Project Role : Custom Software Engineer
Project Role Description : Lead the effort to design, build and configure applications, acting as the primary point of contact.
Must have skills : PySpark
Good to have skills : NA
Minimum
3 year(s) of experience is required
Educational Qualification : 15 years full time education
Summary:
As a Custom Software Engineer, a typical day involves leading the design, development, and configuration of software applications. This role requires taking ownership of the project s technical direction and serving as the main liaison for all application-related activities. The position demands close collaboration with various stakeholders to ensure that the solutions meet business needs while maintaining high standards of quality and performance. The engineer actively manages the application lifecycle, from initial concept through deployment and ongoing support, ensuring seamless integration and functionality within the broader system environment.
Roles & Responsibilities:
ETL/ELT development
Data modeling
Data warehousing concepts
Lakehouse architecture
Batch and streaming pipelines
Performance Optimization
Query optimization
Execution plan analysis
Adaptive Query Execution (AQE)
Join optimization
Partition pruning
Broadcast joins
Skew handling
Caching strategies
File optimization
Cluster resource tuning
Storage Formats
Delta Lake
Apache Iceberg
Parquet
ORC
CI/CD & DevOps
Git
CI/CD pipeline
Docker/Kubernetes (preferred)
Preferred Qualifications
Experience with Databricks, IoMete, EMR, Synapse, or Microsoft Fabric.
Knowledge of Kafka or event-driven architectures.
Experience building data products for AI/ML workloads.
Familiarity with data governance and metadata management.
Experience
5–10 years in Data Engineering.
3+ years of hands-on experience with PySpark and distributed computing.
Proven experience processing large-scale datasets (multi-TB/PB).
Strong track record in Spark performance tuning, query optimization, and building enterprise-grade data products.
Professional & Technical Skills:
- Must To Have Skills: Proficiency in PySpark.
- Strong knowledge of distributed computing principles and big data processing frameworks.
- Experience in writing efficient, reusable, and maintainable code for large-scale data processing.
- Familiarity with data integration techniques and working with various data sources.
- Ability to troubleshoot and optimize complex data workflows and pipelines.
- Understanding of software development lifecycle and version control systems.
Additional Information:
Key Competencies
PySpark development
Spark performance tuning
High-volume data processing
Complex SQL optimization
Distributed systems engineering
Data product engineering
Data modeling
Lakehouse architecture
Cloud data engineering
Production support and troubleshooting
Performance benchmarking and scalability optimization