We are seeking an experienced Senior Big Data Engineer with expertise in Scala, PySpark, AWS Glue, Databricks, Delta Lake, EMR, Snowflake Migration, and Big Data Pipeline Development. The ideal candidate will be responsible for designing, developing, and optimizing scalable data pipelines, modernizing enterprise data platforms, and supporting cloud-based analytics and AI initiatives.
The candidate should have strong experience in distributed data processing, ETL/ELT development, data quality frameworks, performance optimization, and cloud-native data engineering. Experience leveraging AI coding assistants to accelerate development and improve engineering productivity is highly desirable.
Key ResponsibilitiesBig Data Pipeline Development
- Design, develop, and maintain scalable batch and streaming data pipelines.
- Build high-performance ETL/ELT workflows for large-scale data processing.
- Develop reusable and maintainable data engineering frameworks.
- Integrate structured, semi-structured, and unstructured data from multiple sources.
- Ensure high availability, scalability, and reliability of data pipelines.
Databricks & Delta Lake
- Design and implement modern Lakehouse architectures using Databricks and Delta Lake.
- Build and optimize Spark workloads using Scala and PySpark.
- Implement Delta Live Tables, schema evolution, and ACID-compliant data processing.
- Optimize Delta tables through partitioning, indexing, and file compaction.
- Support analytics, reporting, and machine learning workloads.
AWS Data Engineering
- Develop data integration solutions using AWS Glue and Amazon EMR.
- Build serverless ETL jobs and orchestrate large-scale data processing workflows.
- Integrate with Amazon S3, IAM, CloudWatch, and other AWS services.
- Optimize cloud infrastructure for performance and cost efficiency.
- Support cloud migration and modernization initiatives.
Snowflake Migration
- Lead migration of enterprise data warehouses to Snowflake.
- Design migration strategies, data transformation logic, and validation processes.
- Optimize data models and workloads for Snowflake performance.
- Ensure data integrity, security, and minimal downtime during migration.
- Collaborate with business and infrastructure teams throughout migration projects.
Data Validation & Profiling
- Perform data profiling, reconciliation, and validation across enterprise datasets.
- Implement automated validation and reconciliation processes.
- Identify and resolve data inconsistencies and quality issues.
- Develop reusable validation frameworks for production pipelines.
- Support governance and compliance initiatives through data quality monitoring.
Performance Optimization
- Optimize Spark applications, SQL queries, and distributed processing workloads.
- Improve job execution time through partitioning, caching, and resource tuning.
- Monitor pipeline performance and resolve bottlenecks.
- Implement scalable data storage strategies using Parquet and Delta formats.
- Continuously improve system efficiency and operational reliability.
AI-Assisted Development
- Leverage AI Coding Assistants to improve development productivity and code quality.
- Apply AI-assisted code generation, testing, debugging, and documentation practices.
- Evaluate emerging AI tools and integrate them into engineering workflows.
- Promote engineering best practices while maintaining secure and reliable code.
Collaboration & Leadership
- Work closely with Data Architects, Data Scientists, Analysts, and DevOps teams.
- Participate in architecture discussions and technical design reviews.
- Mentor junior engineers and establish data engineering best practices.
- Contribute to Agile/Scrum ceremonies and continuous improvement initiatives.
Required SkillsProgramming
Big Data Technologies
- Apache Spark
- Databricks
- Delta Lake
- AWS Glue
- Amazon EMR
- Hadoop Ecosystem (Preferred)
Data Engineering
- Big Data Pipeline Development
- ETL/ELT Development
- Data Validation
- Data Profiling
- Data Transformation
- Data Integration
Data Warehousing
- Snowflake
- Snowflake Migration
- Data Lakehouse
- Data Warehouse Modernization
File Formats
- Parquet
- Delta Format
- Avro (Preferred)
- ORC (Preferred)
Databases
- PostgreSQL
- SQL Server
- MySQL
- MongoDB (Preferred)
Cloud Platforms
- AWS
- Amazon S3
- IAM
- CloudWatch
- Databricks on AWS
Data Quality
- Data Quality Frameworks
- Data Governance
- Metadata Management (Preferred)
- Data Lineage (Preferred)
DevOps & Tools
- Git
- Jenkins
- Azure DevOps
- GitHub Actions
- CI/CD Pipelines
AI Productivity Tools
- AI Coding Assistants (GitHub Copilot, Amazon Q, Cursor, or similar)
- Prompt Engineering for Development
- AI-Assisted Code Review
Preferred Qualifications
- Bachelor's or Master's degree in Computer Science, Data Engineering, Information Technology, or a related field.
- 5–10 years of experience in Big Data Engineering and cloud-based data platforms.
- Strong expertise in Scala, PySpark, AWS Glue, Databricks, and Delta Lake.
- Experience with Amazon EMR, Snowflake migration, and enterprise-scale ETL development.
- Strong knowledge of distributed computing, Spark optimization, and data quality frameworks.
- Experience working with Agile/Scrum development methodologies.
- Excellent analytical, troubleshooting, and communication skills.
Preferred Certifications
- Databricks Certified Data Engineer Professional
- AWS Certified Data Engineer – Associate
- AWS Certified Solutions Architect – Associate (Preferred)
- Snowflake SnowPro Core Certification
- Microsoft Certified: Azure Data Engineer Associate (DP-203) (Preferred)
- Apache Spark Developer Certification (Preferred)
Work Location: Hybrid remote in Noida, Uttar Pradesh