We are seeking an experienced Data Engineer with strong expertise in Snowflake, Python, Azure Cloud, Spark, SQL, and Kubernetes to design, build, and maintain scalable cloud-based data platforms. The ideal candidate should have hands-on experience developing modern data pipelines, deploying AI/ML models, implementing CI/CD practices, and working with healthcare data. Experience with Databricks is highly preferred.
The role involves collaborating with data scientists, analysts, and business stakeholders to deliver high-quality, secure, and scalable data solutions that support advanced analytics and machine learning initiatives.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines using Python and SQL.
- Build and optimize enterprise data solutions using Snowflake and Azure cloud services.
- Develop and manage big data processing workflows using Apache Spark.
- Design and optimize data models for analytics and reporting.
- Develop, deploy, and monitor AI/ML models in production environments.
- Build and maintain CI/CD pipelines using GitHub Actions.
- Orchestrate and schedule data workflows using Apache Airflow.
- Deploy and manage containerized data applications using Kubernetes.
- Integrate data from multiple structured and unstructured sources.
- Optimize Snowflake performance, query execution, and warehouse utilization.
- Collaborate with Data Scientists, Data Analysts, and business teams to deliver data-driven solutions.
- Ensure data quality, governance, security, and compliance throughout the data lifecycle.
- Support production issues, troubleshoot data pipelines, and implement performance improvements.
- Work with healthcare datasets while ensuring compliance with data privacy and security standards.
Required Skills
- 4–7 years of experience in Data Engineering.
- Strong experience with Snowflake.
- Proficiency in Python programming.
- Hands-on experience with Azure Cloud services.
- Strong knowledge of Apache Spark.
- Excellent SQL development and query optimization skills.
- Experience with Kubernetes for container orchestration.
- Experience building CI/CD pipelines using GitHub Actions.
- Hands-on experience with Apache Airflow.
- Knowledge of data modeling, ETL/ELT development, and data warehousing concepts.
- Experience deploying and monitoring AI/ML models.
- Strong analytical and problem-solving skills.
- Excellent communication and collaboration abilities.
Preferred Skills
- Hands-on experience with Databricks.
- Experience with Delta Lake and Spark optimization.
- Knowledge of Azure Data Factory (ADF).
- Experience with Azure DevOps.
- Familiarity with REST APIs and microservices.
- Experience with Infrastructure as Code (Terraform/Bicep) is a plus.
- Understanding of data governance, security, and compliance frameworks.
Domain Experience
- Experience working in the Healthcare domain is highly preferred.
- Understanding of healthcare data models, interoperability, and compliance standards is an added advantage.
Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, Data Science, or a related field.
Preferred Candidate Profile
- Strong understanding of modern cloud data architectures.
- Ability to design scalable and high-performance data platforms.
- Experience working in Agile/Scrum environments.
- Self-motivated with excellent troubleshooting and problem-solving skills.
- Strong stakeholder management and cross-functional collaboration skills.
- Passion for building secure, reliable, and data-driven solutions
Work Location: Hybrid remote in Noida, Uttar Pradesh