Job Description – Data Engineer
Job Title: Data Engineer
Experience: 5+ Years
Employment Type: Full-Time
Role Summary
We are seeking an experienced Data Engineer with strong expertise in cloud data engineering, distributed processing, and modern data platforms to design, build, and maintain scalable data solutions. The ideal candidate will have hands-on experience with Snowflake, Python, SQL, Azure, Apache Spark, Databricks, Kubernetes, and Airflow, along with exposure to CI/CD practices and AI/ML model deployment workflows.
The candidate will play a key role in developing enterprise data pipelines, enabling analytics capabilities, optimizing data platforms, and supporting healthcare data initiatives with secure, reliable, and scalable solutions.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines and data processing frameworks.
- Build and optimize data solutions using Snowflake, Azure, Spark, and Databricks.
- Develop robust ETL/ELT workflows to ingest, transform, and process large volumes of structured and unstructured data.
- Write efficient and optimized SQL queries for data extraction, transformation, and analytics.
- Develop Python-based data processing applications and automation scripts.
- Work with Apache Spark for large-scale distributed data processing and performance optimization.
- Design and manage data workflows using Apache Airflow.
- Deploy and manage data applications using Kubernetes and cloud-native technologies.
- Implement CI/CD pipelines using GitHub and related DevOps practices.
- Collaborate with data analysts, data scientists, architects, and business teams to deliver analytics solutions.
- Support AI/ML initiatives by enabling data pipelines and assisting with model deployment workflows.
- Build and maintain data models, datasets, and analytical solutions to support business intelligence requirements.
- Perform data quality checks, validation, monitoring, and troubleshooting of production data pipelines.
- Optimize data platform performance, scalability, and reliability.
- Implement security, governance, and compliance best practices for enterprise healthcare data.
- Participate in architecture discussions, code reviews, and continuous improvement initiatives.
Required QualificationsEducation
- Bachelor's degree in Computer Science, Engineering, Information Technology, Data Science, or related field.
Experience
- 5+ years of experience in Data Engineering roles.
- Strong experience building enterprise-scale data pipelines and analytics platforms.
- Experience working with cloud-based data solutions and distributed processing environments.
- Healthcare domain experience is preferred.
Required Technical SkillsData Engineering & Cloud Platforms
- Snowflake
- Microsoft Azure
- Apache Spark
- Databricks (Preferred)
- Data Lake / Lakehouse architectures
- ETL/ELT development
- Data pipeline design and optimization
Programming & Databases
- Strong programming experience with Python
- Advanced SQL skills
- Experience with data transformation, optimization, and performance tuning
Workflow Orchestration
- Apache Airflow
- Workflow scheduling and dependency management
- Pipeline monitoring and troubleshooting
Containerization & DevOps
- Kubernetes
- Docker (Preferred)
- GitHub
- CI/CD pipeline development and automation
- Version control best practices
AI/ML Engineering
- Experience supporting AI/ML workflows
- Understanding of model deployment processes
- Experience integrating ML models into production data platforms is preferred
Data Analytics
- Strong understanding of data analytics concepts
- Ability to support reporting, visualization, and business intelligence requirements
- Experience working with analytical datasets and data models
Preferred Skills
- Experience working with healthcare data platforms, claims, clinical, member, or provider data.
- Knowledge of healthcare data standards and compliance requirements.
- Experience with Azure services such as Azure Data Factory, Azure Databricks, Azure Storage, and Azure Synapse.
- Experience implementing data governance, security, and data quality frameworks.
- Familiarity with modern cloud-native architecture patterns.
- Experience working in Agile/Scrum environments.
Key Competencies
- Data Engineering & Architecture
- Cloud Data Platforms
- Data Pipeline Development
- Distributed Computing
- Data Analytics
- Performance Optimization
- Problem Solving
- Automation & DevOps Practices
- Data Quality Management
- Collaboration & Communication
- Healthcare Data Understanding
Why Join Us?
- Work on enterprise-scale healthcare data engineering initiatives.
- Build modern cloud-native data platforms using cutting-edge technologies.
- Collaborate with engineering, analytics, and AI/ML teams.
- Opportunity to contribute to scalable, secure, and impactful data solutions.
Work Location: Hybrid remote in Noida, Uttar Pradesh (Noida)