Job Summary
We are seeking an experienced Senior Data Engineer with strong expertise in Snowflake, Python, Azure Cloud, Spark, SQL, and Kubernetes to design, develop, and maintain scalable data platforms and analytics solutions. The ideal candidate will have hands-on experience in cloud-native data engineering, data pipeline orchestration, CI/CD automation, and AI/ML model deployments. Experience with Databricks and the Healthcare domain will be highly preferred.
Key Responsibilities
- Design, develop, and optimize scalable ETL/ELT data pipelines using Snowflake, Python, and Apache Spark.
- Build and maintain cloud-based data platforms on Microsoft Azure.
- Develop and manage data ingestion, transformation, and loading processes for structured and semi-structured datasets.
- Implement and optimize data warehouse solutions in Snowflake for high-performance analytics.
- Orchestrate workflows using Apache Airflow for reliable and automated data processing.
- Deploy and manage containerized applications using Kubernetes.
- Build CI/CD pipelines using GitHub Actions to automate code deployment and infrastructure changes.
- Collaborate with Data Scientists to deploy, monitor, and maintain AI/ML models in production environments.
- Implement best practices for data quality, governance, security, and performance optimization.
- Create scalable analytics datasets to support business intelligence and reporting.
- Troubleshoot production issues and optimize system performance across cloud environments.
- Work closely with cross-functional teams, including Data Science, Business Intelligence, Product, and Engineering.
- Ensure compliance with Healthcare data standards, security, and privacy regulations where applicable.
Required Technical Skills
- Snowflake
- Python
- SQL
- Apache Spark
- Microsoft Azure
- Kubernetes
- GitHub (CI/CD & GitHub Actions)
- Apache Airflow
- Data Analytics
- AI/ML Model Deployment
- REST APIs
- Data Warehousing Concepts
- ETL/ELT Pipeline Development
- Version Control (Git)
Preferred Skills
- Databricks
- Delta Lake
- Azure Data Factory (ADF)
- Azure Data Lake Storage (ADLS Gen2)
- MLflow
- Docker
- Terraform
- Azure DevOps
- Healthcare domain experience (HIPAA/FHIR/HL7 knowledge is an added advantage)
Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, Data Engineering, or a related field.
- 5–8 years of experience in Data Engineering, Cloud Data Platforms, or Analytics Engineering.
- Strong problem-solving, analytical, and debugging skills.
- Excellent communication and stakeholder management abilities.
- Experience working in Agile/Scrum environments.
Preferred Experience
- Hands-on experience building enterprise-scale data platforms.
- Experience deploying machine learning models into production.
- Exposure to healthcare data ecosystems and regulatory compliance.
- Experience with large-scale data processing and cloud-native architectures.
Work Location: Remote