We are seeking a highly skilled Senior Data / ML Engineer to join our Data Science & Engineering team. In this role you will design, build, and maintain robust, scalable data pipelines and infrastructure on AWS, powering analytics, machine learning, and business-critical reporting. You will work closely with data scientists, analysts, and product teams to ensure reliable, performant data delivery across the organization.
Kolkata (Rajarhat–Newtown)
We are seeking a highly skilled Senior Data / ML Engineer to join our Data Science & Engineering team. In this role you will design, build, and maintain robust, scalable data pipelines and infrastructure on AWS, powering analytics, machine learning, and business-critical reporting. You will work closely with data scientists, analysts, and product teams to ensure reliable, performant data delivery across the organization.
The ideal candidate combines deep AWS expertise with modern development practices - including AI-assisted coding, containerization, and infrastructure as code - to accelerate delivery without sacrificing quality.
Experience Level: 5+ Years
Location: Kolkata (Rajarhat–Newtown)
Mode of Working: Work from Office
Key Responsibilities:
Design, develop, and optimize large-scale ETL/ELT pipelines using AWS services such as EMR (Spark), Glue, Lambda, and Step Functions.
Orchestrate complex data workflows with Apache Airflow (Amazon MWAA or self-managed), ensuring reliability, observability, and SLA adherence.
Architect and manage data storage solutions across Amazon S3 (data lake), Redshift (data warehouse), and RDS (relational databases), applying best practices for partitioning, compression, and cost optimization.
Build and maintain containerized data applications and microservices using Docker and Amazon ECS/Fargate, including CI/CD automation.
Develop event-driven and serverless data processing solutions with AWS Lambda, SQS, SNS, and EventBridge.
Leverage AI-powered coding assistants and IDE integrations (e.g., Kiro, Cursor, Claude Code) to accelerate development, code review, and documentation.
Implement data quality frameworks, monitoring, and alerting to ensure data integrity across all pipelines.
Collaborate with Data Scientists to productionize ML models and feature pipelines.
Define and enforce data governance, security, and access-control policies in line with organizational and regulatory standards.
Contribute to infrastructure-as-code initiatives using Terraform, CloudFormation, or CDK.
Required Qualifications:
Bachelor's or Master's degree in Computer Science, Engineering, Data Science, or a related field.
5+ years of professional experience in data engineering, with at least 3 years of hands-on AWS production workloads.
3+ years of experience with AWS Services: EMR (Spark/Hadoop), Apache Airflow, S3, Redshift, RDS, Lambda, Sagemaker and ECS.
Solid experience with Docker (building, optimizing, and deploying containers) and container orchestration.
Proficiency in Python and SQL; experience with Typescript is a plus.
Demonstrated ability to use AI-assisted development tools within modern IDEs for rapid prototyping, code generation, testing and debugging.
Strong understanding of data modeling, data governance, and data security best practices.
AWS certifications such as AWS Certified Data Analytics - Specialty or AWS Certified Solutions Architect are a plus.
Experience with infrastructure-as-code tools (Terraform, CloudFormation, CDK).
Exposure to ML Ops workflows, feature stores, or model serving pipelines (e.g., SageMaker).
Knowledge of cost-optimization strategies for large-scale AWS data workloads.
Tech Stack at a Glance:
Compute & Processing: EMR (Spark), Lambda, ECS/Fargate, Glue, Sagemaker
Orchestration: Apache Airflow (MWAA), Step Functions
Storage & Warehousing: S3, Redshift, RDS (PostgreSQL / MySQL)
Containers & DevOps: Docker, ECS, GitHub Actions
AI-Assisted Dev: Cursor, Kiro, Claude Code
Languages: Python, SQL, Bash