Job Title: Lead / Principal MLOps Engineer
Job Location: Remote (Client location Gurugram/Chennai)
Experience Level: 8+ YearsEmployment
Type: Full-time
Department: Artificial Intelligence / Cloud Engineering Infrastructure
About the Role
We are seeking an experienced and battle-tested MLOps Engineer with 8+ years of expertise toown, scale, and maintain our end-to-end Machine Learning and Generative AI infrastructure. Inthis role, you will be responsible for ensuring the high availability, reliability, and continuousperformance of production ML pipelines, automated monitoring systems, and model deploymentWorkflows.You will bridge the gap between Data Science, Data Engineering, and DevOps—building resilient infrastructure that powersproduction AI systems at scale.
Key Responsibilities
1. Production ML Infrastructure & Platform Ownership● Architect, deploy, and manage robust, scalable Machine Learning infrastructure acrossmulti-cloud or hybrid environments (AWS, Azure, or GCP).● Automate CI/CD pipelines for ML models (CT/CD), ensuring seamless modeldeployment, rollback strategies, and zero-downtime releases.● Design and maintain feature stores, model registries, and containerized deploymentruntime environments (Docker, Kubernetes/KServe).2. Pipeline Monitoring & Operational Issue Resolution● Establish real-time telemetry, monitoring, and alerting frameworks to track systemhealth, inference latency, GPU/CPU utilization, and pipeline throughput.● Serve as the primary escalation point for production operational incidents—rapidlydiagnosing and resolving pipeline failures, data drift, and infrastructure bottlenecks.● Conduct root-cause analysis (RCA) for operational outages and implement permanentremediations to safeguard system SLAs.3. System Reliability Performance Optimization● Ensure strict adherence to high-availability (99.9%+ uptime), fault tolerance, and disasterrecovery standards across all ML workflows.● Monitor models in production for concept drift, data drift, and latency degradation,triggering automated retraining and re-deployment workflows.● Optimize resource allocation, cluster auto-scaling, and compute usage to reduce cloudinfrastructure costs without compromising speed or reliability.4. Configuration Management & Security governance● Implement and manage Infrastructure-as-Code (IaC) using Terraform, CloudFormation,or Ansible to enforce configuration consistency across environments.● Manage configuration updates, version control, and secret management for complex,distributed ML and LLM microservices.● Partner with InfoSec teams to enforce data governance, access controls, compliancestandards, and security patches across the ML ecosystem.Requirements & Qualifications
● Experience: 8+ years of professional engineering experience, with at least 4+ yearsdedicated to MLOps, ML Platform Infrastructure, or Cloud Engineering at scale.
● Education: Bachelor’s or Master’s degree in Computer Science, Software Engineering, Information Technology, or a related field.
● Core Technical Expertise:
○ MLOps Frameworks: Hands-on mastery with platforms like MLflow, Kubeflow,Airflow, Weights & Biases, Argo Workflows, or SageMaker.○ Containerization & Orchestration: Advanced expertise in Docker, Kubernetes(EKS/GKE/AKS), Helm charts, and ingress controllers.○ CI/CD & IaC: Strong command of GitHub Actions, GitLab CI, Jenkins, andInfrastructure-as-Code tools like Terraform.○ Monitoring & Observability: Proficiency with Prometheus, Grafana, ELK Stack,Datadog, or specialized ML monitoring tools (Evidently AI, Whylogs, Arize).○ Programming & Scripting: Expert proficiency in Python, Bash, and SQL forautomation, CLI tooling, and service integration.Preferred Qualifications● Experience with LLMOps (deploying, serving, and monitoring Large Language Modelsusing vLLM, Ollama, or Triton Inference Server).● Hands-on experience managing GPU compute clusters, CUDA acceleration, anddistributed inference/training.● Relevant certifications in AWS/GCP/Azure Cloud Architecture or Kubernetes(CKA/CKAD).
Pay: ₹1,200,000.00 - ₹1,600,000.00 per year
Work Location: Remote