We are looking for a highly skilled and experienced Senior DevOps Engineer to design, implement, and manage scalable, secure, and automated infrastructure. The ideal candidate will have strong experience building CI/CD pipelines, managing cloud and containerized environments, and ensuring high availability, security, and performance across development and production systems.
This role is ideal for someone who enjoys solving complex infrastructure challenges, automating processes, supporting machine learning workloads, and collaborating with cross-functional engineering teams to deliver reliable software and ML solutions at scale.
Key Responsibilities
- Design, implement, and maintain robust CI/CD pipelines for automated software delivery.
- Build and manage containerized applications using Docker and Kubernetes.
- Automate infrastructure provisioning, deployment, and configuration management.
- Manage and optimize development, staging, and production environments.
- Support the deployment, scaling, and monitoring of machine learning models and ML platforms.
- Collaborate with data scientists and ML engineers to streamline model training, testing, deployment, and monitoring.
- Monitor system health, performance, and availability using Grafana and other observability tools.
- Configure and maintain local server infrastructure and deployment environments.
- Implement DevOps and MLOps best practices for continuous integration, continuous deployment, infrastructure as code, and automation.
- Ensure infrastructure security, reliability, scalability, and compliance.
- Troubleshoot infrastructure, deployment, networking, database, messaging, and production issues.
- Manage messaging systems and databases supporting high-performance applications and ML workloads.
- Collaborate closely with development, QA, data science, and product teams to improve release cycles and operational efficiency.
Required Experience & Skills
- 7+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or a related field.
- Strong hands-on expertise in:
- CI/CD and test automation
- Infrastructure management and automation
- Cloud and containerized environments
- Infrastructure security
- Monitoring, logging, and observability
- Strong experience with:
- Jenkins
- GitLab and GitHub
- Docker and Kubernetes
- Ansible
- Grafana
- Local server setup and administration
- Kafka and RabbitMQ
- MongoDB
- Strong scripting and automation skills.
- Experience troubleshooting production systems and ensuring high availability.
- Strong experience with machine learning infrastructure, MLOps, or ML model deployment.
- Familiarity with model serving, ML pipelines, experiment tracking, or monitoring ML workloads.
- Excellent problem-solving, analytical, communication, and collaboration skills.
Preferred Qualifications
- Experience with cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform.
- Knowledge of Infrastructure as Code tools such as Terraform.
- Experience implementing security best practices within DevOps and ML pipelines.
- Familiarity with ML frameworks, model serving platforms, or tools such as MLflow, Kubeflow, TensorFlow Serving, or similar technologies.
- Experience with scalable data platforms and distributed systems.
Pay: Up to ₹2,200,000.00 per year
Benefits:
Work Location: In person