AI Powered DevOps Engineer
Location: Remote, India
Job Type: Full-time
Experience: 7+ Years
Department: DevOps / Cloud Engineering / AI
About the Role
We are seeking a senior AI Powered DevOps Engineer to modernize software delivery and infrastructure operations through automation, cloud engineering, and artificial intelligence.
This role is ideal for an experienced DevOps professional who understands traditional CI/CD and cloud infrastructure while also exploring how Generative AI, AI agents, AIOps, and LLM-based automation can improve engineering productivity, observability, incident response, and operational efficiency.
You will work with development, platform, security, and machine learning teams to build resilient infrastructure and intelligent engineering workflows.
Key Responsibilities
- Design, implement, and maintain scalable CI/CD pipelines for enterprise applications.
- Develop automated software delivery workflows using tools such as Jenkins, GitHub Actions, and Azure DevOps.
- Manage infrastructure using Terraform and Terraform Enterprise (TFE).
- Automate configuration and operational tasks using Ansible and Python.
- Build and manage containerized workloads using Docker and Kubernetes.
- Work across AWS, Azure, and GCP environments to support cloud infrastructure and application deployments.
- Develop reusable infrastructure modules and establish infrastructure-as-code standards.
- Improve deployment reliability, speed, security, and operational consistency.
- Integrate AI-powered capabilities into DevOps processes to reduce manual operational effort.
- Explore and implement Generative AI and AI-agent solutions for engineering automation.
- Develop workflows using LLM APIs and frameworks such as LangChain where appropriate.
- Evaluate opportunities for AI-assisted troubleshooting, code analysis, deployment support, and operational automation.
- Contribute to MLOps and LLM deployment practices for AI-based applications.
- Design monitoring and observability solutions using tools such as Prometheus, Grafana, and ELK Stack.
- Analyze infrastructure and application metrics to identify performance and reliability issues.
- Support incident investigation, root-cause analysis, and post-incident improvements.
- Automate repetitive operational tasks using shell scripting and Python.
- Implement appropriate logging, alerting, health checks, and recovery mechanisms.
- Apply security best practices across cloud infrastructure, pipelines, containers, and deployment processes.
- Collaborate with developers and architects to improve application delivery and infrastructure design.
- Document infrastructure architecture, operational processes, and automation solutions.
Required Skills & Experience
- 7+ years of professional experience in DevOps, Cloud Engineering, Platform Engineering, or related disciplines.
- Strong experience designing and managing CI/CD pipelines.
- Hands-on expertise with Jenkins, GitHub Actions, or Azure DevOps.
- Strong knowledge of Terraform, including enterprise infrastructure-management concepts.
- Experience with Ansible and infrastructure automation.
- Strong hands-on experience with Docker and Kubernetes.
- Experience working with one or more major cloud platforms: AWS, Azure, or GCP.
- Strong scripting capabilities using Python and/or Shell.
- Good understanding of Linux systems and cloud infrastructure.
- Experience with monitoring and observability platforms such as Prometheus, Grafana, or ELK.
- Familiarity with Generative AI, LLMs, AI agents, or AI-assisted engineering workflows.
- Understanding of secure software delivery and infrastructure practices.
- Strong troubleshooting and problem-solving skills.
AI & AIOps Responsibilities
The successful candidate will help identify areas where AI can improve DevOps operations, including:
- Intelligent incident analysis.
- Automated log and error investigation.
- AI-assisted root-cause analysis.
- Automated operational recommendations.
- Intelligent infrastructure monitoring.
- LLM-powered engineering assistants.
- AI-supported deployment and release workflows.
- Automated documentation and knowledge retrieval.
- Predictive identification of infrastructure or application issues.
Preferred Skills
- Experience integrating OpenAI APIs or comparable LLM services.
- Knowledge of LangChain or similar AI-agent frameworks.
- Experience deploying LLM-powered services.
- Familiarity with MLOps platforms and machine learning deployment workflows.
- Understanding of Kubernetes-based AI workloads.
- Experience implementing AI-powered observability or AIOps solutions.
- Knowledge of cloud security and identity-management concepts.
- Experience with service meshes or cloud-native networking.
- Familiarity with GitOps methodologies.
Pay: ₹899,000.00 - ₹1,300,000.00 per year
Work Location: Remote