Job Overview
We are looking for an experienced Senior DevOps Engineer to design, implement, and maintain scalable, secure, and highly reliable cloud infrastructure, with a primary focus on Microsoft Azure, Kubernetes, Terraform, CI/CD, and cloud observability.
The ideal candidate will have strong hands-on experience in Infrastructure as Code (IaC), Azure Kubernetes Service (AKS), CI/CD automation, Datadog, cloud security, cost optimization, GitOps, and DevSecOps. You will work closely with development, security, and platform teams to ensure infrastructure and applications meet organizational standards for performance, reliability, security, and compliance.
You will also leverage enterprise-approved AI tools to streamline engineering workflows, automate repetitive tasks, improve operational efficiency, and drive continuous improvement.
Key Responsibilities
- Design, implement, and maintain scalable, secure, and highly available cloud infrastructure primarily on Microsoft Azure.
- Develop and manage Infrastructure as Code (IaC) using Terraform.
- Design, deploy, and manage Azure Kubernetes Service (AKS) and containerized workloads.
- Manage cloud networking, virtual networks, connectivity, security groups, and related infrastructure components.
- Develop, maintain, and optimize CI/CD pipelines using GitHub Actions, Azure DevOps, or similar tools.
- Implement and maintain comprehensive observability solutions using Datadog, Prometheus, or similar platforms.
- Automate the deployment and management of monitoring dashboards, alerts, monitors, and SLOs through Terraform.
- Establish proactive monitoring, logging, alerting, and incident-response capabilities across cloud environments.
- Collaborate with development and security teams to implement DevSecOps practices throughout the software delivery lifecycle.
- Integrate automated security scanning, vulnerability management, compliance validation, and policy-as-code into CI/CD pipelines.
- Implement and maintain GitOps practices using tools such as Argo CD, Flux, or equivalent platforms.
- Ensure cloud infrastructure and applications adhere to organizational security standards and compliance requirements.
- Support compliance initiatives aligned with frameworks such as SOC 2, HIPAA, ISO 27001, or similar standards.
- Monitor cloud infrastructure usage and costs and identify opportunities for cost optimization and FinOps improvements.
- Implement cost-saving strategies while maintaining system performance, availability, security, and reliability.
- Design resilient cloud platforms with appropriate backup, disaster recovery, high availability, and business continuity strategies.
- Define and maintain Service Level Objectives (SLOs) and reliability standards.
- Troubleshoot infrastructure, deployment, networking, performance, and production issues.
- Participate in incident management, root cause analysis, and continuous improvement initiatives.
- Collaborate effectively with engineering, security, product, and other cross-functional teams.
- Utilize approved AI-powered enterprise tools to automate tasks, improve productivity, accelerate troubleshooting, and optimize DevOps processes.
- Maintain technical documentation covering infrastructure, deployment processes, monitoring, security, and operational procedures.
Required Qualifications
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent relevant experience.
- 4+ years of experience in DevOps, Cloud Engineering, Infrastructure Engineering, or a related role.
- Strong hands-on experience with Terraform and Infrastructure as Code (IaC).
- 3+ years of experience working with cloud platforms such as Azure, AWS, or GCP.
- Strong experience with Azure Kubernetes Service (AKS), Kubernetes, container orchestration, and cloud networking.
- 3+ years of experience developing and managing CI/CD pipelines.
- Hands-on experience with GitHub Actions, Azure DevOps, or similar CI/CD platforms.
- Strong understanding of cloud infrastructure architecture, automation, scalability, and reliability.
- Experience with version control systems such as Git.
- Strong troubleshooting, analytical, and problem-solving skills.
- Excellent communication and collaboration skills.
Preferred Qualifications
- Hands-on experience implementing and managing Datadog, Prometheus, Grafana, or similar observability platforms.
- Experience automating monitoring, dashboards, alerts, and infrastructure observability using Terraform.
- Strong understanding of cloud security best practices and compliance frameworks such as SOC 2, HIPAA, and ISO 27001.
- Experience with FinOps and cloud cost optimization.
- Experience implementing GitOps using Argo CD, Flux, or similar tools.
- Strong understanding of DevSecOps, including security scanning, vulnerability management, policy-as-code, and compliance automation.
- Experience designing highly available and resilient cloud platforms.
- Knowledge of disaster recovery, backup strategies, business continuity, and SRE principles.
- Experience defining and managing SLIs, SLOs, and SLAs.
- Experience working in regulated environments such as healthcare, financial services, or other compliance-driven industries.
- Experience with relevant certifications such as:
- HashiCorp Terraform Associate
- Microsoft Azure Solutions Architect
- Certified Kubernetes Administrator (CKA)
- Certified Kubernetes Application Developer (CKAD)
Technical Skills
Cloud: Azure, AWS, GCP
Infrastructure as Code: Terraform
Containers & Orchestration: Kubernetes, AKS, Docker
CI/CD: GitHub Actions, Azure DevOps
Observability: Datadog, Prometheus, Grafana, Monitoring & Alerting
GitOps: Argo CD, Flux
DevSecOps: Security Scanning, Vulnerability Management, Policy-as-Code, Compliance Automation
Cloud Security: IAM, Network Security, Security Controls, Compliance
Networking: Azure Virtual Network, Load Balancing, Connectivity, Network Security
Reliability: High Availability, Disaster Recovery, Backup, SLOs, SRE Practices
Cost Management: FinOps, Cloud Cost Optimization
Version Control: Git
AI: Enterprise-approved AI tools for automation and workflow optimization
Key Competencies
- Cloud Infrastructure & Automation
- Infrastructure as Code
- Kubernetes & Containerization
- CI/CD & Release Automation
- Observability & Monitoring
- DevSecOps & Cloud Security
- GitOps
- Cloud Cost Optimization
- Reliability Engineering
- Disaster Recovery & Business Continuity
- Problem Solving & Troubleshooting
- Cross-functional Collaboration
- Continuous Improvement
Work Location: Remote