Job Description – DevOps / Cloud Infrastructure Engineer
Location: Gurgaon
Experience: 3–6 Years
Role Type: Full-Time | Individual Contributor | Hands-On
Joining: Immediate
About the Role
We are looking for a hands-on DevOps / Cloud Infrastructure Engineer to independently manage our cloud infrastructure, production servers, deployments, and application environments.
The ideal candidate should have strong practical experience in Linux server administration, AWS/Cloud infrastructure, Docker, CI/CD, Auto Scaling, Load Balancing, production deployment, troubleshooting, and infrastructure cost optimization.
This is not a monitoring-only or theoretical DevOps role. The candidate must be capable of setting up servers from scratch, configuring complete application environments, deploying applications, resolving live production issues, scaling infrastructure, and optimizing cloud resources.
Key Responsibilities1. Cloud Infrastructure & Server Management
- Provision and configure Staging, UAT, and Production servers.
- Set up complete application environments from scratch.
- Manage Linux/Ubuntu servers and related infrastructure.
- Configure and maintain Nginx/Apache, PHP/PHP-FPM, MySQL, Redis, SSL, DNS, firewalls/security groups, storage, and backups.
- Maintain secure and properly segregated Staging and Production environments.
- Manage cloud infrastructure using AWS, DigitalOcean, or equivalent cloud platforms.
2. Load Balancing & Auto Scaling
- Configure and manage Load Balancers and Auto Scaling.
- Define scaling thresholds based on CPU, RAM, traffic, connections, and application performance.
- Plan infrastructure capacity for sudden traffic spikes and high-volume events.
- Conduct load/stress testing and optimize infrastructure performance.
- Ensure infrastructure availability, reliability, and scalability.
3. Docker & CI/CD
- Create and maintain Dockerfiles, Docker containers, and Docker Compose configurations.
- Design and maintain CI/CD pipelines for automated Staging and Production deployments.
- Automate application build, testing, Docker image creation, and deployment processes.
- Implement deployment versioning, rollback, and release management mechanisms.
- Troubleshoot deployment and container-related issues.
4. Git & Deployment Automation
- Work extensively with Git and platforms such as Bitbucket, GitHub, or GitLab.
- Configure branch-based deployment workflows.
- Automate repetitive deployment and infrastructure tasks using Shell/Bash scripting.
- Maintain proper version control and deployment practices.
5. Production Troubleshooting & Incident Management
- Independently diagnose and resolve production infrastructure issues, including:
- 502/503/504 errors
- High CPU, RAM, or disk utilization
- Server and network failures
- Nginx/Apache/PHP issues
- Docker/container failures
- Deployment failures
- Database connectivity and connection issues
- Load Balancer and Auto Scaling issues
- Participate in production incident resolution and Root Cause Analysis (RCA).
- Ensure timely resolution of critical infrastructure and deployment issues.
- Monitor system performance and proactively identify potential failures.
6. Cloud Cost Optimization
- Review cloud infrastructure and billing on a monthly basis.
- Identify under-utilized or over-provisioned resources.
- Optimize server, storage, network, and other cloud infrastructure costs.
- Recommend appropriate instance sizing and cost-saving measures.
- Maintain basic monthly infrastructure cost and optimization reports.
Must-Have Skills
- 3–6 years of hands-on DevOps / Cloud Infrastructure experience
- Strong Linux/Ubuntu Server Administration
- Hands-on experience with AWS or equivalent cloud platforms
- Strong knowledge of Docker & Docker Compose
- Experience setting up and managing CI/CD pipelines
- Strong working knowledge of Git
- Experience with Bitbucket, GitHub, or GitLab
- Good knowledge of Nginx/Apache
- Strong Linux Shell/Bash scripting skills
- Experience with Load Balancing and Auto Scaling
- Knowledge of DNS, SSL, Firewalls, and Security Groups
- Basic knowledge of MySQL/database infrastructure
- Strong production troubleshooting and problem-solving skills
- Ability to independently manage production infrastructure
Good to Have
- Terraform / Infrastructure as Code (IaC)
- Ansible
- Kubernetes
- AWS RDS
- AWS CloudFront
- AWS CloudWatch
- Redis
- Prometheus / Grafana
- Load testing and performance optimization
- Disaster Recovery (DR)
- High Availability (HA) architecture
Preferred Candidate Profile
Priority will be given to candidates who have personally handled the complete infrastructure lifecycle, including:
Server Creation → Cloud Configuration → Docker → CI/CD → Deployment → Auto Scaling → Monitoring → Troubleshooting → Cost Optimization
The ideal candidate should be someone who can independently take ownership of a requirement such as:
“Set up a new production environment and make the application production-ready.”
The candidate should be capable of independently planning, configuring, deploying, monitoring, troubleshooting, scaling, securing, and maintaining the complete infrastructure.
Key Competencies
- Strong technical and analytical skills
- Hands-on approach to infrastructure management
- Production-focused troubleshooting ability
- Ability to work independently
- Strong ownership and accountability
- Good understanding of system performance and scalability
- Ability to work under pressure during production incidents
- Good communication and coordination skills
Work Location: In person