About Nerdience
Nerdience Technologies Pvt. Ltd. is a technology company specializing in Platform Engineering, DevOps, Cloud Infrastructure, Site Reliability Engineering (SRE), DevSecOps, FinOps, MLOps, and AI-driven solutions. We help organizations build secure, scalable, and highly reliable cloud-native platforms while accelerating software delivery through automation and modern engineering practices.
Job Summary
We are looking for a passionate Site Reliability Engineer (SRE) with 2 years of hands-on experience in cloud infrastructure, Kubernetes, CI/CD, Linux, and automation. You will work closely with development teams to build reliable, scalable, and secure platforms while ensuring high availability, performance, and operational excellence. Modern SRE roles commonly combine cloud operations, observability, automation, incident management, and infrastructure engineering.
Key Responsibilities
- Design, implement, and maintain highly available cloud infrastructure on AWS, Azure, or Google Cloud.
- Manage and administer Kubernetes clusters in production environments.
- Build, optimize, and maintain CI/CD pipelines using GitHub Actions, GitLab CI, or Jenkins.
- Automate infrastructure provisioning using Terraform, Ansible, or similar Infrastructure as Code (IaC) tools.
- Monitor applications and infrastructure using Prometheus, Grafana, ELK Stack, Datadog, or similar observability platforms.
- Investigate production incidents, perform root cause analysis (RCA), and implement preventive measures.
- Configure monitoring, alerting, logging, and dashboards to improve system reliability.
- Improve system scalability, availability, security, and performance.
- Collaborate with software engineers to ensure reliable application deployments.
- Implement backup, disaster recovery, and high-availability strategies.
- Optimize cloud infrastructure costs and resource utilization.
- Develop automation scripts using Python, Bash, or Go to eliminate repetitive operational tasks.
- Participate in on-call rotations and production support when required.
- Maintain technical documentation, SOPs, and operational runbooks.
Required Skills
- 2 years of experience in Site Reliability Engineering, DevOps, or Platform Engineering.
- Strong knowledge of Linux system administration.
- Experience with AWS, Microsoft Azure, or Google Cloud Platform.
- Hands-on experience with Docker and Kubernetes.
- Experience building and maintaining CI/CD pipelines.
- Good understanding of Infrastructure as Code using Terraform.
- Knowledge of Git and version control workflows.
- Experience with monitoring and observability tools such as Prometheus, Grafana, ELK Stack, or Datadog.
- Experience with scripting using Bash or Python.
- Understanding of networking concepts including DNS, HTTP/HTTPS, Load Balancers, TCP/IP, SSL/TLS, and Firewalls.
- Familiarity with microservices architecture.
- Strong troubleshooting and analytical skills.
- Excellent communication and collaboration skills.
Preferred Skills
- Experience with GitOps tools such as ArgoCD or FluxCD.
- Knowledge of Helm and Kubernetes Operators.
- Experience with DevSecOps practices and security scanning tools.
- Familiarity with SLOs, SLIs, and Error Budgets.
- Experience with PostgreSQL, MySQL, Redis, or MongoDB administration.
- Cloud certifications such as AWS Certified Solutions Architect, AWS SysOps Administrator, Kubernetes (CKA/CKAD), or Terraform Associate are an advantage.
What We Offer
- Opportunity to work on modern cloud-native technologies.
- Exposure to Platform Engineering, DevOps, SRE, FinOps, DevSecOps, and AI infrastructure.
- Collaborative engineering culture with continuous learning opportunities.
- Career growth based on performance.
- Flexible work environment.
- Competitive salary and benefits.
Pay: Up to ₹25,000.00 per month
Benefits:
- Flexible schedule
- Paid sick time
Work Location: In person