Location: Kolkata
Employment Type: Internship
Duration: 3 months
Shift: Rotational (Day/Night)
Department: Engineering / Infrastructure
Role Overview: We are looking for a motivated SRE / DevOps Intern to support the monitoring, performance analysis, and optimization of our APIs and cloud infrastructure. The intern will work closely with the Development, Backend, and DevOps teams to monitor application and infrastructure health, identify performance and capacity requirements, configure alerts, and contribute to automation initiatives.
This role offers hands-on exposure to modern observability and cloud infrastructure technologies, including Prometheus, Grafana, Kubernetes, Docker, cloud platforms, and API monitoring tools.
Key Responsibilities
- Monitor API performance metrics, including Requests Per Second (RPS), latency, throughput, traffic patterns, and error rates.
- Monitor infrastructure and system utilization, including CPU, memory, network, storage, pods, containers, and other resource metrics.
- Build, maintain, and improve Grafana dashboards for API, application, and infrastructure monitoring.
- Configure and maintain alerts for critical events such as traffic spikes, increased latency, elevated error rates, and resource saturation.
- Analyze traffic and utilization patterns to identify capacity requirements and potential scaling needs.
- Coordinate with Development and DevOps teams when infrastructure needs to scale up or scale down based on application demand.
- Identify underutilized infrastructure and provide recommendations for resource right-sizing and cost optimization.
- Monitor API and infrastructure performance during application deployments, releases, and periods of high traffic.
- Assist in investigating performance, availability, and capacity-related issues.
- Automate monitoring, reporting, and alerting processes using Python, Bash, or other scripting tools, wherever applicable.
- Prepare and maintain daily and weekly reports covering API traffic, RPS, system performance, and infrastructure utilization.
- Collaborate closely with Backend Developers, Software Engineers, and DevOps teams to help identify and resolve performance or capacity-related issues.
- Continuously learn and gain practical exposure to modern observability, cloud, and infrastructure technologies.
Preferred Technical Skills:
- Basic understanding of Linux operating systems and command-line environments.
- Basic knowledge of REST APIs, HTTP/HTTPS, API requests, and response status codes.
- Familiarity with cloud infrastructure concepts and platforms such as AWS, Azure, or Google Cloud Platform (GCP).
- Basic understanding of Docker and containerization.
- Exposure to or interest in Kubernetes and container orchestration.
- Familiarity with monitoring and observability tools such as Prometheus and Grafana is preferred.
- Basic knowledge of Python and/or Bash scripting for automation.
- Understanding of fundamental networking concepts such as TCP/IP, DNS, ports, and HTTP would be an advantage.
- Familiarity with API gateways, logging, or infrastructure monitoring tools is a plus.
Preferred Candidate Profile:
- Currently pursuing or recently completed a degree in Computer Science, Information Technology, Software Engineering, or a related technical discipline.
- Strong analytical and problem-solving skills.
- Interest in DevOps, Site Reliability Engineering (SRE), Cloud Infrastructure, and System Monitoring.
- Ability to interpret technical metrics and identify potential performance trends or anomalies.
- Willingness to learn new tools and technologies in a fast-paced engineering environment.
- Good communication and collaboration skills.
Benefits:
Education:
Language:
Shift availability:
- Night Shift (Preferred)
- Day Shift (Preferred)
Work Location: In person