Job Summary
We are looking for an experienced L2 Production Support Engineer to ensure the stability, availability, and performance of production applications. The role involves advanced troubleshooting, incident resolution, and close coordination with L3/development teams to minimize downtime and improve system reliability.
Key Responsibilities
-
Provide Level 2 support for production systems and applications
- Monitor application health, performance, and availability using monitoring tools
- Troubleshoot and resolve complex incidents within SLA timelines
- Perform root cause analysis (RCA) for major incidents and recurring issues
- Analyze logs, application behavior, and database queries to identify problems
- Execute data fixes, configuration changes, and minor enhancements where required
- Coordinate with L3/development teams for code-level fixes and deployments
- Support release activities, patch deployments, and production rollouts
- Maintain and improve runbooks, SOPs, and knowledge base articles
- Ensure compliance with incident, problem, and change management processes
Participate in on-call rotations and handle critical production issues
-
Required Skills & Qualifications
-
4–6 years of experience in production/application support (L2)
- Strong SQL skills for data analysis and troubleshooting
- Experience with relational databases (Oracle, SQL Server, PostgreSQL, etc.)
- Proficiency in log analysis and debugging
- Knowledge of Linux/Unix or Windows environments
- Familiarity with ticketing tools (ServiceNow, Jira, Remedy)
- Understanding of web applications, APIs, and microservices architecture
- Experience with monitoring tools (Splunk, AppDynamics, Dynatrace, etc.)
- Strong analytical, troubleshooting, and problem-solving skills
Good communication skills and ability to work under pressure
-
Preferred Qualifications
-
Experience in 24/7 production support environments
- Knowledge of scripting (Shell, Python, PowerShell)
- Exposure to cloud platforms (AWS, Azure, GCP)
- Familiarity with CI/CD pipelines and DevOps practices
Understanding of ITIL processes or ITIL certification
-
Key Deliverables
-
Timely resolution of production incidents
- Root cause analysis (RCA) reports
- System health and performance reports
Updated documentation (runbooks, SOPs)
-
Key Metrics / KPIs
-
SLA adherence (MTTR, response time)
- Incident resolution rate
- Reduction in recurring incidents
- System uptime/availability
Customer satisfaction
-
Nice-to-Have Skills
-
Experience with batch job scheduling tools (Control-M, Autosys)
- Exposure to middleware systems (Kafka, MQ, etc.)
- Basic understanding of application code (Java, Python etc.)
- Automation of repetitive operational tasks