Role description
Role Title Kubernetes OpenTelemetry OTel Engineer SRE
We are seeking an OpenTelemetry OTel Engineer with a focus on Site Reliability Engineering SRE to support our observability platform running on Kubernetes This is an SREfocused rolethe primary objective is ensuring the reliability availability and performance of the OpenTelemetry collection and pipeline infrastructure A working understanding of OpenTelemetry and Kubernetes is essential to operate monitor and troubleshoot the telemetry platform effectively
Key Responsibilities
SRE Operational Reliability Primary Focus
Monitor the health throughput and performance of OpenTelemetry collectors and pipelines proactively address bottlenecks
Respond to incidents INCs troubleshoot telemetry dataflow issues and drive timely resolution
Participate in oncall rotation and support major incident management
Define and track SLOsSLIs for telemetry pipeline reliability and availability
Contribute to root cause analysis RCA and implement preventive measures
Automate routine operational tasks to reduce manual effort scriptingconfig management
Support capacity planning and forecasting for telemetry data growth
Maintain resilience through failover recovery and pipeline redundancy procedures
OpenTelemetry Platform Support
Deploy configure and maintain OpenTelemetry Collectors receivers processors exporters on IKPKubernetes
Support instrumentation of applications for traces metrics and logs
Manage telemetry pipeline configuration sampling batching and data routing to backends
Operate and maintain OTel components running on Kubernetes pods deployments config maps secrets
Ensure reliable data delivery to observability backends eg Prometheus Grafana Splunk Jaeger
Collaboration Continuous Improvement
Work with development infrastructure and platform teams to onboard new services into observability
Document runbooks standard operating procedures and known error solutions
Identify and implement improvements to enhance stability resilience and efficiency
Required Skills Experience
Solid understanding of SRE principles reliability monitoring incident management automation
Working knowledge of OpenTelemetry concepts traces metrics logs collectors exporters
Experience operating workloads on Kubernetes IKP or equivalent
Familiarity with LinuxUnix systems and commandline operations
Scripting skills Bash Python for automation and operational tooling
Experience with monitoringobservability tools Prometheus Grafana Splunk
Awareness of ITIL processes incident problem and change management
Strong troubleshooting and problemsolving abilities
Good communication and documentation skills.
Skills
Mandatory Skills : Kubernetes
About LTM
LTM is an AI-centric global technology services company and the Business Creativity partner to the world’s largest and most disruptive enterprises. We bring human insights and intelligent systems together to help clients create greater value at the intersection of technology and domain expertise. Our capabilities span integrated operations, transformation, and business AI — enabling new ways of working, new productivity paradigms, and new roads to value. Together with over 87,000 employees across 40 countries and our global network of partners, LTM — a Larsen & Toubro company — owns business outcomes for our clients, helping them not just outperform the market, but to Outcreate it. Please also note that neither LTM nor any of its authorized recruitment agencies/partners charge any candidate registration fee or any other fees from talent (candidates) towards appearing for an interview or securing employment/internship. Candidates shall be solely responsible for verifying the credentials of any agency/consultant that claims to be working with LTM for recruitment. Please note that anyone who relies on the representations made by fraudulent employment agencies does so at their own risk, and LTM disclaims any liability in case of loss or damage suffered as a consequence of the same. Recruitment Fraud Alert - https://www.ltimindtree.com/recruitment-fraud-alert/