Chennai, Tamil Nadu
Job Summary
We are seeking a highly skilled Senior Observability Engineer to lead the evolution of our enterprise observability platform. In this role, you will transition our systems to an open, standards-based observability architecture using OpenTelemetry (OTel). You will design and scale OTel solutions across polyglot microservices (Java, Python, etc.) and Google Cloud Platform (GCP) resources, routing telemetry to Dynatrace, Prometheus, and Jaeger. A major focus of this role is improving the overall developer experience by creating onboarding tools, documentation, migration paths, and automated Quality of Service (QoS) reporting.
Key Responsibilities
Key Responsibilities Expand Observability Coverage: Design and implement standardized OpenTelemetry (OTel) SDK and API configurations for polyglot environments, focusing on Python and Java. Architect and deploy OTel-based collection mechanisms for GCP-native resources, including serverless components (Cloud Functions) and asynchronous messaging (Pub/Sub). Ensure all OTel-ingested metrics, traces, and logs are seamlessly mapped, enriched, and visualized within Dynatrace using native OTLP ingestion. Enhance Interoperability: Design, deploy, and maintain high-availability OpenTelemetry Collector pipelines to receive, process, batch, and export telemetry data. Configure OTel pipelines to dynamically route telemetry to multiple backends (e.g., exporting traces to Dynatrace and Jaeger, and metrics to Prometheus). Establish unified semantic conventions, tagging schemas, and context propagation standards across all services. Improve User Experience & Developer Enablement: Build a frictionless onboarding experience for software engineering teams by creating reusable OTel templates and bootstrap libraries. Create automated migration tools and scripts to help teams transition from legacy/proprietary agents to OpenTelemetry. Design and build automated pipelines to generate customizable Quality of Service (QoS) and Service Level Objective (SLO) reports using Dynatrace and OTel APIs.
Skill Requirements
These are the specific, hands-on technical tools, languages, and frameworks required to perform the daily duties of the role. OpenTelemetry Ecosystem: Deep expertise in the OTel Collector architecture, OTel SDKs/APIs, OTLP protocol, auto-instrumentation agents, and semantic conventions. Observability Backends: Advanced administration of Dynatrace (including Smartscape, Davis AI, and OTLP ingestion), as well as open-source tools like Prometheus and Jaeger. Programming Languages: High proficiency in writing, debugging, and instrumenting code in Python and Java. Cloud Infrastructure: Hands-on experience monitoring Google Cloud Platform (GCP) resources, specifically GKE, Cloud Functions, and Pub/Sub. Infrastructure as Code (IaC) & CI/CD: Experience writing Terraform to deploy observability infrastructure and integrating instrumentation into CI/CD pipelines (e.g., GitHub Actions, GitLab CI, or Jenkins).
Other Requirements
These are the formal credentials, education, certifications, and professional history that validate your technical background. Professional Experience: 5+ years of experience in Site Reliability Engineering (SRE), DevOps, or Platform Engineering with a heavy focus on observability. Education: Bachelor’s degree in Computer Science, Software Engineering, Information Technology, or a related technical field (or equivalent practical experience). Certifications (Highly Preferred): Dynatrace Certified Professional or Associate. GCP Professional Cloud DevOps Engineer or Cloud Architect. Certified Kubernetes Administrator (CKA). Industry Contributions (Plus): Active contributions to the Open Telemetry open-source project or related CNCF communities.
#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-