Role Summary
We are looking for a Senior Observability Architect to lead the design, implementation, and maturity of metrics, alerting, and monitoring platforms across multiple client environments. As a senior, hands-on member of the team, this person will both set architectural direction and build the solution directly. The ideal candidate will own the observability strategy for current and future engagements, standardize how the organization designs, deploys, and operates Prometheus and Grafana stacks at scale, work directly with client stakeholders, and provide technical leadership to delivery teams.
Key Responsibilities
-
Own end-to-end solution architecture for Prometheus and Grafana based observability platforms, from greenfield design through to operational handover.
-
Define reference architectures and reusable patterns that can be applied consistently across multiple clients and deployment environments.
-
Architect alerting and event-correlation pipelines, including complex one-to-many correlation where a single infrastructure failure maps to multiple impacted downstream services or consumers.
-
Design and maintain the mapping between monitored assets and the services they support, including integration with CMDB or equivalent systems of record.
-
Integrate observability platforms with ITSM and ticketing systems using webhooks, APIs, and middleware, ensuring alerts flow reliably into escalation and incident workflows.
-
Lead long-term storage and scalability design using Grafana Mimir, Thanos, or Cortex, including remote-write ingestion and object-storage backends.
-
Deploy and operate observability stacks in Kubernetes environments, including the kube-prometheus-stack, Prometheus Operator, kube-state-metrics, and associated exporters.
-
Establish organizational standards for instrumentation, exporters, service discovery, dashboarding, and alert hygiene.
-
Provide senior technical escalation (L2 and above) and act as the authoritative technical owner on observability questions.
-
Mentor engineers and raise the observability capability of delivery teams across the organization.
Required Skills and Experience
-
12-15 years in infrastructure, platform, or site reliability engineering, with deep hands-on specialization in Prometheus and Grafana based observability.
-
Expert-level Prometheus, including PromQL, recording and alerting rules, TSDB internals, retention, and scaling limitations.
-
Expert-level Grafana, including dashboard design, alerting, data source configuration, and its integration surface (webhooks, alert payload structure, and APIs).
-
Strong command of the exporter ecosystem, including Node Exporter, cAdvisor, Blackbox Exporter, and application instrumentation using Prometheus client libraries.
-
Alertmanager configuration, including deduplication, grouping, silencing, and routing to receivers such as email, Slack, PagerDuty, and Opsgenie.
-
Service discovery across Kubernetes and cloud providers.
-
Long-term storage architecture with Mimir, Thanos, or Cortex, including remote-write and object storage (S3, GCS, Azure Blob, or MinIO).
-
Kubernetes-native deployment and operations, including Helm and the Prometheus Operator.
-
Integration engineering with ITSM platforms, webhooks, and middleware for automated ticketing and escalation.
-
Solid Linux, networking, and general infrastructure fundamentals.
Nice to Have
-
Experience designing observability for managed-service or multi-tenant environments where infrastructure is shared across multiple downstream consumers.
-
Experience with GPU or high-performance compute infrastructure monitoring.
-
Familiarity with hardware-level telemetry, for example Redfish based exporters.
-
Experience with hybrid and distributed cloud platforms, including managed Kubernetes and on-premises or edge Kubernetes distributions.
-
Scripting and automation (Python, Go, or Bash) for tooling and integration.
-
Infrastructure-as-code experience (Terraform or similar).
-
Certifications such as Prometheus Certified Associate, Grafana certifications, or CKA.
Soft Skills
-
Ability to translate ambiguous client requirements into clear, defensible architecture.
-
Strong written communication, including capability notes, reference architectures, and integration documentation.
-
Comfort working directly with client technical teams and presenting to mixed technical and stakeholder audiences.
-
Self-directed ownership of complex technical problems from definition through resolution.