we are looking for the candidates having experience in SRE & Observability Engineer for a high impact enterprise monitoring role. you will be responsible for architecture and managing the entire observability landscape-building unified dashboards, consolidating monitoring tools, implementing automation and ensuring enterprise -wide system reliability across on premise and cloud environment
Key Responsibilities
- Unified Dash boarding: Architect and implement single-pane-of-glass reporting using Grafana/Datadog, integrating data from Splunk, AWS CloudWatch, Azure Monitor, ThousandEyes, Imperva, and Ansible.
- Tools Assessment & Consolidation: Evaluate the current IT landscape, including Datadog, Splunk, Dynatrace, New Relic, AppDynamics, Zabbix, ServiceNow, SMART, SolarWinds, Prometheus, Glowroot, CAMS, etc., to drive consolidation and optimize the target state.
- Enterprise Discovery: Deploy and maintain Device42 for automated infrastructure discovery, application dependency mapping, and continuous CI synchronization with ServiceNow.
- Backend Automation: Implement ticket-based and non-ticket-based routine operational automations leveraging Ansible.
Requirements
- Expert-level knowledge in Monitring Rule setting up and configuring of Dynatrace, App Dynamics, New Relic, Zabbix, Data dogSolarWinds, Prometheus.
- Strong experience building Unified Dashboards using Grafana ,pulling telemetry from diverse sources including Log/SIEM (Splunk), and Infrastructure.
- Deep experience conducting Enterprise Tools Assessments and integrating alerts/metrics from legacy APM, network, and infrastructure monitoring tools
- Enterprise Monitoring & Observability Tools for Onprem & Cloud : (Datadog, New Relic, AppDynamics, Zabbix, SolarWinds Grafana), Log/SIEM (Splunk),
- Broad familiarity supporting and assessing monitoring and observability landscapes (e.g., Device42, SMART, Glowroot, CAMS, Thousand Eyes, Thanos, OpsRamp, AWS CloudWatch, Azure Monitor).
- Hands-on experience with backend automation and orchestration using Ansible, and configuring Major Incident Management (MIM) notifications via PagerDuty.
- ITIL V4 Foundation certification with good process knowledge in Incident, Problem, and Change Management.
- Flexible to work in Off hour / Weekend shift. Excellent verbal and written communication skills.
Pay: ₹1,500,000.00 - ₹3,000,000.00 per year
Benefits:
- Paid sick time
- Paid time off
- Provident Fund
Work Location: In person