Job Title: Site Reliability Engineer
Location: Remote / India
Experience Level: Senior (6–8 years)
Role Type: Contract (Fixed Term)
Working Hours: Primarily in US timezone, details TBD
Position Overview:
We are looking for a Senior Site Reliability Engineer to join our global SRE team. This is a hands-on individual contributor role for an experienced engineer who can take ownership of high-severity incidents, and drive reliability improvements across a fast-moving fintech infrastructure. You will serve in the on-call rotation, work closely with global engineering teams to maintain system health and improve the resilience of our platforms.
Key Responsibilities:
On-Call & Incident Command: Serve in the on-call rotation and act as incident lead for incidents: drive triage, coordinate cross-team response, communicate status to stakeholders, and own the incident through resolution. Engage management or service owners when incidents require architectural decisions or business-level trade-offs. Incident Mitigation: Independently triage and resolve production alerts across AWS infrastructure and application layers. Apply sound judgment to distinguish transient issues from systemic failures and act accordingly.
Monitoring & Alerting Strategy: Own and evolve Datadog monitoring standards: design monitors, dashboards, and SLO/SLI frameworks, and drive signal-to-noise improvements to reduce alert fatigue across the team.
Runbook & SOP Authorship: Create runbooks from scratch for new failure modes, update existing SOPs to reflect operational learnings, and lead post-mortems with structured root cause analysis and actionable follow-ups.
Reliability Initiatives: Lead initiatives to reduce toil and operational inefficiency: identify systemic weaknesses, design and implement automation or process improvements, and see them through to adoption.
Root Cause Analysis: Lead RCA investigations for infrastructure and application-level failures in the AWS environment. Produce clear, action-oriented incident reports that drive lasting fixes.
Mentorship & Review: Guide junior and mid-level on-call engineers through complex investigations, review their runbooks and incident reports, and help raise the team's troubleshooting standards.
Deployment Support: Monitor CI/CD pipelines during deployments, flag reliability risks, and initiate rollbacks following established procedures when stability is at risk.
Qualifications :
Experience: 6–8 years of hands-on experience in SRE, DevOps, or Platform Engineering roles.
Incident Leadership: Proven experience leading the response to high-severity production incidents: coordinating responders, managing stakeholder communication, and driving resolution under pressure.
AWS Expertise: Deep working knowledge of Amazon ECS, IAM, VPC, ALB/NLB, RDS, S3, MSK, Elasti Cache, Lambda, CloudWatch, and an awareness of cost optimization practices. Infrastructure as Code: Proficiency with Terraform for managing and modifying AWS resources; comfortable reading and writing Terraform configurations independently. Observability: Proficiency building and owning Datadog monitors and dashboards, including defining SLOs/SLIs and managing error budgets. Familiarity with Grafana is a plus. Architecture Knowledge: Ability to evaluate architectural patterns (e.g., Microservices, Pub/Sub, Load Balancing) and reason about their reliability trade-offs and failure modes. Linux & Scripting: Strong Linux command-line skills; comfortable writing Python automation scripts for operational tasks and tooling.
CI/CD: Working knowledge of GitHub Actions or GitLab CI for ECS-based deployments. Communication: Excellent written and verbal English communication skills. Able to produce clear async handover notes, post-mortems, and status updates across global time zones.
Ownership Mindset: Demonstrated track record of driving open issues to closure. Verifying root causes rather than assuming silence means resolved, and following up across teams until items are truly done.
Structured Troubleshooting: Systematic, hypothesis-driven approach to diagnosing issues, forming a clear picture of what is known, what to check next, and why, rather than relying on trial and error.
Proactive Teamwork: Posts incremental status updates during investigations, raises blockers early instead of going silent, and asks clarifying questions when expectations or feedback are unclear.
Preferred Skills:
Basic understanding of financial markets and market data (e.g., equities, options, market data feeds).
Database query skills and familiarity with RDS or Cassandra performance metrics. Experience mentoring engineers or establishing operational best practices across a team. Self-directed ramp-up: comfortable learning unfamiliar systems through documentation and runbooks with minimal hand-holding.
Strong documentation habit: proactively leaves high-quality handover notes and improves runbooks as part of daily work.
Pay: ₹600,000.00 - ₹1,200,000.00 per year
Benefits:
- Health insurance
- Leave encashment
- Provident Fund
Work Location: In person