We are hiring for an exciting opportunity for a Site Reliability Engineer (SRE) with 5–7 years of experience to join our team supporting MAM Fund Management Public and Investment Operations.
Position: Site Reliability Engineer (SRE)
Experience: 5–7 Years
Location: Gurugram
Work Model: Hybrid – 3 days/week from office
Domain: Financial Services / Fund Management / Investment Operations
Employment Type: Full-Time
About the Role
We are looking for a skilled and proactive Site Reliability Engineer who will be responsible for ensuring the stability, availability, security, resiliency, and compliance of business-critical applications.The role will involve production support, cloud infrastructure management, monitoring and observability, risk remediation, automation, and implementation of secure and resilient technology solutions.The ideal candidate should have strong hands-on experience with AWS, CloudFormation, CI/CD, monitoring tools, security controls, automation, and production operations, along with a strong risk and compliance mindset.
Key Responsibilities
- Ensure high availability, stability, reliability, and uptime of production applications.
- Monitor application and infrastructure health and proactively identify potential issues.
- Manage secrets, credentials, password rotations, certificates, and access controls.
- Perform vulnerability remediation, security patching, and technology risk mitigation.
- Manage certificate lifecycle activities including renewal, deployment, and expiry prevention.
- Perform software, application, operating system, and infrastructure upgrades.
- Identify and remediate technology obsolescence and unsupported components.
- Support technology resiliency, disaster recovery, high availability, and failover initiatives.
- Automate manual, repetitive, and operational processes using scripting and cloud-native technologies.
- Perform application and infrastructure cleanup and decommissioning of obsolete or unused resources.
- Provide L1/L2 production and on-call support, including incident investigation and resolution.
- Participate in Incident, Problem, and Change Management processes.
- Develop and maintain operational documentation, runbooks, and support procedures.
- Collaborate with global technical and non-technical stakeholders to deliver reliable technology solutions.
Technical Skills RequiredAWS & Infrastructure
- Strong hands-on experience with AWS services.
- Experience building, deploying, and maintaining AWS infrastructure using CloudFormation.
- Experience with:
- EC2
- S3
- IAM
- Lambda
- SQS
- SNS
- RDS
- DynamoDB
Monitoring & Observability
- Hands-on experience with one or more of:
- Splunk
- Grafana
- AWS CloudWatch
- Strong understanding of application/infrastructure monitoring, logging, alerting, metrics, and troubleshooting.
Security & Risk Management
- Experience with Secrets Management.
- Vulnerability assessment and patch remediation.
- Certificate lifecycle management.
- Password and credential rotation.
- User access management and access reviews.
- Technology obsolescence remediation.
- Strong understanding of Secure-by-Design principles.
- Strong risk management and compliance mindset.
DevOps / CI/CD / GitOps
- Experience with Git-based development and deployment practices.
- Hands-on experience with CI/CD tools such as:
- GitHub / Bitbucket
- GitHub Actions
- Bamboo
- ArgoCD
- Understanding of GitOps practices and automated deployment processes.
Automation
- Strong experience automating manual and repetitive operational processes.
- Scripting/programming experience using technologies such as Python, Shell/Bash, or PowerShell would be an advantage.
Production & Service Management
- Strong understanding of Incident Management, Problem Management, and Change Management.
- Experience working in L1/L2 production support environments.
- Comfortable participating in on-call support and handling production incidents.
- Strong troubleshooting and root-cause analysis skills.
- Experience creating and maintaining operational runbooks and technical documentation.
Domain Experience
- Prior experience in Financial Services / BFSI is preferred.
- Experience in Banking, Investment Banking, Asset Management, Fund Management, Capital Markets, Wealth Management, or FinTech will be considered relevant.
- Experience working in highly regulated environments with strong security, risk, audit, and compliance requirements is highly desirable.
Ways of Working
- Ability to work from the Gurugram office 3 days per week.
- Flexibility to work outside standard business hours when required for production support or critical activities.
- Experience working in an Agile delivery environment.
- Hands-on experience with JIRA and Confluence for task management, collaboration, and documentation.
- Ability to work independently with minimal supervision.
- Strong communication and stakeholder management skills, with the ability to work effectively with global teams.
Key Competencies
- Strong analytical and problem-solving skills.
- Excellent troubleshooting and incident-handling capabilities.
- Strong ownership and accountability.
- Risk and security-oriented mindset.
- Ability to prioritize and manage multiple operational activities.
- Strong written and verbal communication skills.
- Ability to collaborate with both technical and business stakeholders.
- Focus on automation, operational excellence, reliability, and continuous improvement.
Work Location: Hybrid remote in Gurugram, Haryana