JOB DESCRIPTION
Join our dynamic team to innovate and refine technology operations, impacting the core of our business services.
As a Technology Support Lead in the Technology Support team, you will play a leadership role in ensuring the operational stability, availability, and performance of our production services.
Job Responsibilities
-
Own end-to-end production support for mission-critical applications and platform components supporting treasury/trading/risk/data/reporting workflows
-
Lead major incident (P1/P2) response: triage, decisioning, restoration, communications, and post-incident review
-
Drive problem management and stability engineering by setting standards for RCA quality and timeliness
-
Define prevention roadmaps, track actions to closure, and reduce repeat incidents while improving MTTR/MTBF
-
Apply an SRE approach: identify operational toil, prioritize automation, and deliver measurable reduction in manual effort and failure rates
-
Build and enforce governance across production operations (runbook standards, operational readiness reviews, change governance, control/evidence routines)
-
Provide KPI/MIS reporting on incidents, availability, batch health, and risk themes
-
Strengthen observability and monitoring across applications and data pipelines (define SLOs/SLIs, alerting standards, improve signal-to-noise, dashboards, proactive detection)
-
Lead batch and data operations governance across AutoSys, Airflow, and data platforms including Databricks/DataLake/ETL
-
Partner with engineering/platform teams to improve resiliency (capacity, performance, HA/DR, error budgets where applicable)
-
Coach and lead a support team (where applicable) and act as a senior stakeholder interface for business, technology, and control partners with clear executive-level communication
Required Qualifications, Capabilities, and Skills
-
8+ years of experience in production/application support and/or SRE/operations for mission-critical platforms in banking/financial services (mandatory), including leadership accountability
-
Demonstrated SRE mindset and execution with proven toil identification and elimination
-
Proven delivery of automation and operational simplification
-
Strong governance and controls ownership (audit-ready)
-
Strong hands-on technical depth in Linux and SQL
-
Automation skills in Python and shell scripting
-
Scheduling/orchestration experience with AutoSys and Airflow
-
Cloud/platform experience with AWS, Kubernetes, and cloud technologies
-
Data platform experience with Databricks, DataLake, and ETL
-
Observability experience with Grafana, Dynatrace, Splunk, and OpenTelemetry
-
Strong incident/problem/change management expertise, including operating under ambiguity and time pressure, plus experience with SLOs/reliability dashboards/alert tuning, DR testing, resiliency reviews, operational risk assessments, and senior stakeholder communication/reporting
Preferred Qualifications, Capabilities, and Skills
-
Experience with containerized microservices, service meshes, and event-driven architectures.
ABOUT US