#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-
Requisition Number: 59454
Job Location: Chennai, IND
Global Grade: Band 6
Work Type: Office Working
Employment Type: Permanent
Posting Start Date: 04/08/2026
Posting End Date: 18/08/2026
:
Job Summary
To own the end-to-end operational health, reliability, and performance of our AI/ML platform running across Azure, AWS, and On-Premises GPU environments. The ideal candidate will have deep expertise in maintaining GPU-accelerated infrastructure, AI/ML application stacks, and hybrid-cloud deployments while ensuring maximum uptime, scalability, and security.
You will play a critical role in supporting production AI workloads, troubleshooting complex platform issues, and continuously optimizing the infrastructure that powers our AI applications (LLMs, ML models, inference services, RAG pipelines, etc.).
Key Responsibilities
Platform Operations & Maintenance
- Manage day-to-day operations, monitoring, and maintenance of the AI Platform across Azure, AWS, and on-premises GPU clusters.
- Ensure high availability, reliability, and performance of AI/ML applications and services (training, inference, fine-tuning, RAG pipelines).
- Perform proactive health checks, capacity planning, patching, and lifecycle management of infrastructure components.
- Handle incident management, root cause analysis (RCA), and problem resolution within defined SLAs.
GPU Infrastructure Management
- Maintain and troubleshoot GPU nodes (NVIDIA A100, H100, V100, L40S, etc.), including driver installations, CUDA updates.
- Optimize GPU utilization, workload scheduling, and resource allocation across multi-tenant environments.
- Manage GPU clusters using Kubernetes (with GPU operators) or similar orchestrators.
Cloud & Hybrid Environment Management
- Administer AI/ML workloads on Azure ML, AWS SageMaker, EKS/AKS, EC2/Azure VMs, and on-prem Kubernetes/OpenShift clusters.
- Manage container-based deployments using Docker, Kubernetes, Helm, and GPU-enabled runtimes.
- Configure and maintain networking, storage (NFS, Blob, S3), and IAM across hybrid environments.
AI Application Support
- Support production AI applications including LLM serving, MLflow, Kubeflow, Ray, and similar platforms.
- Deploy, upgrade, and maintain model registries, vector databases and orchestration tools.
- Collaborate with Data Scientists and ML Engineers to onboard models, optimize inference, and troubleshoot performance bottlenecks.
Automation & CI/CD
- Automate platform provisioning, configuration, and maintenance using Terraform, Ansible, ARM/Bicep, CloudFormation.
- Build and maintain CI/CD pipelines for AI/ML workloads (GitHub Actions, Azure DevOps, Jenkins, GitLab).
- Implement Infrastructure as Code (IaC) practices for reproducible environments.
Monitoring, Logging & Observability
- Implement and manage observability stacks: Prometheus, Grafana, ELK, Azure Monitor, CloudWatch.
- Set up GPU-specific monitoring and alerting for AI workloads.
- Track model performance, drift, latency, and throughput.
Security & Compliance
- Enforce security best practices (identity, network, secrets management, encryption).
- Ensure compliance with organizational and regulatory standards
- Manage RBAC, key vaults, and secure model artifact storage.
Strategy
- To Maintain a highly available, secure, scalable, and cost-optimized AI Platform across Azure, AWS, and On-Premises GPU environments that enables seamless development, deployment, and operation of AI/ML and GenAI workloads with enterprise-grade reliability.
Business
- AI Factory architecture and engineering leadership
- T&A and TTO management
- Chief Data Office and central data governance
- Software Engineering Platform Teams
- Business technology teams across Trade, Cash, Financial Markets, Risk & Compliance, and Retail Operational, Technology & Cyber Risk; Group Internal Audit Compliance, Regulators, and the independent AI governance function
Regulatory & Business Conduct
- Display exemplary conduct and live by the Group’s Values and Code of Conduct.
- Take personal responsibility for embedding the highest standards of ethics, including regulatory and business conduct, across Standard Chartered Bank. This includes understanding and ensuring compliance with, in letter and spirit, all applicable laws, regulations, guidelines and the Group Code of Conduct.
- Effectively and collaboratively identify, escalate, mitigate and resolve risk, conduct and compliance matters.
Key stakeholders
- AI Factory architecture and engineering leadership
- T&A and TTO management
- Chief Data Office and central data governance
- Software Engineering Platform Teams
- Business technology teams across Trade, Cash, Financial Markets, Risk & Compliance, and Retail Operational, Technology & Cyber Risk; Group Internal Audit Compliance, Regulators, and the independent AI governance function
Skills and Experience
- Azure Administrator
- Databricks
- Linux Administrator
- Kubernetes
- GPU
- Python
Qualifications
- Bachelor's/Master's degree in Computer Science, Engineering, or related field.
- 8+ years of experience in Platform/Infrastructure/DevOps/SRE roles, with at least 3+ years in AI/ML platform engineering.
- Strong hands-on experience with Azure and AWS (compute, networking, storage, IAM, GPU instances).
- Solid experience managing on-premises GPU infrastructure (NVIDIA DGX, HPE, Dell, Supermicro GPU servers).
- Expertise in Kubernetes, GPU Operator, container runtimes, and Helm.
- Proficiency in Linux system administration, shell scripting, and Python.
- Experience with IaC tools (Terraform, Ansible) and CI/CD pipelines.
- Familiarity with AI/ML frameworks: PyTorch, TensorFlow, Hugging Face, and inference servers (Triton, vLLM, TGI).
- Strong troubleshooting skills across networking, storage, GPU, and application layers.
- Experience with monitoring/observability tools (Prometheus, Grafana, DCGM).
- Certifications: Azure Solutions Architect, AWS Solutions Architect/DevOps, CKA/CKAD, NVIDIA DLI.
About Standard Chartered
We're an international bank, nimble enough to act, big enough for impact. For more than 170 years, we've worked to make a positive difference for our clients, communities, and each other. We question the status quo, love a challenge and enjoy finding new opportunities to grow and do better than before. If you're looking for a career with purpose and you want to work for a bank making a difference, we want to hear from you. You can count on us to celebrate your unique talents and we can't wait to see the talents you can bring us.
Our purpose, to drive commerce and prosperity through our unique diversity, together with our brand promise, to be here for good are achieved by how we each live our valued behaviours. When you work with us, you'll see how we value difference and advocate inclusion.
Together we:
Do the right thing and are assertive, challenge one another, and live with integrity, while putting the client at the heart of what we do
Never settle, continuously striving to improve and innovate, keeping things simple and learning from doing well, and not so well
Are better together, we can be ourselves, be inclusive, see more good in others, and work collectively to build for the long term
What we offer
In line with our Fair Pay Charter, we offer a competitive salary and benefits to support your mental, physical, financial and social wellbeing.
Core bank funding for retirement savings, medical and life insurance, with flexible and voluntary benefits available in some locations.
Time-off including annual leave, parental/maternity (20 weeks), sabbatical (12 months maximum) and volunteering leave (3 days), along with minimum global standards for annual and public holiday, which is combined to 30 days minimum.
Flexible working options based around home and office locations, with flexible working patterns.
Proactive wellbeing support through Unmind, a market-leading digital wellbeing platform, development courses for resilience and other human skills, global Employee Assistance Programme, sick leave, mental health first-aiders and all sorts of self-help toolkits
A continuous learning culture to support your growth, with opportunities to reskill and upskill and access to physical, virtual and digital learning.
Being part of an inclusive and values driven organisation, one that embraces and celebrates our unique diversity, across our teams, business functions and geographies - everyone feels respected and can realise their full potential.
#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-