Ref#: W181046
Department: Information Technology
City: Bangalore
State/Province: Karnataka
Location: India
Company Description
Ralph Lauren Corporation (NYSE:RL) is a global leader in the design, marketing and distribution of premium lifestyle products in five categories: apparel, accessories, home, fragrances, and hospitality. For more than 50 years, Ralph Lauren's reputation and distinctive image have been consistently developed across an expanding number of products, brands and international markets. The Company's brand names, which include Ralph Lauren, Ralph Lauren Collection, Ralph Lauren Purple Label, Polo Ralph Lauren, Double RL, Lauren Ralph Lauren, Polo Ralph Lauren Children, Chaps, among others, constitute one of the world's most widely recognized families of consumer brands.
At Ralph Lauren, we unite and inspire the communities within our company as well as those in which we serve by amplifying voices and perspectives to create a culture of belonging, ensuring inclusion, and fairness for all. We foster a culture of inclusion through: Talent, Education & Communication, Employee Groups and Celebration.
Position Overview
Purpose & Scope:
The Senior ML Ops Engineer designs, builds, and owns the deployment, monitoring, and automation infrastructure that takes Ralph Lauren’s machine learning and AI models from development into reliable, governed production use — and brings the depth to make independent technical decisions about how that infrastructure is built.
This is a senior individual contributor role. Where a mid-level ML Ops Engineer executes within defined patterns, the Senior ML Ops Engineer shapes those patterns — designing reusable deployment frameworks, setting observability standards, leading incident investigations, and guiding other engineers on the team.
Working closely with Data Science, Data Engineering, Platform, and Governance teams, the role owns the technical quality of ML Ops delivery end-to-end: from pipeline design and environment architecture to monitoring coverage and operational readiness.
As Ralph Lauren scales its AI-driven initiatives across demand forecasting, personalisation, and merchandising, the Senior ML Ops Engineer is accountable for ensuring models reach production reliably, operate safely, and are observable enough to act on.
Essential Duties & Responsibilities
What you will be doing (responsibilities):
1. ML Deployment Architecture & CI/CD
-
Lead the design and build of CI/CD pipelines for ML model packaging, testing, versioning, and deployment across dev, test, and production environments.
-
Define and enforce deployment standards — staged rollouts, canary releases, rollback procedures — as reusable patterns for the team.
-
Architect containerised model serving solutions, making considered decisions on orchestration, resource allocation, and environment parity.
-
Own the reusable deployment framework: templates, shared libraries, and pipeline components that reduce rework across model teams.
2. Monitoring, Observability & Production Reliability
-
Design and implement end-to-end observability for model performance, data drift, pipeline health, and inference quality in production.
-
Set monitoring coverage standards and drive adoption across all deployed models — not just newly deployed ones.
-
Lead root-cause investigation for complex production incidents; own post-incident reviews and structural fixes.
-
Ensure all deployed models have defined SLOs, alerting thresholds, and clear incident ownership before going live.
3. Platform & Infrastructure Design
-
Own training, retraining, and batch-scoring pipeline architecture — including scheduling, dependency management, and failure-recovery design.
-
Make technical decisions on model registry structure, artifact lineage, and versioning conventions across the team.
-
Lead environment architecture decisions: base images, dependency isolation, environment parity between training and serving.
-
Partner with platform and architecture teams to optimise compute resource usage and infrastructure costs for ML workloads.
4. Technical Leadership & Enablement
-
Provide technical direction and design review for other ML Ops engineers on the team.
-
Define and document deployment patterns, runbooks, and operational standards that enable team self-service.
-
Work with Data Scientists to translate model requirements into production-ready, maintainable deployment designs.
-
Partner with Data Engineering and Governance teams to ensure data feeds, lineage, and compliance requirements are met in production.
-
Communicate technical decisions, risks, and trade-offs clearly to cross-functional stakeholders and leadership.
Experience, Skills, and Knowledge
Qualifications
Required
-
5–8+ years of hands-on ML Ops, ML engineering, or platform engineering experience, including a demonstrated track record of owning ML deployment infrastructure in enterprise environments.
-
Strong hands-on expertise in CI/CD design for ML — pipeline architecture, testing strategies, environment promotion, and rollback patterns.
-
Proven experience designing and implementing model monitoring and observability frameworks in production, not just operating existing ones.
-
Experience with containerisation (Docker) and orchestration (Kubernetes or equivalent) at the design level — not just usage.
-
Proficiency in Python and SQL, with the ability to build, debug, and optimise automation and infrastructure code.
-
Working knowledge of Azure and Databricks-based ML environments.
Preferred
-
Experience with model registry and experiment tracking tools (MLflow, Databricks Model Registry, or equivalent) at a standards-setting level.
-
Exposure to LLM or GenAI model deployment patterns, evaluation frameworks, and production monitoring considerations.
-
Experience influencing or defining ML Ops standards across a multi-team or platform-wide scope.
-
Background in retail, consumer goods, or other data-intensive industry environments.
-
Familiarity with data governance and responsible AI requirements as they apply to model deployment.
Success Measures
-
Reusable, standardised deployment patterns adopted across model teams with reduced rework and faster time-to-production.
-
Comprehensive monitoring coverage across all deployed models with defined SLOs and alerting.
-
Faster, more reliable root-cause resolution for production incidents.
-
Clear technical documentation and runbooks that enable team self-service.
-
Positive feedback from Data Science and Engineering peers on technical depth and collaboration