TJX Companies
At TJX Companies, every day brings new opportunities for growth, exploration, and achievement. You’ll be part of our vibrant team that embraces diversity, fosters collaboration, and prioritizes your development. Whether you’re working in our four global Home Offices, Distribution Centers or Retail Stores—TJ Maxx, Marshalls, Homegoods, Homesense, Sierra, Winners, and TK Maxx, you’ll find abundant opportunities to learn, thrive, and make an impact. Come join our TJX family—a Fortune 100 company and the world’s leading off-price retailer.
Job Description:
Location: India IT Office, Hyderabad / India
The Infrastructure and Operations organization embodies the hub of lifecycle engineering at TJX, delivering, maintaining, and optimizing our technology portfolio at cloud scale. We are a service-oriented team aimed at providing extraordinary experiences to thousands of TJX associates, business partners, and application delivery teams across the portfolio.
We are seeking a Site Reliability Engineer – AI Operations to implement Site Reliability Engineering practices with a strong focus on automation, observability, AIOps, Gen AI, and self-healing operations. This role will help improve production reliability, reduce manual toil, accelerate incident response, and enable intelligent operational automation across enterprise applications and platforms.
The ideal candidate combines strong production operations experience with engineering, automation, observability, and AI-enabled operations capabilities. This role will work closely with application, infrastructure, cloud, security, operations, and leadership teams to strengthen reliability, improve operational efficiency, and modernize how production services are supported.
Site Reliability Engineering & Operations
Implement SRE best practices across production support, incident management, problem management, change management, release, and deployment processes.
Support production reliability by improving availability, performance, scalability, and operational readiness of applications and platforms.
Participate in incident response, troubleshooting, root cause analysis, postmortems, and corrective/preventive action planning.
Drive shift-left reliability by embedding monitoring, alerting, automation, and operational readiness into SDLC and CI/CD pipelines.
Automation & Self-Healing
Build and maintain automation scripts, runbooks, self-healing workflows, and operational tools to reduce manual effort and improve MTTR.
Integrate monitoring, ITSM, CI/CD, cloud, and automation platforms using APIs, scripts, and workflows.
Ensure automation activities follow change management, governance, security, and compliance processes.
Observability, Monitoring & Reliability Metrics
Configure and enhance observability across logs, metrics, traces, dashboards, alerts, and synthetic monitoring.
Work with monitoring and APM tools such as Splunk, AppDynamics, Dynatrace, Datadog, New Relic, Azure Monitor, or similar tools.
Support implementation of SLIs, SLOs, SLAs, service health metrics, reliability KPIs, and operational dashboards.
AIOps, Gen AI & Intelligent Operations
Support AI-driven operational capabilities such as incident summarization, alert enrichment, anomaly detection, log analysis, event correlation, and root cause recommendations.
Ensure AI-enabled operations follow enterprise security, compliance, data privacy, and responsible AI guidelines.
Collaboration & Continuous Improvement
Work closely with application, infrastructure, cloud, operations, security, and leadership teams.
Support cross-functional initiatives focused on reliability, automation, service health, and operational excellence.
Contribute to continuous improvement by identifying gaps in process, tooling, monitoring, and documentation.
We seek creative, customer-focused individuals with strong SRE, DevOps, production operations, automation, observability, and AI-enabled operations experience. This role requires a continuous improvement mindset and the ability to balance innovation with reliability, stability, security, and compliance.
Bachelor’s degree in Information Technology, Computer Science, Engineering, or equivalent practical experience.
6+ years of experience in SRE, DevOps, Production Support Engineering, Cloud Operations, Automation Engineering, or related roles.
Hands-on experience in application support, incident management, problem management, change/release management, deployment support, monitoring, and documentation.
Experience working across application, infrastructure, cloud, operations, security, and leadership teams.
Required Technical Skills
Strong experience with observability/APM tools such as Splunk, AppDynamics, Dynatrace, Datadog, New Relic, Azure Monitor, or similar tools.
Experience with SLIs, SLOs, SLAs, operational KPIs, service health dashboards, and reliability metrics.
Strong coding/scripting experience in one or more languages such as Python, Java, PowerShell, Go, or Shell scripting.
Hands-on experience with automation and DevOps tools such as Azure DevOps, GitHub Actions, Jenkins, Ansible, Terraform, Power Automate, Rundeck, or similar tools.
Knowledge of AIOps, Gen AI, or AI-enabled operations use cases such as anomaly detection, log analysis, ticket classification, event correlation, and incident summarization.
Understanding of distributed systems, APIs, microservices, cloud-native applications, and reliability engineering principles.
Experience with Azure OpenAI, Microsoft Copilot Studio, LangChain, Semantic Kernel, vector databases, or RAG-based solutions.
Experience with ITSM and incident response tools such as ServiceNow, Jira Service Management, PagerDuty, xMatters, Opsgenie, or similar platforms.
Experience with Docker, Kubernetes, OpenShift, CI/CD pipelines, Infrastructure as Code, GitOps, or DevSecOps.
Knowledge of machine learning concepts such as anomaly detection, classification, clustering, and time-series analysis.
Understanding of Responsible AI, prompt engineering, model governance, data privacy, and security controls.
Relevant certifications in Azure, DevOps, SRE, AI, or ITIL.
Ability to work across product, engineering, operations, cloud, security, and leadership teams.
This role is based in India and requires collaboration with global teams across U.S., Canada, Europe, India, and Australia. The candidate should be comfortable supporting flexible working hours as needed to engage with global stakeholders, participate in critical meetings, support escalations, and ensure seamless delivery across regions.
Join us and Discover Different at TJX.
In addition to our open door policy and supportive work environment, we also strive to provide a competitive salary and benefits package. TJX considers all applicants for employment without regard to race, color, religion, gender, sexual orientation, national origin, age, disability, gender identity and expression, marital or military status, or based on any individual's status in any group or class protected by applicable federal, state, or local law. TJX also provides reasonable accommodations to qualified individuals with disabilities in accordance with the Americans with Disabilities Act and applicable state and local law.
Address:
Salarpuria Sattva Knowledge City, Inorbit Road
Location:
APAC Home Office Hyderabad IN