This role is for an AI Operations Support Engineer within a centralized team focused on maintaining the reliability of enterprise AI platforms (Azure and AWS) through automation.1
Key responsibilities include:
Incident Management: Performing L2 troubleshooting for incidents not resolved by existing automation.
Automation Development: Writing Python scripts and building "self-healing" utilities to reduce repetitive tasks.
Technical Operations: Supporting Kubernetes-based workloads (Kubeflow, ArgoCD) and managing observability stacks (Grafana, Prometheus).
AI Integration: Leveraging "vibe coding" (AI-assisted development tools like GitHub Copilot) to accelerate development and debugging.
Operations & Strategy: Collaborating with L1 operators, contributing to root-cause analysis, and managing tickets via ServiceNow.
Core Requirements:
Education: Degree in Computer Science or a related quantitative field.
Technical Proficiency: 2–3 years of Python experience, demonstrable Kubernetes skills, and experience with Infrastructure as Code (IaC).
Mindset: A hybrid developer-engineer approach, emphasizing proactive automation over reactive support.
Soft Skills: Strong communication, resilience, and a collaborative team player.
Desirable Additions:
Experience with Agentic AI bot development (e.g., LangChain) and CI/CD pipelines.
Familiarity with ITSM processes and cloud environments (Azure/AWS).