Noida, Uttar Pradesh
Job Summary
Job Description: AI Ops Support Engineer Role Summary The AI Ops Support Engineer is responsible for the daily operational health, monitoring, and L2/L3 support of production Generative AI and Machine Learning solutions. This role focuses on maintaining the stability, availability, and performance of AI/ML platforms by proactively monitoring system health, triaging incidents, executing standard operating procedures (SOPs), and adhering to defined service level agreements (SLAs). Working within an operational rotation, the engineer acts as the first line of defense for AI system anomalies and collaborates closely with MLOps, Platform, and Development teams to ensure seamless service delivery. Key Responsibilities AI/ML Production Operations & Monitoring Health Checks & Monitoring: Perform daily operational health checks, service validations, and active monitoring of production GenAI and ML workloads. Observability Dashboards: Monitor established observability dashboards (Datadog, Dynatrace, etc.) to track model drift, latency, API consumption, token usage, and system costs. Alert Response: Respond immediately to system-generated alerts regarding threshold breaches, data quality issues, and infrastructure failures. Incident Management & Triage L2/L3 Support: Triage, investigate, and resolve incident tickets related to AI/ML services within specified SLAs. Escalation & Collaboration: Follow defined escalation paths to engage L3 engineers, Data Scientists, or Platform teams for complex, systemic issues. Root Cause Analysis (RCA): Assist in gathering logs, metrics, and event timelines to support the senior team in performing Root Cause Analysis (RCA) and documenting incident reports. Model Lifecycle & Maintenance Support Release & Deployment Validation: Support the deployment of model upgrades, prompt adjustments, and system patches. Conduct post-deployment validation and regression testing using operational runbooks. Data & Pipeline Monitoring: Monitor scheduled data pipelines (such as Databricks jobs and SQL warehouses) to ensure continuous training data flow and resolve minor job failures. Continuous Improvement & Automation SOP & Runbook Execution: Execute standard automated scripts to resolve repetitive issues and recover failed services. Operational Scripting: Write and maintain basic Python or Bash scripts to automate routine manual tasks, log parsing, and operational checks. Documentation & Governance Knowledge Base Maintenance: Document incident resolution steps, update troubleshooting runbooks, and draft new standard operating procedures (SOPs) based on resolved issues. Service Desk Logging: Maintain accurate records of all incidents, service requests, and operational actions within the ticketing system. Required Skills & Experience Technical Skills AI/ML Concepts: Fundamental understanding of Generative AI, LLMs, prompt engineering, Retrieval-Augmented Generation (RAG), and MLOps lifecycles. Monitoring & Observability: Experience utilizing APM and observability platforms (e.g., Datadog, Dynatrace, New Relic) to monitor application logs and system metrics. Cloud Infrastructure: Hands-on experience with Google Cloud Platform (GCP) services (Vertex AI, GKE, Cloud Functions, BigQuery, MemoryStore). Scripting: Proficiency in Python and Bash for automating routine operational tasks and querying APIs. Data Platforms: Familiarity with Databricks environments (managing job runs, clusters, and catalog accesses) and basic SQL querying. Operational Skills ITIL Practices: Solid understanding of ITIL-based framework processes, specifically Incident, Problem, and Change Management. SLA Adherence: Proven ability to work under SLA-driven environments, prioritizing tickets and tasks based on business impact. Collaboration & Communication: Excellent written and verbal communication skills to clearly document technical issues and coordinate with cross-functional support tiers. Preferred Experience & Qualifications Experience: 3–5
Key Responsibilities
Job Description: AI Ops Support Engineer Role Summary The AI Ops Support Engineer is responsible for the daily operational health, monitoring, and L2/L3 support of production Generative AI and Machine Learning solutions. This role focuses on maintaining the stability, availability, and performance of AI/ML platforms by proactively monitoring system health, triaging incidents, executing standard operating procedures (SOPs), and adhering to defined service level agreements (SLAs). Working within an operational rotation, the engineer acts as the first line of defense for AI system anomalies and collaborates closely with MLOps, Platform, and Development teams to ensure seamless service delivery. Key Responsibilities AI/ML Production Operations & Monitoring Health Checks & Monitoring: Perform daily operational health checks, service validations, and active monitoring of production GenAI and ML workloads. Observability Dashboards: Monitor established observability dashboards (Datadog, Dynatrace, etc.) to track model drift, latency, API consumption, token usage, and system costs. Alert Response: Respond immediately to system-generated alerts regarding threshold breaches, data quality issues, and infrastructure failures. Incident Management & Triage L2/L3 Support: Triage, investigate, and resolve incident tickets related to AI/ML services within specified SLAs. Escalation & Collaboration: Follow defined escalation paths to engage L3 engineers, Data Scientists, or Platform teams for complex, systemic issues. Root Cause Analysis (RCA): Assist in gathering logs, metrics, and event timelines to support the senior team in performing Root Cause Analysis (RCA) and documenting incident reports. Model Lifecycle & Maintenance Support Release & Deployment Validation: Support the deployment of model upgrades, prompt adjustments, and system patches. Conduct post-deployment validation and regression testing using operational runbooks. Data & Pipeline Monitoring: Monitor scheduled data pipelines (such as Databricks jobs and SQL warehouses) to ensure continuous training data flow and resolve minor job failures. Continuous Improvement & Automation SOP & Runbook Execution: Execute standard automated scripts to resolve repetitive issues and recover failed services. Operational Scripting: Write and maintain basic Python or Bash scripts to automate routine manual tasks, log parsing, and operational checks. Documentation & Governance Knowledge Base Maintenance: Document incident resolution steps, update troubleshooting runbooks, and draft new standard operating procedures (SOPs) based on resolved issues. Service Desk Logging: Maintain accurate records of all incidents, service requests, and operational actions within the ticketing system. Required Skills & Experience Technical Skills AI/ML Concepts: Fundamental understanding of Generative AI, LLMs, prompt engineering, Retrieval-Augmented Generation (RAG), and MLOps lifecycles. Monitoring & Observability: Experience utilizing APM and observability platforms (e.g., Datadog, Dynatrace, New Relic) to monitor application logs and system metrics. Cloud Infrastructure: Hands-on experience with Google Cloud Platform (GCP) services (Vertex AI, GKE, Cloud Functions, BigQuery, MemoryStore). Scripting: Proficiency in Python and Bash for automating routine operational tasks and querying APIs. Data Platforms: Familiarity with Databricks environments (managing job runs, clusters, and catalog accesses) and basic SQL querying. Operational Skills ITIL Practices: Solid understanding of ITIL-based framework processes, specifically Incident, Problem, and Change Management. SLA Adherence: Proven ability to work under SLA-driven environments, prioritizing tickets and tasks based on business impact. Collaboration & Communication: Excellent written and verbal communication skills to clearly document technical issues and coordinate with cross-functional support tiers. Preferred Experience & Qualifications Experience: 3–5
Skill Requirements
Other Requirements
#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-