NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.
We are currently seeking a AI Orchestration / Platform Engineer to join our team in Noida/Gurgaon, Haryāna (IN-HR), India (IN).
4. AI Orchestration / Platform Engineer
Agent runtime | Model routing | LLMOps | CI/CD | Observability and reliability
Number of positions
1
Level
Senior / Lead Platform Individual Contributor
Primary locations
Hyderabad or Noida preferred; exceptional onshore candidates may be considered
Target / alternate titles
AI Platform Engineer; Agent Platform Engineer; LLMOps Engineer; GenAI Platform Engineer; AI DevOps Engineer; AI Site Reliability Engineer; ML Platform Engineer
Core keywords
agent orchestration, workflow engine, multi-agent systems, model gateway, model routing, LLMOps, platform engineering, Kubernetes, containers, CI/CD, infrastructure as code, policy as code, observability, OpenTelemetry, queues, state, retries, resilience, cost telemetry, SRE
Recruiter red flags
Traditional DevOps profile with no AI runtime understanding; framework-only orchestration without production platform depth; manual deployments; no observability or incident ownership; cannot explain model routing, state, retries, evaluation, or AI-specific controls.
Role purpose
Build and operate the orchestration and runtime capabilities that allow agentic applications to move reliably from development into production. The role will integrate agent workflows with existing cloud, security, policy, CI/CD, and observability capabilities, supporting multi-step execution, model routing, state, resiliency, evaluation, and production operations without rebuilding the client's established platform foundations.
Client and delivery context
- The client already has delivery pipelines, policy-based architecture, security controls, and production review processes. The engineer must integrate with and enhance those capabilities.
- The platform should support multiple models, clouds, frameworks, and developer-productivity ecosystems without forcing avoidable lock-in.
- The engineer will support shared patterns used across different business-process agents and may work across two parallel delivery tracks.
- Platform work must be pragmatic and outcome-driven, with emphasis on enabling engineers to release and operate reliable agents quickly.
Primary ownership
- Agent workflow and runtime orchestration, including state, routing, queues, retries, timeouts, scheduling, persistence, and exception management.
- Model gateway and routing patterns, provider abstraction, policy-based selection, fallback, rate limits, quotas, and cost controls.
- CI/CD, environment promotion, configuration, secrets, infrastructure integration, release validation, rollback, and operational readiness.
- Application and platform observability, reliability engineering, incident response, capacity, performance, and production support.
Key responsibilities
- Design and implement orchestration patterns for multi-step agents, multi-agent collaboration, deterministic workflows, long-running tasks, approvals, and event-driven execution.
- Implement durable state, checkpoints, queues, retries, backoff, idempotency, timeouts, compensation, dead-letter handling, and recovery for agent workflows.
- Integrate model gateways and routing logic that select models based on capability, sensitivity, latency, cost, availability, or policy requirements.
- Build deployment and release patterns for agent services, orchestration components, prompts, tool definitions, configurations, and evaluation assets.
- Integrate with existing CI/CD, infrastructure-as-code, policy-as-code, secrets, identity, container, and cloud-runtime capabilities.
- Create observability for agent traces, model calls, tool calls, workflow state, token usage, cost, latency, errors, quality signals, and dependency health.
- Implement automated quality and security gates, including unit and integration tests, evaluation suites, policy checks, vulnerability scans, and rollback criteria.
- Optimize runtime performance, concurrency, throughput, caching, context use, model selection, and infrastructure cost.
- Support production incidents, root-cause analysis, capacity planning, resiliency testing, disaster recovery, and operational runbooks.
- Develop reusable platform templates, SDKs, reference pipelines, dashboards, and onboarding guidance for agent-development teams.
Must-have candidate profile
- 8+ years of platform, DevOps, SRE, cloud, distributed-systems, or software-engineering experience.
- Production experience supporting AI/ML, LLM, agentic, workflow, or high-scale distributed application platforms.
- Strong experience with containers, Kubernetes or equivalent runtimes, CI/CD, infrastructure as code, configuration, secrets, and automated deployment.
- Understanding of agent orchestration concepts including state, checkpoints, retries, timeouts, queues, long-running tasks, human approvals, and failure recovery.
- Experience with logging, metrics, distributed tracing, OpenTelemetry or equivalent observability, alerting, dashboards, and incident response.
- Strong scripting or development skills in Python, Go, Java, TypeScript, or comparable languages.
- Ability to integrate platform capabilities with security, identity, policy, data, network, and enterprise approval requirements.
- Experience balancing reliability, delivery speed, latency, throughput, portability, and operating cost.
Preferred experience
- Experience with model gateways, multi-model routing, provider abstraction, fallback, quotas, or AI cost controls.
- Experience with LangGraph, Temporal, Airflow, Argo Workflows, Durable Functions, Step Functions, Kubernetes operators, or equivalent orchestration technologies.
- Experience with AI tracing and evaluation platforms, prompt/model registries, feature flags, canary releases, and regression gates.
- Experience with Terraform, Pulumi, Helm, GitOps, policy engines, service mesh, event streaming, and API gateways.
- Experience operating platforms across AWS, Azure, GCP, hybrid, or private environments.
- Experience in regulated enterprises or systems with sensitive proprietary data and formal production controls.
Indicative technology exposure
Kubernetes, Docker, serverless or managed container platforms; Terraform/Pulumi/Helm/GitOps; GitHub Actions, GitLab, Jenkins, Azure DevOps, or equivalent; LangGraph, Temporal, Airflow, Argo, Step Functions, Durable Functions, or equivalent orchestration; model gateways and provider APIs; queues and events; OpenTelemetry, Prometheus, Grafana, cloud monitoring, AI tracing and evaluation tools. Exact products are flexible.
Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored to each client’s needs. While many positions offer remote or hybrid work options, these arrangements are subject to change based on client requirements. For employees near an NTT DATA office or client site, in-office attendance may be required for meetings or events, depending on business needs. At NTT DATA, we are committed to staying flexible and meeting the evolving needs of both our clients and employees. NTT DATA recruiters will never ask for payment or banking information and will only use @nttdata.com and @talent.nttdataservices.com email addresses. If you are requested to provide payment or disclose banking information, please submit a contact us form, https://us.nttdata.com/en/contact-us.
NTT DATA endeavors to make https://us.nttdata.com accessible to any and all users. If you would like to contact us regarding the accessibility of our website or need assistance completing the application process, please contact us at https://us.nttdata.com/en/contact-us. This contact information is for accommodation requests only and cannot be used to inquire about the status of applications. NTT DATA is an equal opportunity employer. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status. For our EEO Policy Statement, please click here. If you'd like more information on your EEO rights under the law, please click here. For Pay Transparency information, please click here.