Our Company:
At Teradata, we believe that people thrive when empowered with better information. Teradata Autonomous Knowledge Platform activates enterprise intelligence by unifying data, knowledge and business context to achieve tangible outcomes. With Teradata, organizations can provide agents with full context for impact when it matters. Our solution lets businesses connect and scale on premises, in the cloud, or through a hybrid approach. Teradata delivers real business value with AI.
Ignite the Future of AI at Teradata!
Location: Hybrid – Hyderabad, India
About This Role
Teradata is building the next generation of AI-native enterprise platform, where autonomous agents don't just assist — they reason, plan, act, and learn. We are looking for a Senior Agentic AI Engineer to lead the design and development of production-grade agentic systems: multi-agent pipelines, memory and context management, agent evaluation infrastructure, and the governance framework that makes autonomous AI safe at enterprise scale.
This role is ideal for an engineer who operates at the frontier of LLM capabilities, understands how to harden agents for real-world enterprise workloads, and cares deeply about building systems that are not only capable but trustworthy and measurable.
What You'll Do: Architect the Agentic Layer, Agent Design & Orchestration
- Design and implement production-grade multi-agent architectures, including task decomposition, inter-agent communication, delegation, and coordination patterns.
- Build agentic harnesses: agent loop controllers, tool registries, execution sandboxes, and retry/fallback logic.
- Develop planning and reasoning frameworks (React, chain-of-thought, tree-of-thought) and integrate them with enterprise workflows.
- Implement tool-calling pipelines and dynamic function invocation with schema validation and error recovery.
Agent Evaluation
- Design end-to-end agent evaluation frameworks: task taxonomies, success criteria, ground-truth datasets, and multi-dimensional scoring across task completion, tool accuracy, reasoning quality, safety, and cost.
- Implement LLM-as-judge scoring systems with calibrated rubrics, inter-rater reliability checks, and human-in-the-loop validation.
- Build trajectory-level evaluation tooling to analyze full agent execution traces — tool call sequences, intermediate reasoning steps, and decision points — not just final outputs.
- Design red-teaming and adversarial harnesses to probe agent failure modes: prompt injection, goal misgeneralization, over-refusal, and multi-step reasoning breakdowns.
- Instrument evaluation pipelines with cost, latency, and quality dashboards; establish regression suites that gate every agent or prompt change before production.
Context Management & LLM Memory
- Design context window management systems: dynamic compression, sliding window strategies, and priority-based context eviction.
- Build multi-tier LLM memory architectures:
- In-context working memory for short-horizon task state.
- Episodic memory via vector retrieval for cross-session recall.
- Procedural memory for learned tool-use patterns.
- Implement memory indexing, lifecycle management (retention, decay, invalidation), and privacy-respecting memory scoping across enterprise tenants.
- Optimize prompt assembly pipelines that dynamically compose context from multiple memory tiers under token budget constraints.
Agentic Skill Development & Document Automation
- Design and implement reusable agent skills — self-contained, composable capability modules with metadata schemas, trigger conditions, and runtime discovery — that extend agent capabilities without redeployment.
- Build document generation and file-handling skills covering DOCX, PPTX, XLSX, and PDF, with programmatic layout control, template rendering, and format-aware extraction for heterogeneous input types.
- Create frontend artifact generation skills for producing interactive HTML/React components, data visualizations, and UI prototypes as structured agent outputs.
- Integrate skills with the Model Context Protocol (MCP) and establish skill authoring standards — versioning, deprecation policies, and evaluation harnesses — that ensure reliable agent-to-skill routing at scale.
Vibe Coding & AI-Accelerated Development
- Leverage AI coding assistants (Claude Code, GitHub Copilot, Cursor, or equivalent) as primary development tools — collaborative pair programmers for architecture, implementation, and debugging, not just autocomplete.
- Translate high-level intent into production-quality code through natural-language-driven development loops, maintaining engineering rigor while dramatically compressing cycle time.
- Build and curate prompt libraries and system prompts that reliably steer AI assistants toward enterprise-grade output: typed, tested, documented, and security-conscious.
- Maintain sharp critical judgment over AI-generated code — catching logic errors, security gaps, and architectural drift — and set team norms for code review, testing, and documentation at AI-assisted scale.
Agentic Skill Development & Document Automation
- Design and implement reusable agent skills — self-contained, composable capability modules with metadata schemas, trigger conditions, and runtime discovery — that extend agent capabilities without redeployment.
- Build document generation and file-handling skills covering DOCX, PPTX, XLSX, and PDF, with programmatic layout control, template rendering, and format-aware extraction for heterogeneous input types.
- Create frontend artifact generation skills for producing interactive HTML/React components, data visualizations, and UI prototypes as structured agent outputs.
- Integrate skills with the Model Context Protocol (MCP) and establish skill authoring standards — versioning, deprecation policies, and evaluation harnesses — that ensure reliable agent-to-skill routing at scale.
Vibe Coding & AI-Accelerated Development
- Leverage AI coding assistants (Claude Code, GitHub Copilot, Cursor, or equivalent) as primary development tools — collaborative pair programmers for architecture, implementation, and debugging, not just autocomplete.
- Translate high-level intent into production-quality code through natural-language-driven development loops, maintaining engineering rigor while dramatically compressing cycle time.
- Build and curate prompt libraries and system prompts that reliably steer AI assistants toward enterprise-grade output: typed, tested, documented, and security-conscious.
- Maintain sharp critical judgment over AI-generated code — catching logic errors, security gaps, and architectural drift — and set team norms for code review, testing, and documentation at AI-assisted scale.
Who You'll Work With: Build with the Best
You'll collaborate with AI architects, ML engineers, platform engineers, and domain experts across Teradata's global AI team, including deep technical partnerships with Silicon Valley research and product leadership.
You'll also work cross-functionally with:
- Product and UX to translate enterprise use cases into agent behaviors that are both powerful and intuitive.
- Security and compliance teams to ensure every agent deployment meets enterprise governance standards.
- Infrastructure teams responsible for GPU cluster management, model serving, and cost optimization.
- Customer success teams to validate that agent capabilities solve real problems in regulated industries.
Minimum Requirements
- BS/MS/PhD in Computer Science, AI/ML, or a related field.
- 5+ years of software engineering experience, with 2+ years focused on LLM or agentic systems.
- Hands-on experience building and deploying production agentic systems (not just prototypes).
- Proficiency in Python with strong fundamentals in typing, testing, and observability.
- Familiarity with agent frameworks such as LangChain, LangGraph, AutoGen, CrewAI, or equivalent.
Preferred Qualifications
- Experience publishing or implementing research on agent evaluation, reasoning, or memory-augmented LLMs.
- Familiarity with MLOps tooling: MLflow, Weights & Biases, DVC, or equivalent.
- Experience with model serving infrastructure (vLLM, TGI, Triton) and quantization techniques.
- Exposure to enterprise observability stacks (OpenTelemetry, Datadog, Prometheus/Grafana).
- Experience designing systems for regulated industries (financial services, healthcare) with compliance constraints.
- Familiarity with multi-modal agents (vision, structured data, voice).
- Contributions to open-source agentic or LLM evaluation ecosystem projects.
- Experience with cloud AI platforms (AWS SageMaker, Azure ML, GCP Vertex AI).
- Portfolio of projects built primarily through AI-assisted development, demonstrating velocity, code quality, and critical review practices.
#LI-NM1