Your responsibilities:
▪ Design and implement model evaluation frameworks using RAGAS, custom scorers, and automated accuracy benchmarking to validate agent outputs against structured specifications and Gherkin acceptance criteria.
▪ Develop adversarial prompt testing suites that probe agent behaviour for hallucinations, prompt injection vulnerabilities, and edge-case failures across the Refinement, Decision, and Coding Agents.
▪ Build and maintain regression test suites that continuously verify agent behaviour consistency across model updates, prompt changes, and skill-file modifications within the LangGraph-based architecture.
▪ Integrate AI-specific test tooling into CI/CD pipelines using GitHub Actions, ensuring every agent deployment is gated by automated quality, fairness, and reliability checks aligned with EU AI Act and DORA requirements.
▪ Define and track quantitative quality metrics, confidence thresholds, retrieval precision, response latency, and drift detection, providing actionable dashboards to the engineering team.
Your profile:
▪ Bachelor's degree in Computer Science, Engineering, or a related field, with demonstrated experience in testing AI/ML systems or LLM-based applications.
▪ Strong programming skills in Python with hands-on proficiency in pytest, test automation frameworks, and scripting for evaluation pipelines.
▪ Practical experience with LLM evaluation methodologies and tools such as RAGAS, custom scoring functions, accuracy benchmarking, and adversarial/redteam testing techniques.
▪ Solid understanding of CI/CD integration using GitHub Actions or comparable platforms, with the ability to embed AI test gates into automated deployment workflows.
▪ Familiarity with agentic architectures such as LangGraph or LangChain, including knowledge of retrieval-augmented generation (RAG) patterns and multi-agent orchestration concepts.
Pay: ₹1,000,000.00 - ₹4,000,000.00 per year
Work Location: In person