Project Role : AI Infrastructure Architect
Project Role Description : Architect and build custom Artificial Intelligence (AI) infrastructure/hardware solutions. Optimize AI infrastructure/hardware performance, power consumption, cost and scalability of computational stack. Advise on AI infrastructure technology and vendor evaluation, selection and full stack integration.
Must have skills : Large Language Models (LLMs)
Good to have skills : NA
Minimum
5 year(s) of experience is required
Educational Qualification : 15 years full time education
Summary:
Build the evaluation and quality engineering capability for AI and agentic systems. Establish measurable quality gates that determine whether systems are safe, reliable, effective, and ready for production.
Expert use of LangSmith, Braintrust, Arize Phoenix, Weights & Biases Weave, MLflow and OpenAI Evals, supported by Codex, Claude Code or Cursor, to build automated regression suites, trajectory evaluations, red-team tests, trace analysis and CI/CD quality gates for production agents.
Must have built evaluation or quality systems for production AI, ML, search, or decision systems. Manual prompt testing and subjective review alone are insufficient.
Roles & Responsibilities:
- Design evaluation strategies for agent behavior, task completion, retrieval quality, groundedness, safety, and reliability.
- Build automated evaluation harnesses, regression suites, benchmark datasets, and production quality gates.
- Evaluate multi-step agent trajectories, tool use, planning, recovery, and human escalation.
- Combine deterministic tests, model-based evaluation, human review, and production telemetry.
- Perform failure analysis, red teaming, and root-cause investigation.
- Integrate evaluations into CI/CD, release, monitoring, and incident-management processes.
- Define scorecards for engineering, risk, product, and business stakeholders.
Professional & Technical Skills:
- Python, test automation, data analysis, statistics, and experimentation.
- LLM and RAG evaluation, agent trajectory analysis, benchmark design, and error taxonomy.
- Tracing, observability, adversarial testing, safety testing, and production monitoring.
- Distinguishing model, retrieval, prompt, tool, data, and orchestration failures.