In every role, McCainers are ambitious, curious, and passionate about creating exceptional work experiences - together. With a customer-first mindset, we make doing business with McCain easy.
About the role.
The Responsible-AI / Eval Specialist is responsible for designing, implementing, and operating evaluation frameworks that determine whether AI use cases are ready to progress through governance stage-gates and into production. This role serves as the hands-on execution arm of McCain’s Responsible AI standards, working closely with AI Platform & Engineering, Knowledge Engineering, Forward Deployment Engineers, and Security teams to evaluate accuracy, reliability, fairness, robustness, and overall quality across the McCain AI ecosystem. The role supports AI solutions built on Azure, Anthropic foundation models, Distyl technologies, Databricks, SAP, Salesforce, and other enterprise platforms.
What you’ll be doing.
-
Design and execute evaluation strategies for AI use cases, including assessments of accuracy, reliability, hallucination rates, fairness, and robustness.
-
Build, maintain, and operate evaluation harnesses, benchmark suites, and golden datasets across the McCain AI technology stack.
-
Conduct stage-gate evaluations and provide clear, actionable findings to support governance decisions.
-
Run bias and fairness testing across protected attributes and business-relevant user groups.
-
Identify, quantify, and communicate fairness concerns while recommending practical remediation approaches.
-
Lead red-team and adversarial testing exercises to identify vulnerabilities, failure modes, prompt injection risks, and jailbreak scenarios.
-
Partner with Enterprise Security teams to strengthen AI resilience through broader adversarial testing initiatives.
-
Support risk classification activities by providing evaluation evidence and recommendations that inform governance decisions.
-
Collaborate with AI Platform & Engineering and Knowledge Engineering teams to embed evaluation capabilities into AI development workflows.
-
Maintain model cards, evaluation reports, decision logs, and other governance documentation while contributing to Responsible AI standards and playbooks.
-
Coach AI Engineers and Forward Deployment Engineers on evaluation best practices and quality disciplines.
-
Share evaluation insights, frameworks, and reusable patterns across the AI organization.
What you’ll need to be successful.
-
Bachelor’s degree in Computer Science, Statistics, Engineering, or a related field; advanced degree preferred.
-
10+ years of experience in AI/ML evaluation, model validation, model risk management, or a related discipline.
-
Hands-on experience with AI/ML evaluation methodologies and performance metrics, including accuracy, calibration, fairness, and robustness.
-
Experience evaluating Large Language Models (LLMs) and AI agents using frameworks such as HELM, Eleuther LM Evaluation Harness, or equivalent platforms.
-
Strong Python programming skills and experience with common machine learning evaluation and testing tools.
-
Experience working with Databricks or comparable enterprise data platforms.
-
Familiarity with Responsible AI and AI governance frameworks, including NIST AI RMF, EU AI Act, ISO/IEC 42001, or similar standards.
-
Strong analytical, problem-solving, and critical-thinking skills.
-
Ability to communicate complex technical findings effectively to both technical and non-technical stakeholders.
-
Experience conducting bias testing, red-team exercises, adversarial testing, and AI risk assessments is highly desirable.
About the team.
The Responsible-AI / Eval Specialist is part of McCain’s AI governance and delivery ecosystem and works closely with the Use Case Governance & Operations Lead. The role partners across AI Platform & Engineering, Knowledge Engineering, Forward Deployment Engineering, Enterprise Security, and business stakeholders to ensure AI use cases meet quality, risk, and Responsible AI standards before deployment. This position plays a critical role in advancing safe, reliable, and trustworthy AI solutions across the organization.
About McCain.
Click Here to learn more about McCain and how we provide opportunities to make an impact that matters.
Leadership principles.
At McCain, our leadership principles guide how we engage with customers, collaborate as a team, and achieve success. We focus on understanding customer needs, driving innovation, empowering people, and taking ownership to clear obstacles and deliver results.
AI & Digital Capabilities.
At McCain, AI and digital capabilities are part of how we work and innovate. We value curiosity toward emerging technologies and the thoughtful, responsible use of AI-enabled tools to improve outcomes as part of our future-ready culture.
The McCain experience.
We are McCain. This statement is a testament to our collective strength and our individual value. Your contributions play a vital role in our success. Our winning culture is rooted in authenticity and trust, empowering us to bring out the best in one another. Here, you’ll find opportunities to learn, grow, and thrive. Join us and experience why we’re better together.
Our purpose is grounded in building meaningful relationships. We’re big believers in the power of working together in person—it helps us stay connected, collaborate more effectively, and grow as a team. At the same time, we recognize the importance of flexibility. Most office-based roles follow a hybrid model, with the option to work remotely two days a week. There may be exceptions depending on the role and location, so we encourage you to speak with your recruiter for more details.