NASSCOM Campus, Sector 126, Noida, NCR
Role Summary
The AI Safety Research & Testing Professional will play a key technical role in advancing AI safety research and evaluation capabilities. The role involves designing and conducting structured, code-based evaluations of publicly accessible AI models and AI-enabled products, and developing Python-based evaluation pipelines to support these activities.
The Professional will conduct hands-on technical testing, adversarial evaluation, failure analysis, and comparative benchmarking, while monitoring emerging AI safety and security research. The role will translate empirical findings, published frameworks, and established standards into practical guidance, readiness assessments, and recommendations for industry members and partners.
This is a hands-on, technically focused research and testing role covering model APIs, open-source evaluation frameworks, evaluation datasets, automated testing, RAG, tool use, agents, and other AI-enabled application components.
.career-details-job-brief p{ padding: 0px!important; } .career-details-job-brief ul li span{ width: 92%!important; }
Key Responsibilities
Research & Landscape Analysis
Conduct structured literature reviews of AI safety and security research, including reports from AI safety institutes, evaluation labs, and research organisations.
Track frontier and open-weight model releases and assess their safety and security implications.
Monitor emerging research on evaluation, red-teaming, robustness, interpretability, alignment, AI security, RAG, and agentic systems.
Assess relevant evaluation frameworks, benchmarks, attack methodologies, and testing tools.
Contribute to the Global & Regional AI Safety Landscape briefing and internal knowledge base.
Technical Testing & Evaluation
Design and implement evaluation pipelines using open-source frameworks and custom Python tooling to test publicly accessible models and AI-enabled products.
Build and maintain test scripts, model/API integrations, structured test cases, scoring rubrics, automated grading, and reporting pipelines.
Test key risk areas including prompt injection, indirect injection, jailbreaks, data leakage/PII exposure, hallucination, adversarial robustness, bias/fairness, over- and under-refusal, and agent/tool-use safety.
Evaluate application-level risks across RAG, system prompts, tool/function calling, external APIs, memory/state, and agentic workflows.
Conduct structured adversarial testing and red-teaming, including multi-turn and adaptive attacks where appropriate.
Develop reusable attack libraries, evaluation datasets, edge cases, and regression test suites.
Run comparative benchmarking across open-weight and commercial models and products, combining quantitative scorecards with qualitative failure analysis.
Evaluation Methodology & Failure Analysis
Design rigorous and reproducible evaluations covering objectives, test-set construction, sampling, baselines, scoring criteria, assumptions, and limitations.
Apply appropriate quantitative and statistical methods, including repeated trials, uncertainty analysis, and error analysis.
Assess the reliability and limitations of automated evaluation, including LLM-as-judge approaches and human-evaluation agreement where relevant.
Investigate, reproduce, classify, and document failures and identify likely model-, prompt-, data-, retrieval-, tool-, or application-level causes.
Assess severity and prioritise findings based on likelihood, exploitability, impact, and potential downstream consequences.
Automation, Regression & Mitigation Testing
Develop automated safety regression testing to detect behavioural changes across model versions, prompts, fine-tuning, RAG configurations, tools, and application releases.
Maintain evaluation pipelines and results repositories for longitudinal tracking and benchmarking.
Evaluate proposed safeguards through pre- and post-mitigation testing and verify that fixes do not introduce new failure modes.
Apply appropriate version control, logging, configuration management, and reproducibility practices.
Advisory, Training & Knowledge Sharing
Support AI safety/security readiness assessments, gap analyses, model risk registers, and incident-response playbooks.
Translate testing results into practical recommendations for AI risk identification, evaluation, monitoring, and management.
Contribute technical content to workshops and hands-on AI safety testing modules.
Present methodologies and findings at relevant conferences, working groups, and technical dialogues.
Required Qualifications
2-3 years of relevant experience in AI/ML engineering, applied ML research, AI safety, AI red-teaming, security research, data science, or a related technical field.
Solid Python proficiency, including REST APIs, JSON, pandas/requests, and development of maintainable evaluation scripts and pipelines.
Hands-on experience with an LLM evaluation/red-teaming framework or demonstrable experience building custom evaluation tooling.
Working understanding of modern LLM development and deployment, including pretraining, fine-tuning/RLHF, RAG, system prompting, tool use, and agentic architectures.
Working knowledge of major AI safety/security failure categories and ability to reason about attack surfaces and abuse cases.
Comfortable reading and synthesising technical research papers and evaluation reports.
Familiarity with Git, version control, testing, debugging, and reproducible development practices.
Bachelor's degree in Computer Science, AI/ML, Data Science, Cybersecurity, or a related technical field; Master's/research experience is a strong plus.
Preferred / Nice-to-Have
Experience with AI red-teaming, CTFs, security research challenges, or adversarial evaluation.
Experience developing AI safety benchmarks, evaluation datasets, attack libraries, or reusable evaluation harnesses.
Experience with automated red-teaming, jailbreak research, prompt-injection testing, or adaptive attacks.
Experience with LLM-as-judge evaluation, human evaluation, or evaluation reliability analysis.
Experience with agent frameworks or RAG pipelines such as LangChain, AutoGen, or similar.
Familiarity with AI/ML security risks such as model extraction, membership inference, memorisation, data poisoning, or RAG/agent security.
Experience with Docker, cloud environments, CI/CD, experiment tracking, or evaluation infrastructure.
Public/open-source contributions related to AI safety, ML, security, evaluation tooling, or research.
Participation in an AI safety fellowship, programme, research group, or professional community.
Apply Here