AI Safety Researcher
ROLE OVERVIEW
The AI Safety Researcher will lead adversarial testing, LLM and agentic AI red teaming, model-security assessment, Responsible AI evaluation, and AI-specific threat modelling across government AI programmes. The role combines applied research with repeatable testing, technical advisory, governance evidence, and open-source knowledge transfer.
Educational Qualifications
- B.Tech./B.E. or M.Tech./M.S./M.Sc. in Computer Science, Information Security, AI/ML, or a related quantitative discipline is required.
- An advanced degree with thesis or published work in adversarial ML, AI security, LLM safety, or Responsible AI is highly desirable.
- Desirable certifications include OSCP, GWAPT, CEH with demonstrable AI/ML security work, ML foundations, MLSecOps, or LLM-security credentials.
- Candidates with strong adversarial ML research, safety publications, credible LLM red-team disclosures, or CTF/red-team achievements may be considered in lieu of conventional qualifications.
Experience
- 6+ years in ML, applied AI, security research, or a related discipline, including at least 3 years in adversarial ML, AI red teaming, LLM safety evaluation, or AI/ML security research.
- Demonstrable LLM red-teaming experience covering prompt injection, jailbreaks, indirect prompt injection, data exfiltration, and agentic tool misuse in production or production-like systems.
- Hands-on adversarial ML experience involving evasion, model inversion, membership inference, model extraction, or data poisoning against non-trivial models.
- Experience in government, BFSI, healthcare, or another regulated sector is a strong advantage.
- Experience advising and upskilling engineering, security, compliance, or programme teams on AI-specific risks is desirable.
Key Responsibilities
- Adversarial ML and Model Security: design and execute attack campaigns covering evasion, inversion, membership inference, extraction, and poisoning against document intelligence, predictive analytics, fraud, computer vision, face-embedding, and liveness systems.
- Develop and maintain reusable adversarial ML attack libraries, evaluation harnesses, reproducible notebooks, quantitative robustness measures, and prioritised mitigation guidance.
- LLM, RAG, and Agentic AI Red Teaming: lead structured exercises covering jailbreaks, direct and indirect prompt injection, retrieved-document attacks, data exfiltration, unsafe tool invocation, sandbox escape, multi-turn manipulation, and policy-boundary violations.
- Develop red-team playbooks and anonymised evaluation sets for government AI use cases, and advise teams on guardrails, output filtering, source integrity, retrieval provenance, and agent policy design.
- Hallucination and Responsible AI Evaluation: measure hallucination, groundedness, faithfulness, calibration, uncertainty, retrieval quality, bias, subgroup performance, and fairness.
- Conduct explainability evaluation using SHAP, LIME, or Captum, and perform lightweight privacy impact assessments for model and data pipelines.
- Author safety content for model cards, dataset sheets, bias and hallucination reports, safety evaluations, and Responsible AI evidence aligned with relevant Indian government frameworks.
- AI Security Architecture and Threat Modelling: own AI-specific threat models using STRIDE, MITRE ATT&CK, MITRE ATLAS, OWASP Top 10 for LLMs, and OWASP ML guidance across data, model, retrieval, prompt, tool, and output surfaces.
- Review deployment patterns across sovereign/on-premise GPU environments, IndiaAI Compute, and CSP APIs; recommend defensive controls and architecture patterns.
- Contribute AI-security and Responsible AI requirements to RFQs, RDRs, procurement documents, and agency evaluations.
- Advise AI QA Engineers and delivery teams on safety test design, sampling, metrics, evidence capture, and remediation priorities.
- Track current research and translate relevant findings into practical controls, tests, playbooks, and programme standards.
- Maintain open-source safety toolkits and publish reusable notebooks, evaluation harnesses, playbooks, and datasets to approved government repositories.
Technical Competencies
- Advanced Python; PyTorch or TensorFlow; Hugging Face Transformers; NumPy, pandas, and scikit-learn.
- Adversarial ML frameworks such as ART, Foolbox, CleverHans, or equivalent, including custom attack and defence implementation.
- Production-relevant LLM and agentic AI red teaming, including guardrail frameworks such as NeMo Guardrails, Guardrails AI, or Llama Guard.
- RAG and generative evaluation using RAGAS or equivalent, groundedness, faithfulness, retrieval metrics, calibration, and uncertainty quantification.
- Threat modelling with STRIDE, MITRE ATT&CK, MITRE ATLAS, OWASP Top 10 for LLMs, and OWASP ML Top 10.
- Explainability and fairness using SHAP, LIME, Captum, Fairlearn, and AI Fairness 360.
- Working knowledge of differential privacy, federated learning, PII redaction, anonymisation, and privacy impact assessment.
- Knowledge of IndiaAI Safe & Trusted AI, MeitY Responsible AI and security guidance, CERT-In directions, and DPDPA 2023.
- Ability to produce clear technical reports and brief mixed engineering, architecture, compliance, and executive audiences.
Work Location: In person