AI Engineer – Agentic Cybersecurity
Location: Mumbai
Work Mode: Work From Office (WFO)
Experience: 2–3+ Years
About Redfox Cybersecurity
Redfox Cybersecurity is a cybersecurity company focused on offensive security, penetration testing, adversary simulation, cybersecurity training, and innovative security solutions.
We are building the next generation of AI-powered cybersecurity products that combine AI agents with established penetration-testing capabilities, security tooling, and human expertise.
We are looking for an AI Engineer to join our team and help build the AI layer of our AI-powered penetration-testing platform. This is a hands-on role involving LLMs, AI agents, reasoning systems, tool orchestration, RAG, model evaluation, AI infrastructure, and production deployment.
You will work closely with our Full Stack Developer, penetration testers, and security researchers to build, integrate, and continuously improve the AI capabilities of the platform.
The ideal candidate is someone who is comfortable taking an AI capability from idea → research → prototype → development → product integration → production.
Key ResponsibilitiesAI & Agent Development
- Design and develop AI agents capable of performing complex, multi-step penetration-testing tasks.
- Build the core reasoning and action loop, enabling the AI to understand objectives, determine the next action, execute tools, interpret results, and continue toward an objective.
- Develop systems that maintain context and state across multiple steps, allowing information gathered during an assessment to influence subsequent actions.
- Build AI-driven tool orchestration that allows LLMs to interact with command-line tools, scripts, APIs, and custom cybersecurity utilities.
- Implement mechanisms for self-critique, validation, and verification to distinguish genuine security findings from false positives.
- Develop AI capabilities that can understand and operate within defined scope, methodologies, skills, rules, and engagement-specific requirements.
- Build reusable AI capabilities that can support the platform's different penetration-testing use cases, including Web, API, Internal Network, External Network, Active Directory, Source Code, Cloud, and Firewall assessments.
- Develop mechanisms for the AI to dynamically select and use appropriate tools based on the assessment context.
LLM & Generative AI Engineering
- Work extensively with Large Language Models and Generative AI for cybersecurity applications.
- Design and optimize prompts, system instructions, structured outputs, function calling, and tool-use mechanisms.
- Build LLM-based decision-making and reasoning systems capable of handling long-running, multi-step tasks.
- Implement RAG and retrieval systems to provide the AI with relevant penetration-testing knowledge, methodologies, engagement documentation, and other contextual information.
- Work with embeddings and vector databases such as FAISS, PGVector, Weaviate, Milvus, or equivalent technologies.
- Evaluate and experiment with different foundation models based on capability, accuracy, latency, context window, reliability, and cost.
- Implement model routing strategies that allow different models to be used for different tasks.
- Fine-tune or adapt models using techniques such as LoRA, QLoRA, or other parameter-efficient approaches, where required.
- Build model abstraction layers that allow different models and inference providers to be evaluated and integrated without significant changes to the core platform.
AI + Cybersecurity Integration
- Work closely with Redfox penetration testers and security researchers to understand real-world penetration-testing requirements and incorporate them into the development and continuous improvement of the AI-powered penetration-testing tool.
- Understand real-world penetration-testing techniques, tool usage, assessment requirements, and common challenges faced by security professionals.
- Integrate cybersecurity tools and custom scripts into the AI-powered platform.
- Enable the AI to interpret outputs from tools such as Nmap, web fuzzers, BloodHound, and custom security tooling.
- Develop mechanisms for dynamically loading capabilities and skills based on the assessment type and engagement requirements.
- Work with the cybersecurity team to continuously improve the AI's ability to identify, validate, and reproduce security findings.
- Contribute to the development of Redfox's proprietary AI and cybersecurity capabilities.
AI Evaluation & Quality
- Build an evaluation and benchmarking framework for testing AI models and agents against controlled penetration-testing environments.
- Develop evaluation datasets and scenarios covering reconnaissance, vulnerability discovery, exploitation, multi-step attack chains, and finding validation.
- Measure AI performance using metrics such as findings discovered, true-positive/false-positive rate, finding validation rate, multi-step task completion, tool-selection accuracy, context/state preservation, latency, token usage, and cost per assessment.
- Run regression testing whenever models, prompts, agent behaviour, or other AI components are changed.
- Benchmark different models and configurations to identify the most effective approach for different tasks.
- Identify and address AI failure modes such as hallucination, incorrect tool selection, loss of context, incomplete task execution, and incorrect conclusions.
AI Infrastructure & Production Engineering
- Build scalable infrastructure for AI inference, agent execution, data processing, evaluation, and monitoring.
- Develop production-grade AI services and expose AI capabilities through REST APIs, WebSockets, or other appropriate interfaces.
- Work with the Full Stack Developer to integrate AI capabilities into the platform's backend and user-facing product.
- Build mechanisms for managing long-running AI sessions, execution state, task queues, and intermediate results.
- Work with Docker and cloud infrastructure, particularly GCP, to deploy and operate AI services.
- Optimize AI systems for latency, throughput, reliability, scalability, and inference cost.
- Implement caching, batching, model selection, context management, and other inference optimization techniques.
- Implement logging, tracing, monitoring, and observability across AI components.
- Maintain appropriate model, prompt, agent, and capability versioning to ensure reproducibility.
Product Development & Collaboration
- Work closely with the Full Stack Developer to convert AI capabilities into functional product features.
- Define APIs, data contracts, and interfaces between the AI engine and application layer.
- Participate in architecture and technical design discussions for the AI platform.
- Rapidly prototype new AI capabilities and take successful prototypes through productionisation.
- Research emerging AI technologies and evaluate their practical applicability to Redfox's products.
- Contribute to the development of Redfox's proprietary AI technology and intellectual property.
RequirementsMust Have
- 2–3+ years of hands-on experience in AI/ML, Applied AI, AI Engineering, ML Engineering, or a closely related field.
- Demonstrated experience building and deploying AI/ML or LLM-based applications, preferably with at least one project taken to production.
- Strong programming skills in Python and solid software engineering fundamentals.
- Strong hands-on experience with LLMs and Generative AI.
- Practical experience building AI agents or agentic systems.
- Strong understanding of:
- Prompt engineering
- Function/tool calling
- Structured outputs
- RAG
- Embeddings
- Vector databases
- Context and state management
- LLM evaluation
- Experience integrating LLMs with external tools, APIs, scripts, or execution environments.
- Experience with PyTorch, TensorFlow, or equivalent ML frameworks.
- Experience with one or more vector databases such as FAISS, PGVector, Weaviate, or Milvus.
- Experience with model adaptation/fine-tuning techniques such as LoRA, QLoRA, or equivalent.
- Experience building AI APIs and backend services.
- Strong understanding of REST APIs and service-oriented architectures.
- Experience with Docker and cloud platforms such as GCP, AWS, or Azure.
- Understanding of CI/CD, Git, testing, observability, and production deployment.
- Experience with MLflow or equivalent experiment/model management systems.
- Ability to evaluate and benchmark AI systems using measurable performance criteria.
- Strong problem-solving ability and a product-oriented engineering mindset.
Strongly Preferred
- Experience building autonomous or semi-autonomous AI agents.
- Experience with agent orchestration frameworks or custom agent orchestration systems.
- Experience with tool-use and function-calling architectures.
- Experience building systems involving long-running and multi-step AI tasks.
- Experience with model serving and inference optimization.
- Experience with open-weight LLMs and platforms such as Hugging Face or equivalent.
- Experience with quantization, batching, caching, and inference optimization.
- Experience building evaluation environments for AI agents.
- Experience with cybersecurity, penetration testing, vulnerability assessment, or security tooling.
- Strong understanding of OWASP, penetration-testing methodologies, offensive security tooling, or MITRE ATT&CK.
Good to Have
- Experience building AI-powered cybersecurity products.
- Experience building AI agents for security testing, vulnerability discovery, or security operations.
- Experience working with tools such as Nmap, Burp Suite, BloodHound, web fuzzers, or equivalent security tooling.
- Understanding of Web Application, API, Network, or Active Directory security.
- Experience with source-code analysis or security configuration data.
- Experience with cloud security or cloud configuration analysis.
- Open-source AI/ML contributions, GitHub projects, technical publications, or relevant personal projects.
- Research experience in LLMs, autonomous agents, NLP, or applied machine learning.
The platform architecture places the Agent Core at the centre of this system, responsible for reasoning, model interaction, tool execution, engagement state, self-critique, and evidence generation.
The platform is intended to autonomously perform and assist penetration testing across multiple assessment types, with the AI capable of maintaining context across chained, multi-step activities.
You will work alongside the Full Stack Developer to take these AI capabilities and integrate them into a complete, usable product.
What We Look For
We are looking for an engineer who can build, experiment, evaluate, debug, and continuously improve AI systems.
You should be comfortable going beyond simply calling an LLM API. You should understand how an AI agent can maintain state, make decisions, use tools, interpret outputs, validate its conclusions, and operate reliably across complex multi-step tasks.
You should also be comfortable working with cybersecurity professionals to understand how the tool needs to behave in real-world penetration-testing scenarios and translate those requirements into working AI capabilities within the product.
Most importantly, we are looking for someone who wants to build an AI-powered cybersecurity product from the ground up and contribute to Redfox's proprietary technology and IP.
Work Mode
- Work From Office (WFO)
- Andheri West, Mumbai
Apply Now
If you are excited about AI Agents, LLMs, cybersecurity, and building production-grade AI products from the ground up, we would love to hear from you.
Please share your CV along with relevant:
- AI/ML projects
- AI Agent / LLM projects
- GitHub repositories
- Production AI systems
- Research / publications
- Cybersecurity or offensive-security projects
Pay: ₹500,000.00 - ₹500,001.00 per year
Benefits:
Work Location: In person