Office location : Gurgaon
Role: Data Scientist
Experience: 3–5 Years
Key Responsibilities
Design, develop, and deploy Machine Learning and Generative AI solutions.
Build Retrieval-Augmented Generation (RAG) pipelines using vector databases and enterprise knowledge sources.
Develop AI agents using Agentic AI frameworks such as LangGraph, LangChain, CrewAI, or similar technologies.
Integrate AI agents with enterprise APIs, tools, databases, and external services.
Develop prompts, tool-calling workflows, and structured output pipelines for LLM applications.
Fine-tune, evaluate, and optimize LLM-powered applications for accuracy, latency, and cost.
Implement data preprocessing, feature engineering, and ML model training workflows.
Work with structured and unstructured datasets to solve business problems.
Collaborate with Product Managers, Software Engineers, and Subject Matter Experts to deliver AI-driven features.
Monitor model and agent performance and participate in troubleshooting and continuous improvements.
Write clean, maintainable, and well-tested Python code following engineering best practices.
Stay updated with the latest advancements in Machine Learning, LLMs, and Agentic AI technologies.
Required Technical Skills
Core Skills
Strong proficiency in Python
Machine Learning fundamentals
Natural Language Processing (NLP)
Generative AI and Large Language Models (LLMs)
Prompt Engineering
Retrieval-Augmented Generation (RAG)
Embeddings and semantic search
Model evaluation and validation techniques
Agentic AI Frameworks
Hands-on experience with LangChain and LangGraph
Experience building AI agents with tool calling and workflow orchestration
Familiarity with CrewAI, AutoGen, Semantic Kernel, or similar frameworks
Understanding of agent memory, planning, state management, and multi-step reasoning
ML & AI Libraries
Scikit-learn
XGBoost or LightGBM
PyTorch or TensorFlow
Hugging Face Transformers
OpenAI, Anthropic, Gemini, Bedrock, Azure OpenAI, or similar LLM APIs
Vector databases such as Pinecone, FAISS, ChromaDB, Weaviate, Milvus, or OpenSearch
Data & Cloud
SQL and relational databases
Experience with AWS, Azure, or GCP
Docker and containerized deployments
Basic CI/CD knowledge
MLflow or similar experiment tracking tools
REST APIs/FastAPI for AI model deployment
Good to Have
Experience building production-ready AI or LLM applications.
Exposure to multi-agent systems and workflow orchestration.
Knowledge of Model Context Protocol (MCP).
Experience with AI evaluation frameworks and guardrails.
Understanding of MLOps and model monitoring.
Experience with fine-tuning techniques such as LoRA, PEFT, or QLoRA.
Experience with document processing, OCR, or document intelligence.
Experience in legal, regulatory, financial, healthcare, or publishing domains.