Primary Objective: Focus strictly on the "AI application layer” building prompt templates, tuning RAG (Retrieval-Augmented Generation) pipelines, managing local vector databases, and testing model inference engines.
Target Experience: 1–2+ years in Python development with hands-on GitHub/project experience in local LLMs and open-source AI frameworks.
Key Requirements & Core Competencies:
Python & Open-Source AI Stack:
Strong proficiency in Python and hands-on experience with modern AI orchestration frameworks: LangChain, LlamaIndex, or Haystack.
Familiarity with local model serving engines like vLLM, Ollama, or TGI (Text Generation Inference).
Prompt Engineering & NL-to-SQL Logic:
Demonstrated skill in writing, testing, and refining system prompts, specifically for T-SQL query generation, structured JSON output formatting, and context-grounded Q&A.
Understanding of guardrails to prevent hallucination and SQL injection (blocking DDL/DML execution).
RAG & Vector Database Management:
Hands-on knowledge of embedding models and local vector stores (PGVector, ChromaDB, or FAISS).
Experience chunking, indexing, and retrieving unstructured data (PDFs, clinical notes, medical docs) for context ingestion.
Model Quantization & Edge Optimization:
Basic understanding of model quantization formats (GGUF, AWQ, GPTQ) to run small-to-medium parameter models (3B, 8B, 14B like Qwen or Llama-3) on constrained hardware/VRAM footprints.
Pay: ₹500,000.00 - ₹800,000.00 per year
Benefits:
Work Location: In person