Job Description
AI Lead
Role Overview
We're looking for a hands-on Technical Program Lead to design and deliver enterprise GenAI/SLM solutions, including air-gapped, on-prem, and sovereign deployments. You'll own architecture end-to-end — model selection, infra, deployment, and governance — while leading delivery and the client relationship.
Key Responsibilities
● Architect GenAI/SLM solutions (RAG, agentic workflows, fine-tuning/distillation) suited to customer security and data-sensitivity constraints.
● Evaluate SLMs vs. LLMs (Phi, Mistral, Llama, Qwen, etc.) on cost, latency, and accuracy trade-offs.
● Design air-gapped/offline deployments — local inference, vector stores, and secure model/data update pipelines with no external dependency.
● Architect across hybrid environments: AWS/Azure/GCP, private cloud, and on-prem data centers, optimizing GPU/CPU cost and performance.
● Define AI governance: model evaluation, guardrails, audit logging, and responsible-AI practices — including offline-compatible monitoring for restricted environments.
● Lead client discovery workshops, translate business requirements into a scoped delivery roadmap, and drive the engagement through to shipment/go-live.
● Own planning and task allocation across the team — break architecture into workstreams, assign to the right engineers, and sequence delivery against client timelines.
● Be the primary point of client interaction throughout the engagement — status updates, scope changes, escalations — not just at kickoff/handoff.
● Drive multiple projects/accounts in parallel, balancing priorities across engagements and flagging capacity or scope risk early.
● Lead a team of engineers/data scientists — planning, reviews, and unblocking delivery.
● Support pre-sales: scoping, estimation, and technical proposals.
Required Skills & Experience
● 8+ years in software/data engineering, 3+ years architecting production ML/GenAI solutions.
● Hands-on with SLMs/LLMs, fine-tuning (LoRA/QLoRA), quantization; Python, LangChain/LlamaIndex, vLLM/Ollama.
● Proven experience with air-gapped or on-premise AI deployment.
● Cloud architecture (AWS/Azure/GCP) plus hybrid/private data center deployment.
● Vector DBs deployable offline (FAISS, Milvus, Weaviate, Qdrant).
● Familiarity with AI governance/compliance frameworks (NIST AI RMF, ISO/IEC 42001) and data residency requirements.
● Docker/Kubernetes and infra-as-code (Terraform/Ansible).
● Expert in Claude-driven development — using Claude Code and Claude-based agents as a core part of the build workflow, including authoring custom Skills/MCP tools and agentic coding pipelines to boost team engineering productivity.
● Reviewer-first mindset: with agents doing most of the generation, your value is in specifying correctly, critically reviewing AI-generated architecture/code, catching subtle design and security flaws, and validating trade-offs — not in hand-writing every line yourself.
Behavioural & Leadership Expectations
● Must have: prior experience leading a small team (formally or as a de facto lead) and working across multiple clients/engagements simultaneously — this is not a first team-lead or first multi-client role.
● Leads a team end-to-end; owns the client relationship from requirement gathering through shipment.
● Spends more time planning, allocating, and reviewing than hand-coding — sets direction, defines specs/guardrails for agentic tooling, allocates tasks across the team, and audits output; comfortable being judged on decision quality and delivery outcomes, not lines of code written.
● Able to run multiple projects/accounts simultaneously without losing quality of client interaction on any one of them.
● Fluent in Agile/Scrum ceremonies; hands-on with JIRA/Confluence for backlog and delivery tracking.
● Self-driven, strong client-facing communicator across technical and non-technical stakeholders.
● Preferred: background in an IT/consulting services company over purely captive/product environments.
Good to Have
● Big Data (Spark/Hive/Hadoop), Graph Analytics, or hardware acceleration (GPU/FPGA) experience.
● Regulated-industry (defense, government, BFSI) AI deployment experience.
● Cloud, security, or AI governance certifications.
● Contribution to open source projects, academic papers published, filled patents
Pay: ₹548,257.38 - ₹3,000,000.00 per year
Application Question(s):
- Do you have Cloud, security, or AI governance certifications ?
- What is your Expected CTC in LPA? (30)
- What is your notice period in days? (30)
Experience:
- overall : 6 years (Required)
- Python: 6 years (Required)
- Agile Software Development: 3 years (Required)
- Machine Learning: 5 years (Required)
- Langchain: 3 years (Required)
- Gen-AI: 6 years (Required)
- LLMs & SLMs: 3 years (Required)
- Cloud Architecture: 2 years (Required)
- Artificial Intelligence: 6 years (Required)
Work Location: In person