M U S T - H A V E S K I L L S
Backend Engineering
- Python, FastAPI / Flask
- REST APIs, Async Architecture
- Concurrency, Multiprocessing
- Memory Management & Performance Optimization
- Docker, Kubernetes, CI/CD
- AWS / GCP / Azure
Machine Learning
- 5+ years in ML Development
- Research → Prototype → Production pipeline experience
- PyTorch, TensorFlow
- Feature Engineering & Data Preprocessing
- Model Monitoring & Drift Detection
- MLflow, Weights & Biases, Airflow
LLM / Generative AI
- RAG (Retrieval-Augmented Generation)
- Prompt Engineering
- LoRA / PEFT Fine-Tuning
- LLaMA, Mistral & Open-Source LLMs
- Embeddings, Vector Databases & Similarity Search
- Context Window Optimization
- Cost-Aware LLM Design
- Hallucination Detection, Prompt Testing & Output Validation
VLM / Multimodal AI
- Multimodal Transformers
- CLIP, BLIP, LLaVA
- Image Embeddings
- Cross-Modal Retrieval
- OCR Integration
- Image Preprocessing Pipelines
programming.com | Confidential & Internal Use Only | Confidential Use Only
J O B D E S C R I P T I O N
Senior AI/ML Engineer | LLM, VLM & Generative AI
AI Architecture
- Scalable LLM Systems
- RAG Architecture Design
- Caching Strategies & Token Optimization
- Streaming Responses
- Kafka / RabbitMQ
- Async Task Execution
- Prompt Injection Prevention
- Data Privacy & Rate Limiting
G O O D - T O - H A V E S K I L L S
- Quantization (4-bit / 8-bit)
- Model Distillation
- vLLM / TGI
- AI Agents & Multi-Agent Systems
- GPU Orchestration
- Distributed Training
- Triton Inference Server
- Model Sharding
- GPU Memory Monitoring
- Autoscaling Inference Endpoints
Pay: ₹805,906.54 - ₹1,782,322.22 per year
Benefits:
- Paid sick time
- Paid time off
- Provident Fund
Work Location: In person