Overview
We are looking for an AI Platform Engineer to build, deploy, optimize, and manage production AI inference systems. The ideal candidate has experience deploying machine learning, computer vision, or LLM models at scale and understands GPU acceleration, inference optimization, and cloud-native AI infrastructure. This role bridges Machine Learning and Infrastructure Engineering. You will work closely with AI engineers to deploy models, optimize inference performance, automate model delivery, and build scalable AI serving platforms.
Responsibilities
AI Model Deployment
● Deploy and manage production AI models including Computer Vision, NLP, and Large Language Models (LLMs)
● Build scalable inference services for real-time and batch workloads
● Deploy and manage NVIDIA Triton Inference Server
● Package models using ONNX, TensorRT, TorchScript, or other optimized formats
● Optimize inference latency, throughput, and GPU utilization
AI Infrastructure
● Build and maintain GPU-enabled Kubernetes clusters
● Configure GPU scheduling and resource allocation
● Deploy inference services using Docker, Kubernetes, and Helm
● Manage model repositories and versioning
● Implement scalable APIs for AI inference
Performance Optimization
● Profile inference latency and throughput
● Optimize GPU memory usage
● Benchmark different model formats (PyTorch, ONNX, TensorRT)
● Configure dynamic batching and concurrent inference
● Troubleshoot production inference bottlenecks
ML Platform
● Build CI/CD pipelines for AI models
● Automate model deployment
● Support canary deployments and A/B testing
● Implement rollback strategies
● Support automated model retraining workflows
Monitoring & Operations Monitor production AI systems including:
● GPU utilization
● GPU memory
● CPU utilization
● Model latency
● Request throughput
● Queue length
● Failed inference requests
● Model accuracy drift Build dashboards using Prometheus and Grafana.
Required Skills AI Model Serving
● Triton Inference Server
● TensorRT
● ONNX Runtime
● PyTorch or TensorFlow deployment
● FastAPI or similar inference frameworks
Containerization & Cloud
● Docker
● Kubernetes
● Helm
● GPU-enabled containers
● NVIDIA Container Toolkit
GPU Computing
● CUDA fundamentals
● GPU scheduling
● Multi-GPU inference
● NVIDIA GPU monitoring
● Performance profiling
Programming
● Python
● Bash
● REST APIs
● Linux
Nice to Have
● Experience deploying LLMs using vLLM, TensorRT-LLM, or NVIDIA NIM
● Experience deploying Computer Vision pipelines
● Experience with Ray Serve, KServe, or Seldon Core
● Experience with Kubeflow or MLflow
● Experience with NVIDIA TAO Toolkit
● Experience with distributed inference systems
● Experience optimizing TensorRT engines
● Experience with model quantization (FP16/INT8)
Preferred Background
● AI Platform Engineering
● ML Infrastructure Engineering
● Computer Vision Deployment
● LLM Deployment
● AI Inference Engineering
● GPU Platform Engineering
WhatsApp: +91 9159359391
Mail ID: [email protected]
Pay: Up to ₹900,000.00 per year
Benefits:
Work Location: In person