ROLE AND RESPONSIBILITIES
- 5–8 years of experience in software engineering, platform engineering,
- DevOps, SRE, MLOps, or related infrastructure roles.
- Strong hands-on experience with Kubernetes, including writing Kubernetes Operators and Custom Resource Definitions using frameworks such as Kubebuilder, Operator SDK, or equivalent.
- Experience designing and operating cloud-native infrastructure on AWS, particularly Amazon EKS, EC2, S3, ECR, IAM, VPC, and CloudWatch.
- Experience deploying and operating machine learning, deep learning, or generative AI models in production.
- Experience running and troubleshooting GPU-accelerated workloads on Kubernetes, with an understanding of GPU scheduling, utilization, memory constraints, and performance.
- Familiarity with model-serving frameworks such as KServe, NVIDIA Triton Inference Server, vLLM, Ray Serve, TorchServe, or equivalent technologies.
- Strong programming experience in Go or Python, with experience building production-grade APIs, controllers, or distributed backend services.
- Experience with containers, Helm, CI/CD, infrastructure as code, and observability tools such as Prometheus, OpenTelemetry, and Grafana.
- Strong understanding of Linux, networking, storage, security, and distributed-system fundamentals.
- Strong debugging, problem-solving, communication, and cross-functional collaboration skills.
REQUIRED SKILLS (explicitly call out must-have and nice-to-have skills expectations)
- Experience building an MLOps platform, AI infrastructure platform, internal developer platform, or Kubernetes-based enterprise product.
- Experience writing GPU kernels or performance-critical code using CUDA C/C++ or Triton is a plus.
- Experience with large language model serving, distributed inference, batching, quantization, or inference-performance optimization.
- Experience with NVIDIA GPU Operator, MIG, GPU time-slicing, Dynamic Resource Allocation, or similar GPU-management technologies.
- Familiarity with AWS Inferentia, Trainium, SageMaker, or Amazon Bedrock.
- Experience operating AI platforms across hybrid-cloud, on-premises, air- gapped, or multi-tenant environments.
Pay: ₹1,000,000.00 - ₹3,000,000.00 per year
Application Question(s):
- Immediate to 15 days preferred
Experience:
- MLOps: 5 years (Preferred)
Work Location: In person