We are seeking an experienced AI/ML Architect to lead the design and development of scalable, real-time AI systems. You will work closely with product, data, and engineering teams to architect end-to-end solutions — from model development and deployment to system integration and production monitoring.
Design and architect AI/ML systems that are scalable, low-latency, and production-ready
Lead development of real-time inference pipelines for use cases like voice, vision, or NLP
Select and integrate appropriate tools, frameworks, and infrastructure (e.g., Kubernetes, Kafka, TensorFlow, PyTorch, ONNX, Triton, VLLM etc.)
Collaborate with data scientists and ML engineers to productionize models
Ensure reliability, observability, and performance of deployed systems
Conduct architecture reviews, POCs, and system optimizations
Mentor engineers and help set best practices for ML lifecycle (MLOps)
6+ years of experience building and deploying ML systems in production
Proven expertise in real-time, low-latency system design (e.g., streaming inference, event-driven pipelines)
Strong understanding of scalable architectures — microservices, message queues, distributed training/inference
Proficient in Python and popular ML/DL frameworks (scikit-learn, TensorFlow, PyTorch)
Hands-on experience with LLM inference optimization using frameworks like vLLM, TensorRT-LLM, and SGLang
Familiarity with vector databases, embedding-based retrieval, and RAG pipelines
Experience with containerized environments (Docker, Kubernetes) and managing multi-container applications
Working knowledge of cloud platforms (AWS, GCP, or Azure) and CI/CD practices for ML workflows
Exposure to edge deployments and model compression/optimization techniques
Strong foundation in software engineering principles and system design
Experience in Linux (Ubuntu)
Terminal/Bash Scripting