Lead and grow a team building the on-device AI inference platform for RTX, RTX Pro, and DGX GPUs, with accountability for execution, technical direction, delivery quality, and roadmap alignment.
Drive cross-functional alignment with NVIDIA’s software, research, architecture, and product teams, along with industry partners and open-source communities, to build strategy and strengthen the AI ecosystem across RTX and DGX platforms.
Provide technical leadership for the architecture and evolution of modern inference runtimes and execution stacks across frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX, spanning workloads including LLMs, vision-language models, TTS, ASR, and diffusion models.
Mentor engineers, develop technical leaders, and foster a high-performance team culture centred on innovation, collaboration, and operational excellence.
Coordinate end-to-end optimization of AI models, data pipelines, and inference runtimes to improve performance across current and next-generation GPU architectures.
Drive adoption of model optimization techniques such as quantization, pruning, sparsity, and distillation to enable efficient deployment of large models on local and edge devices.
Establish team processes for system-level debugging, performance optimization, and performance-accuracy trade-off analysis, including infrastructure for performance and accuracy sweeps, gap analysis, and production-readiness improvements.