Hiring: LLM Inference Performance Engineer (CUDA / Triton / GPU Kernel Optimization) – Freelance Job Support
Job Title: LLM Inference Performance Engineering & GPU Kernel Optimization Expert
Job Type: Part-Time | Freelance | Remote
Work Schedule: Monday to Friday – 2 Hours/Day
Job Description
We are looking for an experienced LLM Inference Performance Engineering & GPU Kernel Optimization Expert to provide job support for a client. This is a remote freelance opportunity for professionals with strong expertise in optimizing Large Language Model (LLM) inference workloads.
Required Skills
- Deep expertise in CUDA programming
- Strong hands-on experience with Triton
- Advanced knowledge of NVIDIA GPUs and GPU architecture
- Experience in LLM inference optimization
- GPU Kernel Optimization and Performance Tuning
- CUDA Kernel Development and Debugging
- Memory Optimization, Profiling, and Latency Reduction
- Experience with Transformer-based models and inference frameworks
- Familiarity with PyTorch, TensorRT, vLLM, or similar inference frameworks is an added advantage
- Excellent troubleshooting and problem-solving skills
Responsibilities
- Provide remote job support for ongoing client projects
- Optimize GPU kernels and improve LLM inference performance
- Analyze and resolve performance bottlenecks
- Guide the client on CUDA, Triton, and NVIDIA GPU optimization techniques
- Deliver high-quality technical support during scheduled sessions
Job Details
- Employment Type: Freelance / Part-Time
- Mode: Remote
- Support Hours: 2 Hours per day (Monday to Friday)
- Experience Required: 5+ Years (Preferred)
- Joining: Immediate
If you have strong experience in CUDA, Triton, NVIDIA GPU optimization, and LLM inference performance engineering, we'd love to hear from you.
Pay: ₹25,000.00 - ₹30,000.00 per month
Experience:
- LLM InLLM & GPU Kernel Optimization: 5 years (Required)
Work Location: Remote