Freelance AI Engineer – Vision-Language-Action (VLA) Models
Overview
We are seeking a freelance AI engineer with expertise in Vision-Language-Action (VLA) models for a fixed tabletop humanoid robot featuring two robotic arms, each with a five-finger dexterous hand, and a head-mounted RGB-D camera.
Responsibilities
- Develop, fine-tune, and evaluate VLA models for tabletop manipulation.
- Build perception pipelines using the head-mounted camera.
- Implement coordinated bimanual manipulation for grasping, pick-and-place, assembly, sorting, and tool use.
- Integrate vision, language, and action into an end-to-end robotic system.
- Optimize models for reliable, real-time performance.
- Document code, experiments, and deployment.
Required Qualifications
- Proven experience with VLA models for robotics.
- Strong proficiency in Python and PyTorch.
- Experience with robotic manipulation, imitation learning, reinforcement learning, or diffusion-based policies.
- Familiarity with ROS2 and simulation platforms Isaac Sim and RViz.
Preferred Qualifications
- Experience with models such as OpenVLA, RT-2, Octo, Pi0, or NVIDIA GR00T.
- Experience with bimanual manipulation and humanoid robotic hands.
- Familiarity with model optimization and deployment on robotic hardware.
Deliverables
- End-to-end VLA pipeline for tabletop manipulation.
- Training and evaluation scripts.
- Deployment-ready code with documentation.
Pay: ₹150,000.00 - ₹300,000.00 per year
Work Location: Remote