About O-HIVE
O-HIVE develops Visual Language Model (VLM) technology that enables machines and robots to see, understand, reason, and execute in real-world environments.
Our technology combines multimodal AI, computer vision, spatial intelligence, robotics, and edge computing to build intelligent systems for manufacturing, robotics, visual inspection, construction, and other industrial applications.
We are looking for an AI Language Model Engineer to develop, optimize, and deploy language and multimodal models that power O-HIVE’s VLM platform and robotics products.
Role Overview
As an AI Language Model Engineer, you will work on the development and optimization of Large Language Models (LLMs), Vision-Language Models (VLMs), and multimodal reasoning systems.
You will work closely with our computer vision, robotics, embedded, and semiconductor teams to develop models capable of understanding visual information, reasoning about physical environments, generating structured outputs, and controlling downstream robotic or software systems.
The ideal candidate has strong experience with Transformer-based models, PyTorch, model fine-tuning, inference optimization, and modern multimodal AI architectures.
Key Responsibilities
- Develop and improve O-HIVE's LLM and VLM architectures for industrial and robotic applications.
- Fine-tune and adapt open-source language and multimodal models for O-HIVE-specific use cases.
- Develop multimodal pipelines combining:
- Images and video
- Natural language
- Spatial and 3D information
- Sensor data
- Robot state and control information
- Design prompting, instruction-tuning, and structured-output methods for reliable machine reasoning.
- Develop training and fine-tuning pipelines using techniques such as:
- Supervised Fine-Tuning (SFT)
- LoRA / QLoRA
- Parameter-Efficient Fine-Tuning (PEFT)
- Distillation
- Quantization
- Preference optimization
- Improve model reasoning, grounding, object understanding, spatial reasoning, and task planning.
- Optimize models for low-latency edge inference on GPUs, embedded AI devices, and future O-HIVE custom AI hardware.
- Evaluate model accuracy, latency, memory consumption, hallucination rate, and task reliability.
- Develop datasets and evaluation benchmarks for visual reasoning and industrial AI applications.
- Integrate language models with computer vision, detection, segmentation, depth estimation, 3D mapping, and robotic-control pipelines.
- Build APIs and inference services for O-HIVE's VLM Cloud and Edge AI products.
- Research emerging LLM/VLM architectures and translate relevant research into production systems.
- Collaborate with hardware and ASIC engineers to define model requirements for future O-HIVE VLM accelerator chips.
Required Qualifications
- Bachelor's, Master's, or Ph.D. degree in Computer Science, Artificial Intelligence, Machine Learning, Electrical Engineering, or a related field.
- Strong programming skills in Python.
- Strong experience with PyTorch and modern deep-learning frameworks.
- Hands-on experience with Transformer-based models.
- Experience working with LLMs such as:
- Qwen
- Llama
- Gemma
- Mistral
- or similar open-source architectures
- Experience with Hugging Face Transformers and related ML tooling.
- Experience with model fine-tuning and inference.
- Understanding of:
- Attention mechanisms
- Tokenization
- Transformer architectures
- Embeddings
- KV cache
- Quantization
- Model compression
- Ability to read and implement recent machine-learning research papers.
- Strong analytical and problem-solving skills.
Preferred Qualifications
Experience in one or more of the following areas is highly desirable:
- Vision-Language Models such as Qwen-VL, LLaVA, InternVL, Florence, or similar multimodal architectures.
- Multimodal model training and fine-tuning.
- Visual grounding and spatial reasoning.
- Robotics or embodied AI.
- 3D computer vision.
- Image/video understanding.
- Reinforcement learning or preference optimization.
- CUDA, TensorRT, ONNX, vLLM, or other inference-optimization frameworks.
- FP8 / INT8 / INT4 model optimization.
- NVIDIA Jetson or other edge-AI platforms.
- Distributed training using DeepSpeed, FSDP, or similar frameworks.
- Synthetic-data generation and automated dataset creation.
- RAG, vector databases, and agentic AI architectures.
- Deployment of production AI APIs and inference infrastructure.
- Model-hardware co-design or AI accelerator architecture.
Example Projects
You may work on projects including:
- Developing O-HIVE's proprietary Visual Language Model for robotics.
- Enabling robots to interpret camera input and generate actionable decisions.
- Creating visual inspection models for manufacturing and industrial environments.
- Building multimodal reasoning systems combining RGB, depth, 3D, and natural-language inputs.
- Optimizing multimodal models for real-time inference on edge hardware.
- Developing task-planning and reasoning systems that translate visual observations into robotic actions.
- Compressing VLM models for deployment on O-HIVE's future custom semiconductor platform.
What We Value
We are looking for engineers who:
- Enjoy turning cutting-edge AI research into real products.
- Are comfortable working across research and engineering.
- Take ownership of difficult technical problems.
- Can rapidly prototype, evaluate, and improve AI systems.
- Are excited about the intersection of language models, computer vision, robotics, and semiconductor technology.
- Want to build AI systems that operate in the physical world rather than only inside software applications.
Why O-HIVE
At O-HIVE, you will have the opportunity to work across the full AI stack—from multimodal model architecture and training to edge deployment, robotics integration, and eventually custom AI silicon.
Our goal is to build visual intelligence that allows machines to understand their environment and make intelligent decisions in real time.
Pay: ₹1,000,000.00 - ₹2,000,000.00 per year
Work Location: In person