This role gets AI into production and keeps it running. You own the MLOps backbone and build the agentic systems we deliver for clients: agents, GraphRAG, enterprise copilots, and workflow automation.
The particular challenge here is dependability. An LLM system that dazzles in a demo can behave unpredictably in a client's real workflow, and your job is to make it steady: reliable, affordable to run, guarded, and watched. That's the problem this role turns over, and if it's the kind of problem you enjoy, there's plenty of it.
The work:
- Design and ship agentic systems: agents, RAG and GraphRAG, copilots the client's people actually use
- Own the model-specific side of production: serving logic, evaluation, guardrails, drift and behavior monitoring, and the retraining loop
- Build the MLOps frameworks and reusable AI infrastructure every engagement draws on
- Expose models and agents behind clean endpoints for the applications to consume
- Integrate AI into the client's real operational workflows, not tools that sit unused
Where your work ends. You take models from the Data Scientist and the graph foundation from the Data Engineer and turn them into dependable production AI. You own everything model-specific, but you run it on the platform the Software Engineer provides: you don't own the containers, Kubernetes, cloud provisioning, or the generic observability stack, and you don't build the application front ends. Your line is model behavior; theirs is the platform and the product surface.
What success looks like:
- AI systems stay reliable in a client's real workflow, not just in a demo
- Guardrails, evals, and monitoring are in place before anything goes live
- The cost of running AI stays predictable and defensible
- The AI infrastructure you build makes the next engagement faster, not slower
The stack we work in today: Python for the model and service work, with FastAPI where you're exposing a model or agent behind an endpoint; PyTorch, with real fluency in LLM internals, fine-tuning, and, where it earns its place, mechanistic interpretability; MCP and A2A for agent and tool interop; and Neo4j behind our GraphRAG systems. A depth in a serious slice of this, plus the judgment to learn the rest, matters more than checking every box. For this role, MLOps is the craft itself.
About The Strong AI, and how we work
The Strong AI is an end-to-end AI implementation partner. Clients come to us because most organizations can run an AI experiment, but few can turn it into a system their business depends on. We close that gap. We don't hand over slideware or a notebook; we build systems that work inside a client's business, and where they want it, we run them.
You'll work across engagements and industries, on different problems and often different stacks. We're technology-agnostic: the problem and the client's environment choose the tools, so treat any stack we list as the ground we work on today, not a gate.
Across all roles, we ask for the same way of working:
- Real software. Tested, reviewed, versioned code the next person, or the client's team, can pick up.
- MLOps mindset. A model's life starts at deployment. Monitoring, retraining, drift, and rollback are handled before anything breaks.
- Systems thinking. You see both the value slice and the whole it compounds into.
- Quality and security, owned by you. Designed in from the first decision, not inspected in at the end. Everyone builds to the highest standard.
- Built for handover. Clear code and docs the client's own team can understand, operate, and take over.
Pay: ₹500,000.00 - ₹1,000,000.00 per year
Benefits:
- Flexible schedule
- Work from home
Application Question(s):
- What interests you about working for this company?
Work Location: Remote