Commerce Pundit is a fast-growing technology and services company helping global e-commerce and enterprise clients build intelligent, AI-powered systems. We work across manufacturing, retail, beauty, auto parts, and digital commerce — solving real business problems with production-grade AI. We are a lean, high-output team where engineers own problems end-to-end and ship systems that directly impact client revenue and operations.
We are looking for an Agentic AI Engineer who can design, build, and ship production-grade multi-agent systems — not demos, not POCs, but real systems that run autonomously and deliver measurable business outcomes.
This is a hands-on IC role. You will be in the code daily — architecting agent pipelines, building MCP servers, designing RAG systems, managing LLM costs at scale, and iterating with real users and clients. If you are looking for a role where you can lead a large team without building yourself, this is not it. If you want to own a full agentic system from design to production, this is exactly it.
- Hands-on experience building production multi-agent systems using LangGraph, CrewAI, AutoGen, or equivalent
- Deep understanding of agent orchestration patterns — Supervisor, Planner-Executor, Critic-Generator, and hierarchical delegation
- Experience designing typed state schemas with reducers for parallel agent coordination
- Understanding of human-in-the-loop workflows including interrupt gates and async approval flows
- Experience debugging real production failures in multi-agent systems — loops, state corruption, silent failures
- Experience building or consuming MCP servers
- Understanding of MCP tool schema design, transport mechanisms (stdio / SSE), and security patterns
- Ability to expose existing APIs, databases, or scripts as structured MCP tools for LLM consumption
- Hands-on experience building production RAG pipelines including document ingestion, chunking, embedding, and retrieval
- Understanding of hybrid search (BM25 + vector), reranking (cross-encoder or Cohere), and metadata filtering
- Experience with vector databases — Pinecone, Weaviate, Qdrant, ChromaDB, or pgvector
- Familiarity with RAG evaluation frameworks — RAGAS, TruLens, or DeepEval
- Strong prompt engineering skills including system prompt design, structured output enforcement, and few-shot examples
- Understanding of token cost optimization — prompt caching, model routing, semantic caching, token budget management
- Experience with LLM APIs — OpenAI, Anthropic Claude, AWS Bedrock, Google Gemini
- Ability to design and implement evaluation pipelines for LLM output quality
- Python proficiency — FastAPI, async patterns, Pydantic, background workers
- Experience deploying AI systems to production with proper observability — LangSmith, LangFuse, or equivalent
- Understanding of production failure modes — silent failures, hallucination detection, retry logic, circuit breakers
- Experience with cost tracking and monitoring for LLM API spend at production scale
- AWS or GCP experience — Lambda, ECS, S3, or equivalent
- Docker and basic containerization
- CI/CD pipeline familiarity — GitHub Actions, Jenkins, or equivalent
- Orchestration: LangGraph, CrewAI, AutoGen
- LLM APIs: Anthropic Claude, OpenAI GPT, AWS Bedrock, Google Gemini
- RAG: LangChain, LlamaIndex, hybrid search, cross-encoder reranking
- Vector Databases: Pinecone, Qdrant, ChromaDB, pgvector
- MCP: Custom server development, tool schema design
- Backend: Python, FastAPI, Node.js
- Cloud: AWS (primary), GCP
- Observability: LangSmith, LangFuse, CloudWatch
- Evaluation: RAGAS, custom eval pipelines