Agent architecture & orchestration
- Design and implement agentic systems capable of multi-step reasoning, planning, tool use, and workflow execution for a B2b operational processes.
- Build stateful workflows using frameworks such as LangGraph and LangChain - including branching, retries, self-correction, human-in-the-loop checkpoints, and reusable orchestration patterns.
- Engineer for long-horizon reliability - multi-step task completion, recovery from compounding errors, planning under uncertainty, and robust tool use when individual steps fail.
Retrieval, grounding & context engineering
- Develop end-to-end Retrieval-Augmented Generation (RAG) pipelines: ingestion, chunking, embeddings, vector and hybrid retrieval, reranking, contextual compression, and grounding strategies.
- Engineer memory and context management - conversational state, persistent memory, retrieval-aware context assembly, and token-efficient context selection.
- Apply modern context-delivery patterns (e.g., MCP-style tool/context interfaces) so agents access the right information at the right time.
Reliability, evaluation & safety
- Implement observability and tracing for prompts, tool calls, retrieval quality, agent traces, failures, drift, latency, and production behavior.
- Apply guardrails, safety controls, and failure-handling to reduce hallucinations and unsafe actions.
- Evaluate agents at the trajectory and task level - multi-step task success, failure-mode and regression analysis, and sandboxed test environments - alongside retrieval- and generation-quality metrics, automated checks, and human review.
Integration & production craft
- Build integrations with internal and external tools, APIs, enterprise systems, databases, and model providers so agents operate safely within real business workflows.
- Deliver production-quality code with strong practices in testing, CI/CD, logging, versioning, and documentation; make architecture decisions that balance quality, safety, latency, cost, and model risk.
- Demonstrated depth building and shipping production agentic systems - this is your primary experience, not a recent exploration, strong software/ML fundamentals plus substantial, recent hands-on agentic work.
- Strong, hands-on experience building production agent systems with modern orchestration - LangGraph/LangChain or equivalent, including custom orchestration.
- Experience designing and optimizing end-to-end RAG systems: indexing, retrieval, reranking, grounding, and evaluation.
- Strong understanding of memory and context management, including context windows, retrieval-driven context assembly, persistent memory, and high-signal context selection.
- Deep, practical understanding of LLM behavior - strengths, limitations, hallucination risks, reasoning constraints, and latency/cost trade-offs - and the evaluation methods used to measure them.
- Experience evaluating and debugging agent behavior - task-success and trajectory analysis, not just output quality.
- Strong Python engineering skills and modern software practices: testing, CI/CD, version control, and API integration; experience implementing observability, tracing, and debugging for LLM-based systems in production.
- Hands-on experience with at least one frontier model platform (e.g., Anthropic, Google, OpenAI) and/or open-weight/self-hosted models (e.g., Llama via vLLM), including production tool use and agent capabilities.
Preferred qualifications
- Experience with multi-agent systems and agent collaboration patterns.
- Familiarity with vector databases and retrieval infrastructure such as Pinecone, Weaviate, or Milvus.
- Exposure to model adaptation and fine-tuning techniques such as LoRA or QLoRA.
- Understanding of traditional NLP concepts: tokenization, semantic similarity, entity extraction, summarization, and transformer fundamentals.
- Bachelor's degree in Computer Science, Engineering, Data Science, Computational Linguistics, or a related field.
Please send your resumes to [email protected].