About the Role
We are seeking an execution-focused AI Software Engineer to build and ship production LLM applications, RAG pipelines, and agentic workflows. We want a developer whose daily workflow is already supercharged by AI-assisted coding tools (Claude Code, Cursor, Copilot) to build, test, and ship clean software fast.
What You’ll Do
-
Build, optimize, and maintain end-to-end RAG pipelines (document ingestion, chunking, embeddings, vector indexing, retrieval, and reranking).
-
Integrate LLM APIs (Anthropic Claude, OpenAI, Gemini, open-source models) into scalable backend services using FastAPI or Python.
-
Use AI-native coding setups (Claude Code CLI, Cursor, GitHub Copilot) to explore codebases, write unit tests, and accelerate deployment cycles.
-
Design structured outputs, function/tool calling, and agent workflows using frameworks like LangChain, LlamaIndex, or native SDKs.
-
Implement basic evals, guardrails, and logging to reduce hallucinations, measure retrieval quality, and monitor API costs.
Core Requirements
-
Programming: Solid foundation in Python (async, REST APIs, clean object-oriented code) and Git.
-
GenAI & RAG: Hands-on experience with vector databases (e.g., Chroma, Qdrant, Pinecone, or pgvector) and embedding models.
-
AI Tooling Native: Daily, active user of CLI or IDE agentic coding tools (Claude Code, Cursor, or similar)—you know how to guide AI agents and rigorously inspect generated code.
-
Proof of Work: At least 1 shipped or working project beyond basic chat (e.g., semantic search tool, custom RAG on documents, automated agent, or GitHub repo).
Good to Have
-
Familiarity with Docker, Linux environment, and basic cloud deployment (AWS/GCP).
-
Experience with Model Context Protocol (MCP) or multi-agent orchestration frameworks.
-
Exposure to front-end integration (Streamlit, Next.js, or React basics).