Lexsi Labs is the leading frontier AI lab focused on building aligned, interpretable, and safe superintelligent systems. While that is the vision, the mission to build safety aware autonomous system in the extreme near term. Our research work spans around areas like AI alignment methodologies, interpretability led system design, and foundational model research across structured, tabular, and new autonomous system designs. We published about 25+ papers in the past 15 months across leading conferences ICLR, ICML, WWW, IJCNN, MICCAI, Eurips etc. Our labs are located in India (Mumbai & remote), Paris, & London.
We operate with a flat structure, high autonomy, and a strong bias toward engineers who take full ownership of what they build, from architecture to production behavior.
The Role
Our current sprint on building autonomous system for complex problems, across software engineering, data science, and AI research involves building harness, execution substrate and evaluation system, and each is designed to run inside a customer's environment rather than ours.
This role owns the coding agent: a system that takes an objective, works directly on a repository, verifies its own changes, and produces a record of what it did. The codebases are large, old, thinly tested and load-bearing, and the agent runs inside the customer's network, usually without internet egress.
What'll you work on:
-
The agent architecture and its tool surface for software work, including retrieval, editing, builds, tests, static analysis and version control.
-
Verification, so that a change is established as correct before a human sees it.
-
Long-horizon execution across multi-file refactors and migrations, with checkpointing, recovery and self-correction.
-
The authorisation and rollback model governing what the agent may do and how it is undone.
-
Integrations with the engineering estate the agent operates in, and the observability tooling used to debug and audit it.
-
Evaluation suites built from real repository tasks, and the behavior tuning that follows from them.
You will work closely with our research team on evaluation design, post-training and interpretability of agent behavior.
What We Are Looking For
This is a developer tooling role as much as an AI role. The agent is only as capable as the tools and integrations it is given, so depth here matters more to us than anything else.
-
Deep experience with developer tooling and IDEs. Editor and IDE internals or extensions, language servers and LSP, code intelligence, refactoring engines. You have built things engineers use every day.
-
Code intelligence and program analysis. ASTs and tree-sitter, static and dataflow analysis, symbol indexing and code search, call graph and dependency resolution, codemods and automated migration tooling.
-
Build systems and CI/CD. Bazel, Gradle, Maven, Make and similar, incremental and hermetic builds, artifact and dependency management, pipeline design across Jenkins, GitLab, GitHub Actions or equivalents, including self-hosted and offline deployments.
-
Observability. Distributed tracing, structured logging, OpenTelemetry, and building the replay and inspection tooling that makes non-deterministic systems debuggable.
-
Integrations engineering. Git hosting APIs, ticketing and workflow systems, enterprise auth, and the reality of working against self-hosted software several versions behind current.
-
Agentic systems in production, not prototypes. You know the current framework landscape (ReAct-style agents, LangGraph, LangChain, Semantic Kernel) and where each one stops working.
-
Strong backend fundamentals. Advanced Python, APIs and services, sandboxing and containerisation, databases and data pipelines, cloud and on-premise infrastructure.
-
You reason about performance, reliability, cost, safety and interpretability as one connected set of constraints.
-
You make good calls in the gray area, and you are comfortable when problems are loosely specified and ownership is assumed rather than assigned.
-
Experience building evaluation datasets or frameworks for AI systems is a plus.
-
Compiler background, large monorepo work, or legacy modernisation programs are a plus.
-
Prior exposure to alignment, interpretability or safety tooling in production settings is a plus.
We move quickly and expect candidates to do the same. We value substance over polish and execution over rhetoric.