Lexsi Labs is the leading frontier AI lab focused on building aligned, interpretable, and safe superintelligent systems. While that is the vision, the mission is to build safety aware autonomous systems in the extreme near term. Our research work spans areas like AI alignment methodologies, interpretability-led system design, and foundational model research across structured, tabular, and new autonomous system designs. We published about 25+ papers in the past 15 months across leading conferences including ICLR, ICML, WWW, IJCNN, MICCAI and EurIPS. Our labs are located in India (Mumbai and remote), Paris, and London.
We operate with a flat structure, high autonomy, and a strong bias toward engineers who take full ownership of what they build, from architecture to production behavior.
Our current sprint on building autonomous systems for complex problems, across software engineering, data science, and AI research, involves building the harness, execution substrate and evaluation system, and each is designed to run inside a customer's environment rather than ours.
This role sits underneath all three agents. The coding agent, the data science agent and the AI engineering agent look different from the outside, but they are the same system underneath: a loop that plans, acts, observes, recovers, and knows when to stop. You will build that shared layer, and the evaluation system that tells us whether any change to it made things better. Both halves matter equally. A harness we cannot measure is a harness we cannot improve.
Harness
-
The core agent loop and its execution model, covering orchestration, concurrency, sub-agent coordination, cancellation, timeouts, checkpointing and resume.
-
Context and state management for long-running tasks, including retention, compaction, and how an agent reasons over what it has already done.
-
The shared tool protocol, its schemas, versioning, and the internal libraries every agent team builds against.
-
The execution substrate. Sandboxed environments that are reproducible, snapshot-able and resource-bounded, and that behave identically for a repository task, a training run and a data analysis.
-
Failure semantics. Retry policy, partial failure, idempotency, and the distinction between a recoverable error and a task that should stop and hand back.
-
The trace schema every agent emits, which serves at once as our debugging surface, our training signal, and the audit record our customers keep.
Evals
-
Task suites for each agent type, built from real work rather than synthetic benchmarks, and the infrastructure to run them at scale in parallel.
-
Verifiers and scoring. Deterministic checks where the task allows it, model-graded rubrics where it does not, and calibration of the graders themselves.
-
Regression gates that run on every harness change, with cost and latency accounted alongside quality.
-
The statistics to say a result is real. Seeds, variance, pass@k, confidence, and knowing how many runs are needed before anyone claims an improvement.
-
Turning observed production failures into permanent test cases.
-
Tooling for how we work. Trace inspection, replay, and diffing one run against another.
You will work across all three agent teams and closely with our research team on evaluation design, post-training and interpretability of agent behavior.
This is a systems and infrastructure role. Most of the difficulty here is concurrency, state and measurement, not prompting.
-
Strong software engineering fundamentals and advanced Python. You have shipped and operated production services, not only written them.
-
Concurrency and distributed systems. Async execution, worker pools and queues, retries and idempotency, backpressure, cancellation, and reasoning clearly about what happens when a step fails halfway through.
-
Containers and sandboxing. Docker and OCI internals, resource isolation, reproducible environments, and an understanding of why a job that passes locally fails in a sandbox.
-
Test and CI infrastructure. You have built test harnesses, run large suites in parallel, and dealt with flakiness as an engineering problem rather than an annoyance.
-
Measurement literacy. Comfort with variance, sampling and significance, and healthy skepticism toward benchmark results including your own.
-
Observability instincts. Distributed tracing, structured logging, and building the inspection tooling that makes non-deterministic systems debuggable.
-
You debug systematically. You find out what actually happened rather than adjusting things until the symptom disappears.
-
You are comfortable when problems are loosely specified and ownership is assumed rather than assigned.
-
Experience with evaluation or benchmarking work for AI systems is a strong plus.
-
Experience with agentic systems, LLM inference and serving, or developer tooling is a plus.
-
Open source contributions we can read are a plus.
We are hiring several engineers for this team at a range of experience levels, including engineers early in their careers who have strong fundamentals and want to work on agents from the infrastructure side.
We move quickly and expect candidates to do the same. We value substance over polish and execution over rhetoric.