Data Scientist – Language Models
Role Overview
This is a founding technical hire who owns the AI/ML function end to end. Rather than sitting in one lane of a larger team, this person works across the full lifecycle of a language model project, from raw data to a deployed, monitored model in production. That includes exploring and preparing data, making modeling and training decisions, fine-tuning and compressing models, evaluating them honestly, and getting them served reliably.
A central part of this role is building a small language model (SLM) from the ground up for a specific, focused use case: preparing and curating the training data, making the architecture and training-objective calls, running the fine-tuning and/or pretraining work, distilling or quantizing it down to something efficient, and validating it against a clear evaluation framework. This isn't a one-off project. It's the kind of work this role is expected to repeat and improve on as we take on new use cases.
As the team grows, this person will help shape how the work gets divided into more specialized roles, and may mentor or hand off pieces of the stack to future hires. Until then, they are expected to be comfortable owning a project from a messy dataset to a model running in production.
Key Responsibilities
End-to-end SLM / model development
- Lead the full lifecycle of building a domain-focused small language model: data curation, training/fine-tuning strategy, evaluation design, and deployment.
- Make the core modeling decisions (architecture choices, training objective, and alignment approach such as SFT, LoRA/QLoRA, RLHF, or DPO) appropriate to the use case and available data/compute.
- Run distillation and/or quantization to move from a larger capable model to a smaller, cheaper one without losing too much quality.
- Own the model through to a working, demonstrable result: something that can be shown, benchmarked, and explained end to end.
Data
- Explore and understand raw data (EDA), and turn messy text into clean, usable training data.
- Build instruction/SFT datasets and preference pairs for alignment work.
- Catch and fix language-data-specific problems: duplicates, test-set leakage, bad formatting, and tokenization/context-length issues.
- Version datasets so every model result is traceable to what it was trained on.
Training & modeling
- Build and run fine-tuning and training pipelines, including LoRA/QLoRA and full SFT.
- Debug training runs that stall, diverge, or underperform, and tune hyperparameters.
- Apply distributed training techniques where needed, and manage checkpoints/reproducibility.
- Design evaluation and benchmarking approaches, including for free-text output where there's no single right answer.
Deployment & operations
- Package and serve models, whether self-hosted (e.g., vLLM/Triton-style stacks) or via managed frontier APIs.
- Apply quantization at serving time to manage latency and cost.
- Set up basic monitoring, logging, and cost tracking for models in use.
Communication & growth
- Document methods and results clearly enough that the work can be picked up, scaled, or handed to future specialized hires.
- Be able to explain technical decisions and trade-offs (cost, quality, speed, data needs) in plain terms to non-technical stakeholders.
Qualifications & Experience
- Bachelor's or Master's degree in Computer Science, Data Science, Machine Learning, or a related field. Strong practical/portfolio experience can substitute for formal education.
- Minimum 3 years of hands-on experience in applied ML/data science, with meaningful, demonstrable work on language models specifically, not just classical ML.
- Must be able to show prior end-to-end model development work: fine-tuning, distillation, or training a model (large or small) from data through to a working result. This can be from a job, academic research, or a serious personal/open-source project, but it needs to be real, showable work (portfolio, GitHub, write-up, or similar), not just familiarity with concepts.
Required Skills
- End-to-end model development (primary focus): Has actually taken a language model project from data to a working, evaluated model, not just used pretrained models off the shelf.
- Language models fundamentals: Solid understanding of transformer architecture, tokenization, and what happens during training.
- Fine-tuning & alignment: Hands-on with SFT, LoRA/QLoRA, and ideally RLHF or DPO.
- Distillation & quantization: Can shrink a model and reasonably measure the quality trade-off.
- Data engineering for text: Comfortable cleaning, deduplicating, and structuring large text datasets; understands context-length and tokenization constraints.
- Coding: Strong Python, PyTorch (or JAX), and the Hugging Face ecosystem (Transformers, PEFT, Accelerate, or similar).
- Evaluation: Can design and run benchmarks for language models, including free-text/no-single-right-answer cases.
- Serving basics: Some hands-on experience getting a model into a usable, callable state (self-hosted or via API), and comfort with the trade-offs between the two.
- General ML/DS craft: Pandas, NumPy, SQL, Git, and good collaborative engineering habits.
Good to Have
- Experience with actual pretraining (even small-scale) rather than only fine-tuning existing models.
- Familiarity with distributed training (data/tensor/pipeline parallelism).
- Cloud experience (AWS, Azure, or GCP), particularly GPU-based training/serving.
- Exposure to working with data from a regulated or specialized domain (e.g., healthcare, agriculture, defense-adjacent, legal/compliance).
- Publications, preprints, or open-source contributions in ML/NLP.
- Experience with MLOps basics: CI/CD for models, monitoring, or experiment tracking (W&B, MLflow).
Benefits:
- Flexible schedule
- Health insurance
- Leave encashment
- Paid sick time
- Paid time off
- Provident Fund
Work Location: In person