COMPANY PRFOILE
TRST01 is a ClimateTech and Digital Trust Infrastructure company that builds AI-native platforms for traceability, sustainability, carbon markets, and compliance-driven supply chains. Leveraging Digital Public Infrastructure (DPI), AI, blockchain, geospatial intelligence, and data governance, TRST01 enables trusted, transparent, and interoperable ecosystems across agriculture, commodities, manufacturing, and environmental markets.
Its solutions include end-to-end traceability, EUDR compliance, digital MRV (dMRV), sovereign carbon registries, ESG reporting, and AI-powered decision support. TRST01 also developed the Open Agri Trace Stack (OATS), an open, standards-based blueprint for agricultural traceability.
TRST01 partners with governments, enterprises, and development institutions to build trusted, compliant, and sustainable digital ecosystems that deliver measurable climate and business impact.
Qualifications & Requirements:
Minimum of 2 to 4 years’ experience as LLM Developer
Immediate joiners only
Work Mode:
Remote position
THE ROLE
TRST01 accumulates structured and unstructured data that general-purpose models do not understand well — traceability records, chain of custody documentation, Digital Product Passport data, climate compliance filings, and intelligence from frontier and emerging markets. We want to change that.
Your role is to take our proprietary data and build domain-specific models from it. That means fine-tuning and adapting current open-source LLMs — such as Qwen 3.5 or Qwen 3.8 — to understand our domain accurately, and integrating them into production platforms. This is not research. There is no paper to write. The measure of success is a deployed model that does what our platforms need. TRST01 provides compute budget for training runs.
We are not looking for academics or researchers. We want practitioners who have trained real models on real data and shipped the result.
WHAT YOU WILL WORK FOR?
● Fine-tuning open-source LLMs (Qwen, Llama, Mistral, Phi) on TRST01 proprietary datasets — traceability, DPP, climate, and market intelligence data
● Building data pipelines: cleaning, formatting, and versioning training datasets from domain-specific sources
● Designing domain-specific evaluation benchmarks — not generic metrics, but tests that measure whether the model actually answers our questions correctly
● Building RAG systems for accurate document retrieval across large proprietary document sets
● Integrating trained models into production workflows via REST APIs and N8N
● Advising on model selection, compute cost, and quality trade-offs
NON-NEGOTIABLES
Show evidence of real model training.
A tutorial, Coursera project, or AutoTrain job without understanding the underlying process does not qualify. We want a real fine-tuning run: a base model you selected, a dataset you prepared, a training job you ran, and results you evaluated. Show the HuggingFace model card, GitHub repo, or a written breakdown covering what you did and what the results showed. If the work was part of a team, describe your specific contribution clearly.
Be ready to work on messy, domain-specific data.
Our data is not Wikipedia. It is semi-structured government documents, compliance filings, and supply chain records — some multilingual, some low-quality, all domain-specific. You need experience preparing training data from non-standard sources, not just clean benchmark datasets.
Ship, do not theorise.
Every technical decision you make should be oriented toward a deployed, working model. We will not fund exploratory research without a production path. If that is the kind of work you prefer, this role is not the right fit.
REQUIRED SKILLS
● Solid Python — clean, readable, production-quality code
● Hands-on LLM fine-tuning: LoRA, QLoRA, or full SFT in a real project — not a notebook demo
● Hugging Face ecosystem: Transformers, PEFT, Datasets, TRL
● Training data preparation: cleaning, deduplication, instruction formatting (JSONL, ShareGPT / ChatML templates)
● Evaluation design: task-specific benchmarks and human review — not just perplexity
● GPU training on cloud compute: RunPod, Vast.ai, AWS, or GCP
● Ability to explain model decisions clearly to a non-technical founder
GOOD TO HAVE
● RAG pipeline experience — LangChain, LlamaIndex, or custom vector store
● Model deployment via vLLM, Ollama, or similar inference stacks
● Experience with structured output or function-calling fine-tuning
● Open-source contributions or a public HuggingFace model you trained and can demonstrate
● Background in a regulated or data-intensive industry (supply chain, climate, finance, government)
HOW WE WORK
● Fully remote, async-first
● You own the full pipeline — data preparation through deployed model in production
● Direct access to the founder — fast decisions, no committees
● Short paid technical task required before any offer
HOW TO APPLY
Send your application to [email protected]
Pay: ₹30,000.00 - ₹70,000.00 per month
Work Location: Remote