About the role
We run outbound voice AI agents that call our leads and move them through the loan funnel — from first contact to digital verification and disbursal. The technology works; the job is making it drive *business outcomes*, not just call activity.
We're hiring one person to own this program end to end. You'll be accountable for the number that matters — loans disbursed through bot-led and bot-assisted calling — and for everything upstream of it: pickup rates, early drop-off, interest capture, conversation quality, vendor performance, and the tech decisions that shape all of them.
This is a *business-owning role with a technical spine.* You won't necessarily write production code, but you must be fluent enough in the stack — ASR, TTS, LLMs, latency, telemetry — to challenge a vendor's numbers, diagnose why a metric moved, and decide what to fix. If you can't tell whether a 45% early drop-off is a latency problem, an endpointing problem, or an opening-line problem, this role isn't for you. If you find that question interesting, it is.
---
## What you'll own
### Business outcomes (the core)
- Own the funnel: pickup rate, early drop-off, engagement, interest captured, DV completion, and *disbursal* as the north-star metric.
- Set targets, track them weekly, and be the single person accountable when the number moves — up or down.
- Decide where the bot leads versus where humans take over (e.g. bot for intent capture, humans for complex DV until quality improves), and shift that boundary as the bot improves.
- Prioritise calling by funnel stage and lead value — leads stuck at NACH, eKYC, VKYC, DV, and disbursal — rather than dialling indiscriminately.
### Vendor & technology management
- Own relationships with voice AI vendors (current and pilots). Run structured evaluations on *business outcomes, not demos*.
- Hold vendors accountable on the stack: ASR (e.g. Deepgram), TTS voices, LLM model and version, latency percentiles (p50/p90/p95/p99), and per-stage telemetry.
- Diagnose where problems live — call setup, first-audio delay, endpointing, ASR accuracy, LLM latency, TTS — and drive fixes with the vendor or internal team.
- Own decisions on the conversation experience: latency targets, barge-in handling, backchannel/filler handling (e.g. "haanji", "accha", "ji"), retry logic, and language/code-mixing quality (Hindi / English / Hinglish).
### Product & stack optimization (heavy focus)
- Own the conversation product end to end and continuously optimise it — treat the voice agent as a product with a roadmap, not a fixed vendor deliverable.
- Drive down end-to-end turn latency: profile every stage (endpointing, ASR, orchestration, LLM time-to-first-token, TTS time-to-first-byte, network/jitter) and push each one, understanding which stages stream and overlap and which sit in the blocking path.
- Optimise the ASR layer: benchmark models on your own call audio (not studio demos), tune confidence thresholds, custom vocabulary and keyword boosting for domain terms (EMI, NACH, eKYC, VKYC, disbursal, product names), and code-mixed Hindi/English quality.
- Optimise endpointing and turn-taking: move from fixed silence thresholds toward semantic/model-based endpointing, build backchannel/continuer suppression so fillers like "haanji", "accha", "ji", "hmm" don't derail the bot, and make barge-in reliable.
- Optimise the LLM layer: system-prompt design, prompt/context size and caching, tool/function-calling reliability, RAG/FAQ retrieval quality and freshness, and guardrail behaviour — trading off latency, cost, and answer quality deliberately.
- Optimise TTS: voice selection against retention data, streaming behaviour, prosody on numbers and English loanwords, and the audio-format path into telephony.
- Own the opening experience — connect-to-first-audio delay, the opening line, and everything that decides whether a caller stays past the first ten seconds.
- Turn call recordings and transcripts into a prioritised optimisation backlog: listen to calls, find failure patterns, and translate them into concrete stack and prompt changes.
### Experimentation & analytics
- Design and run controlled experiments — cohort splits, A/Bs, holdouts — to prove what actually moves business results, and to avoid crediting changes that didn't.
- Build and own the reporting: funnel dashboards, drop-off distributions, latency and error telemetry, and outcome tracking that ties calling activity to disbursal.
- Instrument the stack so every optimisation is measurable per stage and per language, and no change ships without a way to read its effect.
### Operations & scale
- Own calling strategy: dialling times, retry cadence, re-contact of "call back later" leads, personalisation by lead attributes (pincode, employment type, funnel stage).
- Manage concurrency, capacity, and the small human team that handles bot-escalated leads.
- Expand channels where they help conversion — WhatsApp journeys, VKYC/link sharing, and similar.
---
## What we're looking for
### Must have
- *A product optimiser at heart.* You see a voice agent as a system to be tuned stage by stage, and you get satisfaction from shaving latency, lifting accuracy, and removing friction one measurable change at a time.
- *Real technical fluency across the stack:* you understand how ASR, TTS, LLMs, endpointing/VAD, and RAG fit together, why each affects the caller's experience, and where the latency and quality levers actually are. You don't need to write production code, but you must be able to profile a pipeline, read a per-stage trace, and tell a vendor exactly what to fix.
- *Genuine comfort with data and telemetry.* You can read latency percentiles, ASR confidence distributions, and drop-off curves, and reason about what they imply. You're fluent in SQL or equivalent, and can build a dashboard without waiting on someone else.
- *Experiment discipline.* You know why an average hides the truth, why a before/after comparison without a control is confounded, and how to design a test that actually isolates a change.
- *Ownership of a hard number.* You've carried a business metric before and can show how your decisions moved it.
- Sharp vendor-management instincts — you can tell a real number from a flattering one.
### Nice to have
- Direct experience with voice AI / conversational AI / IVR / contact-centre tech.
- Familiarity with the Indian telephony and lending context (8kHz call audio, code-mixed speech, NACH/eKYC/VKYC/DV flows).
- Working knowledge of specific vendors in the space (Deepgram, ElevenLabs, Cartesia, Sarvam, Smallest.ai, Bolna, and similar).
- Basic scripting (Python) for pulling data and quick analysis.
- Comfort listening to and QA-ing calls in Hindi and English.
### You'll thrive here if
- You're allergic to vanity metrics and reach instinctively for the business outcome underneath.
- You're equally happy in a vendor call arguing about p95 latency and in a strategy review defending a disbursal target.
- You'd rather ship one controlled experiment than five uncontrolled changes.
Pay: ₹50,000.00 - ₹70,000.00 per month
Work Location: In person