for·mant /ˈfɔː.mənt/ — n. a resonance of the human vocal tract; the bands of energy that make speech intelligible
sovereign voice agents for collections, sales and support — in hindi, english, and the mix of both. on your infrastructure, under india’s rules.
live in production
why voice
for most of human history, it wasn’t written down; it was spoken
india doesn’t call in one language — it switches mid-sentence
the phone line keeps only 300–3400 hz. what survives the wire is the formant
handles the sentence that starts in hindi and ends in english — because that is how the call actually happens.
runs inside dpdp, rbi outsourcing guidelines and dlt-registered outbound from day one.
every call recorded, transcribed and auditable — down to which stage of reasoning produced which answer.
the science we’re built on
every vowel you say is defined by its formants — f1, f2, f3 — the resonant frequencies your vocal tract carves into raw sound. every language arranges its vowels differently in formant space. hindi and english overlap, but they don’t coincide.
when a caller switches language mid-sentence, their formants jump between two maps in milliseconds. systems tuned to one map drop the switch. we are built for exactly that movement — down at the formant level.
the chart is an instrument. press and drag anywhere on it — you are playing the human vocal tract: x is f2, y is f1, and the sound follows your cursor.
instrument 1 — hear the chart · each button synthesises a vowel from nothing but its three formantsthe physics
one model explains every vowel ever spoken — the source–filter model (fant, 1960). your vocal folds make a buzz. your vocal tract filters it. what survives are the formants.
the vocal folds open and close ~120× a second, producing energy at every multiple of that pitch.
throat + mouth form a ~17.5 cm tube. move your tongue, and the tube’s resonant peaks move.
the harmonics that survive the tube are the vowel. the peaks are the formants. that’s the whole trick.
quarter-wave resonances of a tube closed at one end · c = 350 m/s · l = 0.175 m
→ f1 ≈ 500 hz · f2 ≈ 1500 hz · f3 ≈ 2500 hz — the neutral vowel. every other vowel is a deformation of this tube.
a telephone call is band-limited to roughly 300–3400 hz (itu-t g.711). the pitch itself is often cut off entirely — your brain reconstructs it. what actually crosses the wire is the formant structure. telephone speech is formant speech.
which is why an agent built for phone calls has to be trained on real 8 khz telephony audio — noisy lines, regional accents, mid-sentence language switches — not studio recordings. it is why we build — and measure — at the formant level.
the proof you can hear
press play. every saffron word is a switch — the moment the caller’s formants jump from one language’s map to the other, and the agent jumps with them.
illustrative transcript for the site draft. before launch, replace with consented, redacted production recordings and their verbatim transcripts — the module is built to take an audio source per scenario.
platform
single-stage systems guess at intent, then commit. deeprealm chains reasoning and validation stages, so the agent checks what it understood before it acts. that is what keeps a code-switched sentence from becoming a wrong answer — and a wrong answer, in collections, is a compliance event.
hindi, english and the regional languages actually in production — trained on real telephonic audio where speakers switch mid-sentence, not studio recordings.
disputes, hardship, abuse, and anything the agent isn’t sure of go to your team — with the full transcript, the customer’s state, and what was already promised.
sip trunking and byoc, your existing numbers, dlt-registered sender identity, and batch outbound campaigns with concurrency on demand.
rest and webhooks for campaign upload, call events and outcomes. every call exports its recording, transcript and per-stage reasoning trace to your warehouse. model-agnostic — bring your own.
use cases
pre-due, early bucket and hard bucket. ptp capture, dispute routing, and a compliance trail on every call.
[ talk to us → ] [ 02 · lending sales ]work cold leads, verify intent, and book the branch visit — in the customer’s own language.
[ talk to us → ] [ 03 · insurance ]premium reminders, claim status and policy servicing without the hold queue.
[ talk to us → ] [ 04 · onboarding & kyc ]chase documents, confirm details and complete activation — escalating the hard cases to your team.
[ talk to us → ]sovereign voice at scale
tens of thousands of accounts a month worked by agency partners in multiple languages. right-party contact inconsistent, dispute handling varying by agency, and no audit trail a risk committee could read.
deeprealm ran pre-due and early-bucket calling in hindi and regional languages, switching mid-sentence where the borrower did. every promise-to-pay timestamped, every dispute routed to a human with full context — the whole trail inside the lender’s own environment.
inbound volume spiking around sale events, with customers opening in one language and finishing in another. agents re-asking questions the ivr had already asked.
inbound agents that answer in the language the caller opens with, resolve order and billing queries end to end, and hand the rest to the human team with the conversation already summarised.
representative scenarios for this draft — swap in production figures (connect rate, ptp rate, kept-ptp, resolution %) from real deployments before launch.
compliance & data sovereignty
your customer data stays yours. we don’t train on your conversations — ever.
how you deploy it
don’t waste a year on a pilot that doesn’t end up going anywhere.
we listen to your recordings and map the real conversation — including the language mix.
prompts, flows, integrations, escalation rules — built by forward-deployed engineers, with you.
simulated calls against your edge cases before a single customer hears it.
dialer connected, dlt-registered, watched daily by the engineer who built it.
[ built by a team out of iit deep-tech ai research · backed by google for startups ]
this is only the beginning
outbound and inbound voice agents in hindi, english and code-mixed hinglish — in production today.
marathi, tamil, telugu, bengali, gujarati, kannada and more — each with measured code-switch quality, not a language count.
a published accuracy number for code-mixed telephone speech — measured on production calls, open for anyone to beat.
your deeprealm journey starts here
send us fifty of your recordings. we’ll come back with an agent that handles them — in the language they actually happened in.