backed by [ iit deep-tech research ] [ google for startups ]

for·mant /ˈfɔː.mənt/ — n. a resonance of the human vocal tract; the bands of energy that make speech intelligible

the next generation of india’s phone calls starts here

sovereign voice agents for collections, sales and support — in hindi, english, and the mix of both. on your infrastructure, under india’s rules.

live in production

hindi · english · regional
[ named languages, code-mixed ]
outbound & inbound
[ collections · sales · support ]
your infrastructure
[ vpc · private cloud · on-prem ]
never trained on
[ your conversations stay yours ]

why voice

01

for most of human history, it wasn’t written down; it was spoken

02

india doesn’t call in one language — it switches mid-sentence

03

the phone line keeps only 300–3400 hz. what survives the wire is the formant

[ f1 ]
fluent

handles the sentence that starts in hindi and ends in english — because that is how the call actually happens.

[ f2 ]
compliant

runs inside dpdp, rbi outsourcing guidelines and dlt-registered outbound from day one.

[ f3 ]
accountable

every call recorded, transcribed and auditable — down to which stage of reasoning produced which answer.

the science we’re built on

speech is resonance.
we work where it moves.

every vowel you say is defined by its formants — f1, f2, f3 — the resonant frequencies your vocal tract carves into raw sound. every language arranges its vowels differently in formant space. hindi and english overlap, but they don’t coincide.

when a caller switches language mid-sentence, their formants jump between two maps in milliseconds. systems tuned to one map drop the switch. we are built for exactly that movement — down at the formant level.

f2 (hz) → higher = fronter f1 (hz) → lower = closer 2500200015001000500 300500700900 iː · beat ɪ · bit æ · bat ɑː · calm uː · boot
english vowel space hindi vowel space one code-switched sentence, crossing between them

the chart is an instrument. press and drag anywhere on it — you are playing the human vocal tract: x is f2, y is f1, and the sound follows your cursor.

instrument 1 — hear the chart · each button synthesises a vowel from nothing but its three formants
the saffron path, as sound — one sentence’s vowels gliding between the hindi and english maps
instrument 2 — measure yourself · live lpc formant analysis of your own voice
say “aa… ee… oo” and watch your dot cross the map · runs entirely in your browser — nothing is recorded or sent

the physics

source, filter, formant

one model explains every vowel ever spoken — the source–filter model (fant, 1960). your vocal folds make a buzz. your vocal tract filters it. what survives are the formants.

[ 01 · the source ]

a buzz, rich in harmonics

the vocal folds open and close ~120× a second, producing energy at every multiple of that pitch.

[ 02 · the filter — × ]

a tube that resonates

throat + mouth form a ~17.5 cm tube. move your tongue, and the tube’s resonant peaks move.

[ 03 · the speech — = ]

harmonics × resonance

the harmonics that survive the tube are the vowel. the peaks are the formants. that’s the whole trick.

fn = (2n − 1) · c / 4l

quarter-wave resonances of a tube closed at one end · c = 350 m/s · l = 0.175 m
→ f1500 hz · f21500 hz · f32500 hz — the neutral vowel. every other vowel is a deformation of this tube.

why this matters on a phone line

a telephone call is band-limited to roughly 300–3400 hz (itu-t g.711). the pitch itself is often cut off entirely — your brain reconstructs it. what actually crosses the wire is the formant structure. telephone speech is formant speech.

which is why an agent built for phone calls has to be trained on real 8 khz telephony audio — noisy lines, regional accents, mid-sentence language switches — not studio recordings. it is why we build — and measure — at the formant level.

references
  1. g. fant, acoustic theory of speech production, mouton, the hague, 1960.
  2. g. e. peterson & h. l. barney, “control methods used in a study of the vowels,” j. acoust. soc. am. 24(2), 175–184, 1952 — the original f1×f2 vowel chart.
  3. k. n. stevens, acoustic phonetics, mit press, 1998.
  4. p. ladefoged & k. johnson, a course in phonetics, 7th ed., cengage.
  5. itu-t recommendation g.711 — 300–3400 hz telephony pass-band.
  6. vowel positions on the chart above are textbook-typical values (after peterson & barney), not measurements of any one speaker.

the proof you can hear

hinglish is not a setting.
it’s how india answers the phone.

press play. every saffron word is a switch — the moment the caller’s formants jump from one language’s map to the other, and the agent jumps with them.

early-bucket reminder — the borrower already paid
transcript preview · production audio drops in here

illustrative transcript for the site draft. before launch, replace with consented, redacted production recordings and their verbatim transcripts — the module is built to take an audio source per scenario.

platform

reasoning in stages,
so it doesn’t guess

single-stage systems guess at intent, then commit. deeprealm chains reasoning and validation stages, so the agent checks what it understood before it acts. that is what keeps a code-switched sentence from becoming a wrong answer — and a wrong answer, in collections, is a compliance event.

named languages, not a language count

hindi, english and the regional languages actually in production — trained on real telephonic audio where speakers switch mid-sentence, not studio recordings.

it knows when to stop talking

disputes, hardship, abuse, and anything the agent isn’t sure of go to your team — with the full transcript, the customer’s state, and what was already promised.

telephony that reaches the lock screen

sip trunking and byoc, your existing numbers, dlt-registered sender identity, and batch outbound campaigns with concurrency on demand.

your engineers get the keys

rest and webhooks for campaign upload, call events and outcomes. every call exports its recording, transcript and per-stage reasoning trace to your warehouse. model-agnostic — bring your own.

use cases

built for the calls that carry money

sovereign voice at scale

what a deployment looks like

indian nbfc · personal loans, buckets 1–3

challenge

tens of thousands of accounts a month worked by agency partners in multiple languages. right-party contact inconsistent, dispute handling varying by agency, and no audit trail a risk committee could read.

solution

deeprealm ran pre-due and early-bucket calling in hindi and regional languages, switching mid-sentence where the borrower did. every promise-to-pay timestamped, every dispute routed to a human with full context — the whole trail inside the lender’s own environment.

consumer brand · order and service support

challenge

inbound volume spiking around sale events, with customers opening in one language and finishing in another. agents re-asking questions the ivr had already asked.

solution

inbound agents that answer in the language the caller opens with, resolve order and billing queries end to end, and hand the rest to the human team with the conversation already summarised.

representative scenarios for this draft — swap in production figures (connect rate, ptp rate, kept-ptp, resolution %) from real deployments before launch.

compliance & data sovereignty

compliance your risk team
already asked about

[ india ]
dpdp act 2023 rbi outsourcing & localisation irdai trai / dlt-registered outbound in-india data residency
[ global ]
soc 2 type ii — audit in progress, report under nda iso 27001 — in progress gdpr-aligned processing
[ deployment ]
your vpc private cloud on-prem air-gapped

your customer data stays yours. we don’t train on your conversations — ever.

how you deploy it

live in weeks, on your infrastructure

don’t waste a year on a pilot that doesn’t end up going anywhere.

week 1

scope

we listen to your recordings and map the real conversation — including the language mix.

weeks 2–3

build

prompts, flows, integrations, escalation rules — built by forward-deployed engineers, with you.

week 4

test

simulated calls against your edge cases before a single customer hears it.

go live

monitor

dialer connected, dlt-registered, watched daily by the engineer who built it.

[ built by a team out of iit deep-tech ai research · backed by google for startups ]

this is only the beginning

every scheduled call in india,
eventually

live · now

collections, sales & support

outbound and inbound voice agents in hindi, english and code-mixed hinglish — in production today.

next

the regional ten

marathi, tamil, telugu, bengali, gujarati, kannada and more — each with measured code-switch quality, not a language count.

then

the formant benchmark

a published accuracy number for code-mixed telephone speech — measured on production calls, open for anyone to beat.

your deeprealm journey starts here

hear it on your own portfolio

send us fifty of your recordings. we’ll come back with an agent that handles them — in the language they actually happened in.