Live speech-to-speech · 38 languages

Speak your language.
Everyone hears theirs.

Sangam is a meeting room where each person chooses the language they speak and the language they hear. A Hindi sentence reaches one colleague in Tamil, another in Japanese, a third as English captions — in about two seconds, in a spoken voice, while the original plays softly underneath.

23 Indian languages~2s to interpreted voiceSelf-hosted media, DB and storage
One sentence

Priya says it once. Everyone gets it in their own language.

No shared working language, no interpreter on a phone bridge, no “sorry, can you repeat that in English?”. Each listener’s track is synthesised separately, so six people can be hearing six different languages of the same sentence at the same moment.

Hindi · spoken

हमारा बजट पचास लाख रुपये है, और डिलीवरी मार्च तक चाहिए।

Tamil

நமது பட்ஜெட் ஐம்பது லட்சம் ரூபாய், டெலிவரி மார்ச்சுக்குள் வேண்டும்.

Bangla

আমাদের বাজেট পঞ্চাশ লাখ টাকা, আর ডেলিভারি মার্চের মধ্যে চাই।

Telugu

మన బడ్జెట్ యాభై లక్షల రూపాయలు, డెలివరీ మార్చిలోగా కావాలి.

Odia

ଆମ ବଜେଟ୍ ପଚାଶ ଲକ୍ଷ ଟଙ୍କା, ଡେଲିଭରି ମାର୍ଚ୍ଚ ସୁଦ୍ଧା ଦରକାର।

English

Our budget is five million rupees, and we need delivery by March.

38languages, 23 of them Indian
~2send of sentence to interpreted voice
~170msto the first byte of synthesised audio
0third parties inside your meeting
In the meeting

Built for the half-second that decides whether a conversation flows.

01

Interpretation, not subtitles

Everyone hears a spoken voice in their own language while the original is ducked underneath — so tone, pace and who-is-talking survive the trip.

02

Each person picks two languages

One to speak, one to hear. Change either mid-call and the next sentence arrives differently. Nobody has to agree on a common tongue.

03

Captions in the same breath

Text lands about half a second before the voice, in your language, so you can skim while you listen.

04

Routed per language, not per vendor

Each language uses the engine that measured best for it — Google streaming where it wins, Gemini where it wins — and falls through automatically when one fails mid-sentence.

05

Background blur and replacement

Runs on your device with MediaPipe. Your room never leaves your machine.

06

Runs on your own SFU

Self-hosted LiveKit, your Postgres, your storage bucket. No third party sits in the middle of the meeting.

After the meeting

The recap arrives in your language too

One meeting, one timeline — rendered per person. Recordings go to the bucket you nominate, with a path template you control, so the file lands where your compliance team already looks.

01

Recording to your bucket

S3, Google Cloud Storage, Azure Blob, Alibaba OSS or a local directory — you name the path template, we write the file.

02

Transcript in your language

The same meeting reads back in Hindi for one person and Japanese for another, from one shared timeline.

03

Summary and action items

Decisions, owners and dates pulled out after the call — delivered in the language each person chose.

04

Ask the meeting a question

“What did we commit to on pricing?” answered from the transcript, with the moment it was said.

Vendor review · 42 min · 6 participants
SummaryTranscriptRecordingAsk
తెలుగుहिन्दीEnglishதமிழ்日本語

బడ్జెట్ ₹50 లక్షలుగా ఖరారైంది; డెలివరీ మార్చి చివరి వారానికి.

ప్రియా ధరల పట్టికను శుక్రవారం లోపు పంపుతారు; అరుణ్ లీగల్ సమీక్ష మొదలుపెడతారు.

Priya — send revised pricing sheet · Fri
Arun — start legal review · Mon
For companies

Give every client a meeting room with their name on the door.

01

Your brand, not ours

Your logo, colours, product name and a welcome message on the lobby, in the meeting and on every recap link. On Business, “Powered by Sangam” disappears.

02

A role for every seat

Owners, admins, hosts and members in the workspace; hosts, co-hosts, participants and webinar viewers in the room. Guests never need an account.

03

Gates that hold

Passcodes, waiting rooms, members-only meetings and a lock — checked on the server, not in the browser.

04

Built to resell

One workspace per client, plans with limits and usage metering, an audit log, a REST API and signed webhooks, and a console to run them all.

Why it feels different

Most tools translate the text. Sangam translates the conversation.

SangamCaption-only tools
What you hearA voice in your language, liveThe original audio only
What you readCaptions in your languageCaptions in the speaker’s language
Indian languages23, including ones most tools skipUsually Hindi, sometimes none
Where it runsYour servers, your database, your bucketTheir cloud
Languages

Every language in the Eighth Schedule we could do honestly.

Each language is routed to the engine that measured best for it, and the ones we can only caption are marked as such rather than quietly mis-voiced. Fifteen more languages cover the rest of the world.

Questions

Straight answers, including the unflattering ones.

How fast is it, really?

In our own end-to-end runs, captions arrive roughly a second after a sentence ends and the interpreted voice follows at about 1.4–2.6 seconds, depending on the language pair and sentence length. We publish the numbers rather than a round marketing figure.

Is every language equally good?

No, and we say so. Some languages have strong streaming recognition and a natural voice; a few are caption-only today because no voice we trust exists for them yet. The per-language table in the repo lists exactly what each one does.

Do I need an account to try it?

No. Start a meeting, share the link, pick your language. Accounts add history, recordings and recaps.

Where does my audio go?

To your self-hosted media server, and to the speech providers you configure, for as long as it takes to interpret a sentence. Nothing is kept unless you turn recording on and point it at your own bucket.

Why “Sangam”?

सङ्गम — the confluence, where rivers meet and become one. Which is roughly the job description.

सङ्गम — where the rivers meet.
Start talking.

No download, no account, no agreeing on a common language first.

Start a meeting