You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Access to Echo Omni is provided through Zero Runtime. Visit https://zeroruntime.ai/?utm_source=huggingface&utm_medium=referral&utm_campaign=huggingface to get started.

Log in or Sign Up to review the conditions and access this model content.

Echo Omni. 95.39% accuracy, 27 languages, 4 states, verdict under 80 ms.

Echo Omni hears what was said and how it was said, at the same time.
A verdict in under 80 ms, faster than the pause it is judging.

Get Access   Documentation

🏆 Ranked 4th of 18 on end-of-turn and 4th of 16 on interruption on TurnBench, Sesame AI Labs' public turn-taking benchmark.

Most turn detectors wait for silence and hope. A 500 ms gap looks identical whether someone finished a sentence, paused to think, or simply took a breath, so agents built on timers end up talking over people or leaving them hanging.

Echo Omni does not guess. It reads the user's audio together with the transcript of that same turn, so it catches the meaning of the words and the sound of the delivery at once: the trailing pitch, the hesitation, the clipped "mm-hm" that was never a turn at all. One prediction per turn, with a confidence score, fast enough to live inside a real conversation.

It is the flagship of the Echo family, and the broadest of the three: 27 languages, four conversational states, and accuracy that holds across every one of them. It is also independently ranked on TurnBench, scored blind against a held-out test set.


🎯 What it does

Echo Omni classifies every user turn into one of four states. Each state tells the agent exactly what to do next.

Every turn resolves to one of four states: Incomplete, Complete, Backchannel or Wait, each with the action the agent should take.

State What it means What your agent should do
Complete The user has finished their turn Hand the turn to the LLM and respond
Incomplete The user is mid-sentence, just pausing Keep listening. Do not take the floor
Backchannel A short acknowledgement: "uh-huh", "okay okay" Keep speaking. This is not an interruption
Wait The user is asking you to hold: "wait a minute", "hold on" Stop speaking immediately

Most turn detectors only answer the first two. Backchannel and Wait are the states that make an agent feel polite instead of oblivious: not stopping every time someone says "mm-hm", and stopping the instant someone says "hold on".


🔌 Input and output

Echo Omni is multimodal, and it needs both signals at once.

Audio and the transcript of the same turn go in together, and one classified state comes back.

Input, per turn:

Field Description
Audio The speech segment for the turn
Transcript The finalized transcript of that same segment
Language One of the 27 supported codes

Output, one prediction per turn:

Field Description
State Complete, Incomplete, Backchannel or Wait
Confidence A score for the prediction, so you can tune how decisive your agent is

Both the audio and the transcript are required, and they must describe the same turn. Text alone cannot tell a thinking pause from a finished thought, and audio alone cannot tell you what was actually said. Echo Omni is built to use both together, which is where its accuracy comes from.


🌍 Supported languages

27 languages, spanning Indian, European and East Asian language families.

Code Language Code Language Code Language
ar 🇸🇦 Arabic bn 🇧🇩 Bengali da 🇩🇰 Danish
de 🇩🇪 German en 🇺🇸 English es 🇪🇸 Spanish
fi 🇫🇮 Finnish fr 🇫🇷 French gu 🇮🇳 Gujarati
hi 🇮🇳 Hindi id 🇮🇩 Indonesian it 🇮🇹 Italian
ja 🇯🇵 Japanese ko 🇰🇷 Korean mr 🇮🇳 Marathi
nl 🇳🇱 Dutch no 🇳🇴 Norwegian pl 🇵🇱 Polish
pt 🇵🇹 Portuguese ru 🇷🇺 Russian ta 🇮🇳 Tamil
te 🇮🇳 Telugu tr 🇹🇷 Turkish uk 🇺🇦 Ukrainian
ur Urdu vi 🇻🇳 Vietnamese zh 🇨🇳 Chinese

This is not a model that works well in English and degrades everywhere else. Accuracy holds across the full set.


📊 Performance

Measured on a held-out multilingual test set of 25,000+ utterances covering all four labels.

Metric Echo Omni
Accuracy 95.39%
Macro F1 0.9712
F1 Score (Complete) 0.9508

Accuracy holds up across all 27 supported languages, with a mean macro-F1 of 0.9712.

The macro F1 is the number worth looking at. It weights all four states equally, including Backchannel and Wait, which are rare in real traffic and are exactly where lesser turn detectors quietly fall apart. Echo Omni does not trade those away to flatter its headline accuracy.

Results are measured on the benchmark described above. Performance may vary depending on language, deployment configuration, user behaviour and application requirements.


🏆 TurnBench

TurnBench is Sesame AI Labs' public benchmark for conversational turn-taking. Systems are scored blind against a held-out test set they never see. We submitted Echo Omni on 10 September 2026.

Task Rank Recall False-positive rate
End-of-turn 4th of 18 0.839 0.071
Interruption 4th of 16 0.904 0.129

Echo Omni is one of only three systems on the board to place top five on both tasks, which is what a real conversation needs: knowing when someone has finished and when they have cut in, from the same model.

Its end-of-turn false-positive rate is 0.071, so it wrongly claims the floor on seven of every hundred mid-turn pauses. That is the number users feel, because each one is the agent cutting somebody off mid-sentence.

Filtered by conversation type, it ranks 3rd of 18 on end-of-turn in casual, collaborative and narrative conversation.

See Echo Omni's full TurnBench scorecard

Scored by Sesame AI Labs. Ranks as of 11 September 2026.


🧠 Trained on data built for this problem

Echo Omni is trained on a proprietary, closed-source dataset built in-house, purpose-made for conversational turn-taking across all 27 languages and all four states, with real acoustic variety rather than clean read speech.

That dataset is the reason the rare states hold up. Backchannels and hold requests barely appear in off-the-shelf speech corpora, so a model trained on public data has almost nothing to learn them from.


🎛️ The Echo family

Echo Omni is the multimodal flagship. Two text-only siblings cover the rest of the latency and accuracy curve, and all three return the same four states, so you can move between them without changing your agent logic.

Model Modality Latency Best for
echo-small Semantic 5 to 10 ms The lowest latency. The default when responsiveness matters most
echo-large Semantic 10 to 20 ms Higher accuracy, when it matters more than raw speed
echo-omni Audio + semantic 60 to 80 ms The widest language coverage, with acoustic understanding on top of the semantics

🚀 Get access

Echo Omni is served for you. There is nothing to download, host, or keep running. Plug it straight into your voice pipeline alongside your existing STT, LLM and TTS.

# Set ZERORUNTIME_AUTH_TOKEN in your environment.

from zeroruntime.inference import TurnDetector

# Multimodal: audio and transcript together, 27 languages
turn_detector = TurnDetector(model="echo-omni")

Get Access   Documentation


Echo Omni

Echo Omni by Zero Runtime

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including zero-runtime/echo-omni

Evaluation results