Cascaded vs Omni
What the published model sets can tell us—and what their architecture labels leave unresolved.
These are published model results. VoiceAgent.best ran no model calls or voice-agent tests. They do not measure a platform, phone call, production latency, voice quality or handoff reliability.
Two source-defined groups
VoiceBench classifies cascaded rows as a separate speech-recognition and language-model pipeline. Its Omni category represents turn-based speech input and output. These are community labels; the table separately classifies S2S / Full-Duplex, so Omni does not establish simultaneous listening and speaking.
The capture contains 5 cascaded rows and 21 omni rows. Model sizes, access, prompt choices and run settings differ or remain unspecified.
Observed metric ranges
| Metric / source scale | Cascaded / n | Omni / n |
|---|---|---|
| AlpacaEvalsource 5-point scale | 4.45–4.80 / 5 | 1.90–4.78 / 21 |
| CommonEvalsource 5-point scale | 3.82–4.47 / 5 | 1.79–4.54 / 21 |
| WildVoicesource 5-point scale | 4.04–4.62 / 5 | 1.60–4.58 / 21 |
| SD-QAsource score /100 | 47.47–75.77 / 5 | 4.16–76.90 / 21 |
| MMSUsource score /100 | 51.37–81.69 / 5 | 24.27–80.25 / 21 |
| OBQAsource score /100 | 60.66–92.97 / 5 | 25.27–89.70 / 21 |
| BBHsource score /100 | 63.90–87.20 / 5 | 46.30–84.10 / 21 |
| IFEvalsource score /100 | 69.53–78.99 / 5 | 11.56–77.80 / 21 |
| AdvBenchsource score /100 | 98.08–99.81 / 5 | 11.35–100.00 / 21 |
Interpretation limits
A range describes this selected set of published rows. It is not an estimate of an architecture’s expected performance and cannot establish a causal advantage. The groups are unequal, not randomly sampled and include model/configuration differences. VoiceBench Overall is deliberately excluded from this derived summary.
Evaluate the deployment choices separately: audio handling, interruption behavior, integration ownership, telephony and handoff need evidence beyond this table.