Explore voiceagent.best
Estimate costOur independent methodology

Start here

Comparison LabCompare 27 voice AI platforms and components across seven buying contexts, with sourced capabilities, current price captures and visible unknowns.Voice AI Price Index: prices, evidence and historyUnderstand the complete voice-agent cost stack: models, voice, transcription, telephony, subscriptions, capacity and add-ons.Know the cost before the first callEstimate monthly voice-agent costs, AI minutes, setup expenses and volume scenarios using your own transparent assumptions.AI voice agents for dentalAppointments, rescheduling and after-hours calls—with a person ready for clinical questions. Explore call flows, safe automation, integrations, tests and cost considerations.

Cascaded vs Omni

What the published model sets can tell us—and what their architecture labels leave unresolved.

Third-party Benchmark · VoiceBench · Captured 2026-09-27

These are published model results. VoiceAgent.best ran no model calls or voice-agent tests. They do not measure a platform, phone call, production latency, voice quality or handoff reliability.

Leaderboard · Repository · Paper · Dataset card · Apache-2.0 repository terms

Citation: Chen, Yue, Zhang, Gao, Tan and Li (2024). VoiceBench: Benchmarking LLM-Based Voice Assistants. arXiv:2410.17196.

Two source-defined groups

VoiceBench classifies cascaded rows as a separate speech-recognition and language-model pipeline. Its Omni category represents turn-based speech input and output. These are community labels; the table separately classifies S2S / Full-Duplex, so Omni does not establish simultaneous listening and speaking.

The capture contains 5 cascaded rows and 21 omni rows. Model sizes, access, prompt choices and run settings differ or remain unspecified.

Observed metric ranges

VoiceAgent.best derived summary of public VoiceBench rows. For each metric and architecture, range = min(non-null published values) to max(non-null published values); n = number of populated rows. No averages, normalized score or architecture winner.

Derived ranges · VoiceBench source · 2026-09-27
Metric / source scaleCascaded / nOmni / n
AlpacaEvalsource 5-point scale4.45–4.80 / 51.90–4.78 / 21
CommonEvalsource 5-point scale3.82–4.47 / 51.79–4.54 / 21
WildVoicesource 5-point scale4.04–4.62 / 51.60–4.58 / 21
SD-QAsource score /10047.47–75.77 / 54.16–76.90 / 21
MMSUsource score /10051.37–81.69 / 524.27–80.25 / 21
OBQAsource score /10060.66–92.97 / 525.27–89.70 / 21
BBHsource score /10063.90–87.20 / 546.30–84.10 / 21
IFEvalsource score /10069.53–78.99 / 511.56–77.80 / 21
AdvBenchsource score /10098.08–99.81 / 511.35–100.00 / 21

Interpretation limits

A range describes this selected set of published rows. It is not an estimate of an architecture’s expected performance and cannot establish a causal advantage. The groups are unequal, not randomly sampled and include model/configuration differences. VoiceBench Overall is deliberately excluded from this derived summary.

Evaluate the deployment choices separately: audio handling, interruption behavior, integration ownership, telephony and handoff need evidence beyond this table.