Explore voiceagent.best
Estimate costOur independent methodology

Start here

Comparison LabCompare 27 voice AI platforms and components across seven buying contexts, with sourced capabilities, current price captures and visible unknowns.Voice AI Price Index: prices, evidence and historyUnderstand the complete voice-agent cost stack: models, voice, transcription, telephony, subscriptions, capacity and add-ons.Know the cost before the first callEstimate monthly voice-agent costs, AI minutes, setup expenses and volume scenarios using your own transparent assumptions.AI voice agents for dentalAppointments, rescheduling and after-hours calls—with a person ready for clinical questions. Explore call flows, safe automation, integrations, tests and cost considerations.

Public evidence, inspectable limits.

L0 Benchmark Intelligence aggregates published metadata and explains its boundaries.

Third-party Benchmark · VoiceBench · Captured 2026-09-27

These are published model results. VoiceAgent.best ran no model calls or voice-agent tests. They do not measure a platform, phone call, production latency, voice quality or handoff reliability.

Leaderboard · Repository · Paper · Dataset card · Apache-2.0 repository terms

Citation: Chen, Yue, Zhang, Gao, Tan and Li (2024). VoiceBench: Benchmarking LLM-Based Voice Assistants. arXiv:2410.17196.

Evidence labels

Official
Specifications attached to an exact public model-card record.
Third-party Benchmark
Values published in the VoiceBench leaderboard, preserved with their source scales. Some rows are imported there from official model results.
Vendor Claim
A vendor’s performance assertion. Such claims are not our measurements and are not merged into the score table.
Derived from public benchmark
Explicit calculations over captured rows, with a formula and denominator beside the output.
Unknown
Information the cited record does not establish. Missing scores remain null rather than zero.

Capture and validation

We capture public leaderboard metadata, retain a SHA-256 of the source artifact and store a dated JSON snapshot. The adapter rejects unexpected columns, duplicate model identities, out-of-scale scores and missing provenance. New captures use exclusive file creation: an existing date cannot be overwritten.

The capture fingerprint versions our source artifact. It is not a VoiceBench protocol release, a model release date or the date experiments occurred. A repository revision is recorded as supporting provenance; it does not prove that a deployed GitHub Pages artifact corresponds to that revision.

Comparison rules

Pair pages require two model rows in the same source capture and nine populated shared metric columns. Same table and scales permit a descriptive comparison; identical protocols and checkpoints are not assumed. We do not average incompatible metrics or produce a VoiceAgent.best universal score.

Models and platforms

A model score cannot become a Retell, Vapi or other platform score. Transcription, speech generation, voice activity detection, prompts, tools, network conditions, telephony and human handoff all affect the deployed system. Official model-support documentation may support contextual links only, with this limitation visible.

Rights and attribution

The repository license and dataset card were checked on 2026-09-27; both identify Apache-2.0. We retain factual result metadata and the paper citation. We do not mirror audio, download model weights or reproduce the leaderboard design. These source terms are not a blanket license for every third-party dataset or model linked by the project.

A source with unclear terms stays citation-only until review. Dataset license status does not establish an evaluation result’s reproducibility.

Refresh policy

Review public source metadata at least monthly. Append a new dated capture when the source changes; preserve prior files and review row differences before publication. This module has no scheduled job, paid API integration or simulated self-test path.

Known limits

Public submissions may have different evaluation settings and unverified checkpoints. Architecture and open/closed classifications are community-maintained. The nine columns cover benchmark tasks rather than the full operating behavior of a voice agent. No uncertainty intervals or independently repeated runs are available in this capture.