VoiceAgent Benchmark Intelligence
Inspect published voice-model evidence with its original dimensions, provenance and limitations.
These are published model results. VoiceAgent.best ran no model calls or voice-agent tests. They do not measure a platform, phone call, production latency, voice quality or handoff reliability.
Choose the question, then inspect the evidence
Architecture context
Inspect group coverage and metric ranges, with unequal samples and selection limits visible.
Cascaded vs Omni →Comparable published rows
Inspect the public result rows
Filter by the classifications published by VoiceBench. “Open” describes its weight-availability label; it does not establish a commercial license or a platform’s availability.
46 model records
| Model | Architecture | Weights | AlpacaEvalsource 5-point scale | CommonEvalsource 5-point scale | WildVoicesource 5-point scale | SD-QAsource score /100 | MMSUsource score /100 | OBQAsource score /100 | BBHsource score /100 | IFEvalsource score /100 | AdvBenchsource score /100 | VoiceBench OverallSource reported /100 | Provenance |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Baichuan-Audio | OmniCommunity label | openSource classification | 4.41 | 4.08 | 3.92 | 45.84 | 53.19 | 71.65 | 54.80 | 50.31 | 99.42 | 69.27VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Baichuan-Omni-1.5 | OmniCommunity label | openSource classification | 4.50 | 4.05 | 4.06 | 43.40 | 57.25 | 74.51 | 62.70 | 54.54 | 97.31 | 71.32VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| BR-Voice-Reasoner | Vision + Audio + LLMCommunity label | openSource classification | 4.87 | 4.71 | 4.66 | 74.30 | 85.70 | 96.00 | 90.40 | 83.20 | 99.80 | 90.46VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| DiVA | Audio-LLMCommunity label | openSource classification | 3.67 | 3.54 | 3.74 | 57.05 | 25.76 | 25.49 | 51.80 | 39.15 | 98.27 | 57.39VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Freeze-Omni | S2S / Full-DuplexCommunity label | openSource classification | 4.03 | 3.46 | 3.15 | 53.45 | 28.14 | 30.98 | 50.70 | 23.40 | 97.30 | 55.20VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| GLM-4-Voice | OmniCommunity label | openSource classification | 3.97 | 3.42 | 3.18 | 36.98 | 39.75 | 53.41 | 52.80 | 25.92 | 88.08 | 56.48VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| GPT-4o-Audio | OmniCommunity label | closedSource classification | 4.78 | 4.49 | 4.58 | 75.50 | 80.25 | 89.23 | 84.10 | 76.02 | 98.65 | 86.75VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| GPT-4o-mini-Audio | OmniCommunity label | closedSource classification | 4.75 | 4.24 | 4.40 | 67.36 | 72.90 | 84.84 | 81.50 | 72.90 | 98.27 | 82.84VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Ichigo | OmniCommunity label | openSource classification | 3.79 | 3.17 | 2.83 | 36.53 | 25.63 | 26.59 | 46.50 | 21.59 | 57.50 | 45.57VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Kimi-Audio | OmniCommunity label | openSource classification | 4.46 | 3.97 | 4.20 | 63.12 | 62.17 | 83.52 | 69.70 | 61.10 | 100.00 | 76.91VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| LFG-1 | Audio-LLMCommunity label | openSource classification | 4.60 | 3.71 | 4.01 | 62.39 | 75.60 | 82.42 | 83.90 | 74.85 | 91.15 | 79.63VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| LFG-2 | Audio-LLMCommunity label | openSource classification | 4.67 | 4.26 | 4.23 | 73.42 | 85.20 | 93.63 | 87.10 | 84.54 | 95.77 | 86.98VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| LFG-3 | Audio-LLMCommunity label | openSource classification | 4.73 | 4.40 | 4.45 | 78.12 | 85.52 | 94.73 | 92.20 | 88.54 | 98.27 | 89.88VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| LLaMA-Omni | OmniCommunity label | openSource classification | 3.70 | 3.46 | 2.92 | 39.69 | 25.93 | 27.47 | 49.20 | 14.87 | 11.35 | 41.12VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Lyra-Base | OmniCommunity label | openSource classification | 3.85 | 3.50 | 3.42 | 38.25 | 49.74 | 72.75 | 59.00 | 36.28 | 59.62 | 59.00VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Lyra-Mini | OmniCommunity label | openSource classification | 2.99 | 2.69 | 2.58 | 19.89 | 31.42 | 41.54 | 48.40 | 20.91 | 80.00 | 45.26VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Mair-hub-0.5B-Omni | OmniCommunity label | openSource classification | 3.06 | 2.87 | 2.48 | 21.70 | 25.60 | 25.27 | 50.90 | 14.85 | 94.81 | 44.59VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Megrez-3B-Omni | OmniCommunity label | openSource classification | 3.50 | 2.95 | 2.34 | 25.95 | 27.03 | 28.35 | 50.30 | 25.71 | 87.69 | 46.76VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| MERaLiON | Audio-LLMCommunity label | openSource classification | 4.50 | 3.77 | 4.12 | 55.06 | 34.95 | 27.23 | 62.60 | 62.93 | 94.81 | 65.04VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Mini-Omni | OmniCommunity label | openSource classification | 1.95 | 2.02 | 1.61 | 13.92 | 24.69 | 26.59 | 46.30 | 13.58 | 37.12 | 30.42VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Mini-Omni2 | OmniCommunity label | openSource classification | 2.32 | 2.18 | 1.79 | 9.31 | 24.27 | 26.59 | 46.40 | 11.56 | 57.50 | 33.49VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| MiniCPM-o | OmniCommunity label | openSource classification | 4.42 | 4.15 | 3.94 | 50.72 | 54.78 | 78.02 | 60.40 | 49.25 | 97.69 | 71.23VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Moshi | S2S / Full-DuplexCommunity label | openSource classification | 2.01 | 1.60 | 1.30 | 15.64 | 24.04 | 25.93 | 47.40 | 10.12 | 44.23 | 29.51VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Nemotron 3 VoiceChat (V1) | S2S / Full-DuplexCommunity label | openSource classification | 3.59 | 3.12 | 3.06 | 43.40 | 50.46 | 64.40 | 52.70 | 17.12 | 99.62 | 58.10VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| NVIDIA Nemotron 3 Nano Omni 30B A3B | Vision + Audio + LLMCommunity label | openSource classification | 4.75 | 4.57 | 4.58 | 71.43 | 82.30 | 92.97 | 91.10 | 88.66 | 100.00 | 89.39VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Ola | OmniCommunity label | openSource classification | 4.12 | 2.97 | 3.19 | 33.82 | 45.97 | 67.91 | 51.10 | 39.57 | 90.77 | 59.42VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Parakeet-TDT-0.6b-V2 + Qwen3-8B | CascadedCommunity label | openSource classification | 4.68 | 4.46 | 4.35 | 47.47 | 59.10 | 80.00 | 77.90 | 78.99 | 99.81 | 79.23VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Phi-4-multimodal | Vision + Audio + LLMCommunity label | openSource classification | 3.81 | 3.82 | 3.56 | 39.78 | 42.19 | 65.93 | 61.80 | 45.35 | 100.00 | 64.32VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Qwen2-Audio | Audio-LLMCommunity label | openSource classification | 3.74 | 3.43 | 3.01 | 35.71 | 35.72 | 49.45 | 54.70 | 26.33 | 96.73 | 55.80VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Qwen3-Omni-30B-A3B-Instruct | OmniCommunity label | openSource classification | 4.74 | 4.54 | 4.58 | 76.90 | 68.10 | 89.70 | 80.40 | 77.80 | 99.30 | 85.50VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Qwen3-Omni-30B-A3B-Thinking | Vision + Audio + LLMCommunity label | openSource classification | 4.82 | 4.53 | 4.53 | 78.10 | 83.00 | 94.30 | 88.90 | 80.60 | 97.20 | 88.80VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| SLAM-Omni | OmniCommunity label | openSource classification | 1.90 | 1.79 | 1.60 | 4.16 | 26.06 | 25.27 | 48.80 | 13.38 | 94.23 | 35.30VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Step-Audio | OmniCommunity label | openSource classification | 4.13 | 3.09 | 2.93 | 44.21 | 28.33 | 33.85 | 50.60 | 27.96 | 69.62 | 50.84VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Ultravox-GLM-4P6 | Audio-LLMCommunity label | openSource classification | 4.93 | 4.42 | 4.57 | 84.24 | 80.40 | 89.45 | 81.48 | 75.59 | 99.23 | 87.05VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Ultravox-GLM-4P7 | Audio-LLMCommunity label | openSource classification | 4.87 | 4.30 | 4.55 | 84.15 | 83.82 | 94.72 | 87.16 | 76.26 | 99.23 | 88.86VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Ultravox-GLM-4P7 (thinking) | Audio-LLMCommunity label | openSource classification | 4.69 | 4.05 | 4.21 | 88.57 | 87.15 | 95.01 | 89.78 | 80.99 | 98.46 | 88.79VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Ultravox-v0.4.1-LLaMA-3.1-8B | Audio-LLMCommunity label | openSource classification | 4.55 | 3.90 | 4.12 | 53.35 | 47.17 | 65.27 | 66.30 | 66.88 | 98.46 | 72.09VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Ultravox-v0.5-LLaMA-3.1-8B | Audio-LLMCommunity label | openSource classification | 4.59 | 4.11 | 4.28 | 58.68 | 54.16 | 68.35 | 67.80 | 66.51 | 98.65 | 74.86VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Ultravox-v0.5-LLaMA-3.2-1B | Audio-LLMCommunity label | openSource classification | 4.04 | 3.57 | 3.47 | 34.72 | 30.03 | 35.60 | 52.70 | 45.56 | 96.92 | 57.46VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Ultravox-v0.6-LLaMA-3.3-70B | Audio-LLMCommunity label | openSource classification | 4.69 | 4.26 | 4.38 | 82.60 | 69.20 | 86.40 | 78.80 | 61.50 | 91.20 | 81.81VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| VITA-1.0 | OmniCommunity label | openSource classification | 3.38 | 2.15 | 1.87 | 27.94 | 25.70 | 29.01 | 47.70 | 22.82 | 26.73 | 36.43VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| VITA-1.5 | OmniCommunity label | openSource classification | 4.21 | 3.66 | 3.48 | 38.88 | 52.15 | 71.65 | 55.30 | 38.14 | 97.69 | 64.53VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Whisper-v3-large + GPT-4o | CascadedCommunity label | closedSource classification | 4.80 | 4.47 | 4.62 | 75.77 | 81.69 | 92.97 | 87.20 | 76.51 | 98.27 | 87.80VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Whisper-v3-large + LLaMA-3.1-8B | CascadedCommunity label | openSource classification | 4.53 | 4.04 | 4.16 | 70.43 | 62.43 | 72.53 | 69.70 | 69.53 | 98.08 | 77.48VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Whisper-v3-turbo + LLaMA-3.1-8B | CascadedCommunity label | openSource classification | 4.55 | 4.02 | 4.12 | 58.23 | 62.04 | 72.09 | 69.10 | 71.12 | 98.46 | 76.09VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| Whisper-v3-turbo + LLaMA-3.2-3B | CascadedCommunity label | openSource classification | 4.45 | 3.82 | 4.04 | 49.28 | 51.37 | 60.66 | 63.90 | 69.71 | 98.08 | 71.02VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
Official model context stays separate
Two Qwen variants have source-linked input, output and component specifications. Other rows retain unknown official specifications until exact model evidence is attached. A published family name is not enough to infer pricing, latency or platform support.