GPT-4o-Audio vs GPT-4o-mini-Audio
Compare their published VoiceBench columns before drawing conclusions about a deployment.
These are published model results. VoiceAgent.best ran no model calls or voice-agent tests. They do not measure a platform, phone call, production latency, voice quality or handoff reliability.
Shared published evidence
Both rows have 9 populated dimensions in the same captured source table. Their metric columns and display scales match. Exact evaluation configurations, API checkpoints and run dates remain unknown, so this is a comparison of published rows rather than a controlled experiment.
| Model | Architecture | Weights | AlpacaEvalsource 5-point scale | CommonEvalsource 5-point scale | WildVoicesource 5-point scale | SD-QAsource score /100 | MMSUsource score /100 | OBQAsource score /100 | BBHsource score /100 | IFEvalsource score /100 | AdvBenchsource score /100 | VoiceBench OverallSource reported /100 | Provenance |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-4o-Audio | OmniCommunity label | closedSource classification | 4.78 | 4.49 | 4.58 | 75.50 | 80.25 | 89.23 | 84.10 | 76.02 | 98.65 | 86.75VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
| GPT-4o-mini-Audio | OmniCommunity label | closedSource classification | 4.75 | 4.24 | 4.40 | 67.36 | 72.90 | 84.84 | 81.50 | 72.90 | 98.27 | 82.84VoiceBench Overall | VoiceBenchThird-party Benchmark 2026-09-27 |
Reasoning, instruction following and safety
Reasoning
GPT-4o-Audio: 84.10
GPT-4o-mini-Audio: 81.50
Instruction following
GPT-4o-Audio: 76.02
GPT-4o-mini-Audio: 72.90
Safety
GPT-4o-Audio: 98.65
GPT-4o-mini-Audio: 98.27
How to use this comparison
Use the dimensions to identify questions for a real pilot: which instructions fail, which reasoning tasks matter and which unsafe requests should be escalated. The table has no measurements of call interruption, network behavior, cost or customer outcomes.
Differences cannot be attributed to architecture alone. Other model, prompt and evaluation choices are uncontrolled in this public snapshot. Supporting one of these model families does not transfer these scores to a platform.