Compare Lab
OpenAI Realtime API vs Speechmatics
OpenAI Realtime API is positioned as speech-to-speech API for developers implementing their own voice application. Speechmatics is positioned as recognition APIs with cloud and private-deployment choices for transcription-heavy pr
OpenAI Realtime API is positioned as speech-to-speech API for developers implementing their own voice application. Speechmatics is positioned as recognition APIs with cloud and private-deployment choices for transcription-heavy products. Start by deciding who owns the missing layers, then inspect the documented differences below.
Developer Frameworks and STT are different product categories. An underlying engine or framework is not a complete substitute for a deployed agent workflow.
7 jointly documented catalog fields. This measures evidence coverage, not equivalent quality or an overall winner.
Test opening hours, booking conflicts and an unanswered transfer.
Start where the records differ.
Values and evidence states are compared together. Shared facts follow the differences; Unknown never means No.
| Decision point | OpenAI Realtime API | Speechmatics |
|---|---|---|
| Billing unitDifferent records | 1 million text/audio/image tokens; transcription separately metered | Credits (1 credit = $1); product-specific per-hour rates, calculated to the second |
| Bring your telephonyDifferent records | Yes | Unknown |
| Bring your modelDifferent records | Unknown | Not applicable |
| HIPAA pathDifferent records | Unknown | Yes |
| Billing modelShared state | component-stack | component-stack |
| Warm transferShared state | Unknown | Unknown |
Your requirements & next steps
Your requirements
Ready-made / AI receptionist · Strict evidence mode
No personal requirements yet. Add criteria in Advanced Filter to prioritize this comparison.
Realtime model tokens plus input transcription, tools and carrier/hosting; not GPT-Live session-minute billing.
Credit billing: one credit equals $1; usage measured by product hours/seconds. This does not mean a character-priced TTS rate or flat agent minute rate.
Cached audio/text and image inputs have separate rates. Rate is not per-minute.
Price page exposes a Pro 0.129 figure without a clear unit in captured content; exact model rate and enterprise commitment require confirmation.
Model scope fixed to GPT-Realtime-2, not all voice models.
Billing currency versus usage unit kept distinct.
Published Realtime token rates do not establish these account/contract terms. GPT-Live duration billing and unrelated fine-tuning discounts must not be imported.
No fixed subscription number inferred from credit billing.
Published Realtime token rates do not establish these account/contract terms. GPT-Live duration billing and unrelated fine-tuning discounts must not be imported.
Existing pre-August accounts have different transition grant.
Published Realtime token rates do not establish these account/contract terms. GPT-Live duration billing and unrelated fine-tuning discounts must not be imported.
No forced per-minute bundle conversion.
Repeated context, cache behavior and turn count affect token charges.
Buyer-selected component costs sit outside speech API usage.
Not yet researched against this buyer-critical field standard.
Not yet researched against this buyer-critical field standard.
Full bill requires measured model/token/cache/tool/carrier usage; conversation minutes alone are insufficient.
Speech component credits and contract-specific workloads do not determine full voice-agent cost; external LLM/hosting/telephony unknown.
Full bill requires measured model/token/cache/tool/carrier usage; conversation minutes alone are insufficient.
Speech component credits and contract-specific workloads do not determine full voice-agent cost; external LLM/hosting/telephony unknown.
Anyone with this link can view the encoded decision criteria.
Vendor documentation establishes a claim. Public model benchmarks are separate context, not measurements of these platforms.
Open saved comparison ↗What to verify together.
OpenAI Realtime API
A token rate cannot be compared directly with a bundled call-minute price.
GPT-Realtime-2 audio input $32/M, audio output $64/M; text input $4/M, output $24/M · V3.3 evidence review · GPT-Realtime-2 input audio tokens only; output/audio/text/cache have different rates; not 2.1 rate evidence
Excluded / confirm: Complete cost cannot be established from this reference. Full bill requires measured model/token/cache/tool/carrier usage; conversation minutes alone are insufficient. · Input transcription, tool/backend calls, carrier and application infrastructure
Realtime model tokens plus input transcription, tools and carrier/hosting; not GPT-Live session-minute billing.
Speechmatics
The lowest advertised batch price is not a real-time price.
Pro self-serve PAYG; Enterprise contract billing · V3.3 evidence review · STT exact-second metering billed in audio hours; no verified current complete rate schedule
Excluded / confirm: Complete cost cannot be established from this reference. Speech component credits and contract-specific workloads do not determine full voice-agent cost; external LLM/hosting/telephony unknown. · External LLM, orchestrator, telephony and application hosting
Credit billing: one credit equals $1; usage measured by product hours/seconds. This does not mean a character-priced TTS rate or flat agent minute rate.
The evidence overlap.
Both records have documented values for Product type, Technical setup, Pricing model, Published numeric pricing, Realtime / streaming audio, Checked within 30 days, Retained price snapshot. Matching availability does not measure behavior under your call conditions.
Run the same booking, tool failure and unsuccessful human-transfer cases. Record both outcomes and billable units before signing a deployment agreement.
Build a buyer-operated pilot →