Compare Lab
Deepgram vs OpenAI Realtime API
Deepgram is positioned as speech recognition models with separate speech generation and managed voice-agent APIs. OpenAI Realtime API is positioned as speech-to-speech API for developers implementing their own voice application. S
Deepgram is positioned as speech recognition models with separate speech generation and managed voice-agent APIs. OpenAI Realtime API is positioned as speech-to-speech API for developers implementing their own voice application. Start by deciding who owns the missing layers, then inspect the documented differences below.
STT and Developer Frameworks are different product categories. An underlying engine or framework is not a complete substitute for a deployed agent workflow.
11 jointly documented catalog fields. This measures evidence coverage, not equivalent quality or an overall winner.
Test opening hours, booking conflicts and an unanswered transfer.
Start where the records differ.
Values and evidence states are compared together. Shared facts follow the differences; Unknown never means No.
| Decision point | Deepgram | OpenAI Realtime API |
|---|---|---|
| Billing unitDifferent records | STT audio minute; TTS thousand characters; agent WebSocket connection minute | 1 million text/audio/image tokens; transcription separately metered |
| Bring your modelDifferent records | Yes | Unknown |
| HIPAA pathDifferent records | Yes | Unknown |
| Billing modelShared state | component-stack | component-stack |
| Bring your telephonyShared state | Yes | Yes |
| Warm transferShared state | Unknown | Unknown |
Your requirements & next steps
Your requirements
Ready-made / AI receptionist · Strict evidence mode
No personal requirements yet. Add criteria in Advanced Filter to prioritize this comparison.
STT audio minutes, TTS characters and Voice Agent connection minutes are distinct.
Realtime model tokens plus input transcription, tools and carrier/hosting; not GPT-Live session-minute billing.
External provider and telephony bills excluded where applicable. Do not infer a complete monthly ceiling.
Cached audio/text and image inputs have separate rates. Rate is not per-minute.
Billable open connection time is not just spoken audio.
Model scope fixed to GPT-Realtime-2, not all voice models.
Growth is annual prepaid commitment, not invented monthly subscription.
Published Realtime token rates do not establish these account/contract terms. GPT-Live duration billing and unrelated fine-tuning discounts must not be imported.
No fixed included agent minutes inferred from credit dollars.
Published Realtime token rates do not establish these account/contract terms. GPT-Live duration billing and unrelated fine-tuning discounts must not be imported.
Initial credits, Growth commitments and tier rates are documented; no recurring free plan or universal rounding/contract terms inferred.
Published Realtime token rates do not establish these account/contract terms. GPT-Live duration billing and unrelated fine-tuning discounts must not be imported.
Agent pricing covers selected tier and components; full workload determines bill.
Repeated context, cache behavior and turn count affect token charges.
Not yet researched against this buyer-critical field standard.
Not yet researched against this buyer-critical field standard.
Carrier, hosting, model tier and BYO usage prevent a universal full monthly or per-minute ceiling.
Full bill requires measured model/token/cache/tool/carrier usage; conversation minutes alone are insufficient.
Carrier, hosting, model tier and BYO usage prevent a universal full monthly or per-minute ceiling.
Full bill requires measured model/token/cache/tool/carrier usage; conversation minutes alone are insufficient.
Anyone with this link can view the encoded decision criteria.
Vendor documentation establishes a claim. Public model benchmarks are separate context, not measurements of these platforms.
Open saved comparison ↗What to verify together.
Deepgram
The current Nova-3 streaming rate is promotional.
Voice Agent Standard $0.075/min; BYO TTS $0.065/min; BYO LLM + TTS $0.050/min; Advanced $0.163/min (PAYG) · V3.3 evidence review · Voice Agent Standard PAYG WebSocket connection minutes; other tiers/BYO variants differ
Excluded / confirm: Complete cost cannot be established from this reference. Carrier, hosting, model tier and BYO usage prevent a universal full monthly or per-minute ceiling. · BYO model/TTS charges, carrier and customer transport hosting
STT audio minutes, TTS characters and Voice Agent connection minutes are distinct.
OpenAI Realtime API
A token rate cannot be compared directly with a bundled call-minute price.
GPT-Realtime-2 audio input $32/M, audio output $64/M; text input $4/M, output $24/M · V3.3 evidence review · GPT-Realtime-2 input audio tokens only; output/audio/text/cache have different rates; not 2.1 rate evidence
Excluded / confirm: Complete cost cannot be established from this reference. Full bill requires measured model/token/cache/tool/carrier usage; conversation minutes alone are insufficient. · Input transcription, tool/backend calls, carrier and application infrastructure
Realtime model tokens plus input transcription, tools and carrier/hosting; not GPT-Live session-minute billing.
The evidence overlap.
Both records have documented values for Product type, Technical setup, Pricing model, Published numeric pricing, Realtime / streaming audio, OpenAI models, Function / tool calling, Interruption controls, Official product documentation, Checked within 30 days, Retained price snapshot. Matching availability does not measure behavior under your call conditions.
Run the same booking, tool failure and unsuccessful human-transfer cases. Record both outcomes and billable units before signing a deployment agreement.
Build a buyer-operated pilot →