Compare Lab
AssemblyAI vs Deepgram
AssemblyAI is positioned as transcription APIs for live and recorded audio, with model-specific language and analysis options. Deepgram is positioned as speech recognition models with separate speech generation and managed voice-a
AssemblyAI is positioned as transcription APIs for live and recorded audio, with model-specific language and analysis options. Deepgram is positioned as speech recognition models with separate speech generation and managed voice-agent APIs. Start by deciding who owns the missing layers, then inspect the documented differences below.
17 jointly documented catalog fields. This measures evidence coverage, not equivalent quality or an overall winner.
Test opening hours, booking conflicts and an unanswered transfer.
Start where the records differ.
Values and evidence states are compared together. Shared facts follow the differences; Unknown never means No.
| Decision point | AssemblyAI | Deepgram |
|---|---|---|
| Billing unitDifferent records | Async audio hours; streaming connection hours; agent session minutes; gateway tokens | STT audio minute; TTS thousand characters; agent WebSocket connection minute |
| Bring your telephonyDifferent records | Yes | Yes |
| HIPAA pathDifferent records | Yes | Yes |
| Billing modelShared state | component-stack | component-stack |
| Bring your modelShared state | Yes | Yes |
| Warm transferShared state | Unknown | Unknown |
Your requirements & next steps
Your requirements
Ready-made / AI receptionist · Strict evidence mode
No personal requirements yet. Add criteria in Advanced Filter to prioritize this comparison.
STT hours/session duration, agent minutes, gateway tokens and add-ons remain distinct.
STT audio minutes, TTS characters and Voice Agent connection minutes are distinct.
Model rates are scoped; add-ons stack; reference states updated May 29, 2026 and conflicts are retained.
External provider and telephony bills excluded where applicable. Do not infer a complete monthly ceiling.
Idle streaming connection duration is billable. No fabricated all-in phone cost per minute.
Billable open connection time is not just spoken audio.
SSO costs $199 per connection per month; do not mistake no base fee for no fixed optional charges.
Growth is annual prepaid commitment, not invented monthly subscription.
Not a fixed monthly audio allowance; credits do not expire.
No fixed included agent minutes inferred from credit dollars.
No invented bundle overage.
Initial credits, Growth commitments and tier rates are documented; no recurring free plan or universal rounding/contract terms inferred.
Component choices and call mix must be specified.
Agent pricing covers selected tier and components; full workload determines bill.
Not yet researched against this buyer-critical field standard.
Not yet researched against this buyer-critical field standard.
No complete phone bill ceiling from minutes alone; carrier, model and addon scope vary.
Carrier, hosting, model tier and BYO usage prevent a universal full monthly or per-minute ceiling.
No complete phone bill ceiling from minutes alone; carrier, model and addon scope vary.
Carrier, hosting, model tier and BYO usage prevent a universal full monthly or per-minute ceiling.
Anyone with this link can view the encoded decision criteria.
Vendor documentation establishes a claim. Public model benchmarks are separate context, not measurements of these platforms.
Open saved comparison ↗What to verify together.
AssemblyAI
A transcription hour is not a complete agent hour.
Universal-2 async / Universal-Streaming $0.15/hour; U3.5 Pro async $0.21/hour; U3 realtime $0.45/hour; Voice Agent $0.075/min · V3.3 evidence review · Voice Agent API reference; standalone STT is hourly; transport/add-ons separate
Excluded / confirm: Complete cost cannot be established from this reference. No complete phone bill ceiling from minutes alone; carrier, model and addon scope vary. · Twilio carrier; optional custom/Gateway LLM; standalone STT add-ons and per-channel billing
STT hours/session duration, agent minutes, gateway tokens and add-ons remain distinct.
Deepgram
The current Nova-3 streaming rate is promotional.
Voice Agent Standard $0.075/min; BYO TTS $0.065/min; BYO LLM + TTS $0.050/min; Advanced $0.163/min (PAYG) · V3.3 evidence review · Voice Agent Standard PAYG WebSocket connection minutes; other tiers/BYO variants differ
Excluded / confirm: Complete cost cannot be established from this reference. Carrier, hosting, model tier and BYO usage prevent a universal full monthly or per-minute ceiling. · BYO model/TTS charges, carrier and customer transport hosting
STT audio minutes, TTS characters and Voice Agent connection minutes are distinct.
The evidence overlap.
Both records have documented values for Product type, Technical setup, Pricing model, Published numeric pricing, Free trial, Twilio, Realtime / streaming audio, OpenAI models, Anthropic models, Google models, Bring your model, Function / tool calling and additional fields shown in Compare. Matching availability does not measure behavior under your call conditions.
Run the same booking, tool failure and unsuccessful human-transfer cases. Record both outcomes and billable units before signing a deployment agreement.
Build a buyer-operated pilot →