Compare Lab
Deepgram vs Fish Audio
Deepgram is positioned as speech recognition models with separate speech generation and managed voice-agent APIs. Fish Audio is positioned as a developer speech API and creative voice workspace; paid TTS usage is metered in UTF-8
Deepgram is positioned as speech recognition models with separate speech generation and managed voice-agent APIs. Fish Audio is positioned as a developer speech API and creative voice workspace; paid TTS usage is metered in UTF-8 bytes. Start by deciding who owns the missing layers, then inspect the documented differences below.
STT and Voice Engines are different product categories. An underlying engine or framework is not a complete substitute for a deployed agent workflow.
14 jointly documented catalog fields. This measures evidence coverage, not equivalent quality or an overall winner.
Test opening hours, booking conflicts and an unanswered transfer.
Start where the records differ.
Values and evidence states are compared together. Shared facts follow the differences; Unknown never means No.
| Decision point | Deepgram | Fish Audio |
|---|---|---|
| Billing unitDifferent records | STT audio minute; TTS thousand characters; agent WebSocket connection minute | TTS: million UTF-8 bytes; ASR: audio hours; Agents: conversation minutes |
| Bring your telephonyDifferent records | Yes | Yes |
| Warm transferDifferent records | Unknown | Yes |
| HIPAA pathDifferent records | Yes | Unknown |
| WebhooksDifferent records | YesUnknown | YesAPI / webhook |
| Billing modelShared state | component-stack | component-stack |
Your requirements & next steps
Your requirements
Ready-made / AI receptionist · Strict evidence mode
No personal requirements yet. Add criteria in Advanced Filter to prioritize this comparison.
STT audio minutes, TTS characters and Voice Agent connection minutes are distinct.
TTS uses UTF-8 bytes, ASR audio hours, Agents minutes plus carrier/model/add-ons. Multiple products must not share a fabricated normalized unit.
External provider and telephony bills excluded where applicable. Do not infer a complete monthly ceiling.
Free TTS model and Agents promotional basic LLM scope are separate.
Billable open connection time is not just spoken audio.
Each product keeps its own billing unit.
Growth is annual prepaid commitment, not invented monthly subscription.
API access; managed numbers have separate rental and enterprise agreement may differ.
No fixed included agent minutes inferred from credit dollars.
Real sessions consume balance including relevant carrier/model charges; no monthly included minutes.
Initial credits, Growth commitments and tier rates are documented; no recurring free plan or universal rounding/contract terms inferred.
Destination, transfer and selected model affect cost; no complete ceiling.
Agent pricing covers selected tier and components; full workload determines bill.
Rates have explicit event/destination scope, not a single maximum.
Not yet researched against this buyer-critical field standard.
Not yet researched against this buyer-critical field standard.
Carrier, hosting, model tier and BYO usage prevent a universal full monthly or per-minute ceiling.
Call mix, destination, transfer time, selected model/token usage and separate speech products prevent a universal complete monthly/minute ceiling.
Carrier, hosting, model tier and BYO usage prevent a universal full monthly or per-minute ceiling.
Call mix, destination, transfer time, selected model/token usage and separate speech products prevent a universal complete monthly/minute ceiling.
Anyone with this link can view the encoded decision criteria.
Vendor documentation establishes a claim. Public model benchmarks are separate context, not measurements of these platforms.
Open saved comparison ↗What to verify together.
Deepgram
The current Nova-3 streaming rate is promotional.
Voice Agent Standard $0.075/min; BYO TTS $0.065/min; BYO LLM + TTS $0.050/min; Advanced $0.163/min (PAYG) · V3.3 evidence review · Voice Agent Standard PAYG WebSocket connection minutes; other tiers/BYO variants differ
Excluded / confirm: Complete cost cannot be established from this reference. Carrier, hosting, model tier and BYO usage prevent a universal full monthly or per-minute ceiling. · BYO model/TTS charges, carrier and customer transport hosting
STT audio minutes, TTS characters and Voice Agent connection minutes are distinct.
Fish Audio
UTF-8 bytes and characters are not interchangeable for multilingual scripts.
TTS $15/M UTF-8 bytes; ASR $0.36/audio hour; Agents $0.06/min + applicable surcharges · V3.3 evidence review · TTS API reference; Voice Agents beta has separate minute-based billing
Excluded / confirm: Complete cost cannot be established from this reference. Call mix, destination, transfer time, selected model/token usage and separate speech products prevent a universal complete monthly/minute ceiling. · Phone destination surcharge, $1.20/month managed number, cold $0.015/min or warm $0.025/min post-transfer, non-basic LLM tokens; BYO carrier separate
TTS uses UTF-8 bytes, ASR audio hours, Agents minutes plus carrier/model/add-ons. Multiple products must not share a fabricated normalized unit.
The evidence overlap.
Both records have documented values for Product type, Technical setup, Pricing model, Published numeric pricing, Free trial, Twilio, Realtime / streaming audio, Bring your model, Function / tool calling, Interruption controls, Published included concurrency at least, Official product documentation and additional fields shown in Compare. Matching availability does not measure behavior under your call conditions.
Run the same booking, tool failure and unsuccessful human-transfer cases. Record both outcomes and billable units before signing a deployment agreement.
Build a buyer-operated pilot →