Compare Lab
Fish Audio vs OpenAI Realtime API
Fish Audio is positioned as a developer speech API and creative voice workspace; paid TTS usage is metered in UTF-8 bytes. OpenAI Realtime API is positioned as speech-to-speech API for developers implementing their own voice appli
Fish Audio is positioned as a developer speech API and creative voice workspace; paid TTS usage is metered in UTF-8 bytes. OpenAI Realtime API is positioned as speech-to-speech API for developers implementing their own voice application. Start by deciding who owns the missing layers, then inspect the documented differences below.
Voice Engines and Developer Frameworks are different product categories. An underlying engine or framework is not a complete substitute for a deployed agent workflow.
10 jointly documented catalog fields. This measures evidence coverage, not equivalent quality or an overall winner.
Test opening hours, booking conflicts and an unanswered transfer.
Start where the records differ.
Values and evidence states are compared together. Shared facts follow the differences; Unknown never means No.
| Decision point | Fish Audio | OpenAI Realtime API |
|---|---|---|
| Billing unitDifferent records | TTS: million UTF-8 bytes; ASR: audio hours; Agents: conversation minutes | 1 million text/audio/image tokens; transcription separately metered |
| Bring your telephonyDifferent records | Yes | Yes |
| Bring your modelDifferent records | Yes | Unknown |
| Warm transferDifferent records | Yes | Unknown |
| WebhooksDifferent records | YesAPI / webhook | YesUnknown |
| Billing modelShared state | component-stack | component-stack |
Your requirements & next steps
Your requirements
Ready-made / AI receptionist · Strict evidence mode
No personal requirements yet. Add criteria in Advanced Filter to prioritize this comparison.
TTS uses UTF-8 bytes, ASR audio hours, Agents minutes plus carrier/model/add-ons. Multiple products must not share a fabricated normalized unit.
Realtime model tokens plus input transcription, tools and carrier/hosting; not GPT-Live session-minute billing.
Free TTS model and Agents promotional basic LLM scope are separate.
Cached audio/text and image inputs have separate rates. Rate is not per-minute.
Each product keeps its own billing unit.
Model scope fixed to GPT-Realtime-2, not all voice models.
API access; managed numbers have separate rental and enterprise agreement may differ.
Published Realtime token rates do not establish these account/contract terms. GPT-Live duration billing and unrelated fine-tuning discounts must not be imported.
Real sessions consume balance including relevant carrier/model charges; no monthly included minutes.
Published Realtime token rates do not establish these account/contract terms. GPT-Live duration billing and unrelated fine-tuning discounts must not be imported.
Destination, transfer and selected model affect cost; no complete ceiling.
Published Realtime token rates do not establish these account/contract terms. GPT-Live duration billing and unrelated fine-tuning discounts must not be imported.
Rates have explicit event/destination scope, not a single maximum.
Repeated context, cache behavior and turn count affect token charges.
Not yet researched against this buyer-critical field standard.
Not yet researched against this buyer-critical field standard.
Call mix, destination, transfer time, selected model/token usage and separate speech products prevent a universal complete monthly/minute ceiling.
Full bill requires measured model/token/cache/tool/carrier usage; conversation minutes alone are insufficient.
Call mix, destination, transfer time, selected model/token usage and separate speech products prevent a universal complete monthly/minute ceiling.
Full bill requires measured model/token/cache/tool/carrier usage; conversation minutes alone are insufficient.
Anyone with this link can view the encoded decision criteria.
Vendor documentation establishes a claim. Public model benchmarks are separate context, not measurements of these platforms.
Open saved comparison ↗What to verify together.
Fish Audio
UTF-8 bytes and characters are not interchangeable for multilingual scripts.
TTS $15/M UTF-8 bytes; ASR $0.36/audio hour; Agents $0.06/min + applicable surcharges · V3.3 evidence review · TTS API reference; Voice Agents beta has separate minute-based billing
Excluded / confirm: Complete cost cannot be established from this reference. Call mix, destination, transfer time, selected model/token usage and separate speech products prevent a universal complete monthly/minute ceiling. · Phone destination surcharge, $1.20/month managed number, cold $0.015/min or warm $0.025/min post-transfer, non-basic LLM tokens; BYO carrier separate
TTS uses UTF-8 bytes, ASR audio hours, Agents minutes plus carrier/model/add-ons. Multiple products must not share a fabricated normalized unit.
OpenAI Realtime API
A token rate cannot be compared directly with a bundled call-minute price.
GPT-Realtime-2 audio input $32/M, audio output $64/M; text input $4/M, output $24/M · V3.3 evidence review · GPT-Realtime-2 input audio tokens only; output/audio/text/cache have different rates; not 2.1 rate evidence
Excluded / confirm: Complete cost cannot be established from this reference. Full bill requires measured model/token/cache/tool/carrier usage; conversation minutes alone are insufficient. · Input transcription, tool/backend calls, carrier and application infrastructure
Realtime model tokens plus input transcription, tools and carrier/hosting; not GPT-Live session-minute billing.
The evidence overlap.
Both records have documented values for Product type, Technical setup, Pricing model, Published numeric pricing, Realtime / streaming audio, Function / tool calling, Interruption controls, Official product documentation, Checked within 30 days, Retained price snapshot. Matching availability does not measure behavior under your call conditions.
Run the same booking, tool failure and unsuccessful human-transfer cases. Record both outcomes and billable units before signing a deployment agreement.
Build a buyer-operated pilot →