Compare Lab
Fish Audio vs LiveKit
Fish Audio is positioned as a developer speech API and creative voice workspace; paid TTS usage is metered in UTF-8 bytes. LiveKit is positioned as realtime agent infrastructure for teams that own the application and model stack.
Fish Audio is positioned as a developer speech API and creative voice workspace; paid TTS usage is metered in UTF-8 bytes. LiveKit is positioned as realtime agent infrastructure for teams that own the application and model stack. Start by deciding who owns the missing layers, then inspect the documented differences below.
Voice Engines and Developer Frameworks are different product categories. An underlying engine or framework is not a complete substitute for a deployed agent workflow.
16 jointly documented catalog fields. This measures evidence coverage, not equivalent quality or an overall winner.
Test opening hours, booking conflicts and an unanswered transfer.
Start where the records differ.
Values and evidence states are compared together. Shared facts follow the differences; Unknown never means No.
| Decision point | Fish Audio | LiveKit |
|---|---|---|
| Billing unitDifferent records | TTS: million UTF-8 bytes; ASR: audio hours; Agents: conversation minutes | Plan/month + agent minutes + telephony minutes + inference tokens/characters/audio + observability entries |
| Bring your telephonyDifferent records | Yes | Yes |
| HIPAA pathDifferent records | Unknown | Yes |
| WebhooksDifferent records | YesAPI / webhook | YesUnknown |
| Billing modelShared state | component-stack | component-stack |
| Bring your modelShared state | Yes | Yes |
Your requirements & next steps
Your requirements
Ready-made / AI receptionist · Strict evidence mode
No personal requirements yet. Add criteria in Advanced Filter to prioritize this comparison.
TTS uses UTF-8 bytes, ASR audio hours, Agents minutes plus carrier/model/add-ons. Multiple products must not share a fabricated normalized unit.
Plan plus session, telephony, inference, observability and other usage; example calculator minute total is not fixed price.
Free TTS model and Agents promotional basic LLM scope are separate.
Agent session component only; not complete call cost.
Each product keeps its own billing unit.
No universal conversion of model inference to minute rate.
API access; managed numbers have separate rental and enterprise agreement may differ.
Usage overages and model inference additional.
Real sessions consume balance including relevant carrier/model charges; no monthly included minutes.
Separate allowance buckets; no implied complete free call quantity.
Destination, transfer and selected model affect cost; no complete ceiling.
Build allowances are hard caps; paid usage beyond allowances is incremental, not a complete price/minute.
Rates have explicit event/destination scope, not a single maximum.
Usage mix and selected providers required.
Not yet researched against this buyer-critical field standard.
Not yet researched against this buyer-critical field standard.
Call mix, destination, transfer time, selected model/token usage and separate speech products prevent a universal complete monthly/minute ceiling.
The calculator example $0.0479/min is assumption-specific; no maximum complete bill from minutes alone.
Call mix, destination, transfer time, selected model/token usage and separate speech products prevent a universal complete monthly/minute ceiling.
The calculator example $0.0479/min is assumption-specific; no maximum complete bill from minutes alone.
Anyone with this link can view the encoded decision criteria.
Vendor documentation establishes a claim. Public model benchmarks are separate context, not measurements of these platforms.
Open saved comparison ↗What to verify together.
Fish Audio
UTF-8 bytes and characters are not interchangeable for multilingual scripts.
TTS $15/M UTF-8 bytes; ASR $0.36/audio hour; Agents $0.06/min + applicable surcharges · V3.3 evidence review · TTS API reference; Voice Agents beta has separate minute-based billing
Excluded / confirm: Complete cost cannot be established from this reference. Call mix, destination, transfer time, selected model/token usage and separate speech products prevent a universal complete monthly/minute ceiling. · Phone destination surcharge, $1.20/month managed number, cold $0.015/min or warm $0.025/min post-transfer, non-basic LLM tokens; BYO carrier separate
TTS uses UTF-8 bytes, ASR audio hours, Agents minutes plus carrier/model/add-ons. Multiple products must not share a fabricated normalized unit.
LiveKit
Agent-session pricing excludes a complete inference and telephony stack.
Paid Cloud agent sessions $0.01/min beyond allowance · V3.3 evidence review · Cloud Ship starting subscription; session overage, inference, transport and artifacts extra
Excluded / confirm: Complete cost cannot be established from this reference. The calculator example $0.0479/min is assumption-specific; no maximum complete bill from minutes alone. · Inference, phone number/transport, recording and observability, outbound carrier, bandwidth and optional private links
Plan plus session, telephony, inference, observability and other usage; example calculator minute total is not fixed price.
The evidence overlap.
Both records have documented values for Product type, Technical setup, Pricing model, Published numeric pricing, Free tier, Twilio, Call recording, Realtime / streaming audio, Bring your model, Function / tool calling, Interruption controls, Published included concurrency at least and additional fields shown in Compare. Matching availability does not measure behavior under your call conditions.
Run the same booking, tool failure and unsuccessful human-transfer cases. Record both outcomes and billable units before signing a deployment agreement.
Build a buyer-operated pilot →