Text-to-speech pricing compared
Synthetic voices priced by characters, tokens or audio time. Listen for pronunciation and check latency and voice rights.
16 services · sources last checked 2026-10-06 · how we compare
What the most-compared services cost
Estimated yearly bill at 10M characters / month. Enter your own usage below.
- ElevenLabs$9,590/yrFree / Pay as you go · $80 per M chars · Published rates
- Cartesia$4,500/yrScale · $38 per M chars · Published rates
- Google Cloud Text-to-Speech$3,240/yr$30 per M chars · Published rates
- Amazon Polly$1,920/yr$16 per M chars · Published rates
- OpenAI TTS$1,800/yr$15 per M chars · Modeled rates
Enter your usage to estimate each bill. Services differ in what they include; check the rates, assumptions and extra charges below.
The estimates below use the default usage. Choose your current provider to compare costs.
| Provider | Parent jurisdiction | Features / pricing source | Est. monthly | Est. yearly |
|---|---|---|---|---|
| Kokoro-82M on DeepInfraOpen-weight Kokoro-82M TTS served pay-per-use by US inference host DeepInfra, plus other open TTS models such as Chatterbox and Orpheus. | 🇺🇸 United States | Published rates | $6.20$0.62 per M chars | $74 |
| Speechify APISpeechify's developer TTS API (Simba models) with monthly plans that include usage and a published per-character rate for overage. | 🇺🇸 United States | GDPR DPAPublished rates | $91Starter · $10 per M chars | $1,092 |
| Murf APIPay-as-you-go TTS API from Murf with the low-latency Falcon 2 model and regional endpoints in 11 regions, including Germany and the UK. | 🇺🇸 United States | EU residencyPublished rates | $100$10 per M chars | $1,200 |
| Azure AI SpeechMicrosoft's speech service with hundreds of neural and HD voices, EU regions, a free F0 tier and monthly commitment tiers. | 🇺🇸 United States | EU residencyGDPR DPAPublished rates | $150Standard (S0) pay-as-you-go · $15 per M chars | $1,800 |
| Fish Audio APIPay-as-you-go TTS and voice-cloning API from Fish Audio (S2.1 Pro model), billed per UTF-8 byte with no subscription or minimum. | 🇺🇸 United States | GDPR DPAPublished rates | $150$15 per M chars | $1,800 |
| OpenAI TTSOpenAI speech API with the steerable gpt-4o-mini-tts model plus legacy tts-1 and tts-1-hd; pay-as-you-go, no EU region on self-serve. | 🇺🇸 United States | GDPR DPAModeled ratesConverts estimated audio-minute costs to characters at 1,000 English characters per minute. Speed and language change the bill. | $150$15 per M chars | $1,800 |
| Amazon PollyAWS text-to-speech with Standard, Neural, Long-Form and Generative voices; pay-as-you-go per character in EU and other regions. | 🇺🇸 United States | EU residencyGDPR DPAPublished rates | $160$16 per M chars | $1,920 |
| Mistral Voxtral TTSVoxtral TTS from Paris-based Mistral AI, pay-as-you-go per character with zero-shot voice cloning and EU hosting by default. | 🇫🇷 France | EU residencyGDPR DPAPublished rates | $160$16 per M chars | $1,920 |
| Inworld TTSRealtime TTS-2 voice API from US-based Inworld AI, pay-as-you-go or monthly credit plans with lower per-character rates. | 🇺🇸 United States | Published rates | $175Builder · $17.5 per M chars | $2,101 |
| Google Cloud Text-to-SpeechGoogle Cloud speech synthesis with Chirp 3 HD, Neural2, WaveNet and Gemini-TTS voices, a recurring free allowance and an EU endpoint. | 🇺🇸 United States | EU residencyGDPR DPAPublished rates | $270$30 per M chars | $3,240 |
| Deepgram AuraDeepgram's low-latency Aura-2 text-to-speech for voice agents, pay-as-you-go per character with an EU endpoint. | 🇺🇸 United States | EU residencyGDPR DPAPublished rates | $300$30 per M chars | $3,600 |
| CartesiaLow-latency Sonic text-to-speech for voice agents, sold as monthly credit plans with tiered overage rates. | 🇺🇸 United States | GDPR DPAPublished rates | $375Scale · $38 per M chars | $4,500 |
| ElevenLabsPremium voice platform with expressive multilingual models, voice cloning and a large voice library; API billed per character in USD. | 🇺🇸 United States | GDPR DPAPublished rates | $799Free / Pay as you go · $80 per M chars | $9,590 |
| Hume AI (Octave TTS)Expressive Octave TTS from New York-based Hume AI; self-serve plan prices are no longer published, as Hume now focuses on voice AI evaluation. | 🇺🇸 United States | contractNo listed rates | —Contract only · no public price | — |
Standard voices and quality vary a lot. Listen before you switch. Voice cloning and commercial-use rights differ by plan.
Need EU residency? Turn on the EU filter to pick a plan that includes it, which may cost more. “EU plan available” alone does not mean the price shown includes residency. Shared links save your inputs and filters, but prices can change.
Self-hosted options
Operating cost not estimated · servers, licenses and engineering time are yours. We list these separately because there’s no calculated bill to compare with managed services.
- Chatterbox (self-hosted) — Resemble AI's open-source (MIT) Chatterbox TTS models with zero-shot voice cloning, run on your own GPUs. Resemble no longer sells paid TTS.
- Kokoro-82M (self-hosted) — Open-weight (Apache-2.0) 82M-parameter TTS model with 54 voices in 8 languages, small enough to run on CPU or a modest GPU.
No self-hosted options match these filters.
Head-to-head comparisons
Text-to-speech pricing explained
What drives the bill
Text-to-speech is billed per character, or per minute of generated audio. Premium voice platforms sell subscriptions with character credits, and their effective rate can be many times the price of hyperscaler voices or small open models such as Kokoro on a cheap inference host. Voice quality varies more than in most categories here, so price alone doesn’t decide.
Costs the calculator leaves out
- Commercial-use rights, which some plans only include at higher tiers.
- Voice cloning and custom voices, often limited to specific plans.
- Credits that expire, or plan overage rates higher than the advertised per-character price.
- Real-time streaming latency, which matters for voice agents and differs a lot between vendors.
When paying more makes sense
A more expressive voice may be worth paying for in audiobooks, characters or branded assistants. For prompts, notifications and accessibility, pronunciation and clarity may matter more. Generate the same sample text with each vendor and ask people who use the product to compare it.