Kokoro-82M (self-hosted)
Open-weight (Apache-2.0) 82M-parameter TTS model with 54 voices in 8 languages, small enough to run on CPU or a modest GPU.
With only 82M parameters, the model runs on CPU or on a single consumer GPU with about 1-2 GB of VRAM, so a small instance can serve hundreds of millions of characters a month. The model card puts serving costs under $1 per million characters, and DeepInfra's hosted $0.62 is a useful benchmark for whether self-hosting is worth the effort. Its 54 voices in 8 languages suit narration, but there is no voice cloning; unlike Voxtral TTS's non-commercial weights, the licence is Apache-2.0.
- Headquarters
- 🇺🇸 United States
- Parent jurisdiction
- 🇺🇸 United States
- EU self-serve plan
- Available; check eligible plans and regions
- GDPR DPA
- No / not documented
- Billing
- self-hosted
- Free tier
- Free and open weights (Apache-2.0).
- Open source
- Yes · self-hostable
- Website
- huggingface.co
- Source last checked
- 2026-10-06
- Pricing evidence
- Operating cost not estimated
Pricing
Operating cost not estimated · servers, licenses and engineering time are yours. See notes for software editions and license terms.
Self-hosted: budget for hardware, storage and maintenance. The model card cites serving costs of "under $1 per million characters of text input" or "under $0.06 per hour of audio output"; DeepInfra's hosted price of $0.62 per million characters is a good market reference. At 82M parameters it runs on CPU (a browser WASM build reports roughly real-time speed on one laptop thread) and on a single consumer GPU with ~1-2 GB of VRAM, so a small GPU instance can serve hundreds of millions of characters a month. v1.0 (January 2025) has 54 voices in 8 languages; quality is good for narration but below the expressive premium models, and there is no voice cloning. Developed by the independent hexgrad project; data stays wherever you run it.
Managed alternatives
At 10M characters / month. Adjust usage and requirements.
Kokoro-82M has no bill to estimate: you pay for the servers, storage and time to run it. The cheapest managed option here is Kokoro-82M on DeepInfra at $74/yr, so running Kokoro-82M yourself saves money only if hosting and upkeep stay below about $6.20 a month.
- Kokoro-82M on DeepInfra: $74/yr estimatedPublished rates
- Speechify API: $1,092/yr estimatedPublished rates
- Murf API: $1,200/yr estimatedPublished rates
- Azure AI Speech: $1,800/yr estimatedPublished rates
- Fish Audio API: $1,800/yr estimatedPublished rates
No cheaper managed estimate matches your filters. Adjust your usage or compare with your actual contract bill.
The price difference leaves out migration costs, operating costs and any features you’d need to replace.
Self-hosted options
Operating cost is not estimated. These options are not counted as savings against managed services.
Other text-to-speech comparisons
Found a wrong rate or plan restriction? Report a correction with a current source.