Whisper (self-hosted)
OpenAI's open Whisper models run on your own GPUs via faster-whisper or whisper.cpp; no per-minute fees, you pay for compute.
Running Whisper yourself removes per-minute fees, which pays off once volume is high enough to keep GPUs busy and someone can own the operations. Based on faster-whisper's benchmark of about 45x real-time and an assumed $0.50-1.00 per GPU-hour, that works out to roughly $0.01-0.03 per audio hour at full utilisation, before idle capacity, storage and ops time. Diarization is not built in, and data stays wherever you run it.
- Headquarters
- 🇫🇷 France
- Parent jurisdiction
- 🇫🇷 France
- EU self-serve plan
- Available; check eligible plans and regions
- GDPR DPA
- No / not documented
- Billing
- self-hosted
- Free tier
- Free and open source (Whisper weights and faster-whisper under MIT; whisper.cpp MIT).
- Open source
- Yes · self-hostable
- Website
- github.com
- Source last checked
- 2026-10-06
- Pricing evidence
- Operating cost not estimated
Pricing
Operating cost not estimated · servers, licenses and engineering time are yours. See notes for software editions and license terms.
Self-hosted: budget for hardware, storage and maintenance. faster-whisper's README benchmark transcribes 13 minutes of audio with large-v2 in 17 seconds (fp16, batch size 8) on a consumer RTX 3070 Ti, about 45x real-time, i.e. roughly 0.02 GPU-hours per audio hour (large-v3 is similar in size; turbo is faster). At an assumed $0.50-1.00 per GPU-hour for a cloud L4/A10-class GPU that is about $0.01-0.03 per audio hour at full utilisation, before idle capacity, storage and ops time. whisper.cpp runs on CPUs and Apple Silicon as well. No diarization built in (pair with e.g. pyannote). hq/jurisdiction reflect faster-whisper's maintainer SYSTRAN (France); the Whisper weights are from OpenAI (US). Data stays wherever you run it.
Managed alternatives
At 1,000 audio hours / month. Adjust usage and requirements.
Whisper has no bill to estimate: you pay for the servers, storage and time to run it. The cheapest managed option here is DeepInfra at $324/yr, so running Whisper yourself saves money only if hosting and upkeep stay below about $27 a month.
- DeepInfra: $324/yr estimatedPublished rates
- Groq: $1,332/yr estimatedPublished rates
- Azure AI Speech: $2,160/yr estimatedPublished rates
- Mistral AI (Voxtral): $2,160/yr estimatedPublished rates
- Rev.ai: $2,400/yr estimatedPublished rates
No cheaper managed estimate matches your filters. Adjust your usage or compare with your actual contract bill.
The price difference leaves out migration costs, operating costs and any features you’d need to replace.
Self-hosted options
Operating cost is not estimated. These options are not counted as savings against managed services.
Other speech-to-text comparisons
Found a wrong rate or plan restriction? Report a correction with a current source.