Speech-to-text pricing compared
Transcription prices for recorded audio. Streaming, speaker labels and specialist vocabulary may cost extra.
14 services · sources last checked 2026-10-06 to 2026-10-08 · how we compare
What the most-compared services cost
Estimated yearly bill at 1,000 audio hours / month. Enter your own usage below.
Enter your usage to estimate each bill. Services differ in what they include; check the rates, assumptions and extra charges below.
The estimates below use the default usage. Choose your current provider to compare costs.
| Provider | Parent jurisdiction | Features / pricing source | Est. monthly | Est. yearly |
|---|---|---|---|---|
| DeepInfraPay-per-minute hosting of open speech models (Whisper large-v3, Qwen3-ASR, Voxtral, Nemotron ASR) at very low per-hour rates. | 🇺🇸 United States | Published rates | $27$0.027/hour | $324 |
| GroqGroqCloud API serving OpenAI's open Whisper large-v3 and large-v3-turbo models on Groq LPUs, billed per audio hour. | 🇺🇸 United States | GDPR DPAPublished rates | $111$0.111/hour | $1,332 |
| Azure AI SpeechMicrosoft Azure speech-to-text with cheap batch transcription, pricier real-time and fast modes, and commitment tiers; EU regions available. | 🇺🇸 United States | EU residencyGDPR DPAPublished rates | $180$0.18/hour | $2,160 |
| Mistral AI (Voxtral)Mistral's Voxtral Transcribe 2 speech-to-text API with diarization included, from a French company hosting in the EU. | 🇫🇷 France | EU residencyGDPR DPAPublished rates | $180$0.18/hour | $2,160 |
| Rev.aiRev's speech-to-text API with the Reverb ASR model plus Whisper and human transcription; 15-second minimum per file. | 🇺🇸 United States | GDPR DPAPublished rates | $200$0.2/hour | $2,400 |
| AssemblyAIUS speech AI API with Universal-3.5 Pro pre-recorded and Universal streaming models, pay-as-you-go; EU endpoint in Dublin. | 🇺🇸 United States | EU residencyGDPR DPAPublished rates | $210$0.21/hour | $2,520 |
| ElevenLabs ScribeElevenLabs' Scribe v2 speech-to-text API, billed per hour pay-as-you-go or from hours included in monthly subscription plans. | 🇺🇸 United States | GDPR DPAPublished rates | $219Free / Pay as you go · $0.22/hour | $2,628 |
| SpeechmaticsUK speech recognition vendor with Melia 1 multilingual, Enhanced and Standard models; choice of EU, US or Australia processing. | 🇬🇧 United Kingdom | EU residencyGDPR DPAPublished rates | $240$0.24/hour | $2,880 |
| DeepgramUS speech AI vendor with Nova-3 pre-recorded and Flux/Nova-3 streaming STT, pay-as-you-go with no minimums; EU endpoint available. | 🇺🇸 United States | EU residencyGDPR DPAPublished rates | $258$0.258/hour | $3,096 |
| OpenAIOpenAI's transcription API (gpt-transcribe, gpt-4o-transcribe, whisper-1), billed per audio minute or token. US company; EU residency is sales-gated. | 🇺🇸 United States | GDPR DPAEstimated ratesToken-billed transcription uses the vendor's estimated per-minute cost, not a fixed hourly rate. | $270$0.27/hour | $3,240 |
| Amazon TranscribeAWS speech-to-text with batch and streaming APIs, per-second billing, diarization included; runs in several EU regions. | 🇺🇸 United States | EU residencyGDPR DPAPublished rates | $360$0.36/hour | $4,320 |
| GladiaFrench speech-to-text API (Solaria models) with async and real-time transcription, diarization included, and EU hosting by default. | 🇫🇷 France | EU residencyGDPR DPAPublished rates | $610$0.61/hour | $7,320 |
| Google Cloud Speech-to-TextGoogle Cloud's Speech-to-Text V2 API with Chirp models, volume tiers and a cheaper dynamic batch mode; EU regions available. | 🇺🇸 United States | EU residencyGDPR DPAPublished rates | $960$0.96/hour | $11,520 |
Pre-recorded audio, standard model, English. Real-time streaming, diarization and other add-ons often cost extra. Accuracy varies by language and audio quality.
Need EU residency? Turn on the EU filter to pick a plan that includes it, which may cost more. “EU plan available” alone does not mean the price shown includes residency. Shared links save your inputs and filters, but prices can change.
Looking for a European-owned provider? See European speech-to-text providers.
Self-hosted options
Operating cost not estimated · servers, licenses and engineering time are yours. We list these separately because there’s no calculated bill to compare with managed services.
- Whisper (self-hosted) — OpenAI's open Whisper models run on your own GPUs via faster-whisper or whisper.cpp; no per-minute fees, you pay for compute.
No self-hosted options match these filters.
Head-to-head comparisons
Speech-to-text pricing explained
What drives the bill
Transcription is billed per minute or hour of audio. The model, hosting provider and choice of batch or streaming affect the rate. Hosted open models can have much lower API prices, but compare accuracy on your recordings, especially names, accents and noisy audio.
Costs the calculator leaves out
- Real-time streaming, which usually costs more than transcribing pre-recorded files.
- Speaker diarization, word timestamps, PII redaction and summaries, often billed as add-ons.
- Billing increments. Some vendors round each file up to the nearest 15 seconds or minute.
- Language coverage. Accuracy in non-English languages varies more between vendors than English does.
When paying more makes sense
Specialist vendors may be worth it for low-latency streaming, difficult audio, medical or legal vocabulary, and contractual data handling. For batch podcasts, meetings or archives, test a hosted open model. Include corrections, retries and any separate diarization service when comparing the cost.