LLM APIs pricing compared
Input and output token rates across model providers. A small model may be enough for your task; test it on real inputs.
14 services · sources last checked 2026-10-06 to 2026-10-09 · how we compare
Enter your usage to estimate each bill. Services differ in what they include; check the rates, assumptions and extra charges below.
The estimates below use the default usage. Choose your current provider to compare costs.
| Provider / model | Parent jurisdiction | Features / pricing source | Est. monthly | Est. yearly |
|---|---|---|---|---|
| DeepInfragpt-oss-120b | 🇺🇸 United States | open weightsmidPublished rates | $27gpt-oss-120b · $0.037 in / $0.17 out per M | $324 |
| DeepInfraDeepSeek-V4-Flash-0731 | 🇺🇸 United States | open weightsmidPublished rates | $39DeepSeek-V4-Flash-0731 · $0.06 in / $0.18 out per M | $468 |
| Groqgpt-oss-20b | 🇺🇸 United States | GDPR DPAopen weightssmallPublished rates | $53gpt-oss-20b · $0.075 in / $0.3 out per M | $630 |
| OpenAIGPT-6 Luna | 🇺🇸 United States | EU residencyGDPR DPAsmallPublished rates | $75GPT-6 Luna · $0.1 in / $0.5 out per M | $900 |
| Mistral AIMinistral 3 8B | 🇫🇷 France | EU residencyGDPR DPAopen weightssmallPublished rates | $83Ministral 3 8B · $0.15 in / $0.15 out per M | $990 |
| Nebius Token FactoryDeepSeek-V4-Flash-0731 | 🇳🇱 Netherlands | EU residencyGDPR DPAopen weightsmidPublished rates | $84DeepSeek-V4-Flash-0731 · $0.14 in / $0.28 out per M | $1,008 |
| Alibaba Cloud Model Studio (Qwen)Qwen3.8-Flash | 🇨🇳 China | EU residencyGDPR DPAsmallPublished rates | $99Qwen3.8-Flash · $0.15 in / $0.47 out per M | $1,182 |
| Fireworks AIGLM 5.3 Flash | 🇺🇸 United States | open weightssmallPublished rates | $100GLM 5.3 Flash · $0.15 in / $0.5 out per M | $1,200 |
| Z.ai (Zhipu GLM)GLM-5.3-Flash | 🇨🇳 China | open weightssmallPublished rates | $100GLM-5.3-Flash · $0.15 in / $0.5 out per M | $1,200 |
| Fireworks AIgpt-oss-120b | 🇺🇸 United States | open weightsmidPublished rates | $105gpt-oss-120b · $0.15 in / $0.6 out per M | $1,260 |
| Groqgpt-oss-120b | 🇺🇸 United States | GDPR DPAopen weightsmidPublished rates | $105gpt-oss-120b · $0.15 in / $0.6 out per M | $1,260 |
| Mistral AIMistral Small 4 | 🇫🇷 France | EU residencyGDPR DPAopen weightssmallPublished rates | $105Mistral Small 4 · $0.15 in / $0.6 out per M | $1,260 |
| Nebius Token Factorygpt-oss-120b | 🇳🇱 Netherlands | EU residencyGDPR DPAopen weightsmidPublished rates | $105gpt-oss-120b · $0.15 in / $0.6 out per M | $1,260 |
| Together AIgpt-oss-120b | 🇺🇸 United States | EU residencyopen weightsmidPublished rates | $105gpt-oss-120b · $0.15 in / $0.6 out per M | $1,260 |
| Scaleway Generative APIsMistral Small 3.2 24B | 🇫🇷 France | EU residencyGDPR DPAopen weightssmallPublished rates | $111Mistral Small 3.2 24B · $0.18 in / $0.41 out per M | $1,326 |
| Scaleway Generative APIsgpt-oss-120b | 🇫🇷 France | EU residencyGDPR DPAopen weightsmidPublished rates | $125gpt-oss-120b · $0.18 in / $0.7 out per M | $1,500 |
| Google Gemini APIGemini 3.1 Flash-Lite | 🇺🇸 United States | GDPR DPAsmallPublished rates | $200Gemini 3.1 Flash-Lite · $0.25 in / $1.5 out per M | $2,400 |
| DeepSeekDeepSeek V4 Flash | 🇨🇳 China | open weightsmidPublished rates | $210DeepSeek V4 Flash · $0.3 in / $1.2 out per M | $2,520 |
| Fireworks AIDeepSeek V4.1 Flash | 🇺🇸 United States | open weightsmidPublished rates | $210DeepSeek V4.1 Flash · $0.3 in / $1.2 out per M | $2,520 |
| Together AIDeepSeek V4.1 Flash | 🇺🇸 United States | EU residencyopen weightsmidPublished rates | $210DeepSeek V4.1 Flash · $0.3 in / $1.2 out per M | $2,520 |
| Z.ai (Zhipu GLM)GLM-5.3-FlashX | 🇨🇳 China | midPublished rates | $248GLM-5.3-FlashX · $0.37 in / $1.25 out per M | $2,970 |
| Google Gemini APIGemini 3.5 Flash-Lite | 🇺🇸 United States | GDPR DPAsmallPublished rates | $275Gemini 3.5 Flash-Lite · $0.3 in / $2.5 out per M | $3,300 |
| Alibaba Cloud Model Studio (Qwen)Qwen3.7-Plus | 🇨🇳 China | EU residencyGDPR DPAmidPublished rates | $280Qwen3.7-Plus · $0.4 in / $1.6 out per M | $3,360 |
| Scaleway Generative APIsDeepSeek-V4-Flash-0731 | 🇫🇷 France | EU residencyGDPR DPAopen weightsmidPublished rates | $282DeepSeek-V4-Flash-0731 · $0.47 in / $0.94 out per M | $3,384 |
| Mistral AIMistral Large 3 | 🇫🇷 France | EU residencyGDPR DPAopen weightsmidPublished rates | $325Mistral Large 3 · $0.5 in / $1.5 out per M | $3,900 |
| DeepInfraGLM-5.3 | 🇺🇸 United States | open weightsfrontierPublished rates | $407GLM-5.3 · $0.563 in / $2.5 out per M | $4,878 |
| Google Gemini APIGemini 3.8 Flash | 🇺🇸 United States | GDPR DPAmidPublished rates | $563Gemini 3.8 Flash · $0.75 in / $3.75 out per M | $6,750 |
| Scaleway Generative APIsQwen3-235B-A22B-Instruct-2507 | 🇫🇷 France | EU residencyGDPR DPAopen weightsmidPublished rates | $572Qwen3-235B-A22B-Instruct-2507 · $0.88 in / $2.63 out per M | $6,858 |
| GroqQwen3.8-27B (preview) | 🇺🇸 United States | GDPR DPAopen weightsmidPublished rates | $600Qwen3.8-27B (preview) · $0.8 in / $4 out per M | $7,200 |
| OpenAIGPT-5.4 mini | 🇺🇸 United States | EU residencyGDPR DPAsmallPublished rates | $600GPT-5.4 mini · $0.75 in / $4.5 out per M | $7,200 |
| Moonshot AI (Kimi)Kimi K2.6 | 🇨🇳 China | open weightsmidPublished rates | $675Kimi K2.6 · $0.95 in / $4 out per M | $8,100 |
| Alibaba Cloud Model Studio (Qwen)Qwen3-Coder-Plus | 🇨🇳 China | EU residencyGDPR DPAmidPublished rates | $750Qwen3-Coder-Plus · $1 in / $5 out per M | $9,000 |
| AnthropicClaude Haiku 4.5 | 🇺🇸 United States | GDPR DPAsmallPublished rates | $750Claude Haiku 4.5 · $1 in / $5 out per M | $9,000 |
| DeepInfraDeepSeek-V4-Pro | 🇺🇸 United States | open weightsfrontierPublished rates | $780DeepSeek-V4-Pro · $1.3 in / $2.6 out per M | $9,360 |
| DeepSeekDeepSeek V4 Pro | 🇨🇳 China | open weightsfrontierPublished rates | $858DeepSeek V4 Pro · $1.32 in / $3.96 out per M | $10,296 |
| Together AIDeepSeek V4 Pro 0813 | 🇺🇸 United States | EU residencyopen weightsfrontierPublished rates | $858DeepSeek V4 Pro 0813 · $1.32 in / $3.96 out per M | $10,296 |
| Fireworks AIGLM 5.3 | 🇺🇸 United States | open weightsfrontierPublished rates | $920GLM 5.3 · $1.4 in / $4.4 out per M | $11,040 |
| Nebius Token FactoryGLM-5.3 | 🇳🇱 Netherlands | EU residencyGDPR DPAopen weightsfrontierPublished rates | $920GLM-5.3 · $1.4 in / $4.4 out per M | $11,040 |
| Together AIGLM-5.3 | 🇺🇸 United States | EU residencyopen weightsfrontierPublished rates | $920GLM-5.3 · $1.4 in / $4.4 out per M | $11,040 |
| Z.ai (Zhipu GLM)GLM-5.3 | 🇨🇳 China | open weightsfrontierPublished rates | $920GLM-5.3 · $1.4 in / $4.4 out per M | $11,040 |
| Nebius Token FactoryDeepSeek-V4-Pro | 🇳🇱 Netherlands | EU residencyGDPR DPAopen weightsfrontierPublished rates | $1,050DeepSeek-V4-Pro · $1.75 in / $3.5 out per M | $12,600 |
| Mistral AIMistral Medium 3.5 | 🇫🇷 France | EU residencyGDPR DPAopen weightsmidPublished rates | $1,125Mistral Medium 3.5 · $1.5 in / $7.5 out per M | $13,500 |
| Alibaba Cloud Model Studio (Qwen)Qwen3.8-Max | 🇨🇳 China | EU residencyGDPR DPAfrontierPublished rates | $1,300Qwen3.8-Max · $2 in / $6 out per M | $15,600 |
| Scaleway Generative APIsGLM-5.2 | 🇫🇷 France | EU residencyGDPR DPAopen weightsfrontierPublished rates | $1,377GLM-5.2 · $2.11 in / $6.44 out per M | $16,524 |
| AnthropicClaude Sonnet 5.5 | 🇺🇸 United States | GDPR DPAmidPublished rates | $1,500Claude Sonnet 5.5 · $2 in / $10 out per M | $18,000 |
| OpenAIGPT-6.1 Sol | 🇺🇸 United States | EU residencyGDPR DPAfrontierPublished rates | $1,500GPT-6.1 Sol · $2 in / $10 out per M | $18,000 |
| Google Gemini APIGemini 3.1 Pro (preview) | 🇺🇸 United States | GDPR DPAfrontierPublished rates | $1,600Gemini 3.1 Pro (preview) · $2 in / $12 out per M | $19,200 |
| Together AIKimi K3 | 🇺🇸 United States | EU residencyopen weightsfrontierPublished rates | $2,025Kimi K3 · $2.7 in / $13.5 out per M | $24,300 |
| DeepInfraKimi-K3 | 🇺🇸 United States | open weightsfrontierPublished rates | $2,138Kimi-K3 · $2.85 in / $14.25 out per M | $25,650 |
| Fireworks AIKimi K3 | 🇺🇸 United States | open weightsfrontierPublished rates | $2,250Kimi K3 · $3 in / $15 out per M | $27,000 |
| Moonshot AI (Kimi)Kimi K3 | 🇨🇳 China | open weightsfrontierPublished rates | $2,250Kimi K3 · $3 in / $15 out per M | $27,000 |
| Nebius Token FactoryKimi-K3 | 🇳🇱 Netherlands | EU residencyGDPR DPAopen weightsfrontierPublished rates | $2,250Kimi-K3 · $3 in / $15 out per M | $27,000 |
| AnthropicClaude Opus 5.5 | 🇺🇸 United States | GDPR DPAfrontierPublished rates | $3,000Claude Opus 5.5 · $4 in / $20 out per M | $36,000 |
| AnthropicClaude Fable 5.1 | 🇺🇸 United States | GDPR DPAfrontierPublished rates | $7,500Claude Fable 5.1 · $10 in / $50 out per M | $90,000 |
| OpenAIGPT-6 Astra | 🇺🇸 United States | EU residencyGDPR DPAfrontierPublished rates | $7,500GPT-6 Astra · $10 in / $50 out per M | $90,000 |
Price per token isn’t price per task. Cheaper models can need more retries or longer outputs. Benchmark on your own workload.
Need EU residency? Turn on the EU filter to pick a plan that includes it, which may cost more. “EU plan available” alone does not mean the price shown includes residency. Shared links save your inputs and filters, but prices can change.
Looking for a European-owned provider? See European LLM APIs.
LLM APIs pricing explained
What drives the bill
You pay separately for input and output tokens, with output usually costing 3–5× more. Listed prices vary by more than 100× across models. Smaller models may be enough for classification, extraction and summaries. Test them on your actual inputs rather than assuming a higher price means better results for every task.
Costs the calculator leaves out
- Reasoning tokens. Thinking models bill their hidden reasoning as output, which can multiply the cost of a task.
- Cached input discounts. Repeated system prompts and documents can be billed at a fraction of the normal input price.
- Batch APIs, often about 50% off for work that can wait hours.
- Long-context surcharges on some models, and per-call fees for tools such as web search.
When paying more makes sense
Use the strongest model where errors are expensive: complex coding, multi-step agents, difficult reasoning. Elsewhere, route by task. Price per token isn’t price per task, so build a small evaluation set from your real inputs and compare models on accuracy and total cost before switching. An LLM gateway can route traffic between models.