Fireworks AI
Serverless and dedicated inference platform for open-weight models with Standard, Priority and Fast serving paths.
Teams that want to trade cost against latency get that choice here: listed prices are for the Standard serving path, and the Priority and Fast paths cost more. On shared models Standard pricing matches Together AI and Nebius for GLM 5.3 ($1.40 / $4.40) and gpt-oss-120b ($0.15 / $0.60), while Kimi K3 at $3.00 / $15.00 sits at the top of the range. Compliance paperwork is thin, with SCCs and an EU GDPR representative referenced but no public DPA or EU region found.
- Headquarters
- 馃嚭馃嚫 United States
- Parent jurisdiction
- 馃嚭馃嚫 United States
- EU self-serve plan
- No / not documented
- GDPR DPA
- No / not documented
- Billing
- payg
- Open source
- No
- Website
- fireworks.ai
- Source last checked
- 2026-10-06
- Pricing evidence
- Published rates
Pricing
Calculated from the listed vendor rates; exclusions still apply. How to read these estimates.
| Model | Input / 1M tokens | Output / 1M tokens | Cached / 1M tokens | Context | Details |
|---|---|---|---|---|---|
| Kimi K3 | $3 | $15 | $0.3 | 1048k | frontieropen weights |
| GLM 5.3 | $1.4 | $4.4 | $0.26 | 1048k | frontieropen weights |
| DeepSeek V4.1 Flash | $0.3 | $1.2 | $0.006 | 1000k | midopen weights |
| gpt-oss-120b | $0.15 | $0.6 | $0.015 | 131k | midopen weights |
| GLM 5.3 Flash | $0.15 | $0.5 | $0.03 | 1048k | smallopen weights |
Standard serving-path serverless prices; Priority and Fast paths cost more. Context lengths are not on the pricing page and were taken from the same models' published specs (assumed full context). Privacy policy references SCCs and an EU GDPR representative, but no public DPA or EU region was found.
Managed alternatives
At 500M input tokens / month, 50M output tokens / month. Adjust usage and requirements.
Fireworks AI 路 GLM 5.3 Flash: $1,200/yr estimated.
- Groq 路 gpt-oss-20b: $630/yr estimated 路 $570/yr bill differencePublished rates
- OpenAI 路 GPT-6 Luna: $900/yr estimated 路 $300/yr bill differencePublished rates
- Mistral AI 路 Ministral 3 8B: $990/yr estimated 路 $210/yr bill differencePublished rates
- Alibaba Cloud Model Studio (Qwen) 路 Qwen3.8-Flash: $1,182/yr estimated 路 $18/yr bill differencePublished rates
No cheaper managed estimate matches your filters. Adjust your usage or compare with your actual contract bill.
The price difference leaves out migration costs, operating costs and any features you鈥檇 need to replace. We compare the selected model or the cheapest listed model in the chosen capability tier. Sharing a tier does not mean equal quality.
Self-hosted options
Operating cost is not estimated. These options are not counted as savings against managed services.
Want a European-owned provider? See European alternatives to Fireworks AI among european LLM APIs.
Found a wrong rate or plan restriction? Report a correction with a current source.