Groq
GroqCloud low-latency inference API running open-weight models on Groq's own LPU chips.
Groq focuses on fast inference, with a small public catalog: gpt-oss-120b, gpt-oss-20b and a Qwen3.8-27B preview, all with a 131K context. Llama pricing now goes through enterprise sales. The gpt-oss-120b rates match several other hosts, so compare latency on your own requests. A Helsinki data center opened in 2025, but we found no documented self-serve EU residency option.
- Headquarters
- 🇺🇸 United States
- Parent jurisdiction
- 🇺🇸 United States
- EU self-serve plan
- No / not documented
- GDPR DPA
- Yes
- Billing
- payg
- Open source
- No
- Website
- groq.com
- Source last checked
- 2026-10-06
- Pricing evidence
- Published rates
Pricing
Calculated from the listed vendor rates; exclusions still apply. How to read these estimates.
| Model | Input / 1M tokens | Output / 1M tokens | Cached / 1M tokens | Context | Details |
|---|---|---|---|---|---|
| gpt-oss-120b | $0.15 | $0.6 | — | 131k | midopen weights |
| gpt-oss-20b | $0.075 | $0.3 | — | 131k | smallopen weights |
| Qwen3.8-27B (preview) | $0.8 | $4 | — | 131k | midopen weights |
Groq Inc. remains independent: in Dec 2025 Nvidia took a non-exclusive ~$20B license to Groq's LPU technology and hired founder Jonathan Ross and other staff, but did not acquire the company (new CEO Simon Edwards; raised $650M in June 2026). The public catalog is small: Llama 3.1 8B / 3.3 70B now show "enterprise pricing" only, and Qwen3.8-27B is a preview model. A Helsinki (Equinix) data center opened in 2025, but no self-serve EU data residency option is documented. Customer DPA published at console.groq.com/docs/legal.
Managed alternatives
At 500M input tokens / month, 50M output tokens / month. Adjust usage and requirements.
Groq · gpt-oss-20b: $630/yr estimated.
No cheaper managed estimate matches your filters. Adjust your usage or compare with your actual contract bill.
The price difference leaves out migration costs, operating costs and any features you’d need to replace. We compare the selected model or the cheapest listed model in the chosen capability tier. Sharing a tier does not mean equal quality.
Self-hosted options
Operating cost is not estimated. These options are not counted as savings against managed services.
Want a European-owned provider? See European alternatives to Groq among european LLM APIs.
Found a wrong rate or plan restriction? Report a correction with a current source.