SV Premium

LLM APIs

DeepInfra

Pay-per-token inference API for open-weight models including DeepSeek, Kimi, GLM, Qwen, Llama and gpt-oss.

Among the hosts reselling open-weight models, DeepInfra lists the lowest per-token price on several shared models: DeepSeek-V4-Flash-0731 is $0.06 / $0.18 against $0.14 / $0.28 on Nebius, and gpt-oss-120b is $0.037 / $0.17. A cheaper Flex tier exists for some models. That suits cost-sensitive workloads without compliance requirements, because although inputs are stated not to be stored to disk, no DPA, GDPR terms or EU region are documented.

Headquarters
🇺🇸 United States
Parent jurisdiction
🇺🇸 United States
EU self-serve plan
No / not documented
GDPR DPA
No / not documented
Billing
payg
Open source
No
Website
deepinfra.com
Source last checked
2026-10-06
Pricing evidence
Published rates

Pricing

Calculated from the listed vendor rates; exclusions still apply. How to read these estimates.

ModelInput / 1M tokensOutput / 1M tokensCached / 1M tokensContextDetails
Kimi-K3$2.85$14.25$0.2851024kfrontieropen weights
DeepSeek-V4-Pro$1.3$2.6$0.11024kfrontieropen weights
GLM-5.3$0.563$2.5$0.1251048kfrontieropen weights
DeepSeek-V4-Flash-0731$0.06$0.18$0.0151024kmidopen weights
gpt-oss-120b$0.037$0.17—131kmidopen weights

Standard-tier prices; a cheaper Flex tier exists for some models (e.g. GLM-5.3 $0.45/$2.00). Data-privacy docs state inputs are not stored to disk, but no DPA, GDPR terms or EU region are documented.

Managed alternatives

At 500M input tokens / month, 50M output tokens / month. Adjust usage and requirements.

DeepInfra · gpt-oss-120b: $324/yr estimated.

No cheaper managed estimate matches your filters. Adjust your usage or compare with your actual contract bill.

The price difference leaves out migration costs, operating costs and any features you’d need to replace. We compare the selected model or the cheapest listed model in the chosen capability tier. Sharing a tier does not mean equal quality.

Want a European-owned provider? See European alternatives to DeepInfra among european LLM APIs.

Found a wrong rate or plan restriction? Report a correction with a current source.