Together AI
Serverless and dedicated inference for open-weight models (DeepSeek, Kimi, GLM, Qwen, Llama, gpt-oss) via an OpenAI-compatible API.
Serverless users after the lowest Kimi K3 price will find it here at $2.70 / $13.50, below Moonshot's own $3.00 / $15.00, while GLM-5.3 and gpt-oss-120b match Fireworks AI and Nebius. The EU residency flag needs reading carefully: EU data centres are documented only for dedicated endpoints on Scale and Enterprise plans, and serverless traffic is not guaranteed to stay in the EU. No public self-serve DPA was found.
- Headquarters
- 🇺🇸 United States
- Parent jurisdiction
- 🇺🇸 United States
- EU self-serve plan
- Available; check eligible plans and regions
- GDPR DPA
- No / not documented
- Billing
- payg
- Open source
- No
- Website
- www.together.ai
- Source last checked
- 2026-10-06
- Pricing evidence
- Published rates
Pricing
Calculated from the listed vendor rates; exclusions still apply. How to read these estimates.
| Model | Input / 1M tokens | Output / 1M tokens | Cached / 1M tokens | Context | Details |
|---|---|---|---|---|---|
| Kimi K3 | $2.7 | $13.5 | $0.27 | 1048k | frontieropen weights |
| GLM-5.3 | $1.4 | $4.4 | $0.26 | 1048k | frontieropen weights |
| DeepSeek V4 Pro 0813 | $1.32 | $3.96 | $0.13 | 1048k | frontieropen weights |
| DeepSeek V4.1 Flash | $0.3 | $1.2 | $0.006 | 1000k | midopen weights |
| gpt-oss-120b | $0.15 | $0.6 | — | 131k | midopen weights |
Serverless per-token prices shown. EU data centers are documented only for dedicated endpoints on Scale/Enterprise plans; serverless traffic is not guaranteed to stay in the EU. Privacy policy cites Standard Contractual Clauses but no public self-serve DPA was found. Context lengths from docs.together.ai serverless model list.
Managed alternatives
At 500M input tokens / month, 50M output tokens / month. Adjust usage and requirements.
Together AI · gpt-oss-120b: $1,260/yr estimated.
- DeepInfra · gpt-oss-120b: $324/yr estimated · $936/yr bill differencePublished rates
- Nebius Token Factory · DeepSeek-V4-Flash-0731: $1,008/yr estimated · $252/yr bill differencePublished rates
No cheaper managed estimate matches your filters. Adjust your usage or compare with your actual contract bill.
The price difference leaves out migration costs, operating costs and any features you’d need to replace. We compare the selected model or the cheapest listed model in the chosen capability tier. Sharing a tier does not mean equal quality.
Self-hosted options
Operating cost is not estimated. These options are not counted as savings against managed services.
Want a European-owned provider? See European alternatives to Together AI among european LLM APIs.
Found a wrong rate or plan restriction? Report a correction with a current source.