Z.ai (Zhipu GLM)
Z.ai (Zhipu AI) international API serving the GLM-5.x model family, including open-weight GLM-5.3 and cheaper Flash variants.
Two models, GLM-4.7-Flash and GLM-4.5-Flash, are free via the API, which makes Z.ai a cheap place to experiment. GLM-5.3 costs $1.40 / $4.40, the same as at Fireworks AI, Together AI and Nebius, so going direct buys no discount on the flagship. The cheaper GLM-5.3-FlashX is only sold here, while the MIT-licensed GLM-5.3-Flash is also listed at Fireworks AI at the same $0.15 / $0.50. Data falls under Chinese jurisdiction with no EU residency or GDPR DPA.
- Headquarters
- 馃嚚馃嚦 China
- Parent jurisdiction
- 馃嚚馃嚦 China
- EU self-serve plan
- No / not documented
- GDPR DPA
- No / not documented
- Billing
- payg
- Free tier
- GLM-4.7-Flash and GLM-4.5-Flash are free to use via the API.
- Open source
- No
- Website
- z.ai
- Source last checked
- 2026-10-09
- Pricing evidence
- Published rates
Pricing
Calculated from the listed vendor rates; exclusions still apply. How to read these estimates.
| Model | Input / 1M tokens | Output / 1M tokens | Cached / 1M tokens | Context | Details |
|---|---|---|---|---|---|
| GLM-5.3 | $1.4 | $4.4 | $0.26 | 1000k | frontieropen weights |
| GLM-5.3-FlashX | $0.37 | $1.25 | $0.075 | 1000k | mid |
| GLM-5.3-Flash | $0.15 | $0.5 | $0.03 | 1000k | smallopen weights |
GLM-5.3 weights (744B MoE) were released on Hugging Face in late August 2026. Cached-input storage is listed as "limited-time free". GLM-5.3-Flash weights are published under the MIT licence (huggingface.co/zai-org/GLM-5.3-Flash, Aug 2026).
Managed alternatives
At 500M input tokens / month, 50M output tokens / month. Adjust usage and requirements.
Z.ai (Zhipu GLM) 路 GLM-5.3-Flash: $1,200/yr estimated.
- Groq 路 gpt-oss-20b: $630/yr estimated 路 $570/yr bill differencePublished rates
- OpenAI 路 GPT-6 Luna: $900/yr estimated 路 $300/yr bill differencePublished rates
- Mistral AI 路 Ministral 3 8B: $990/yr estimated 路 $210/yr bill differencePublished rates
- Alibaba Cloud Model Studio (Qwen) 路 Qwen3.8-Flash: $1,182/yr estimated 路 $18/yr bill differencePublished rates
No cheaper managed estimate matches your filters. Adjust your usage or compare with your actual contract bill.
The price difference leaves out migration costs, operating costs and any features you鈥檇 need to replace. We compare the selected model or the cheapest listed model in the chosen capability tier. Sharing a tier does not mean equal quality.
Self-hosted options
Operating cost is not estimated. These options are not counted as savings against managed services.
Found a wrong rate or plan restriction? Report a correction with a current source.