GLM-5.3-Flash
Native multimodal GLM model for efficient coding and long-horizon agent tasks.
- Capability
- 151.9ECI · #42 of 148
- Input
- $0.15per 1M tokens
- Output
- $0.50per 1M tokens
- Context
- 1M131K max output
What it costs
- Input
- $0.15per million tokens
- Output
- $0.50per million tokens
- Cached input
- $0.03per million tokens
- Blended (3:1)
- $0.237cheaper than 76% of priced models
Typical monthly bills
| Monthly usage | Estimated cost |
|---|---|
| Side project2M input + 0.5M output tokens | $0.55 |
| Team assistant25M input + 5M output tokens | $6.25 |
| Production app250M input + 50M output tokens | $62.50 |
Official Z.AI API price, as listed on models.dev.
Independent benchmarks
#42 of 148 scored models
The shaded band is Epoch AI’s confidence range (149.4–154.3); the tick marks the median scored model.
Scores from Epoch AI, run independently of Z.ai (Zhipu).
The details
- Lab
- Z.ai (Zhipu)
- Released
- Aug 26, 2026
- Context window
- 1,000,000 tokens
- Max output
- 131,072 tokens
- Inputs
- Text, Images, PDFs, Video
- Output
- Text
- Reasoning
- Adjustable effort low · high · max
- Tool calling
- Yes
- Structured output
- Yes
- Weights
- Open weights
- API model ID
glm-5.3-flashon Z.AI- Availability
- 65 API providerslisted on models.dev
What Z.ai (Zhipu) claims
Published by the lab at launch. Settings vary, so compare these only with care.
More from Z.ai (Zhipu)
About GLM-5.3-Flash
How much does GLM-5.3-Flash cost?
GLM-5.3-Flash costs $0.15 input / $0.50 output per million tokens (official Z.AI API price). Cached input is $0.03 per million tokens. At a 3:1 input-to-output mix that is $0.237 per million tokens, cheaper than 76% of the 360 priced models we track.
What is the context window of GLM-5.3-Flash?
GLM-5.3-Flash accepts up to 1,000,000 tokens per request and can write up to 131,072 tokens in one response.
How good is GLM-5.3-Flash?
Epoch AI gives GLM-5.3-Flash a Capabilities Index score of 151.9 (likely range 149.4–154.3), ranking it #42 of 148 models Epoch has scored. Epoch AI benchmark results: GPQA Diamond 90.2%, FrontierMath Tiers 1–3 55.8%, OTIS Mock AIME 2024–2025 93.9%.
Is GLM-5.3-Flash open source?
Yes. Z.ai (Zhipu) publishes the weights, so it can be downloaded and self-hosted.
What inputs does GLM-5.3-Flash support?
GLM-5.3-Flash accepts text, images, PDFs and video and replies in text. It is a reasoning model with low, high and max effort settings, supports tool calling and can return structured JSON output.
When was GLM-5.3-Flash released?
Z.ai (Zhipu) released GLM-5.3-Flash on Aug 26, 2026.
What are the best alternatives to GLM-5.3-Flash?
The closest current models from other labs on capability, price and release date are GPT-6 Luna (OpenAI, $0.10 / $0.50), DeepSeek V4.1 Flash (DeepSeek, ECI 155.0, $0.15 / $0.60), Inkling Small (Thinking Machines, ECI 150.2, $0.50 / $1.20) and Qwen3.8 27B (Alibaba (Qwen), ECI 149.4, $0.40 / $2.50).