GLM-4.5-Flash
Efficient GLM model for fast reasoning, coding, and agent workflows.
- Capability
- —No ECI score yet
- Input
- Freeper 1M tokens
- Output
- Freeper 1M tokens
- Context
- 131K98K max output
What it costs
- Input
- Freeper million tokens
- Output
- Freeper million tokens
Listed as free on Z.AI’s own API.
Independent benchmarks
Epoch AI has not published a Capabilities Index score for GLM-4.5-Flash yet.
The details
- Lab
- Z.ai (Zhipu)
- Released
- Jul 28, 2025
- Knowledge cutoff
- Apr 2025
- Context window
- 131,072 tokens
- Max output
- 98,304 tokens
- Inputs
- Text
- Output
- Text
- Reasoning
- Can be switched on or off
- Tool calling
- Yes
- Structured output
- No
- Weights
- Proprietary
- API model ID
glm-4.5-flashon Z.AI- Availability
- 4 API providerslisted on models.dev
More from Z.ai (Zhipu)
About GLM-4.5-Flash
How much does GLM-4.5-Flash cost?
GLM-4.5-Flash is listed as free on Z.AI.
What is the context window of GLM-4.5-Flash?
GLM-4.5-Flash accepts up to 131,072 tokens per request and can write up to 98,304 tokens in one response.
How good is GLM-4.5-Flash?
Epoch AI has not published a Capabilities Index score for GLM-4.5-Flash yet.
Is GLM-4.5-Flash open source?
No. GLM-4.5-Flash is proprietary; you use it through Z.AI’s API or partner platforms.
What inputs does GLM-4.5-Flash support?
GLM-4.5-Flash accepts text and replies in text. It is a reasoning model, supports tool calling.
When was GLM-4.5-Flash released?
Z.ai (Zhipu) released GLM-4.5-Flash on Jul 28, 2025. Its training data runs to Apr 2025.
What are the best alternatives to GLM-4.5-Flash?
The closest current models from other labs on capability, price and release date are Llama 3.1 Nemotron Ultra 253B (NVIDIA, Free / Free), Nova 2 Pro (Amazon, Free / Free), Laguna M.1 (Poolside, Free / Free) and North Mini Code (Cohere, Free / Free).