ZMIME
Z.ai (Zhipu) · Released Jul 28, 2025

GLM-4.5-Flash

Efficient GLM model for fast reasoning, coding, and agent workflows.

  • Proprietary
  • Reasoning
  • Tool calling
Capability
—No ECI score yet
Input
Freeper 1M tokens
Output
Freeper 1M tokens
Context
131K98K max output
Pricing

What it costs

Input
Freeper million tokens
Output
Freeper million tokens

Listed as free on Z.AI’s own API.

Capability

Independent benchmarks

Epoch AI has not published a Capabilities Index score for GLM-4.5-Flash yet.

Specifications

The details

Lab
Z.ai (Zhipu)
Released
Jul 28, 2025
Knowledge cutoff
Apr 2025
Context window
131,072 tokens
Max output
98,304 tokens
Inputs
Text
Output
Text
Reasoning
Can be switched on or off
Tool calling
Yes
Structured output
No
Weights
Proprietary
API model ID
glm-4.5-flashon Z.AI
Availability
4 API providerslisted on models.dev

Z.AI model documentation

Same lab

More from Z.ai (Zhipu)

Questions

About GLM-4.5-Flash

How much does GLM-4.5-Flash cost?

GLM-4.5-Flash is listed as free on Z.AI.

What is the context window of GLM-4.5-Flash?

GLM-4.5-Flash accepts up to 131,072 tokens per request and can write up to 98,304 tokens in one response.

How good is GLM-4.5-Flash?

Epoch AI has not published a Capabilities Index score for GLM-4.5-Flash yet.

Is GLM-4.5-Flash open source?

No. GLM-4.5-Flash is proprietary; you use it through Z.AI’s API or partner platforms.

What inputs does GLM-4.5-Flash support?

GLM-4.5-Flash accepts text and replies in text. It is a reasoning model, supports tool calling.

When was GLM-4.5-Flash released?

Z.ai (Zhipu) released GLM-4.5-Flash on Jul 28, 2025. Its training data runs to Apr 2025.

What are the best alternatives to GLM-4.5-Flash?

The closest current models from other labs on capability, price and release date are Llama 3.1 Nemotron Ultra 253B (NVIDIA, Free / Free), Nova 2 Pro (Amazon, Free / Free), Laguna M.1 (Poolside, Free / Free) and North Mini Code (Cohere, Free / Free).