ZMIME
Z.ai (Zhipu) · Released Aug 26, 2026

GLM-5.3-Flash

Native multimodal GLM model for efficient coding and long-horizon agent tasks.

  • Open weights
  • Reasoning
  • Tool calling
  • Vision
  • Video input
Capability
151.9ECI · #42 of 148
Input
$0.15per 1M tokens
Output
$0.50per 1M tokens
Context
1M131K max output
Pricing

What it costs

Input
$0.15per million tokens
Output
$0.50per million tokens
Cached input
$0.03per million tokens
Blended (3:1)
$0.237cheaper than 76% of priced models

Typical monthly bills

Monthly usageEstimated cost
Side project2M input + 0.5M output tokens$0.55
Team assistant25M input + 5M output tokens$6.25
Production app250M input + 50M output tokens$62.50

Official Z.AI API price, as listed on models.dev.

Capability

Independent benchmarks

151.9Epoch Capabilities Index
#42 of 148 scored models
88Median 146167

The shaded band is Epoch AI’s confidence range (149.4–154.3); the tick marks the median scored model.

  • GPQA DiamondGraduate-level science questions · max setting90.2%
  • FrontierMath Tiers 1–3Research-level mathematics · max setting55.8%
  • OTIS Mock AIME 2024–2025Competition mathematics · max setting93.9%

Scores from Epoch AI, run independently of Z.ai (Zhipu).

Specifications

The details

Lab
Z.ai (Zhipu)
Released
Aug 26, 2026
Context window
1,000,000 tokens
Max output
131,072 tokens
Inputs
Text, Images, PDFs, Video
Output
Text
Reasoning
Adjustable effort low · high · max
Tool calling
Yes
Structured output
Yes
Weights
Open weights
API model ID
glm-5.3-flashon Z.AI
Availability
65 API providerslisted on models.dev

Z.AI model documentation

Lab-reported

What Z.ai (Zhipu) claims

Published by the lab at launch. Settings vary, so compare these only with care.

BenchmarkScoreSettingSource
Terminal-Bench v2.184.3max effortSource
DeepSWE v1.163.4 resolvedmax effortSource
Agents' Last Exam26.3max effortSource
AutomationBench v1.0.648.8max effortSource
Humanity's Last Exam55.3max effort; with toolsSource
GDPval-AA v21773 Elomax effortSource
Same lab

More from Z.ai (Zhipu)

Questions

About GLM-5.3-Flash

How much does GLM-5.3-Flash cost?

GLM-5.3-Flash costs $0.15 input / $0.50 output per million tokens (official Z.AI API price). Cached input is $0.03 per million tokens. At a 3:1 input-to-output mix that is $0.237 per million tokens, cheaper than 76% of the 360 priced models we track.

What is the context window of GLM-5.3-Flash?

GLM-5.3-Flash accepts up to 1,000,000 tokens per request and can write up to 131,072 tokens in one response.

How good is GLM-5.3-Flash?

Epoch AI gives GLM-5.3-Flash a Capabilities Index score of 151.9 (likely range 149.4–154.3), ranking it #42 of 148 models Epoch has scored. Epoch AI benchmark results: GPQA Diamond 90.2%, FrontierMath Tiers 1–3 55.8%, OTIS Mock AIME 2024–2025 93.9%.

Is GLM-5.3-Flash open source?

Yes. Z.ai (Zhipu) publishes the weights, so it can be downloaded and self-hosted.

What inputs does GLM-5.3-Flash support?

GLM-5.3-Flash accepts text, images, PDFs and video and replies in text. It is a reasoning model with low, high and max effort settings, supports tool calling and can return structured JSON output.

When was GLM-5.3-Flash released?

Z.ai (Zhipu) released GLM-5.3-Flash on Aug 26, 2026.

What are the best alternatives to GLM-5.3-Flash?

The closest current models from other labs on capability, price and release date are GPT-6 Luna (OpenAI, $0.10 / $0.50), DeepSeek V4.1 Flash (DeepSeek, ECI 155.0, $0.15 / $0.60), Inkling Small (Thinking Machines, ECI 150.2, $0.50 / $1.20) and Qwen3.8 27B (Alibaba (Qwen), ECI 149.4, $0.40 / $2.50).