ZMIME
Google · Released Jun 17, 2025

Gemini 2.5 Flash-Lite

Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents.

  • Proprietary
  • Reasoning
  • Tool calling
  • Vision
  • Audio input
  • Video input
Capability
133.9ECI · #118 of 148
Input
$0.10per 1M tokens
Output
$0.40per 1M tokens
Context
1.05M66K max output
Pricing

What it costs

Input
$0.10per million tokens
Output
$0.40per million tokens
Cached input
$0.01per million tokens
Blended (3:1)
$0.175cheaper than 79% of priced models

Typical monthly bills

Monthly usageEstimated cost
Side project2M input + 0.5M output tokens$0.40
Team assistant25M input + 5M output tokens$4.50
Production app250M input + 50M output tokens$45.00

Official Google API price, as listed on models.dev.

Capability

Independent benchmarks

133.9Epoch Capabilities Index
#118 of 148 scored models
88Median 146167

The shaded band is Epoch AI’s confidence range (129.8–136.3); the tick marks the median scored model.

Scores from Epoch AI, run independently of Google.

Specifications

The details

Lab
Google
Released
Jun 17, 2025
Knowledge cutoff
Jan 2025
Context window
1,048,576 tokens
Max output
65,536 tokens
Inputs
Text, Images, PDFs, Audio, Video
Output
Text
Reasoning
Thinking budget
Tool calling
Yes
Structured output
Yes
Weights
Proprietary
API model ID
gemini-2.5-flash-liteon Google
Availability
20 API providerslisted on models.dev

Google model documentation

Lab-reported

What Google claims

Published by the lab at launch. Settings vary, so compare these only with care.

BenchmarkScoreSettingSource
Artificial Analysis Coding Index9.5 index—Source
SciCode19.3 percent correct—Source
Terminal-Bench Hard4.5 success rate—Source
Same lab

More from Google

Questions

About Gemini 2.5 Flash-Lite

How much does Gemini 2.5 Flash-Lite cost?

Gemini 2.5 Flash-Lite costs $0.10 input / $0.40 output per million tokens (official Google API price). Cached input is $0.01 per million tokens. At a 3:1 input-to-output mix that is $0.175 per million tokens, cheaper than 79% of the 360 priced models we track.

What is the context window of Gemini 2.5 Flash-Lite?

Gemini 2.5 Flash-Lite accepts up to 1,048,576 tokens per request and can write up to 65,536 tokens in one response.

How good is Gemini 2.5 Flash-Lite?

Epoch AI gives Gemini 2.5 Flash-Lite a Capabilities Index score of 133.9 (likely range 129.8–136.3), ranking it #118 of 148 models Epoch has scored.

Is Gemini 2.5 Flash-Lite open source?

No. Gemini 2.5 Flash-Lite is proprietary; you use it through Google’s API or partner platforms.

What inputs does Gemini 2.5 Flash-Lite support?

Gemini 2.5 Flash-Lite accepts text, images, PDFs, audio and video and replies in text. It is a reasoning model, supports tool calling and can return structured JSON output.

When was Gemini 2.5 Flash-Lite released?

Google released Gemini 2.5 Flash-Lite on Jun 17, 2025. Its training data runs to Jan 2025.

What are the best alternatives to Gemini 2.5 Flash-Lite?

The closest current models from other labs on capability, price and release date are Mistral Small 3.2 (Mistral AI, ECI 131.7, $0.10 / $0.30), Qwen3 30B A3B (Alibaba (Qwen), ECI 136.2, $0.114 / $0.50), GPT OSS 20B (OpenAI, ECI 137.8, $0.07 / $0.295) and DeepSeek V3 0324 (DeepSeek, ECI 135.9, $0.24 / $0.90).