ZMIME
Google · Released Jul 21, 2026

Gemini 3.5 Flash Lite

Fast Gemini model balancing multimodal reasoning, tool use, and cost.

  • Proprietary
  • Reasoning
  • Tool calling
  • Vision
  • Audio input
  • Video input
Capability
145.1ECI · #79 of 148
Input
$0.30per 1M tokens
Output
$2.50per 1M tokens
Context
1.05M66K max output
Pricing

What it costs

Input
$0.30per million tokens
Output
$2.50per million tokens
Cached input
$0.03per million tokens
Blended (3:1)
$0.85pricier than 51% of priced models

Typical monthly bills

Monthly usageEstimated cost
Side project2M input + 0.5M output tokens$1.85
Team assistant25M input + 5M output tokens$20.00
Production app250M input + 50M output tokens$200.00

Official Google API price, as listed on models.dev.

Capability

Independent benchmarks

145.1Epoch Capabilities Index
#79 of 148 scored models
88Median 146167

The shaded band is Epoch AI’s confidence range (142.5–146.7); the tick marks the median scored model.

  • GPQA DiamondGraduate-level science questions · high setting83.3%
  • FrontierMath Tiers 1–3Research-level mathematics · high setting26.0%
  • OTIS Mock AIME 2024–2025Competition mathematics · high setting71.1%

Scores from Epoch AI, run independently of Google.

Specifications

The details

Lab
Google
Released
Jul 21, 2026
Knowledge cutoff
Mar 2026
Context window
1,048,576 tokens
Max output
65,536 tokens
Inputs
Text, Images, PDFs, Audio, Video
Output
Text
Reasoning
Adjustable effort minimal · low · medium · high
Tool calling
Yes
Structured output
Yes
Weights
Proprietary
API model ID
gemini-3.5-flash-liteon Google
Availability
26 API providerslisted on models.dev

Google model documentation

Lab-reported

What Google claims

Published by the lab at launch. Settings vary, so compare these only with care.

BenchmarkScoreSettingSource
SWE-Bench Pro54.2 resolve rate—Source
Terminal-Bench v2.154—Source
MLE-Bench39.2—Source
GDPval-AA vv21140 Elo—Source
OSWorld-Verified74 success rate—Source
CharXiv Reasoning74.5no toolsSource
CharXiv Reasoning76.5with toolsSource
GDM-MRCR vv272.2128k average, 8-needleSource
GDM-MRCR vv221.31M pointwise, 8-needleSource
Same lab

More from Google

Questions

About Gemini 3.5 Flash Lite

How much does Gemini 3.5 Flash Lite cost?

Gemini 3.5 Flash Lite costs $0.30 input / $2.50 output per million tokens (official Google API price). Cached input is $0.03 per million tokens. At a 3:1 input-to-output mix that is $0.85 per million tokens, more expensive than 51% of the 360 priced models we track.

What is the context window of Gemini 3.5 Flash Lite?

Gemini 3.5 Flash Lite accepts up to 1,048,576 tokens per request and can write up to 65,536 tokens in one response.

How good is Gemini 3.5 Flash Lite?

Epoch AI gives Gemini 3.5 Flash Lite a Capabilities Index score of 145.1 (likely range 142.5–146.7), ranking it #79 of 148 models Epoch has scored. Epoch AI benchmark results: GPQA Diamond 83.3%, FrontierMath Tiers 1–3 26.0%, OTIS Mock AIME 2024–2025 71.1%.

Is Gemini 3.5 Flash Lite open source?

No. Gemini 3.5 Flash Lite is proprietary; you use it through Google’s API or partner platforms.

What inputs does Gemini 3.5 Flash Lite support?

Gemini 3.5 Flash Lite accepts text, images, PDFs, audio and video and replies in text. It is a reasoning model with minimal, low, medium and high effort settings, supports tool calling and can return structured JSON output.

When was Gemini 3.5 Flash Lite released?

Google released Gemini 3.5 Flash Lite on Jul 21, 2026. Its training data runs to Mar 2026.

What are the best alternatives to Gemini 3.5 Flash Lite?

The closest current models from other labs on capability, price and release date are Nemotron 3 Ultra 550B A55B (NVIDIA, ECI 146.2, $0.50 / $2.50), Qwen3.7 Plus (Alibaba (Qwen), ECI 147.4, $0.40 / $1.60), MiniMax-M3 (MiniMax, ECI 147.0, $0.30 / $1.20) and GPT-5.4 nano (OpenAI, ECI 145.8, $0.20 / $1.25).