Gemini 3.5 Flash Lite
Fast Gemini model balancing multimodal reasoning, tool use, and cost.
- Capability
- 145.1ECI · #79 of 148
- Input
- $0.30per 1M tokens
- Output
- $2.50per 1M tokens
- Context
- 1.05M66K max output
What it costs
- Input
- $0.30per million tokens
- Output
- $2.50per million tokens
- Cached input
- $0.03per million tokens
- Blended (3:1)
- $0.85pricier than 51% of priced models
Typical monthly bills
| Monthly usage | Estimated cost |
|---|---|
| Side project2M input + 0.5M output tokens | $1.85 |
| Team assistant25M input + 5M output tokens | $20.00 |
| Production app250M input + 50M output tokens | $200.00 |
Official Google API price, as listed on models.dev.
Independent benchmarks
#79 of 148 scored models
The shaded band is Epoch AI’s confidence range (142.5–146.7); the tick marks the median scored model.
Scores from Epoch AI, run independently of Google.
The details
- Lab
- Released
- Jul 21, 2026
- Knowledge cutoff
- Mar 2026
- Context window
- 1,048,576 tokens
- Max output
- 65,536 tokens
- Inputs
- Text, Images, PDFs, Audio, Video
- Output
- Text
- Reasoning
- Adjustable effort minimal · low · medium · high
- Tool calling
- Yes
- Structured output
- Yes
- Weights
- Proprietary
- API model ID
gemini-3.5-flash-liteon Google- Availability
- 26 API providerslisted on models.dev
What Google claims
Published by the lab at launch. Settings vary, so compare these only with care.
| Benchmark | Score | Setting | Source |
|---|---|---|---|
| SWE-Bench Pro | 54.2 resolve rate | — | Source |
| Terminal-Bench v2.1 | 54 | — | Source |
| MLE-Bench | 39.2 | — | Source |
| GDPval-AA vv2 | 1140 Elo | — | Source |
| OSWorld-Verified | 74 success rate | — | Source |
| CharXiv Reasoning | 74.5 | no tools | Source |
| CharXiv Reasoning | 76.5 | with tools | Source |
| GDM-MRCR vv2 | 72.2 | 128k average, 8-needle | Source |
| GDM-MRCR vv2 | 21.3 | 1M pointwise, 8-needle | Source |
More from Google
About Gemini 3.5 Flash Lite
How much does Gemini 3.5 Flash Lite cost?
Gemini 3.5 Flash Lite costs $0.30 input / $2.50 output per million tokens (official Google API price). Cached input is $0.03 per million tokens. At a 3:1 input-to-output mix that is $0.85 per million tokens, more expensive than 51% of the 360 priced models we track.
What is the context window of Gemini 3.5 Flash Lite?
Gemini 3.5 Flash Lite accepts up to 1,048,576 tokens per request and can write up to 65,536 tokens in one response.
How good is Gemini 3.5 Flash Lite?
Epoch AI gives Gemini 3.5 Flash Lite a Capabilities Index score of 145.1 (likely range 142.5–146.7), ranking it #79 of 148 models Epoch has scored. Epoch AI benchmark results: GPQA Diamond 83.3%, FrontierMath Tiers 1–3 26.0%, OTIS Mock AIME 2024–2025 71.1%.
Is Gemini 3.5 Flash Lite open source?
No. Gemini 3.5 Flash Lite is proprietary; you use it through Google’s API or partner platforms.
What inputs does Gemini 3.5 Flash Lite support?
Gemini 3.5 Flash Lite accepts text, images, PDFs, audio and video and replies in text. It is a reasoning model with minimal, low, medium and high effort settings, supports tool calling and can return structured JSON output.
When was Gemini 3.5 Flash Lite released?
Google released Gemini 3.5 Flash Lite on Jul 21, 2026. Its training data runs to Mar 2026.
What are the best alternatives to Gemini 3.5 Flash Lite?
The closest current models from other labs on capability, price and release date are Nemotron 3 Ultra 550B A55B (NVIDIA, ECI 146.2, $0.50 / $2.50), Qwen3.7 Plus (Alibaba (Qwen), ECI 147.4, $0.40 / $1.60), MiniMax-M3 (MiniMax, ECI 147.0, $0.30 / $1.20) and GPT-5.4 nano (OpenAI, ECI 145.8, $0.20 / $1.25).