Gemini 3.5 Flash
Fast Gemini model balancing multimodal reasoning, tool use, and cost.
- Capability
- 154.5ECI · #33 of 148
- Input
- $1.50per 1M tokens
- Output
- $9.00per 1M tokens
- Context
- 1.05M66K max output
What it costs
- Input
- $1.50per million tokens
- Output
- $9.00per million tokens
- Cached input
- $0.15per million tokens
- Blended (3:1)
- $3.38pricier than 77% of priced models
Typical monthly bills
| Monthly usage | Estimated cost |
|---|---|
| Side project2M input + 0.5M output tokens | $7.50 |
| Team assistant25M input + 5M output tokens | $82.50 |
| Production app250M input + 50M output tokens | $825.00 |
Official Google API price, as listed on models.dev.
Independent benchmarks
#33 of 148 scored models
The shaded band is Epoch AI’s confidence range (152.5–156.6); the tick marks the median scored model.
Scores from Epoch AI, run independently of Google.
The details
- Lab
- Released
- May 19, 2026
- Knowledge cutoff
- Jan 2025
- Context window
- 1,048,576 tokens
- Max output
- 65,536 tokens
- Inputs
- Text, Images, PDFs, Audio, Video
- Output
- Text
- Reasoning
- Adjustable effort minimal · low · medium · high
- Tool calling
- Yes
- Structured output
- Yes
- Weights
- Proprietary
- API model ID
gemini-3.5-flashon Google- Availability
- 32 API providerslisted on models.dev
What Google claims
Published by the lab at launch. Settings vary, so compare these only with care.
| Benchmark | Score | Setting | Source |
|---|---|---|---|
| Terminal-Bench v2.1 | 76.2 success rate | — | Source |
| SWE-Bench Pro | 55.1 resolve rate | single attempt | Source |
| MCP Atlas | 83.6 success rate | — | Source |
| Toolathlon | 56.5 success rate | — | Source |
| OSWorld-Verified | 78.4 success rate | — | Source |
| MMMU Pro | 83.6 | no tools | Source |
| CharXiv Reasoning | 84.2 | no tools | Source |
| Humanity's Last Exam | 40.2 | — | Source |
| ARC-AGI-2 | 72.1 | — | Source |
| GDPval-AA | 1656 Elo | — | Source |
More from Google
About Gemini 3.5 Flash
How much does Gemini 3.5 Flash cost?
Gemini 3.5 Flash costs $1.50 input / $9.00 output per million tokens (official Google API price). Cached input is $0.15 per million tokens. At a 3:1 input-to-output mix that is $3.38 per million tokens, more expensive than 77% of the 360 priced models we track.
What is the context window of Gemini 3.5 Flash?
Gemini 3.5 Flash accepts up to 1,048,576 tokens per request and can write up to 65,536 tokens in one response.
How good is Gemini 3.5 Flash?
Epoch AI gives Gemini 3.5 Flash a Capabilities Index score of 154.5 (likely range 152.5–156.6), ranking it #33 of 148 models Epoch has scored. Epoch AI benchmark results: GPQA Diamond 92.8%, FrontierMath Tiers 1–3 62.8%, OTIS Mock AIME 2024–2025 95.6%, SWE-bench Verified 79.3%, SimpleQA Verified 66.2%.
Is Gemini 3.5 Flash open source?
No. Gemini 3.5 Flash is proprietary; you use it through Google’s API or partner platforms.
What inputs does Gemini 3.5 Flash support?
Gemini 3.5 Flash accepts text, images, PDFs, audio and video and replies in text. It is a reasoning model with minimal, low, medium and high effort settings, supports tool calling and can return structured JSON output.
When was Gemini 3.5 Flash released?
Google released Gemini 3.5 Flash on May 19, 2026. Its training data runs to Jan 2025.
What are the best alternatives to Gemini 3.5 Flash?
The closest current models from other labs on capability, price and release date are Qwen3.7 Max (Alibaba (Qwen), ECI 153.7, $2.50 / $7.50), Grok 4.5 (xAI, ECI 154.0, $2.00 / $6.00), Claude Sonnet 5 (Anthropic, ECI 156.3, $2.00 / $10.00) and Muse Spark 1.1 (Meta, ECI 154.3, $1.25 / $4.25).