ZMIME
Google · Released May 7, 2026

Gemini 3.1 Flash Lite

Low-latency Gemini model for high-volume multimodal and agent workloads.

  • Proprietary
  • Reasoning
  • Tool calling
  • Vision
  • Audio input
  • Video input
Capability
—No ECI score yet
Input
$0.25per 1M tokens
Output
$1.50per 1M tokens
Context
1.05M66K max output
Pricing

What it costs

Input
$0.25per million tokens
Output
$1.50per million tokens
Cached input
$0.025per million tokens
Blended (3:1)
$0.563cheaper than 57% of priced models

Typical monthly bills

Monthly usageEstimated cost
Side project2M input + 0.5M output tokens$1.25
Team assistant25M input + 5M output tokens$13.75
Production app250M input + 50M output tokens$137.50

Official Google API price, as listed on models.dev.

Capability

Independent benchmarks

Epoch AI has not published a Capabilities Index score for Gemini 3.1 Flash Lite yet.

Specifications

The details

Lab
Google
Released
May 7, 2026
Knowledge cutoff
Jan 2025
Context window
1,048,576 tokens
Max output
65,536 tokens
Inputs
Text, Images, PDFs, Audio, Video
Output
Text
Reasoning
Adjustable effort minimal · low · medium · high
Tool calling
Yes
Structured output
Yes
Weights
Proprietary
API model ID
gemini-3.1-flash-liteon Google
Availability
23 API providerslisted on models.dev

Google model documentation

Same lab

More from Google

Questions

About Gemini 3.1 Flash Lite

How much does Gemini 3.1 Flash Lite cost?

Gemini 3.1 Flash Lite costs $0.25 input / $1.50 output per million tokens (official Google API price). Cached input is $0.025 per million tokens. At a 3:1 input-to-output mix that is $0.563 per million tokens, cheaper than 57% of the 360 priced models we track.

What is the context window of Gemini 3.1 Flash Lite?

Gemini 3.1 Flash Lite accepts up to 1,048,576 tokens per request and can write up to 65,536 tokens in one response.

How good is Gemini 3.1 Flash Lite?

Epoch AI has not published a Capabilities Index score for Gemini 3.1 Flash Lite yet.

Is Gemini 3.1 Flash Lite open source?

No. Gemini 3.1 Flash Lite is proprietary; you use it through Google’s API or partner platforms.

What inputs does Gemini 3.1 Flash Lite support?

Gemini 3.1 Flash Lite accepts text, images, PDFs, audio and video and replies in text. It is a reasoning model with minimal, low, medium and high effort settings, supports tool calling and can return structured JSON output.

When was Gemini 3.1 Flash Lite released?

Google released Gemini 3.1 Flash Lite on May 7, 2026. Its training data runs to Jan 2025.

What are the best alternatives to Gemini 3.1 Flash Lite?

The closest current models from other labs on capability, price and release date are MiMo-V2.5-Pro (Xiaomi, $0.435 / $0.87), Ornith 1.5 35B A3B (DeepReinforce, $0.20 / $1.70), Nemotron 3 Nano Omni 30B A3B Reasoning (NVIDIA, $0.25 / $0.85) and Solar Pro 4 (Upstage, $0.30 / $1.20).