ZMIME
NVIDIA · Released Aug 11, 2026

Nemotron 3.5 Lightning 30B A3B

Fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads.

  • Open weights
  • Reasoning
  • Tool calling
Capability
—No ECI score yet
Input
$0.05per 1M tokens
Output
$0.20per 1M tokens
Context
262K262K max output
Pricing

What it costs

Input
$0.05per million tokens
Output
$0.20per million tokens
Blended (3:1)
$0.087cheaper than 92% of priced models

Typical monthly bills

Monthly usageEstimated cost
Side project2M input + 0.5M output tokens$0.20
Team assistant25M input + 5M output tokens$2.25
Production app250M input + 50M output tokens$22.50

Median of 9 paid API providers on models.dev; NVIDIA also offers it free on Nvidia.

Capability

Independent benchmarks

Epoch AI has not published a Capabilities Index score for Nemotron 3.5 Lightning 30B A3B yet.

Specifications

The details

Lab
NVIDIA
Released
Aug 11, 2026
Context window
262,144 tokens
Max output
262,144 tokens
Inputs
Text
Output
Text
Reasoning
Can be switched on or off
Tool calling
Yes
Structured output
Yes
Weights
Open weights
API model ID
nvidia/nemotron-3.5-lightning-30b-a3bon Nvidia
Availability
12 API providerslisted on models.dev

Nvidia model documentation

Lab-reported

What NVIDIA claims

Published by the lab at launch. Settings vary, so compare these only with care.

BenchmarkScoreSettingSource
MMLU-Pro81.94BF16; reasoningSource
AA-Omniscience17.5BF16; reasoningSource
SciCode32.6BF16; reasoningSource
PinchBench85.37BF16; reasoningSource
BrowseComp36.97BF16; reasoningSource
AA-LCR52BF16; reasoningSource
GPQA Diamond75.44BF16; reasoning; no toolsSource
Humanity's Last Exam11.72BF16; reasoning; no toolsSource
SWE-Bench Verified51.56BF16; reasoningSource
SWE-Bench Multilingual39.33BF16; reasoningSource
Terminal-Bench v2.124.58BF16; reasoningSource
tau3-bench9.28BF16; reasoning; bankingSource
Same lab

More from NVIDIA

Questions

About Nemotron 3.5 Lightning 30B A3B

How much does Nemotron 3.5 Lightning 30B A3B cost?

Nemotron 3.5 Lightning 30B A3B costs $0.05 input / $0.20 output per million tokens (median across 9 API providers; free on Nvidia). At a 3:1 input-to-output mix that is $0.087 per million tokens, cheaper than 92% of the 360 priced models we track.

What is the context window of Nemotron 3.5 Lightning 30B A3B?

Nemotron 3.5 Lightning 30B A3B accepts up to 262,144 tokens per request and can write up to 262,144 tokens in one response.

How good is Nemotron 3.5 Lightning 30B A3B?

Epoch AI has not published a Capabilities Index score for Nemotron 3.5 Lightning 30B A3B yet.

Is Nemotron 3.5 Lightning 30B A3B open source?

Yes. NVIDIA publishes the weights, so it can be downloaded and self-hosted.

What inputs does Nemotron 3.5 Lightning 30B A3B support?

Nemotron 3.5 Lightning 30B A3B accepts text and replies in text. It is a reasoning model, supports tool calling and can return structured JSON output.

When was Nemotron 3.5 Lightning 30B A3B released?

NVIDIA released Nemotron 3.5 Lightning 30B A3B on Aug 11, 2026.

What are the best alternatives to Nemotron 3.5 Lightning 30B A3B?

The closest current models from other labs on capability, price and release date are Laguna XS 2.1 (Poolside, $0.06 / $0.12), Ling 3.0 Flash Fin (inclusionAI, $0.075 / $0.22), Mercury 2.5 (Inception, $0.04 / $0.15) and Gemma 4 12B IT (Google, $0.075 / $0.275).