ZMIME
NVIDIA · Released Jun 4, 2026

Nemotron 3 Ultra 550B A55B

Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy.

  • Open weights
  • Reasoning
  • Tool calling
Capability
146.2ECI · #70 of 148
Input
$0.50per 1M tokens
Output
$2.50per 1M tokens
Context
1M128K max output
Pricing

What it costs

Input
$0.50per million tokens
Output
$2.50per million tokens
Cached input
$0.15per million tokens
Blended (3:1)
$1.00pricier than 55% of priced models

Typical monthly bills

Monthly usageEstimated cost
Side project2M input + 0.5M output tokens$2.25
Team assistant25M input + 5M output tokens$25.00
Production app250M input + 50M output tokens$250.00

Official Nvidia API price, as listed on models.dev.

Capability

Independent benchmarks

146.2Epoch Capabilities Index
#70 of 148 scored models
88Median 146167

The shaded band is Epoch AI’s confidence range (143.9–148.1); the tick marks the median scored model.

  • GPQA DiamondGraduate-level science questions85.4%
  • OTIS Mock AIME 2024–2025Competition mathematics86.7%

Scores from Epoch AI, run independently of NVIDIA.

Specifications

The details

Lab
NVIDIA
Released
Jun 4, 2026
Context window
1,000,000 tokens
Max output
128,000 tokens
Inputs
Text
Output
Text
Reasoning
Can be switched on or off
Tool calling
Yes
Structured output
Yes
Weights
Open weights
API model ID
nvidia/nemotron-3-ultra-550b-a55bon Nvidia
Availability
21 API providerslisted on models.dev

Nvidia model documentation

Lab-reported

What NVIDIA claims

Published by the lab at launch. Settings vary, so compare these only with care.

BenchmarkScoreSettingSource
SWE-Bench Verified70.7 resolved—Source
SWE-Bench Multilingual67.7 resolve rate—Source
Terminal-Bench v2.156.4 success rate—Source
GPQA87no toolsSource
Humanity's Last Exam26.7no toolsSource
Humanity's Last Exam37.4with toolsSource
LiveCodeBench vv689—Source
MMLU-Pro86.8—Source
BrowseComp44.4—Source
IFBench81.7prompt looseSource
GDPval46.7 wins or ties—Source
Same lab

More from NVIDIA

Questions

About Nemotron 3 Ultra 550B A55B

How much does Nemotron 3 Ultra 550B A55B cost?

Nemotron 3 Ultra 550B A55B costs $0.50 input / $2.50 output per million tokens (official Nvidia API price). Cached input is $0.15 per million tokens. At a 3:1 input-to-output mix that is $1.00 per million tokens, more expensive than 55% of the 360 priced models we track.

What is the context window of Nemotron 3 Ultra 550B A55B?

Nemotron 3 Ultra 550B A55B accepts up to 1,000,000 tokens per request and can write up to 128,000 tokens in one response.

How good is Nemotron 3 Ultra 550B A55B?

Epoch AI gives Nemotron 3 Ultra 550B A55B a Capabilities Index score of 146.2 (likely range 143.9–148.1), ranking it #70 of 148 models Epoch has scored. Epoch AI benchmark results: GPQA Diamond 85.4%, OTIS Mock AIME 2024–2025 86.7%.

Is Nemotron 3 Ultra 550B A55B open source?

Yes. NVIDIA publishes the weights, so it can be downloaded and self-hosted.

What inputs does Nemotron 3 Ultra 550B A55B support?

Nemotron 3 Ultra 550B A55B accepts text and replies in text. It is a reasoning model, supports tool calling and can return structured JSON output.

When was Nemotron 3 Ultra 550B A55B released?

NVIDIA released Nemotron 3 Ultra 550B A55B on Jun 4, 2026.

What are the best alternatives to Nemotron 3 Ultra 550B A55B?

The closest current models from other labs on capability, price and release date are Qwen3.6 27B (Alibaba (Qwen), ECI 146.5, $0.60 / $3.60), Gemini 3.5 Flash Lite (Google, ECI 145.1, $0.30 / $2.50), Kimi K2 Thinking (Moonshot AI, ECI 146.0, $0.60 / $2.50) and MiniMax-M3 (MiniMax, ECI 147.0, $0.30 / $1.20).