Nemotron 3 Ultra 550B A55B
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy.
- Capability
- 146.2ECI · #70 of 148
- Input
- $0.50per 1M tokens
- Output
- $2.50per 1M tokens
- Context
- 1M128K max output
What it costs
- Input
- $0.50per million tokens
- Output
- $2.50per million tokens
- Cached input
- $0.15per million tokens
- Blended (3:1)
- $1.00pricier than 55% of priced models
Typical monthly bills
| Monthly usage | Estimated cost |
|---|---|
| Side project2M input + 0.5M output tokens | $2.25 |
| Team assistant25M input + 5M output tokens | $25.00 |
| Production app250M input + 50M output tokens | $250.00 |
Official Nvidia API price, as listed on models.dev.
Independent benchmarks
#70 of 148 scored models
The shaded band is Epoch AI’s confidence range (143.9–148.1); the tick marks the median scored model.
Scores from Epoch AI, run independently of NVIDIA.
The details
- Lab
- NVIDIA
- Released
- Jun 4, 2026
- Context window
- 1,000,000 tokens
- Max output
- 128,000 tokens
- Inputs
- Text
- Output
- Text
- Reasoning
- Can be switched on or off
- Tool calling
- Yes
- Structured output
- Yes
- Weights
- Open weights
- API model ID
nvidia/nemotron-3-ultra-550b-a55bon Nvidia- Availability
- 21 API providerslisted on models.dev
What NVIDIA claims
Published by the lab at launch. Settings vary, so compare these only with care.
| Benchmark | Score | Setting | Source |
|---|---|---|---|
| SWE-Bench Verified | 70.7 resolved | — | Source |
| SWE-Bench Multilingual | 67.7 resolve rate | — | Source |
| Terminal-Bench v2.1 | 56.4 success rate | — | Source |
| GPQA | 87 | no tools | Source |
| Humanity's Last Exam | 26.7 | no tools | Source |
| Humanity's Last Exam | 37.4 | with tools | Source |
| LiveCodeBench vv6 | 89 | — | Source |
| MMLU-Pro | 86.8 | — | Source |
| BrowseComp | 44.4 | — | Source |
| IFBench | 81.7 | prompt loose | Source |
| GDPval | 46.7 wins or ties | — | Source |
More from NVIDIA
About Nemotron 3 Ultra 550B A55B
How much does Nemotron 3 Ultra 550B A55B cost?
Nemotron 3 Ultra 550B A55B costs $0.50 input / $2.50 output per million tokens (official Nvidia API price). Cached input is $0.15 per million tokens. At a 3:1 input-to-output mix that is $1.00 per million tokens, more expensive than 55% of the 360 priced models we track.
What is the context window of Nemotron 3 Ultra 550B A55B?
Nemotron 3 Ultra 550B A55B accepts up to 1,000,000 tokens per request and can write up to 128,000 tokens in one response.
How good is Nemotron 3 Ultra 550B A55B?
Epoch AI gives Nemotron 3 Ultra 550B A55B a Capabilities Index score of 146.2 (likely range 143.9–148.1), ranking it #70 of 148 models Epoch has scored. Epoch AI benchmark results: GPQA Diamond 85.4%, OTIS Mock AIME 2024–2025 86.7%.
Is Nemotron 3 Ultra 550B A55B open source?
Yes. NVIDIA publishes the weights, so it can be downloaded and self-hosted.
What inputs does Nemotron 3 Ultra 550B A55B support?
Nemotron 3 Ultra 550B A55B accepts text and replies in text. It is a reasoning model, supports tool calling and can return structured JSON output.
When was Nemotron 3 Ultra 550B A55B released?
NVIDIA released Nemotron 3 Ultra 550B A55B on Jun 4, 2026.
What are the best alternatives to Nemotron 3 Ultra 550B A55B?
The closest current models from other labs on capability, price and release date are Qwen3.6 27B (Alibaba (Qwen), ECI 146.5, $0.60 / $3.60), Gemini 3.5 Flash Lite (Google, ECI 145.1, $0.30 / $2.50), Kimi K2 Thinking (Moonshot AI, ECI 146.0, $0.60 / $2.50) and MiniMax-M3 (MiniMax, ECI 147.0, $0.30 / $1.20).