Llama 3.1 Nemotron Ultra 253B
Flagship Nemotron model for high-throughput reasoning and complex agents.
- Capability
- —No ECI score yet
- Input
- Freeper 1M tokens
- Output
- Freeper 1M tokens
- Context
- 128K8K max output
What it costs
- Input
- Freeper million tokens
- Output
- Freeper million tokens
Listed as free on Nvidia’s own API.
Independent benchmarks
Epoch AI has not published a Capabilities Index score for Llama 3.1 Nemotron Ultra 253B yet.
The details
- Lab
- NVIDIA
- Released
- Apr 7, 2025
- Context window
- 128,000 tokens
- Max output
- 8,192 tokens
- Inputs
- Text
- Output
- Text
- Reasoning
- Can be switched on or off
- Tool calling
- Yes
- Structured output
- No
- Weights
- Open weights
- API model ID
nvidia/llama-3.1-nemotron-ultra-253b-v1on Nvidia- Availability
- 1 API providerlisted on models.dev
More from NVIDIA
About Llama 3.1 Nemotron Ultra 253B
How much does Llama 3.1 Nemotron Ultra 253B cost?
Llama 3.1 Nemotron Ultra 253B is listed as free on Nvidia.
What is the context window of Llama 3.1 Nemotron Ultra 253B?
Llama 3.1 Nemotron Ultra 253B accepts up to 128,000 tokens per request and can write up to 8,192 tokens in one response.
How good is Llama 3.1 Nemotron Ultra 253B?
Epoch AI has not published a Capabilities Index score for Llama 3.1 Nemotron Ultra 253B yet.
Is Llama 3.1 Nemotron Ultra 253B open source?
Yes. NVIDIA publishes the weights, so it can be downloaded and self-hosted.
What inputs does Llama 3.1 Nemotron Ultra 253B support?
Llama 3.1 Nemotron Ultra 253B accepts text and replies in text. It is a reasoning model, supports tool calling.
When was Llama 3.1 Nemotron Ultra 253B released?
NVIDIA released Llama 3.1 Nemotron Ultra 253B on Apr 7, 2025.
What are the best alternatives to Llama 3.1 Nemotron Ultra 253B?
The closest current models from other labs on capability, price and release date are GLM-4.5-Flash (Z.ai (Zhipu), Free / Free), Nova 2 Pro (Amazon, Free / Free), Laguna M.1 (Poolside, Free / Free) and Aya Vision 32B (Cohere).