ZMIME
NVIDIA · Released Apr 7, 2025

Llama 3.3 Nemotron Super 49B v1

Nemotron model for efficient reasoning, coding, and specialized AI agents.

  • Open weights
  • Reasoning
  • Tool calling
  • Deprecated
Capability
—No ECI score yet
Input
$0.15per 1M tokens
Output
$0.15per 1M tokens
Context
131K131K max output
Pricing

What it costs

Input
$0.15per million tokens
Output
$0.15per million tokens
Blended (3:1)
$0.15cheaper than 83% of priced models

Typical monthly bills

Monthly usageEstimated cost
Side project2M input + 0.5M output tokens$0.375
Team assistant25M input + 5M output tokens$4.50
Production app250M input + 50M output tokens$45.00

Median of 1 paid API provider on models.dev; NVIDIA also offers it free on Nvidia.

Capability

Independent benchmarks

Epoch AI has not published a Capabilities Index score for Llama 3.3 Nemotron Super 49B v1 yet.

Specifications

The details

Lab
NVIDIA
Released
Apr 7, 2025
Context window
131,072 tokens
Max output
131,072 tokens
Inputs
Text
Output
Text
Reasoning
Can be switched on or off
Tool calling
Yes
Structured output
No
Weights
Open weights
API model ID
nvidia/llama-3.3-nemotron-super-49b-v1on Nvidia
Availability
2 API providerslisted on models.dev
Status
Deprecated by the provider

Nvidia model documentation

Same lab

More from NVIDIA

Questions

About Llama 3.3 Nemotron Super 49B v1

How much does Llama 3.3 Nemotron Super 49B v1 cost?

Llama 3.3 Nemotron Super 49B v1 costs $0.15 input / $0.15 output per million tokens (median across 1 API provider; free on Nvidia). At a 3:1 input-to-output mix that is $0.15 per million tokens, cheaper than 83% of the 360 priced models we track.

What is the context window of Llama 3.3 Nemotron Super 49B v1?

Llama 3.3 Nemotron Super 49B v1 accepts up to 131,072 tokens per request and can write up to 131,072 tokens in one response.

How good is Llama 3.3 Nemotron Super 49B v1?

Epoch AI has not published a Capabilities Index score for Llama 3.3 Nemotron Super 49B v1 yet.

Is Llama 3.3 Nemotron Super 49B v1 open source?

Yes. NVIDIA publishes the weights, so it can be downloaded and self-hosted.

What inputs does Llama 3.3 Nemotron Super 49B v1 support?

Llama 3.3 Nemotron Super 49B v1 accepts text and replies in text. It is a reasoning model, supports tool calling.

When was Llama 3.3 Nemotron Super 49B v1 released?

NVIDIA released Llama 3.3 Nemotron Super 49B v1 on Apr 7, 2025.

What are the best alternatives to Llama 3.3 Nemotron Super 49B v1?

The closest current models from other labs on capability, price and release date are Voxtral Small 24B 2507 (Mistral AI, $0.10 / $0.30), Qwen Flash (Alibaba (Qwen), $0.05 / $0.40), Phi-4-mini (Microsoft, $0.075 / $0.30) and Solar Pro 2 (Upstage, $0.25 / $0.25).