ZMIME
NVIDIA · Released Jul 25, 2025

Llama 3.3 Nemotron Super 49B v1.5

Nemotron model for efficient reasoning, coding, and specialized AI agents.

  • Open weights
  • Reasoning
  • Tool calling
  • Deprecated
Capability
—No ECI score yet
Input
$0.40per 1M tokens
Output
$0.40per 1M tokens
Context
131K131K max output
Pricing

What it costs

Input
$0.40per million tokens
Output
$0.40per million tokens
Blended (3:1)
$0.40cheaper than 66% of priced models

Typical monthly bills

Monthly usageEstimated cost
Side project2M input + 0.5M output tokens$1.00
Team assistant25M input + 5M output tokens$12.00
Production app250M input + 50M output tokens$120.00

Median of 1 paid API provider on models.dev; NVIDIA also offers it free on Nvidia.

Capability

Independent benchmarks

Epoch AI has not published a Capabilities Index score for Llama 3.3 Nemotron Super 49B v1.5 yet.

Specifications

The details

Lab
NVIDIA
Released
Jul 25, 2025
Context window
131,072 tokens
Max output
131,072 tokens
Inputs
Text
Output
Text
Reasoning
Can be switched on or off
Tool calling
Yes
Structured output
No
Weights
Open weights
API model ID
nvidia/llama-3.3-nemotron-super-49b-v1.5on Nvidia
Availability
2 API providerslisted on models.dev
Status
Deprecated by the provider

Nvidia model documentation

Same lab

More from NVIDIA

Questions

About Llama 3.3 Nemotron Super 49B v1.5

How much does Llama 3.3 Nemotron Super 49B v1.5 cost?

Llama 3.3 Nemotron Super 49B v1.5 costs $0.40 input / $0.40 output per million tokens (median across 1 API provider; free on Nvidia). At a 3:1 input-to-output mix that is $0.40 per million tokens, cheaper than 66% of the 360 priced models we track.

What is the context window of Llama 3.3 Nemotron Super 49B v1.5?

Llama 3.3 Nemotron Super 49B v1.5 accepts up to 131,072 tokens per request and can write up to 131,072 tokens in one response.

How good is Llama 3.3 Nemotron Super 49B v1.5?

Epoch AI has not published a Capabilities Index score for Llama 3.3 Nemotron Super 49B v1.5 yet.

Is Llama 3.3 Nemotron Super 49B v1.5 open source?

Yes. NVIDIA publishes the weights, so it can be downloaded and self-hosted.

What inputs does Llama 3.3 Nemotron Super 49B v1.5 support?

Llama 3.3 Nemotron Super 49B v1.5 accepts text and replies in text. It is a reasoning model, supports tool calling.

When was Llama 3.3 Nemotron Super 49B v1.5 released?

NVIDIA released Llama 3.3 Nemotron Super 49B v1.5 on Jul 25, 2025.

What are the best alternatives to Llama 3.3 Nemotron Super 49B v1.5?

The closest current models from other labs on capability, price and release date are GLM-4.5-Air (Z.ai (Zhipu), $0.20 / $1.10), Seed 1.6 Vision (ByteDance Seed, $0.119 / $1.19), Gemma-SEA-LION-v4-27B-IT (AI Singapore, $0.351 / $0.555) and Qwen3 Coder Flash (Alibaba (Qwen), $0.30 / $1.50).