ZMIME
NVIDIA · Released Apr 15, 2025

Llama 3.1 Nemotron 70B Instruct

Nemotron model for efficient reasoning, coding, and specialized AI agents.

  • Open weights
  • Tool calling
Capability
—No ECI score yet
Input
$0.478per 1M tokens
Output
$0.504per 1M tokens
Context
128K8K max output
Pricing

What it costs

Input
$0.478per million tokens
Output
$0.504per million tokens
Blended (3:1)
$0.485cheaper than 62% of priced models

Typical monthly bills

Monthly usageEstimated cost
Side project2M input + 0.5M output tokens$1.21
Team assistant25M input + 5M output tokens$14.48
Production app250M input + 50M output tokens$144.82

Median of 2 paid API providers on models.dev; NVIDIA also offers it free on Nvidia.

Capability

Independent benchmarks

Epoch AI has not published a Capabilities Index score for Llama 3.1 Nemotron 70B Instruct yet.

Specifications

The details

Lab
NVIDIA
Released
Apr 15, 2025
Context window
128,000 tokens
Max output
8,192 tokens
Inputs
Text
Output
Text
Reasoning
No
Tool calling
Yes
Structured output
No
Weights
Open weights
API model ID
nvidia/llama-3.1-nemotron-70b-instructon Nvidia
Availability
3 API providerslisted on models.dev

Nvidia model documentation

Same lab

More from NVIDIA

Questions

About Llama 3.1 Nemotron 70B Instruct

How much does Llama 3.1 Nemotron 70B Instruct cost?

Llama 3.1 Nemotron 70B Instruct costs $0.478 input / $0.504 output per million tokens (median across 2 API providers; free on Nvidia). At a 3:1 input-to-output mix that is $0.485 per million tokens, cheaper than 62% of the 360 priced models we track.

What is the context window of Llama 3.1 Nemotron 70B Instruct?

Llama 3.1 Nemotron 70B Instruct accepts up to 128,000 tokens per request and can write up to 8,192 tokens in one response.

How good is Llama 3.1 Nemotron 70B Instruct?

Epoch AI has not published a Capabilities Index score for Llama 3.1 Nemotron 70B Instruct yet.

Is Llama 3.1 Nemotron 70B Instruct open source?

Yes. NVIDIA publishes the weights, so it can be downloaded and self-hosted.

What inputs does Llama 3.1 Nemotron 70B Instruct support?

Llama 3.1 Nemotron 70B Instruct accepts text and replies in text. It is not a dedicated reasoning model, supports tool calling.

When was Llama 3.1 Nemotron 70B Instruct released?

NVIDIA released Llama 3.1 Nemotron 70B Instruct on Apr 15, 2025.

What are the best alternatives to Llama 3.1 Nemotron 70B Instruct?

The closest current models from other labs on capability, price and release date are Qwen3-VL 30B-A3B (Alibaba (Qwen), $0.20 / $0.80), GLM-4.5-Air (Z.ai (Zhipu), $0.20 / $1.10), Seed 1.6 Vision (ByteDance Seed, $0.119 / $1.19) and Gemma-SEA-LION-v4-27B-IT (AI Singapore, $0.351 / $0.555).