Llama 3.1 Nemotron 70B Instruct
Nemotron model for efficient reasoning, coding, and specialized AI agents.
- Capability
- —No ECI score yet
- Input
- $0.478per 1M tokens
- Output
- $0.504per 1M tokens
- Context
- 128K8K max output
What it costs
- Input
- $0.478per million tokens
- Output
- $0.504per million tokens
- Blended (3:1)
- $0.485cheaper than 62% of priced models
Typical monthly bills
| Monthly usage | Estimated cost |
|---|---|
| Side project2M input + 0.5M output tokens | $1.21 |
| Team assistant25M input + 5M output tokens | $14.48 |
| Production app250M input + 50M output tokens | $144.82 |
Median of 2 paid API providers on models.dev; NVIDIA also offers it free on Nvidia.
Independent benchmarks
Epoch AI has not published a Capabilities Index score for Llama 3.1 Nemotron 70B Instruct yet.
The details
- Lab
- NVIDIA
- Released
- Apr 15, 2025
- Context window
- 128,000 tokens
- Max output
- 8,192 tokens
- Inputs
- Text
- Output
- Text
- Reasoning
- No
- Tool calling
- Yes
- Structured output
- No
- Weights
- Open weights
- API model ID
nvidia/llama-3.1-nemotron-70b-instructon Nvidia- Availability
- 3 API providerslisted on models.dev
More from NVIDIA
About Llama 3.1 Nemotron 70B Instruct
How much does Llama 3.1 Nemotron 70B Instruct cost?
Llama 3.1 Nemotron 70B Instruct costs $0.478 input / $0.504 output per million tokens (median across 2 API providers; free on Nvidia). At a 3:1 input-to-output mix that is $0.485 per million tokens, cheaper than 62% of the 360 priced models we track.
What is the context window of Llama 3.1 Nemotron 70B Instruct?
Llama 3.1 Nemotron 70B Instruct accepts up to 128,000 tokens per request and can write up to 8,192 tokens in one response.
How good is Llama 3.1 Nemotron 70B Instruct?
Epoch AI has not published a Capabilities Index score for Llama 3.1 Nemotron 70B Instruct yet.
Is Llama 3.1 Nemotron 70B Instruct open source?
Yes. NVIDIA publishes the weights, so it can be downloaded and self-hosted.
What inputs does Llama 3.1 Nemotron 70B Instruct support?
Llama 3.1 Nemotron 70B Instruct accepts text and replies in text. It is not a dedicated reasoning model, supports tool calling.
When was Llama 3.1 Nemotron 70B Instruct released?
NVIDIA released Llama 3.1 Nemotron 70B Instruct on Apr 15, 2025.
What are the best alternatives to Llama 3.1 Nemotron 70B Instruct?
The closest current models from other labs on capability, price and release date are Qwen3-VL 30B-A3B (Alibaba (Qwen), $0.20 / $0.80), GLM-4.5-Air (Z.ai (Zhipu), $0.20 / $1.10), Seed 1.6 Vision (ByteDance Seed, $0.119 / $1.19) and Gemma-SEA-LION-v4-27B-IT (AI Singapore, $0.351 / $0.555).