ZMIME
Meta · Released Sep 25, 2024

Llama-3.2-11B-Vision-Instruct

Open multimodal Llama model for image understanding, captioning, and visual QA.

  • Open weights
  • Tool calling
  • Vision
Capability
—No ECI score yet
Input
$0.197per 1M tokens
Output
$0.51per 1M tokens
Context
128K4K max output
Pricing

What it costs

Input
$0.197per million tokens
Output
$0.51per million tokens
Blended (3:1)
$0.275cheaper than 72% of priced models

Typical monthly bills

Monthly usageEstimated cost
Side project2M input + 0.5M output tokens$0.649
Team assistant25M input + 5M output tokens$7.47
Production app250M input + 50M output tokens$74.70

Median of 2 paid API providers on models.dev.

Capability

Independent benchmarks

Epoch AI has not published a Capabilities Index score for Llama-3.2-11B-Vision-Instruct yet.

Specifications

The details

Lab
Meta
Released
Sep 25, 2024
Knowledge cutoff
Dec 2023
Context window
128,000 tokens
Max output
4,096 tokens
Inputs
Text, Images
Output
Text
Reasoning
No
Tool calling
Yes
Structured output
No
Availability
2 API providerslisted on models.dev
Same lab

More from Meta

Questions

About Llama-3.2-11B-Vision-Instruct

How much does Llama-3.2-11B-Vision-Instruct cost?

Llama-3.2-11B-Vision-Instruct costs $0.197 input / $0.51 output per million tokens (median across 2 API providers). At a 3:1 input-to-output mix that is $0.275 per million tokens, cheaper than 72% of the 360 priced models we track.

What is the context window of Llama-3.2-11B-Vision-Instruct?

Llama-3.2-11B-Vision-Instruct accepts up to 128,000 tokens per request and can write up to 4,096 tokens in one response.

How good is Llama-3.2-11B-Vision-Instruct?

Epoch AI has not published a Capabilities Index score for Llama-3.2-11B-Vision-Instruct yet.

Is Llama-3.2-11B-Vision-Instruct open source?

Yes. Meta publishes the weights, so it can be downloaded and self-hosted.

What inputs does Llama-3.2-11B-Vision-Instruct support?

Llama-3.2-11B-Vision-Instruct accepts text and images and replies in text. It is not a dedicated reasoning model, supports tool calling.

When was Llama-3.2-11B-Vision-Instruct released?

Meta released Llama-3.2-11B-Vision-Instruct on Sep 25, 2024. Its training data runs to Dec 2023.

What are the best alternatives to Llama-3.2-11B-Vision-Instruct?

The closest current models from other labs on capability, price and release date are Command R (Cohere, $0.15 / $0.60), Qwen-MT Turbo (Alibaba (Qwen), $0.16 / $0.49), Pixtral 12B (Mistral AI, $0.15 / $0.15) and Solar Pro 2 (Upstage, $0.25 / $0.25).