ZMIME
Vispark · Released May 15, 2024

Vision Small

Fast, low-cost multimodal model for understanding text, images, audio, video, and PDFs, with tool calling and a 1M-token context window.

  • Proprietary
  • Reasoning
  • Tool calling
  • Vision
  • Audio input
  • Video input
Capability
—No ECI score yet
Input
—No listed price
Output
—No listed price
Context
1M66K max output
Pricing

What it costs

No per-token price is published for this model yet.

No per-token price is listed for this model on models.dev.

Capability

Independent benchmarks

Epoch AI has not published a Capabilities Index score for Vision Small yet.

Specifications

The details

Lab
Vispark
Released
May 15, 2024
Context window
1,000,000 tokens
Max output
65,536 tokens
Inputs
Text, Images, PDFs, Audio, Video
Output
Text
Reasoning
Always on
Tool calling
Yes
Structured output
Yes
Weights
Proprietary
Same lab

More from Vispark

Questions

About Vision Small

How much does Vision Small cost?

Vision Small has no published per-token API price on models.dev yet.

What is the context window of Vision Small?

Vision Small accepts up to 1,000,000 tokens per request and can write up to 65,536 tokens in one response.

How good is Vision Small?

Epoch AI has not published a Capabilities Index score for Vision Small yet.

Is Vision Small open source?

No. Vision Small is proprietary; you use it through an API or partner platforms.

What inputs does Vision Small support?

Vision Small accepts text, images, PDFs, audio and video and replies in text. It is a reasoning model, supports tool calling and can return structured JSON output.

When was Vision Small released?

Vispark released Vision Small on May 15, 2024.

What are the best alternatives to Vision Small?

The closest current models from other labs on capability, price and release date are Codestral (Mistral AI, $0.30 / $0.90), solar-mini (Upstage, $0.15 / $0.15), Qwen-VL Max (Alibaba (Qwen), $0.80 / $3.20) and o3-deep-research (OpenAI).