ZMIME
Thinking Machines · Released Jul 15, 2026

Inkling

Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio.

  • Open weights
  • Reasoning
  • Tool calling
  • Vision
  • Audio input
Capability
148.6ECI · #57 of 148
Input
$3.74per 1M tokens
Output
$9.36per 1M tokens
Context
1.05M1.05M max output
Pricing

What it costs

Input
$3.74per million tokens
Output
$9.36per million tokens
Cached input
$0.748per million tokens
Blended (3:1)
$5.14pricier than 89% of priced models

Typical monthly bills

Monthly usageEstimated cost
Side project2M input + 0.5M output tokens$12.16
Team assistant25M input + 5M output tokens$140.30
Production app250M input + 50M output tokens$1,403

Official Thinking Machines API price, as listed on models.dev.

Capability

Independent benchmarks

148.6Epoch Capabilities Index
#57 of 148 scored models
88Median 146167

The shaded band is Epoch AI’s confidence range (145.7–150.6); the tick marks the median scored model.

  • GPQA DiamondGraduate-level science questions · xhigh setting88.3%
  • FrontierMath Tiers 1–3Research-level mathematics · xhigh setting33.3%
  • OTIS Mock AIME 2024–2025Competition mathematics · xhigh setting88.9%
  • SimpleQA VerifiedShort factual questions · xhigh setting40.3%

Scores from Epoch AI, run independently of Thinking Machines.

Specifications

The details

Lab
Thinking Machines
Released
Jul 15, 2026
Context window
1,048,576 tokens
Max output
1,048,576 tokens
Inputs
Text, Images, Audio
Output
Text
Reasoning
Adjustable effort low · medium · high · xhigh · max
Tool calling
Yes
Structured output
No
API model ID
thinkingmachines/Inkling:peft:262144on Thinking Machines
Availability
23 API providerslisted on models.dev

Thinking Machines model documentation

Same lab

More from Thinking Machines

Questions

About Inkling

How much does Inkling cost?

Inkling costs $3.74 input / $9.36 output per million tokens (official Thinking Machines API price). Cached input is $0.748 per million tokens. At a 3:1 input-to-output mix that is $5.14 per million tokens, more expensive than 89% of the 360 priced models we track.

What is the context window of Inkling?

Inkling accepts up to 1,048,576 tokens per request and can write up to 1,048,576 tokens in one response.

How good is Inkling?

Epoch AI gives Inkling a Capabilities Index score of 148.6 (likely range 145.7–150.6), ranking it #57 of 148 models Epoch has scored. Epoch AI benchmark results: GPQA Diamond 88.3%, FrontierMath Tiers 1–3 33.3%, OTIS Mock AIME 2024–2025 88.9%, SimpleQA Verified 40.3%.

Is Inkling open source?

Yes. Thinking Machines publishes the weights under the Apache-2.0 licence, so it can be downloaded and self-hosted.

What inputs does Inkling support?

Inkling accepts text, images and audio and replies in text. It is a reasoning model with low, medium, high, xhigh and max effort settings, supports tool calling.

When was Inkling released?

Thinking Machines released Inkling on Jul 15, 2026.

What are the best alternatives to Inkling?

The closest current models from other labs on capability, price and release date are Qwen3.6 Max Preview (Alibaba (Qwen), ECI 149.2, $1.30 / $7.80), Claude Sonnet 4.6 (Anthropic, ECI 152.3, $3.00 / $15.00), GPT-5.1 (OpenAI, ECI 149.7, $1.25 / $10.00) and GLM-5.1 (Z.ai (Zhipu), ECI 149.9, $1.40 / $4.40).