Wait Which Model?
← Back to directory
Thinking Machines Lab

Inkling

Released Jul 15, 2026 · knowledge cutoff unpublished

Status
Superseded
Location
United States
Modality
Multimodal
Context window
1M
Max output
Speed
Price ($/MTok in / out)
$1.87 / $4.68
Cost per task
Open weights
Yes
Benchmarks7 of 8 reported

GPQA Diamond

87.2%

SWE-bench Verified

77.6%

Terminal-Bench 2.1

63.8%

AIME (latest)

97.1%

Humanity's Last Exam

29.7%

LMArena Elo

1441

ARC-AGI-2

36.5%

MMLU-Pro

Strengths & weaknesses

Strengths

  • Built to be fine-tuned — same-day tuning support and an unrestricted licence
  • Takes images and long audio natively, still unusual for an open-weights model
  • Open at frontier scale from a US lab, with no regional restrictions

Weaknesses

  • Weak factuality — it hallucinates at a high rate and needs grounding for anything factual
  • Degrades in long multi-turn conversations
  • Positioned as a base to customise, not a model to deploy as it ships

Thinking Machines Lab's first production model, released 2026-07-15. All benchmarks are the lab's own model-card figures at effort=0.99; HLE is 29.7% text-only / 46.0% with tools, and the model card reports Global-MMLU-Lite 88.7% rather than MMLU-Pro, so MMLU-Pro, LMArena Elo and ARC-AGI-2 are left null. Pricing is Tinker's 64K-context tier under a limited-time 50% discount ($1.87/$4.68 per MTok, $0.374 cached); the 256K tier is $3.74/$9.36, and OpenRouter hosts list about $1.00/$4.05. Context is 1M with the open weights but capped at 256K on the Tinker API; no max output figure is published. Also reported: SWE-Bench Pro Public 54.3%, Terminal-Bench 2.1 63.8%, GDPval-AA v2 Elo 1238, MMMU Pro 73.5%. ARC-AGI-2 36.5% (and ARC-AGI-1 79.5%) and LMArena Elo 1441 were added later, from the Inkling-Small comparison table and the LMArena text leaderboard respectively.

For developers

API model strings

Not researched

Licence

Not researched

Retirement

No retirement announced

Lineage

No recorded predecessor

News