Wait Which Model?
← Back to directory
Thinking Machines Lab

Inkling-Small

Released Jul 30, 2026 · knowledge cutoff unpublished

Status
Superseded
Location
United States
Modality
Multimodal
Context window
1M
Max output
Speed
Price ($/MTok in / out)
$0.50 / $1.20
Cost per task
$0.073
Open weights
Yes
Benchmarks7 of 8 reported

GPQA Diamond

89.5%

SWE-bench Verified

80.2%

Terminal-Bench 2.1

64.7%

AIME (latest)

95.5%

Humanity's Last Exam

31.6%

LMArena Elo

1431

ARC-AGI-2

40.1%

MMLU-Pro

Strengths & weaknesses

Strengths

  • Thinking effort sweeps from minimal to xhigh, and the whole performance curve sits above its much larger sibling's
  • Reads audio and images through the same encoder-free stack, and will drive Python to crop and zoom into a chart it cannot read directly
  • Overtook the full-size Inkling on reasoning and agentic coding after two extra weeks of RL on a distilled checkpoint
  • Runs cleanly across third-party coding and agent harnesses, not just its own

Weaknesses

  • Factually thin — it covers less ground than Inkling and still needs retrieval and human review for anything consequential
  • Degrades in long multi-turn conversations, and the model card concedes it sometimes ignores instructions outright
  • Multi-step tool use went backwards versus Inkling even as the reasoning scores went up

Thinking Machines Lab's second release, a 276B/12B MoE trained on GB300 systems. All benchmarks are the lab's own model-card figures at effort 0.99; SWE-bench Verified uses a bash-only harness and Terminal-Bench 2.1 an internal one with web-contaminated solutions scored zero. HLE 31.6% is text-only no tools / 47.8% with tools. MMLU-Pro is not reported (Global-MMLU-Lite 86.7% instead). Also reported: SWE-bench Pro public 55.9%, GDPval-AA v2 Elo 1269, MMMU Pro 74.0%, IFBench 82.2%, ARC-AGI-1 84.0%. Context is 1M with the open weights but OpenRouter serves 512K; Thinking Machines publishes only the $1.20 output price, and $0.50 input is the OpenRouter listing (Artificial Analysis lists $0.30). LMArena Elo 1431 is an AutoEval rating.

For developers

API model strings

Not researched

Licence

Not researched

Retirement

No retirement announced

Lineage

No recorded predecessor

News