Wait Which Model?
← Back to directory
Sakana AI

Fugu Ultra v2.0

Released Sep 11, 2026 · knowledge cutoff 2026-08

Status
Unknown
Location
Japan
Modality
Multimodal
Context window
1M
Max output
128K
Speed
Price ($/MTok in / out)
$5 / $30
Cost per task
Open weights
No
Benchmarks0 of 10 reported

MMLU-Proretired

GPQA Diamond

SWE-bench Verified

SWE-bench Pro

Terminal-Bench 2.1

AIME (latest)retired

Humanity's Last Exam

LMArena Elo

GDPval-AA v2

ARC-AGI-2

Strengths & weaknesses

Strengths

  • Configurable reasoning effort (high/xhigh/max) plus function calling, structured outputs and built-in web search/fetch in one endpoint, aimed at autonomous research and full-stack agentic work
  • Sakana reports it best or joint-best on five of its own eight tracked evals against single frontier models on multi-step reasoning and document-heavy tasks
  • A smaller, more selective model pool than the original Fugu Ultra, traded for a reported score increase rather than broader coverage

Weaknesses

  • Orchestration tax is severe at this tier: one first-touch coding test logged roughly nine times as many hidden orchestration tokens as visible output, and took several minutes
  • Routing decisions are opaque end to end — no way to see or audit which pool model handled a given part of a request
  • A heavy user can exhaust a fixed monthly budget in a few hours of sustained use; cost is unpredictable by design rather than merely high
  • Too new for independent benchmarking — no Artificial Analysis or third-party leaderboard coverage exists yet, only Sakana's own reported figures

Not a single foundation model: the higher-capability tier of Sakana's TRINITY/Conductor-derived (ICLR 2026) orchestration architecture, routing each request across a fixed, undisclosed pool of open-weight and specialized models and recursively calling itself, versioned v2.0 as the successor line to the original Fugu Ultra (launched 2026-06-22, later v1.1) — neither earlier build is tracked in this dataset, so predecessorId stays null. Closed/proprietary: no weights published, model-pool composition undisclosed. Pricing rises to $10/$45 per million tokens (and $1.00/M cached input, from $0.50/M) above 272,000 tokens of context. Sakana's launch post reports it best or joint-best on five of eight of its own benchmarks, including 74.3 on DeepSWE and 48.3 on Chartography (vs. 27.3 for Claude Opus 5 and 29.5 for Claude Fable 5 on the latter) — neither DeepSWE nor Chartography maps to a benchmark tracked on this site, and this site's ten tracked keys (GPQA Diamond, Terminal-Bench 2.1, SWE-bench Pro, HLE, etc., all reportedly among Fugu Ultra's strong evals per Sakana) appear only in the announcement's chart images, not as extractable text, so all stay null here. No Artificial Analysis coverage exists for any Fugu model, hence costPerTask and speed are null. Knowledge cutoff (2026-08) is sourced from OpenRouter's model listing, not an official Sakana page — treat as third-party-sourced.

For developers

API model strings

  • sakanafugu-ultra-v2
  • openroutersakana/fugu-ultra-v2

Licence

Proprietary — weights not released

Retirement

No retirement announced

Lineage

No recorded predecessor

News