Fugu Max
Released Sep 11, 2026 · knowledge cutoff 2026-08
- Status
- Unknown
- Location
- Japan
- Modality
- Multimodal
- Context window
- 1M
- Max output
- 128K
- Speed
- —
- Price ($/MTok in / out)
- $2 / $6
- Cost per task
- —
- Open weights
- No
Benchmarks0 of 10 reported
MMLU-Proretired
GPQA Diamond
SWE-bench Verified
SWE-bench Pro
Terminal-Bench 2.1
AIME (latest)retired
Humanity's Last Exam
LMArena Elo
GDPval-AA v2
ARC-AGI-2
Strengths & weaknesses
Strengths
- Genuinely cost-competitive on output pricing against same-class flagship single models rather than just against other orchestrators
- One OpenAI-compatible endpoint handles function calling, structured outputs and built-in web search/fetch without extra integration work
- Sakana reports it as best-in-class on several of its own coding and document-reasoning evals against single frontier models, not just against its own prior tier
Weaknesses
- Routing is a black box: which underlying model answered a given request is not exposed or auditable, a real gap for compliance-sensitive use
- Orchestration overhead shows up as billed tokens the user never sees — one hands-on test logged the large majority of a coding task's billed tokens as back-channel coordination rather than visible output
- Latency is unpredictable rather than merely high: simple questions have taken well over a minute end to end while the coordinator deliberates regardless of task difficulty
- Too new for independent benchmark verification — every headline number so far is Sakana's own vendor-reported evaluation
Not a single foundation model: a language-model orchestrator (built on Sakana's TRINITY and Conductor research, ICLR 2026) trained to route each request across a fixed, undisclosed pool of open-weight and specialized models — including NVIDIA's Nemotron family — and to recursively call itself, then assemble a final answer behind one API. Closed/proprietary: no weights are published and the composition of the model pool is not disclosed. Cache-hit input is $0.25/M; web_search and web_fetch tool calls are billed separately, reported as $0.007/call by some trackers and as $10/1,000 calls ($0.01/call) by OpenRouter's listing — this entry could not reconcile the two. Sakana's launch post claims best-overall score on six of its own benchmarks (Terminal-Bench 2.1, GPQA Diamond, AA-LCR, GDP.pdf, AutomationBench and SWEFish) against single frontier models, but the actual percentages are rendered only in chart images in the announcement, not as text, so none of this site's tracked benchmark fields could be filled with a specific number; GDP.pdf (a Surge AI PDF/document-reasoning benchmark) is a different, unrelated benchmark from this site's gdpvalAA (Artificial Analysis' GDPval-AA v2), so no mapping was made between them. No Artificial Analysis Intelligence Index coverage exists for any Fugu model as of launch, hence costPerTask and speed are null. Knowledge cutoff (2026-08) is sourced from OpenRouter's model listing, not confirmed on an official Sakana page — treat as third-party-sourced.
For developers
API model strings
- sakana
fugu-max - openrouter
sakana/fugu-max
Licence
Proprietary — weights not released
Retirement
No retirement announced
Lineage
No recorded predecessor