Laguna XS 2.1
Released Jul 2, 2026 · knowledge cutoff unpublished
- Status
- Unknown
- Location
- United States
- Modality
- Text
- Context window
- 262K
- Max output
- 33K
- Speed
- —
- Price ($/MTok in / out)
- $0.10 / $0.20
- Cost per task
- —
- Open weights
- Yes
Benchmarks1 of 8 reported
SWE-bench Verified
MMLU-Pro
GPQA Diamond
Terminal-Bench 2.1
AIME (latest)
Humanity's Last Exam
LMArena Elo
ARC-AGI-2
Strengths & weaknesses
Strengths
- Fits on a 36 GB Mac, so a long-horizon coding agent can stay entirely local
- Thinks between tool calls and lets you switch that off per request
- Ships FP8, NVFP4 and INT4 checkpoints alongside BF16, plus speculator models that roughly double achieved tokens/sec
- Licence permits commercial use and modification
Weaknesses
- A coding specialist — text-only, and neither built nor measured for knowledge, maths or general reasoning
- Ollama on Apple Silicon can return empty output, a Metal overflow in the MoE down-projection with the fix still pending
- Evaluated only in poolside's own harness with thinking on, so behavior in other agent frameworks is unmeasured
Smaller sibling of Laguna S 2.1, released three weeks earlier; it replaced Laguna XS.2, which poolside sunset on its API a week after launch. SWE-bench Verified 70.9% is poolside's own figure, mean pass@1 over four attempts with thinking enabled; also reported are SWE-bench Multilingual 63.1%, SWE-Bench Pro public 47.6% and Terminal-Bench 2.0 37.5% — the 2.0 score is not comparable to the 2.1 figures tracked here, so terminalBench stays null. Pricing is poolside's list rate; OpenRouter was serving it at a 40% discount and also offers a free tier. No MMLU-Pro, GPQA, AIME, HLE, LMArena or ARC-AGI-2 score has been published, and knowledge cutoff is undisclosed.
For developers
API model strings
Not researched
Licence
Not researched
Retirement
No retirement announced
Lineage
No recorded predecessor