Wait Which Model?
← Back to directory
poolside

Laguna XS 2.1

Released Jul 2, 2026 · knowledge cutoff unpublished

Status
Unknown
Location
United States
Modality
Text
Context window
262K
Max output
33K
Speed
Price ($/MTok in / out)
$0.10 / $0.20
Cost per task
Open weights
Yes
Benchmarks1 of 8 reported

SWE-bench Verified

70.9%

MMLU-Pro

GPQA Diamond

Terminal-Bench 2.1

AIME (latest)

Humanity's Last Exam

LMArena Elo

ARC-AGI-2

Strengths & weaknesses

Strengths

  • Fits on a 36 GB Mac, so a long-horizon coding agent can stay entirely local
  • Thinks between tool calls and lets you switch that off per request
  • Ships FP8, NVFP4 and INT4 checkpoints alongside BF16, plus speculator models that roughly double achieved tokens/sec
  • Licence permits commercial use and modification

Weaknesses

  • A coding specialist — text-only, and neither built nor measured for knowledge, maths or general reasoning
  • Ollama on Apple Silicon can return empty output, a Metal overflow in the MoE down-projection with the fix still pending
  • Evaluated only in poolside's own harness with thinking on, so behavior in other agent frameworks is unmeasured

Smaller sibling of Laguna S 2.1, released three weeks earlier; it replaced Laguna XS.2, which poolside sunset on its API a week after launch. SWE-bench Verified 70.9% is poolside's own figure, mean pass@1 over four attempts with thinking enabled; also reported are SWE-bench Multilingual 63.1%, SWE-Bench Pro public 47.6% and Terminal-Bench 2.0 37.5% — the 2.0 score is not comparable to the 2.1 figures tracked here, so terminalBench stays null. Pricing is poolside's list rate; OpenRouter was serving it at a 40% discount and also offers a free tier. No MMLU-Pro, GPQA, AIME, HLE, LMArena or ARC-AGI-2 score has been published, and knowledge cutoff is undisclosed.

For developers

API model strings

Not researched

Licence

Not researched

Retirement

No retirement announced

Lineage

No recorded predecessor

News