Wait Which Model?
← Back to directory
poolside

Laguna S 2.1

Released Jul 21, 2026 · knowledge cutoff 2025-11

Status
Unknown
Location
United States
Modality
Text
Context window
1M
Max output
131K
Speed
Price ($/MTok in / out)
$0.10 / $0.20
Cost per task
Open weights
Yes
Benchmarks1 of 8 reported

Terminal-Bench 2.1

70.2%

MMLU-Pro

GPQA Diamond

SWE-bench Verified

AIME (latest)

Humanity's Last Exam

LMArena Elo

ARC-AGI-2

Strengths & weaknesses

Strengths

  • Trained to keep checking its work and retry failed approaches instead of giving up mid-run
  • Small enough to serve on a single workstation-class box while doing long-horizon coding
  • Full evaluation trajectories published, so its behavior can be inspected rather than taken on trust
  • Licence permits commercial use and modification

Weaknesses

  • A coding specialist — text-only, and it was neither built nor measured for knowledge, maths or general reasoning
  • Emits invalid JSON on nested tool calls inside some third-party agent scaffolds
  • Tuned closely to its own harness; behavior shifts when moved to a different agent framework

Terminal-Bench 2.1 is the only tracked benchmark reported: poolside published agentic coding evals only, all with thinking enabled — Terminal-Bench 2.1 70.2%, SWE-bench Multilingual 78.5%, SWE-Bench Pro public 59.4%, DeepSWE v1.1 40.4%, SWE Atlas 46.2%, Toolathlon Verified 49.7% (pass@1 over 3-4 attempts). Without thinking, Terminal-Bench 2.1 drops to 60.4% and DeepSWE to 16.5%. SWE-bench Verified is absent from poolside's blog and Hugging Face model card — the Multilingual and Pro figures are different splits and are not recorded as a Verified score. Pricing is the OpenRouter listing served by poolside ($0.10/$0.20 per MTok, $0.01 cached read); poolside publishes no separate per-token price sheet and also offers a free 256K-context tier. Knowledge cutoff of November 2025 is poolside's own, stated in the launch post while ruling out contamination on an Erdős-problem result; the model shares XS 2.1's pre-training data. Max output 131,072 tokens per the OpenRouter endpoint spec. Artificial Analysis does not cover poolside models, so there is no cost-per-task figure.

For developers

API model strings

Not researched

Licence

Not researched

Retirement

No retirement announced

Lineage

No recorded predecessor

News