Laguna S 2.1
Released Jul 21, 2026 · knowledge cutoff 2025-11
- Status
- Unknown
- Location
- United States
- Modality
- Text
- Context window
- 1M
- Max output
- 131K
- Speed
- —
- Price ($/MTok in / out)
- $0.10 / $0.20
- Cost per task
- —
- Open weights
- Yes
Benchmarks1 of 8 reported
Terminal-Bench 2.1
MMLU-Pro
GPQA Diamond
SWE-bench Verified
AIME (latest)
Humanity's Last Exam
LMArena Elo
ARC-AGI-2
Strengths & weaknesses
Strengths
- Trained to keep checking its work and retry failed approaches instead of giving up mid-run
- Small enough to serve on a single workstation-class box while doing long-horizon coding
- Full evaluation trajectories published, so its behavior can be inspected rather than taken on trust
- Licence permits commercial use and modification
Weaknesses
- A coding specialist — text-only, and it was neither built nor measured for knowledge, maths or general reasoning
- Emits invalid JSON on nested tool calls inside some third-party agent scaffolds
- Tuned closely to its own harness; behavior shifts when moved to a different agent framework
Terminal-Bench 2.1 is the only tracked benchmark reported: poolside published agentic coding evals only, all with thinking enabled — Terminal-Bench 2.1 70.2%, SWE-bench Multilingual 78.5%, SWE-Bench Pro public 59.4%, DeepSWE v1.1 40.4%, SWE Atlas 46.2%, Toolathlon Verified 49.7% (pass@1 over 3-4 attempts). Without thinking, Terminal-Bench 2.1 drops to 60.4% and DeepSWE to 16.5%. SWE-bench Verified is absent from poolside's blog and Hugging Face model card — the Multilingual and Pro figures are different splits and are not recorded as a Verified score. Pricing is the OpenRouter listing served by poolside ($0.10/$0.20 per MTok, $0.01 cached read); poolside publishes no separate per-token price sheet and also offers a free 256K-context tier. Knowledge cutoff of November 2025 is poolside's own, stated in the launch post while ruling out contamination on an Erdős-problem result; the model shares XS 2.1's pre-training data. Max output 131,072 tokens per the OpenRouter endpoint spec. Artificial Analysis does not cover poolside models, so there is no cost-per-task figure.
For developers
API model strings
Not researched
Licence
Not researched
Retirement
No retirement announced
Lineage
No recorded predecessor