Soofi S 30B-A3B
Released Jul 13, 2026 · knowledge cutoff unpublished
- Status
- Unknown
- Location
- Germany
- Modality
- Text
- Context window
- —
- Max output
- —
- Speed
- —
- Price ($/MTok in / out)
- — / —
- Cost per task
- —
- Open weights
- Yes
Benchmarks2 of 8 reported
MMLU-Pro
GPQA Diamond
SWE-bench Verified
Terminal-Bench 2.1
AIME (latest)
Humanity's Last Exam
LMArena Elo
ARC-AGI-2
Strengths & weaknesses
Strengths
- Throughput stays essentially flat as context grows, where dense-attention models slow down
- German treated as a first-class language rather than a translation target
- Few active parameters per token, so it runs quickly on modest hardware
- Training data and process published in full, which matters for regulated deployments
Weaknesses
- An unaligned base checkpoint — no instruction tuning, no safety tuning, not usable as an assistant without fine-tuning
- Licence terms were not final at release
- A foundation to build on rather than a model that competes on capability
No SWE-bench/AIME/HLE/LMArena/ARC-AGI-2 scores have been published; reported metrics instead cover HumanEval (73.8%), MBPP (70.2%, 84.2% for the German variant) and aggregate English/German scores. Coordinated by the KI Bundesverband and funded under Germany's IPCEI-CIS program; arXiv preprint posted 2026-07-10, press release and Hugging Face weights 2026-07-13. MMLU-Pro 51.4 and GPQA Diamond 43.4 corrected/added from the SOOFI consortium's own technical report (arXiv:2607.09424, Tables 4-5 and Figure 7), for the released base checkpoint iter_1056000 — not published in initial press coverage. The paper's Phase 3 training extends usable context to 1M tokens, but a separate long-context checkpoint is evaluated apart from the base-model benchmark table, so a single released context-window figure could not be confirmed.
For developers
API model strings
Not researched
Licence
Not researched
Retirement
No retirement announced
Lineage
No recorded predecessor