Wait Which Model?
← Back to directory
Soofi

Soofi S 30B-A3B

Released Jul 13, 2026 · knowledge cutoff unpublished

Status
Unknown
Location
Germany
Modality
Text
Context window
Max output
Speed
Price ($/MTok in / out)
— / —
Cost per task
Open weights
Yes
Benchmarks2 of 8 reported

MMLU-Pro

51.4%

GPQA Diamond

43.4%

SWE-bench Verified

Terminal-Bench 2.1

AIME (latest)

Humanity's Last Exam

LMArena Elo

ARC-AGI-2

Strengths & weaknesses

Strengths

  • Throughput stays essentially flat as context grows, where dense-attention models slow down
  • German treated as a first-class language rather than a translation target
  • Few active parameters per token, so it runs quickly on modest hardware
  • Training data and process published in full, which matters for regulated deployments

Weaknesses

  • An unaligned base checkpoint — no instruction tuning, no safety tuning, not usable as an assistant without fine-tuning
  • Licence terms were not final at release
  • A foundation to build on rather than a model that competes on capability

No SWE-bench/AIME/HLE/LMArena/ARC-AGI-2 scores have been published; reported metrics instead cover HumanEval (73.8%), MBPP (70.2%, 84.2% for the German variant) and aggregate English/German scores. Coordinated by the KI Bundesverband and funded under Germany's IPCEI-CIS program; arXiv preprint posted 2026-07-10, press release and Hugging Face weights 2026-07-13. MMLU-Pro 51.4 and GPQA Diamond 43.4 corrected/added from the SOOFI consortium's own technical report (arXiv:2607.09424, Tables 4-5 and Figure 7), for the released base checkpoint iter_1056000 — not published in initial press coverage. The paper's Phase 3 training extends usable context to 1M tokens, but a separate long-context checkpoint is evaluated apart from the base-model benchmark table, so a single released context-window figure could not be confirmed.

For developers

API model strings

Not researched

Licence

Not researched

Retirement

No retirement announced

Lineage

No recorded predecessor

News