Muse Glimmer
Released Aug 10, 2026 · knowledge cutoff 2026-01
- Status
- Superseded
- Location
- United States
- Modality
- Multimodal
- Context window
- 131K
- Max output
- —
- Speed
- —
- Price ($/MTok in / out)
- $0.35 / $1.50
- Cost per task
- —
- Open weights
- Yes
Benchmarks5 of 8 reported
GPQA Diamond
SWE-bench Verified
Terminal-Bench 2.1
AIME (latest)
Humanity's Last Exam
MMLU-Pro
LMArena Elo
ARC-AGI-2
Strengths & weaknesses
Strengths
- Runs entirely offline on a single 24–32GB consumer GPU or Apple Silicon Mac, agent loop and all — no account or network connection required
- DFlash speculative decoding proposes 16-token blocks at once, more than tripling raw throughput on an RTX 5090 over standard decoding
- Diagnoses a failed tool call and retries rather than stalling out, trained explicitly for failure recovery across multi-turn agent loops
- Reads interleaved text, screenshots, charts and documents through a dedicated perception encoder, across 100+ languages
Weaknesses
- Trails Qwen3.6-27B on OSWorld-Verified and Terminal-Bench 2.1 — weaker at precise GUI and command-line work than its size-class rival
- No audio input or output, and video is sampled as individual frames rather than processed natively
- Launched compared only against Meta's own chosen size-class peers (Gemma4-31B, Qwen3.6-27B) — no independent LMArena or Artificial Analysis leaderboard placement yet
Distilled from Muse Spark via logit distillation, but Meta's model card doesn't state it replaces Muse Spark, which stays on as Meta's separate, closed flagship. Benchmarks are Meta's own reported 'High Reasoning' figures except GPQA Diamond (83.5) and HLE Text-only (22.0), which the model card itself attributes to Artificial Analysis measurement rather than Meta's own eval. Terminal-Bench 2.1 (51.7%) was run 'with terminus2' scaffold; SWE-bench Pro (51.2%) was also reported but isn't the tracked metric. MMLU-Pro, LMArena Elo and ARC-AGI-2 are unpublished; Artificial Analysis hadn't indexed the model for its Intelligence Index or speed benchmarks as of this check (released 2026-08-10), so costPerTask and speed are null. Pricing shown ($0.35/$1.50 per MTok) is OpenRouter's hosted rate — Meta offers no official API for this open-weight release; Together AI and Fireworks AI also host it.
For developers
API model strings
Not researched
Licence
Apache License 2.0 · Permissive · commercial use permitted
Retirement
No retirement announced
Lineage
No recorded predecessor