Model news
The frontier log.
Releases, benchmark milestones, company moves, research results and policy — everything that shifts the frontier, in one dated record.
July 2026
- Jul 16releaseMoonshot AI
Moonshot AI launches Kimi K3
Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-weight MoE model with a 1M-token context window and native visual understanding, priced at $3/$15 per million input/output tokens; full weights are due July 27, 2026.
VentureBeat ↗ - Jul 13releaseSoofi
German SOOFI consortium releases Soofi S, a sovereign open-weight MoE model
A consortium led by the KI Bundesverband — with Fraunhofer, DFKI, TU Darmstadt and others — released Soofi S 30B-A3B, a hybrid Mamba-2/MoE model trained on 27T tokens with up-weighted German, as a preview checkpoint on Hugging Face.
Fraunhofer IAIS ↗ - Jul 9releaseOpenAI
OpenAI ships GPT-5.6 (Sol, Terra, Luna) to the public
OpenAI lifted the government-gated preview and launched GPT-5.6 broadly; flagship Sol posts record GPQA Diamond (94.6%) and ARC-AGI-2 (92.5% at max effort) while OpenAI calls it its strongest cybersecurity model yet.
TechCrunch ↗ - Jul 1policyxAI
Colorado guts its AI Act as xAI and DOJ challenge proceeds
After xAI sued to block SB24-205 and the DOJ intervened — its first move against a state AI law — Governor Polis signed SB 26-189, stripping algorithmic-discrimination duties and bias audits in favor of consumer disclosure.
US Department of Justice ↗
June 2026
- Jun 26releaseOpenAI
OpenAI previews GPT-5.6 Sol, Terra and Luna under government-gated release
GPT-5.6 opens to ~20 government-approved companies at the administration's behest: Sol is the flagship, Terra runs ~2x cheaper than GPT-5.5, and Luna is the low-cost tier. General availability is promised 'in coming weeks.'
Axios ↗ - Jun 15policy
EU appoints scientific panel ahead of AI Act penalty powers
The European Commission named 60 independent frontier-AI experts to support the AI Office before enforcement penalties activate on August 2, 2026, shifting the AI Act from infrastructure-building to active enforcement.
European Commission ↗ - Jun 15benchmarkAnthropicOpenAIGoogle DeepMind
LMArena top tier compresses to its tightest spread on record
The top five models — Claude Opus 4.8 (~1510), GPT-5.5 Pro, Gemini 3.1 Pro, Claude Opus 4.7 and GPT-5.5 — sit within ~55 Elo, with the top ten inside ~20 points; task fit now matters more than leaderboard rank.
Swfte AI Leaderboard ↗ - Jun 13releaseZhipu AI
Zhipu AI releases GLM-5.2, an MIT-licensed open-weights coding flagship
Z.ai's GLM-5.2 (753B/40B-active MoE) shipped with a 1M-token context window and beat GPT-5.5 on SWE-bench Pro (62.1 vs 58.6) at roughly one-sixth the API cost, topping the Artificial Analysis open-weights index.
VentureBeat ↗ - Jun 13policyAnthropic
US bars foreign nationals from accessing Fable 5 and Mythos 5
Two days after Anthropic's distillation-throttling apology, the administration barred foreign nationals from Anthropic's two newest frontier models, citing national security.
Council on Foreign Relations ↗ - Jun 11companyAnthropic
Anthropic apologizes for Fable 5 silently throttling suspected distillation
Anthropic apologized after Claude Fable 5 was found silently limiting responses to users suspected of trying to replicate its capabilities, and faced criticism for over-refusing cyber-related queries.
Council on Foreign Relations ↗ - Jun 9releaseAnthropic
Anthropic launches Claude Fable 5 and restricted Claude Mythos 5
Fable 5 becomes the first generally available Mythos-class model ($10/$50 per MTok, state of the art on nearly all benchmarks); Mythos 5 — the same weights with safeguards lifted — goes to vetted cyberdefenders via Project Glasswing.
Anthropic ↗ - Jun 2policyOpenAIAnthropicGoogle DeepMindxAIMeta AI
White House issues executive order on frontier AI innovation and security
The order creates a voluntary pre-release federal review process for frontier models with 30-day government access, plus early access for critical-infrastructure entities, framed around cybersecurity.
The White House ↗
May 2026
- May 28releaseAnthropic
Anthropic releases Claude Opus 4.8
Opus 4.8 ships at $5/$25 per MTok with 88.6% on SWE-bench Verified, a USAMO jump to 96.7%, a new fast mode, and the #1 LMArena spot (~1510 Elo). Reviewers call it a modest but tangible improvement.
Anthropic ↗ - May 28companyAnthropicOpenAI
Anthropic raises $65B at a $965B valuation, overtaking OpenAI as most valuable AI startup
The Series H led by Altimeter, Dragoneer, Greenoaks and Sequoia nearly triples February's $380B valuation; Anthropic reports a $47B revenue run rate driven largely by Claude Code.
CNBC ↗ - May 22companyOpenAI
OpenAI files confidential S-1 for IPO
OpenAI submits a confidential draft registration to the SEC after a record $122B March round at ~$852B; later reports suggest it may wait until 2027 to list.
TechJournal ↗ - May 22researchOpenAI
OpenAI reasoning model disproves Erdős-linked conjecture
A general-purpose OpenAI reasoning model disproved a central conjecture tied to Erdős's 1946 planar unit-distance problem, finding an infinite family of counterexample point arrangements verified by outside mathematicians.
Forbes ↗ - May 20releaseAlibaba (Qwen)
Alibaba unveils Qwen3.7 Max, its first closed-weight flagship
Announced at the Alibaba Cloud Summit, Qwen3.7 Max posts the highest Artificial Analysis index score ever for a Chinese model (56.6) — but ships without open weights, a notable strategy shift.
Digital Applied ↗ - May 19releaseGoogle DeepMind
Google launches Gemini 3.5 Flash at I/O 2026
Gemini 3.5 Flash ($1.50/$9 per MTok) beats Gemini 3.1 Pro on coding and agentic benchmarks at ~25% lower cost and up to 4x the speed.
MarkTechPost ↗ - May 11release
Thinking Machines Lab previews near-realtime voice-and-video 'interaction models'
Mira Murati's lab unveils TML-Interaction-Small (276B MoE, 12B active) with 0.4s turn-taking latency that listens while it talks — its first model release, in limited research preview.
TechCrunch ↗ - May 1companyDeepSeekAlibaba (Qwen)
DeepSeek V4 triggers scramble for Huawei Ascend 950 chips
V4's optimization for Huawei Ascend rather than Nvidia hardware prompts ByteDance, Tencent and Alibaba to rush Ascend 950 orders; SMIC shares jump 10% while US export controls constrain supply.
Capacity ↗
April 2026
- Apr 30releasexAI
xAI launches Grok 4.3 with native video input and a 40% price cut
Grok 4.3 hits the public API at $1.25/$2.50 per MTok — the cheapest near-frontier flagship — while Grok 5, the 6T-parameter Colossus 2 model, slips again with no release date.
Artificial Analysis ↗ - Apr 24releaseDeepSeek
DeepSeek releases open-weights V4 family, tying Gemini 3.1 Pro on SWE-bench
V4-Pro-Max posts 80.6% on SWE-bench Verified — the best open-weights score — with a 1M context and 384K max output at $0.435/$0.87 per MTok under MIT license.
MorphLLM ↗ - Apr 23releaseOpenAI
OpenAI ships GPT-5.5 with an 85% ARC-AGI-2 score
GPT-5.5 posts an 11.7-point ARC-AGI-2 jump over GPT-5.4 and takes Terminal-Bench 2.0 state of the art; the $30/$180 GPT-5.5 Pro variant targets long-horizon research.
OpenAI ↗ - Apr 8releaseMeta AI
Meta releases open-weights Llama 5 and closed flagship Muse Spark on the same day
Llama 5 (600B+, 5M-token context, trained on 500K+ B200s) renews Meta's open-weights push, while Muse Spark — the first closed model from Meta Superintelligence Labs — becomes its consumer flagship.
Meta AI ↗ - Apr 2releaseGoogle DeepMind
Google DeepMind releases Gemma 4, its first multimodal open-weights line
Gemma 4 ships in five Apache 2.0-licensed sizes (E2B to 31B), adding native image/video (and audio on the smaller tiers) input; the 31B flagship ranked #3 on the LMArena text leaderboard at release.
Google ↗
February 2026
- Feb 11releaseGoogle DeepMind
Google previews Gemini 3.1 Pro with a 77% ARC-AGI-2 score
Gemini 3.1 Pro more than doubles its predecessor's abstract-reasoning performance and takes the GPQA Diamond lead at unchanged $2/$12 pricing.
Google DeepMind ↗
November 2025
- Nov 24releaseAnthropic
Anthropic releases Claude Opus 4.5, first model past 80% on SWE-bench Verified
Opus 4.5 tops coding benchmarks while cutting flagship pricing 3x to $5/$25 per million tokens, capping a three-week stretch in which GPT-5.1, Gemini 3 Pro and Grok 4.1 all shipped.
Anthropic ↗ - Nov 18releaseGoogle DeepMind
Google ships Gemini 3 Pro with record HLE and ARC-AGI-2 scores
Gemini 3 Pro posts 37.5% on Humanity's Last Exam and becomes the first model rated above 1500 Elo on LMArena; it rolls out to Search's AI Mode on day one.
Google DeepMind ↗ - Nov 17releasexAI
xAI's Grok 4.1 takes the top LMArena spot
Grok 4.1 debuts at #1 on the LMArena text leaderboard with sharply reduced hallucination rates and improved writing over Grok 4.
xAI ↗ - Nov 12releaseOpenAI
OpenAI releases GPT-5.1 with adaptive reasoning
GPT-5.1 Instant and Thinking replace GPT-5 in ChatGPT, spending reasoning tokens only when a task needs them and shipping a warmer default personality after months of user complaints.
OpenAI ↗ - Nov 6releaseMoonshot AI
Kimi K2 Thinking sets open-weights records on agentic benchmarks
Moonshot AI's 1T-parameter open reasoning model beats GPT-5 on agentic Humanity's Last Exam and holds 200+ sequential tool calls, reportedly trained for ~$4.6M.
Moonshot AI ↗
October 2025
- Oct 15releaseAnthropic
Claude Haiku 4.5 brings near-frontier coding to the small-model tier
Anthropic's small model matches Sonnet 4's coding performance at a third of the cost and more than twice the speed, at $1/$5 per million tokens.
Anthropic ↗
September 2025
- Sep 29releaseAnthropic
Claude Sonnet 4.5 claims best-coding-model title
Sonnet 4.5 posts 77.2% on SWE-bench Verified and sustains 30+ hour autonomous coding sessions, shipping alongside the Claude Agent SDK.
Anthropic ↗ - Sep 29releaseDeepSeek
DeepSeek-V3.2 halves long-context API prices with sparse attention
DeepSeek's experimental sparse-attention release cuts API prices to $0.28/$0.42 per million tokens, keeping open-weights pressure on frontier pricing.
DeepSeek ↗
August 2025
- Aug 7releaseOpenAI
OpenAI launches GPT-5 as a unified system
GPT-5 routes between fast and reasoning modes automatically and reaches 74.9% on SWE-bench Verified; a bumpy rollout forces OpenAI to restore GPT-4o for paying users within days.
OpenAI ↗
July 2025
- Jul 11releaseMoonshot AI
Moonshot AI open-sources trillion-parameter Kimi K2
Kimi K2 becomes the strongest open non-reasoning model, purpose-built for agentic tool use, intensifying the China open-weights wave.
Moonshot AI ↗ - Jul 9releasexAI
xAI debuts Grok 4 with leading HLE and ARC-AGI-2 scores
Grok 4 and multi-agent Grok 4 Heavy lead Humanity's Last Exam and ARC-AGI-2 at launch, days after high-profile Grok chatbot safety failures on X.
xAI ↗
January 2025
- Jan 20releaseDeepSeek
DeepSeek-R1 shocks the market with open o1-class reasoning
The MIT-licensed reasoning model matches OpenAI's o1 on math and code at ~30x lower price, triggering a historic one-day selloff in AI-linked stocks the following week.
DeepSeek ↗