FrontierObservatory

Model news

The frontier log.

Releases, benchmark milestones, company moves, research results and policy — everything that shifts the frontier, in one dated record.

July 2026

  1. Jul 16
    releaseMoonshot AI

    Moonshot AI launches Kimi K3

    Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-weight MoE model with a 1M-token context window and native visual understanding, priced at $3/$15 per million input/output tokens; full weights are due July 27, 2026.

    VentureBeat
  2. Jul 13
    releaseSoofi

    German SOOFI consortium releases Soofi S, a sovereign open-weight MoE model

    A consortium led by the KI Bundesverband — with Fraunhofer, DFKI, TU Darmstadt and others — released Soofi S 30B-A3B, a hybrid Mamba-2/MoE model trained on 27T tokens with up-weighted German, as a preview checkpoint on Hugging Face.

    Fraunhofer IAIS
  3. Jul 9
    releaseOpenAI

    OpenAI ships GPT-5.6 (Sol, Terra, Luna) to the public

    OpenAI lifted the government-gated preview and launched GPT-5.6 broadly; flagship Sol posts record GPQA Diamond (94.6%) and ARC-AGI-2 (92.5% at max effort) while OpenAI calls it its strongest cybersecurity model yet.

    TechCrunch
  4. Jul 1
    policyxAI

    Colorado guts its AI Act as xAI and DOJ challenge proceeds

    After xAI sued to block SB24-205 and the DOJ intervened — its first move against a state AI law — Governor Polis signed SB 26-189, stripping algorithmic-discrimination duties and bias audits in favor of consumer disclosure.

    US Department of Justice

June 2026

  1. Jun 26
    releaseOpenAI

    OpenAI previews GPT-5.6 Sol, Terra and Luna under government-gated release

    GPT-5.6 opens to ~20 government-approved companies at the administration's behest: Sol is the flagship, Terra runs ~2x cheaper than GPT-5.5, and Luna is the low-cost tier. General availability is promised 'in coming weeks.'

    Axios
  2. Jun 15
    policy

    EU appoints scientific panel ahead of AI Act penalty powers

    The European Commission named 60 independent frontier-AI experts to support the AI Office before enforcement penalties activate on August 2, 2026, shifting the AI Act from infrastructure-building to active enforcement.

    European Commission
  3. Jun 15
    benchmarkAnthropicOpenAIGoogle DeepMind

    LMArena top tier compresses to its tightest spread on record

    The top five models — Claude Opus 4.8 (~1510), GPT-5.5 Pro, Gemini 3.1 Pro, Claude Opus 4.7 and GPT-5.5 — sit within ~55 Elo, with the top ten inside ~20 points; task fit now matters more than leaderboard rank.

    Swfte AI Leaderboard
  4. Jun 13
    releaseZhipu AI

    Zhipu AI releases GLM-5.2, an MIT-licensed open-weights coding flagship

    Z.ai's GLM-5.2 (753B/40B-active MoE) shipped with a 1M-token context window and beat GPT-5.5 on SWE-bench Pro (62.1 vs 58.6) at roughly one-sixth the API cost, topping the Artificial Analysis open-weights index.

    VentureBeat
  5. Jun 13
    policyAnthropic

    US bars foreign nationals from accessing Fable 5 and Mythos 5

    Two days after Anthropic's distillation-throttling apology, the administration barred foreign nationals from Anthropic's two newest frontier models, citing national security.

    Council on Foreign Relations
  6. Jun 11
    companyAnthropic

    Anthropic apologizes for Fable 5 silently throttling suspected distillation

    Anthropic apologized after Claude Fable 5 was found silently limiting responses to users suspected of trying to replicate its capabilities, and faced criticism for over-refusing cyber-related queries.

    Council on Foreign Relations
  7. Jun 9
    releaseAnthropic

    Anthropic launches Claude Fable 5 and restricted Claude Mythos 5

    Fable 5 becomes the first generally available Mythos-class model ($10/$50 per MTok, state of the art on nearly all benchmarks); Mythos 5 — the same weights with safeguards lifted — goes to vetted cyberdefenders via Project Glasswing.

    Anthropic
  8. Jun 2
    policyOpenAIAnthropicGoogle DeepMindxAIMeta AI

    White House issues executive order on frontier AI innovation and security

    The order creates a voluntary pre-release federal review process for frontier models with 30-day government access, plus early access for critical-infrastructure entities, framed around cybersecurity.

    The White House

May 2026

  1. May 28
    releaseAnthropic

    Anthropic releases Claude Opus 4.8

    Opus 4.8 ships at $5/$25 per MTok with 88.6% on SWE-bench Verified, a USAMO jump to 96.7%, a new fast mode, and the #1 LMArena spot (~1510 Elo). Reviewers call it a modest but tangible improvement.

    Anthropic
  2. May 28
    companyAnthropicOpenAI

    Anthropic raises $65B at a $965B valuation, overtaking OpenAI as most valuable AI startup

    The Series H led by Altimeter, Dragoneer, Greenoaks and Sequoia nearly triples February's $380B valuation; Anthropic reports a $47B revenue run rate driven largely by Claude Code.

    CNBC
  3. May 22
    companyOpenAI

    OpenAI files confidential S-1 for IPO

    OpenAI submits a confidential draft registration to the SEC after a record $122B March round at ~$852B; later reports suggest it may wait until 2027 to list.

    TechJournal
  4. May 22
    researchOpenAI

    OpenAI reasoning model disproves Erdős-linked conjecture

    A general-purpose OpenAI reasoning model disproved a central conjecture tied to Erdős's 1946 planar unit-distance problem, finding an infinite family of counterexample point arrangements verified by outside mathematicians.

    Forbes
  5. May 20
    releaseAlibaba (Qwen)

    Alibaba unveils Qwen3.7 Max, its first closed-weight flagship

    Announced at the Alibaba Cloud Summit, Qwen3.7 Max posts the highest Artificial Analysis index score ever for a Chinese model (56.6) — but ships without open weights, a notable strategy shift.

    Digital Applied
  6. May 19
    releaseGoogle DeepMind

    Google launches Gemini 3.5 Flash at I/O 2026

    Gemini 3.5 Flash ($1.50/$9 per MTok) beats Gemini 3.1 Pro on coding and agentic benchmarks at ~25% lower cost and up to 4x the speed.

    MarkTechPost
  7. May 11
    release

    Thinking Machines Lab previews near-realtime voice-and-video 'interaction models'

    Mira Murati's lab unveils TML-Interaction-Small (276B MoE, 12B active) with 0.4s turn-taking latency that listens while it talks — its first model release, in limited research preview.

    TechCrunch
  8. May 1
    companyDeepSeekAlibaba (Qwen)

    DeepSeek V4 triggers scramble for Huawei Ascend 950 chips

    V4's optimization for Huawei Ascend rather than Nvidia hardware prompts ByteDance, Tencent and Alibaba to rush Ascend 950 orders; SMIC shares jump 10% while US export controls constrain supply.

    Capacity

April 2026

  1. Apr 30
    releasexAI

    xAI launches Grok 4.3 with native video input and a 40% price cut

    Grok 4.3 hits the public API at $1.25/$2.50 per MTok — the cheapest near-frontier flagship — while Grok 5, the 6T-parameter Colossus 2 model, slips again with no release date.

    Artificial Analysis
  2. Apr 24
    releaseDeepSeek

    DeepSeek releases open-weights V4 family, tying Gemini 3.1 Pro on SWE-bench

    V4-Pro-Max posts 80.6% on SWE-bench Verified — the best open-weights score — with a 1M context and 384K max output at $0.435/$0.87 per MTok under MIT license.

    MorphLLM
  3. Apr 23
    releaseOpenAI

    OpenAI ships GPT-5.5 with an 85% ARC-AGI-2 score

    GPT-5.5 posts an 11.7-point ARC-AGI-2 jump over GPT-5.4 and takes Terminal-Bench 2.0 state of the art; the $30/$180 GPT-5.5 Pro variant targets long-horizon research.

    OpenAI
  4. Apr 8
    releaseMeta AI

    Meta releases open-weights Llama 5 and closed flagship Muse Spark on the same day

    Llama 5 (600B+, 5M-token context, trained on 500K+ B200s) renews Meta's open-weights push, while Muse Spark — the first closed model from Meta Superintelligence Labs — becomes its consumer flagship.

    Meta AI
  5. Apr 2
    releaseGoogle DeepMind

    Google DeepMind releases Gemma 4, its first multimodal open-weights line

    Gemma 4 ships in five Apache 2.0-licensed sizes (E2B to 31B), adding native image/video (and audio on the smaller tiers) input; the 31B flagship ranked #3 on the LMArena text leaderboard at release.

    Google

February 2026

  1. Feb 11
    releaseGoogle DeepMind

    Google previews Gemini 3.1 Pro with a 77% ARC-AGI-2 score

    Gemini 3.1 Pro more than doubles its predecessor's abstract-reasoning performance and takes the GPQA Diamond lead at unchanged $2/$12 pricing.

    Google DeepMind

November 2025

  1. Nov 24
    releaseAnthropic

    Anthropic releases Claude Opus 4.5, first model past 80% on SWE-bench Verified

    Opus 4.5 tops coding benchmarks while cutting flagship pricing 3x to $5/$25 per million tokens, capping a three-week stretch in which GPT-5.1, Gemini 3 Pro and Grok 4.1 all shipped.

    Anthropic
  2. Nov 18
    releaseGoogle DeepMind

    Google ships Gemini 3 Pro with record HLE and ARC-AGI-2 scores

    Gemini 3 Pro posts 37.5% on Humanity's Last Exam and becomes the first model rated above 1500 Elo on LMArena; it rolls out to Search's AI Mode on day one.

    Google DeepMind
  3. Nov 17
    releasexAI

    xAI's Grok 4.1 takes the top LMArena spot

    Grok 4.1 debuts at #1 on the LMArena text leaderboard with sharply reduced hallucination rates and improved writing over Grok 4.

    xAI
  4. Nov 12
    releaseOpenAI

    OpenAI releases GPT-5.1 with adaptive reasoning

    GPT-5.1 Instant and Thinking replace GPT-5 in ChatGPT, spending reasoning tokens only when a task needs them and shipping a warmer default personality after months of user complaints.

    OpenAI
  5. Nov 6
    releaseMoonshot AI

    Kimi K2 Thinking sets open-weights records on agentic benchmarks

    Moonshot AI's 1T-parameter open reasoning model beats GPT-5 on agentic Humanity's Last Exam and holds 200+ sequential tool calls, reportedly trained for ~$4.6M.

    Moonshot AI

October 2025

  1. Oct 15
    releaseAnthropic

    Claude Haiku 4.5 brings near-frontier coding to the small-model tier

    Anthropic's small model matches Sonnet 4's coding performance at a third of the cost and more than twice the speed, at $1/$5 per million tokens.

    Anthropic

September 2025

  1. Sep 29
    releaseAnthropic

    Claude Sonnet 4.5 claims best-coding-model title

    Sonnet 4.5 posts 77.2% on SWE-bench Verified and sustains 30+ hour autonomous coding sessions, shipping alongside the Claude Agent SDK.

    Anthropic
  2. Sep 29
    releaseDeepSeek

    DeepSeek-V3.2 halves long-context API prices with sparse attention

    DeepSeek's experimental sparse-attention release cuts API prices to $0.28/$0.42 per million tokens, keeping open-weights pressure on frontier pricing.

    DeepSeek

August 2025

  1. Aug 7
    releaseOpenAI

    OpenAI launches GPT-5 as a unified system

    GPT-5 routes between fast and reasoning modes automatically and reaches 74.9% on SWE-bench Verified; a bumpy rollout forces OpenAI to restore GPT-4o for paying users within days.

    OpenAI

July 2025

  1. Jul 11
    releaseMoonshot AI

    Moonshot AI open-sources trillion-parameter Kimi K2

    Kimi K2 becomes the strongest open non-reasoning model, purpose-built for agentic tool use, intensifying the China open-weights wave.

    Moonshot AI
  2. Jul 9
    releasexAI

    xAI debuts Grok 4 with leading HLE and ARC-AGI-2 scores

    Grok 4 and multi-agent Grok 4 Heavy lead Humanity's Last Exam and ARC-AGI-2 at launch, days after high-profile Grok chatbot safety failures on X.

    xAI

January 2025

  1. Jan 20
    releaseDeepSeek

    DeepSeek-R1 shocks the market with open o1-class reasoning

    The MIT-licensed reasoning model matches OpenAI's o1 on math and code at ~30x lower price, triggering a historic one-day selloff in AI-linked stocks the following week.

    DeepSeek