Wait Which Model?

Info

How this site works

What “frontier” means here, why some figures are blank, and how the data stays up to date.

What counts as “frontier”

A model is labeled "frontier" one of two ways: it's the newest flagship-caliber release from one of a handful of labs whose releases are trusted to be near state-of-the-art, or it's genuinely competitive with the best currently available in its class by the numbers.

  • Major-lab recency: the newest release from OpenAI, Anthropic, Google, or Meta in a given tier is frontier automatically, as long as it shipped within the last 3 months — no benchmark data required. These labs' flagship releases are trusted on reputation, since official benchmark figures are often published weeks after launch and a model shouldn't lose frontier status just because that data isn't out yet.
  • Everyone else — recency: it must have been released within the last 9 months. Older models automatically lose frontier status even if nothing has replaced them yet.
  • Everyone else — tier: models are grouped into Flagship, Balanced, and Fast tiers (see below) so a small, cheap model is judged against other small, cheap models — not against a lab's biggest release.
  • Everyone else — capability bar: within its tier, a model's benchmark scores must land within about 15% of the best score any recent model in that tier has posted. Being newest isn't enough on its own — a fresh release with weaker benchmarks than the current leader doesn't qualify.
  • Everyone else — verified data required: a model needs scores on at least 3 of the tracked benchmarks to be evaluated at all. A freshly announced model without enough published benchmark data yet is labeled "Unknown" rather than guessed into Frontier or Superseded.

Definition last reviewed Aug 5, 2026.

Tiers

Every model belongs to one tier, so it’s only ever compared against peers built for the same job.

Flagship
A lab's top-of-line model, built for maximum capability regardless of cost.
Balanced
A mid-tier model trading some capability for lower cost and latency (e.g. Claude Sonnet, Mistral Medium).
Fast
A small, cheap, low-latency model, optimized for speed and cost over raw capability (e.g. Haiku, Flash, Mini variants).

Status labels

Frontier
Currently near the top of its tier by the criteria above.
Superseded
No longer near the top of its tier — either aged out, or outperformed by a newer model.
Unknown
Too new or under-covered to judge yet — not enough benchmark data has been published to place it. This is an honest gap, not a verdict; it isn't the same as Superseded.
Deprecated
Officially retired or shut down by the lab that made it (e.g. removed from their API). This is set by hand when a lab announces it, not inferred automatically.

Benchmarks

Scores come from each benchmark’s own methodology — they measure different things and aren’t directly interchangeable.

MMLU-ProHarder multi-task knowledge and reasoning benchmark (10-choice questions across 14 domains).
GPQA DiamondGraduate-level, Google-proof science questions (hardest split). PhD experts score ~65%.
SWE-bench VerifiedReal GitHub issues resolved end-to-end in real repositories; the standard agentic coding benchmark.
Terminal-Bench 2.1Long-horizon agentic tasks executed in a real terminal; increasingly reported in place of SWE-bench Verified. Version 2.1 only — 2.0 scores are not comparable and are left unset.
AIME (latest)American Invitational Mathematics Examination — competition math, no tools unless noted.
Humanity's Last Exam2,500 expert-written frontier questions across dozens of fields; designed to resist saturation.
LMArena EloCrowd-sourced head-to-head preference rating from LMArena (formerly Chatbot Arena).
ARC-AGI-2Abstraction-and-reasoning puzzles that are easy for humans and hard for models.

Cost per task

Cost per task is what a model actually costs to do one unit of work, rather than what its tokens list for. The figure is Artificial Analysis' "cost per Intelligence Index task": they run every model through the same suite of evaluations and record the weighted average USD spent on a single task — input, cached reads and writes, reasoning tokens, and the answer itself. It's the number the model card leads with because a cheap per-token price stops meaning much once a model thinks for thirty thousand tokens before answering.

  • Sticker price and cost per task can rank models very differently. A model with low token prices that reasons at length can cost more per task than a pricier model that answers quickly — the per-million-token rates are still on every model's detail panel, they just aren't the headline any more.
  • Reasoning models are measured at a chosen effort level, and cost scales steeply with it. Where Artificial Analysis publishes a medium-effort figure, that's the one shown; otherwise it's whichever effort they measured, and the model's detail panel names it.
  • Because the effort level isn't identical across every model, treat these as an order-of-magnitude comparison, not a precise ranking — a 10x gap is real, a 10% gap may just be a difference in effort setting.
  • Blank means Artificial Analysis has not run the Intelligence Index on that model — common for older or newly released models — not that it is free or cheap.
Further reading

Price per 1M tokens is meaningless Jan Iłowski

Why per-token rates don't compare across labs — different tokenizers split the same text into different token counts, and models burn wildly different amounts of hidden reasoning to reach the same answer.

Why some numbers are missing

Any figure you see as "—" is genuinely unpublished or unverified, never a guess. Numbers are only added once they can be traced to an official announcement, model card, or a reputable third-party benchmark tracker.

Where the numbers come from

All benchmark scores are launch-time reported figures from the model's own announcement or model card, not scores re-run independently by this site — so numbers across labs aren't always measured under identical conditions.

How current this data is

This site has no backend — everything is static data, refreshed by periodic research passes rather than a live feed. Roughly weekly, and whenever a major lab ships something new, the model directory, news log, and "frontier" statuses get re-checked against current web sources and updated. There's no fixed daily schedule; treat the data as accurate as of its most recent update, not real-time.

How long models held the frontier

How long each model held the top of its tier — reconstructed from release dates and benchmark scores, not recorded as it happened. Read the notes below the chart before drawing conclusions from it: the timeline starts later than the real frontier did, because models that never reported three benchmarks are missing from it entirely.

Reconstructed, not observed — and incomplete before November 2023.

Flagship

  • GPT-4 Turbo101 days
  • Gemini 1.5 Pro88 days
  • GPT-4o71 days
  • Llama 3.1 405B135 days
  • o174 days
  • Grok 3219 days
  • Qwen3-Max55 days
  • Gemini 3 Pro85 days
  • Gemini 3.1 Pro118 days
  • Claude Fable 5— · current

Balanced

  • Claude 3.5 Sonnet249 days
  • Claude 3.7 Sonnet87 days
  • Claude Sonnet 4130 days
  • Claude Sonnet 4.5141 days
  • Claude Sonnet 4.6133 days
  • Claude Sonnet 5— · current

Fast

  • Gemini 2.0 Flash237 days
  • gpt-oss-20b71 days
  • Claude Haiku 4.5169 days
  • Gemma 447 days
  • Gemini 3.5 Flash63 days
  • Gemini 3.6 Flash23 days
  • Gemini 3.7 Flash— · current
  • This is a reconstruction, not an observed record. The site recomputes frontier status from scratch and keeps no history, so reigns are inferred after the fact.
  • A model takes the crown on its release date if its composite score beats the incumbent's, and holds it until a later release scores higher.
  • Only models reporting at least three benchmarks can be crowned — one lucky score is not enough to claim a reign.
  • The composite is min-max normalised across each tier and averaged over whichever benchmarks a model reports, so two models reporting different benchmark sets are not compared on identical ground.
  • Deprecated models are included: a retired model still held the frontier while it was alive.
  • A model that never reported three benchmarks is absent from the chart altogether, not merely denied a reign. The flagship timeline therefore begins in November 2023 with GPT-4 Turbo — not because nothing led the field before then, but because earlier models never published enough verified scores to be ranked here.

Benchmark coverage

Not every model reports every benchmark. A dash means the score was never published or could not be verified from a primary source — it does not mean zero.

  • MMLU-Pro44 of 98
  • GPQA Diamond80 of 98
  • SWE-bench Verified56 of 98
  • Terminal-Bench 2.127 of 98
  • AIME (latest)41 of 98
  • Humanity's Last Exam60 of 98
  • LMArena Elo59 of 98
  • ARC-AGI-223 of 98

Cost calculator

Monthly cost is tasks per day x 30 x the model's measured cost per task.

  • A task is one Artificial Analysis Intelligence Index task — one self-contained question or job.
  • Cost per task is taken at whatever reasoning effort Artificial Analysis publishes for that model — the effort varies model to model, and is unrecorded for many of them. Figures measured at different efforts are not strictly comparable, even when they're shown side by side here.
  • A 30-day month is used regardless of the calendar month.
  • Models with no measured cost-per-task figure are excluded and named beneath the results rather than silently dropped.
  • Every figure is measured. Nothing here is estimated, so no uncertainty range is shown.

In the spec comparison on the Compare page, each model's value is coloured against the baseline: green means better, red means worse — following each field's own direction. Lower is better for price and time to first answer token, so a cheaper or faster figure is green even though it is the smaller number; higher is better for benchmarks, context window and output speed. Green means better, never merely bigger. The baseline's own value turns blue on any row where it beat every model it was compared against. Fields with no meaningful ordering, such as licence, are never coloured. Speed and cost-per-task figures are each measured at a reasoning-effort setting, and the setting dominates the result, so a value is only coloured when both models were measured at the same recorded effort; where the effort isn't recorded for one or both models, no comparison is drawn — the figures are shown side by side, uncoloured, with a warning instead.