AI Model Timeline

How the timeline is curated: admission criteria, benchmark eras, source priority and how leadership is decided.

Methodology

AI Model Timeline tracks the release of frontier AI models since 2019 and records, for each one, what it scored on the benchmarks that mattered at the time, what it costs to call, and, for open-weight models, what it takes to run it yourself. Four views sit side by side: Cloud for proprietary and API-served models, Local for open-weight models you can download, The Model Race for the frontier of a single benchmark over time, and Reliability for the properties that decide whether a model survives production. The catalog runs through 2026-09-10 and is edited by hand; the checks described below keep the editing honest.

Why this exists

I built this because I kept looking for it and it was not there. Not another leaderboard — those exist, and they answer a different question. What I wanted to see was the movement: when a company stepped on the gas, when release intervals that used to be measured in years dropped to months, and who quietly stopped keeping up. A ranking shows you one moment. A timeline shows you the shape of the race.

That is why the first thing you see is a horizontal timeline with one row per company rather than a table sorted by score. Read it left to right and the acceleration is not a claim you have to take on trust, it is a pattern in the spacing: rows that thicken, rows that thin out, gaps that close. The leadership lines running across it answer the second question — who actually held the lead, and for how long, rather than who announced it loudest.

And the race is worth watching. Not as sport: what these models can do, how fast that changes and which of it is real decides a great deal about the next few years, and it is being reported almost entirely by the people selling the models. Every vendor publishes a benchmark table on launch day, and every one of those tables is drawn so its own model wins. The comparison columns are picked after the results are known, the baselines are whichever competitor version is furthest behind, and six months later nobody can reconstruct which number came from where. That is not a conspiracy, it is marketing working as intended — but it means the honest question, which model should I actually use for this, today, has no honest place to be answered.

So this is an attempt to show the field as it is, past the hype and the launch-day framing. It is deliberately boring about it: one catalog, every score attached to the source it came from, the ranking rule written down before the ranking is run, and the uncertainty shown rather than smoothed away. When two models are inside the measurement noise the timeline says so instead of picking a winner. When a benchmark stops separating models, it is retired on a rule, not on a hunch. When a recommendation rests on two data points, the card admits it.

One more thing changed the design as it went along: the interesting question stopped being who is ahead and became what can I run. An open-weight model that fits on hardware you own is a different kind of answer than an API endpoint, and the two are not comparable on a single leaderboard. So the Local view ranks by what a machine can actually hold, not by what looks best in a headline.

What earns a model an entry

A model gets an entry if it meets at least one of two tests: it is a milestone release from a tracked company (a new generation, a first of its kind, or a shift in what its class of model can do), or it carries an externally verifiable signal, meaning an LMArena Elo rating, an Epoch Capabilities Index entry, or a reported score on a tracked benchmark. Everything else stays out. Without that rule the Local view would drift into a mirror of Hugging Face: there are tens of thousands of fine-tunes, and a derivative with no independent signal adds noise, not coverage.

Benchmark eras

Benchmarks saturate. When the top models bunch at the ceiling, a benchmark stops telling them apart and a harder successor takes over. The timeline records this as eras. Era 1 (2024 to mid 2025) is MMLU, GPQA Diamond, SWE-bench Verified and Humanity’s Last Exam. Era 2 (from Q3 2025) replaces MMLU with MMLU-Pro and SWE-bench Verified with SWE-bench Pro. Era 3 (from 2026) adds Terminal-Bench 2.x for agentic coding in real terminal environments and MMMU-Pro for multimodal reasoning. Terminal-Bench 3.0 is tracked as its own axis because it runs on a different scale than 2.x; a 3.0 score is never entered in the 2.x field.

An era only hands over once at least three models carry both the old and the new value, so the two generations can be compared across the handover. A saturation check and a coverage check run on a schedule: the first asks whether a benchmark is topped out, the second whether the field still reports it. The second exists because in August 2026 the leadership line froze, not because progress stopped, but because the benchmark the race was gated on had been quietly abandoned by new releases.

Who still reports what
BenchmarkScoredLast 90 daysLast score
Terminal-Bench 2.xAgentic work in a real terminal5187% · 34/39Sep 2026
GPQA DiamondGraduate-level science questions11879% · 31/39Sep 2026
Humanity's Last ExamExpert questions across domains, no tools9169% · 27/39Sep 2026
MMLU-ProKnowledge, ten options per question5846% · 18/39Sep 2026
SWE-bench ProReal GitHub issues, harder split3641% · 16/39Sep 2026
MMMU-ProMultimodal college-level reasoning2233% · 13/39Sep 2026
SWE-bench VerifiedReal GitHub issues, human-validated7323% · 9/39Aug 2026
Terminal-Bench 4.0Same name, rewritten again818% · 7/39Sep 2026
DeepSWE v1.1Long-horizon agentic coding on hand-written tasks718% · 7/39Sep 2026
Terminal-Bench 3.0Same name, rewritten task set613% · 5/39Sep 2026
ARC-AGI-1Abstract visual reasoning23% · 1/39Sep 2026
ARC-AGI-2Abstract visual reasoning, harder23% · 1/39Sep 2026
MMLUMultiple-choice knowledge, the 2020 standard690/39Feb 2026

Coverage is the share of the 39 benchmarked models released in the 90 days before 10 Sep 2026 that carry a score. A row at zero means no recent release reported it. Computed from the catalog on every build, not maintained by hand.

Read the last column and the difference between the two failure modes is visible: a benchmark that saturates keeps being reported, its scores just stop separating models, while one that is abandoned simply stops appearing on new releases. MMLU is the second kind. Nobody announced its retirement; the scores just stopped arriving, which is precisely the thing a leaderboard cannot notice about itself.

Where the numbers come from

Every benchmark value is numeric or absent, never a string, and should carry a source. Sources are ranked: independent evaluations first (Vals AI, SWE-rebench, Epoch AI, LiveBench), official vendor announcements second, aggregators last. Vendor tables are read with cherry-picking in mind; when a vendor figure and an independent measurement disagree by more than three percent, the independent value is recorded and the vendor figure kept as a note. The current top value on every active benchmark must be sourced, and the data validation fails otherwise: an unsourced record is the signature of a data error.

How thick the evidence is

45%of the 543 benchmark values in the catalog carry a source entry (244 of 543). Where one exists, it is more often the vendor’s own announcement than an independent evaluation.

  • Independent56
  • Vendor134
  • Aggregator54
  • No source recorded299
Thickest
Terminal-Bench 2.x 86% (44/51)
Thinnest
MMLU 1% (1/69)

Best and worst are taken over axes with at least 20 values, so a two-model benchmark cannot top the list. Sourcing is enforced where a wrong number would change what the site says — the current record on an active benchmark, and any value that decides a hardware recommendation — and burned down elsewhere.

That ranking says which source wins when several exist. It does not say that an independent measurement always exists, and the bar above shows how often it does not: for most values the only published number is the one the vendor put in its own launch table. Source entries are required where a wrong value would change what the site tells you — the current record on an active benchmark, and any value that decides a hardware recommendation — and are being filled in elsewhere as models are revisited. The honest reading of this figure is that the catalog is a record of what the field published, not a re-run of it.

Comparative vendor claims (“beats X”, “matches frontier models”) are never copied into descriptions. The underlying score is entered with its source and the comparison is left to the reader.

Who leads, and when it is too close to call

The leadership lines on the Cloud timeline show which model held the overall and the coding lead at each point in time. From 2026 on a model enters the race with a coding signal (SWE-bench Pro or Terminal-Bench 2.x), GPQA Diamond, and at least three of the four era benchmarks. Scores are min-max normalized over the race pool before they are averaged, so a model does not gain by omitting the benchmark it would score worst on, and a benchmark only enters the composite once three models in the pool report it. The coding line is the same construction over the two coding signals alone, either one of which suffices.

A challenger needs a lead of one normalized point over the incumbent to take the crown, and only a model released after the incumbent can challenge — adding a model rescales everyone, and without that rule an older model could win a handover dated before the one it replaces. Inside that band the crown does not move, and the chart says so: the near-tie is drawn as a faint dashed stub to a hollow dot at the end of the line, which reads as unresolved rather than as a handover. Because the crown only moves once the band is cleared, past handovers on the line are never contested; the live near-tie, when there is one, is always at the end.

Image and video models carry no academic benchmark. Their crowns come from the LMArena text-to-image and text-to-video boards. Robotics models likewise get no scores and no leadership tracking: no tracked benchmark applies to a robot control policy, and vendor task-success rates are not comparable across fleets. World models are kept out of the video type and get the same treatment for a different reason: a video model goes from prompt to clip, a world model from an image or scene to a navigable space with camera control, and no established, comparable metric exists for that yet. The timeline says so where a crown would sit, rather than rendering an empty one.

The Model Race

The Model Race carries a second view for the problem every benchmark chip eventually has. A single benchmark saturates or gets abandoned, and the frontier line stops moving for reasons that have nothing to do with progress. The Capability Index view plots the Epoch Capabilities Index instead, which fits roughly 50 benchmarks onto one difficulty scale so models stay comparable across that handover. It is a separate view rather than another chip because it is not a measurement and not a percentage: the number is fitted rather than run, it sits on a scale around 110 to 160, and Epoch refits the whole scale when new models land, so a record drawn there can shift between snapshots. Two scales never share one axis.

The Model Race view answers a narrower question than the leadership lines: on one benchmark, how far apart are two blocs, and is the gap closing? Pick a benchmark and a lens — US/West against China, or closed-weight against open-weight — and each bloc draws its running best: a point appears only when a model beats every earlier model in that bloc, so the line is monotone by construction and never falls. Nine benchmarks with enough cross-bloc, multi-point coverage are offered; GPQA Diamond is the default because it spans 2022 to 2026 with entries on both sides.

The construction is what makes the gap readable and also what limits it. A frontier line is the best score anyone in the bloc has ever posted, not the state of the median model, and it inherits every weakness of the value behind it — a bloc can appear to leap because one lab published a vendor-measured figure. The same lens runs as a weekly check that opens an issue when an open-weight model leads, or comes within two points of, the proprietary top on a benchmark that is not saturated. That check compares values only within the same evaluator, because a one-sided coverage gap can otherwise look exactly like a frontier crossing.

Reliability

Benchmarks measure how clever a model is on a good day. The Reliability board measures the two properties that decide whether it survives contact with production. Grounding is 100 minus the Vectara HHEM hallucination rate on grounded summarization: how often the model asserts something its source document does not say — the failure mode that breaks RAG systems. Tool-Use is τ-bench, multi-turn agentic tool use against a simulated user, which is a different and harder thing than a single well-formed function call. τ³-Banking is that axis’s successor: the agent has to locate the governing policy in roughly 700 interconnected documents and then run the multi-step tool calls it prescribes — retrieval and tool use in one task, which is what an enterprise deployment actually asks for. All read higher-is-better and each carries its own source per cell.

Where a model has no entry on a column’s canonical source, a comparable alternative is used and the cell is marked with an asterisk naming it, rather than left blank or silently mixed in. Each row also carries the vendor’s country of origin, because for enterprise readers that is frequently the first filter, not the last. Berkeley Function-Calling v4 is collected in the catalog but has no column: it stopped scoring the frontier, and standalone call accuracy largely overlaps with the τ-bench axis. A withheld column is not advertised — the page title used to name it while the table did not draw it.

The board also states how old it is, and why. Reliability is measured by a handful of small teams whose leaderboards stop and restart: Vectara published no HHEM snapshot after May 2026, and the canonical τ-bench repository now declares its tasks frozen in favour of τ³-bench. Those two facts explain most of the blank cells on recent models — the figure does not exist, rather than being un-entered. Because a source going quiet leaves every existing value correct and sourced, nothing in the data can reveal it; the state of each upstream source is therefore declared in the repository and printed under the table, and npm run check-reliability-health reports when the board falls behind the catalog.

The τ³ axis is what a benchmark succession looks like when it is done honestly. τ-bench, τ²-bench and τ³-bench score different task sets and domains, so their numbers are not interchangeable — the same trap as Terminal-Bench 2.x versus 3.0. The successor therefore got its own column rather than new values in the old one, and the τ-bench column stays frozen on what it measured, because it is the only tool-use figure the 2025 and early-2026 models carry. Reading across the two columns is the one thing this board asks you not to do: τ³-Banking is deliberately the hardest domain in the suite, and its numbers sit far below the retail and airline scores beside them. The values are Artificial Analysis’s own runs of Sierra’s benchmark, which is not the same harness as Sierra’s own leaderboard; every cell names which run it came from.

Hardware classes for open-weight models

The Local view sorts every open-weight model into a hardware class from the memory it needs at 4-bit quantization plus a 15 percent framework overhead: Mobile/NPU up to 4 GB, Edge up to 24 GB, Workstation up to 80 GB, Multi-GPU up to 300 GB, Frontier above that. The class is computed from total parameters, never authored, and mixture-of-experts models are measured on all their weights: every expert has to be resident before the first token, so a sparse model needs its full footprint regardless of how few experts fire per token. Active parameters are shown separately because they predict throughput once the weights are loaded, which is a different question from what can load them.

Within each class, one pick per task (all-round, coding and agents, reasoning) is computed from the benchmarks that task ranks on. A class is a budget, not a size bracket: a workstation can run everything an edge GPU runs, so the candidates for a larger class include every model a smaller one holds, and where the larger budget buys nothing the card says so instead of inventing a difference. Picks that rest on two candidates, on one, or on a margin inside the measurement noise are labelled rather than hidden. A model that currently wins a pick must have a source on the benchmark that decided it, or the data validation fails — the Local analogue of the rule for frontier records.

Derived models and shared lineage

Open weights get post-trained by other labs, and a leaderboard that ignores this counts one set of weights twice: once as the original, once as the derivative that inherited most of its capability. A model built from another lab’s open release carries its lineage on the card as a “derived from” tag, with the base model named and a note on how the provenance is established. Where a vendor names the base only at family level — “built on Qwen3.5” without saying which size — the tag says family-level rather than implying a precision the vendor did not publish. Variants inside a lab’s own family are not tagged; that belongs in the description.

A derivative’s license covers only the weights that lab published; the upstream license of the base still applies, which is why lineage is a visible field rather than a footnote. Benchmark rules do not change for derivatives: vendor values stay vendor values, and comparative claims stay out.

Arena crowns and open weights

The open-weight LLM crown on the Local macro timeline follows the Epoch Capabilities Index, a capability measure that reacts faster than community preference. The coding crown follows the LMArena WebDev board. The board’s license field is not trusted: a model counts as open weight only if its Arena name is mapped to a catalog entry whose weights are actually published, and unmapped entries are surfaced for curation by the same health check that watches benchmark coverage.

The checks that keep it honest

Four checks run against the catalog, and they answer four different questions. The saturation check asks whether a benchmark is topped out — scores bunching at the ceiling. The benchmark health check asks the opposite: whether the field still reports it, how thin the current lead is, and how many candidates each task and hardware-class cell has left to rank. That second check exists because a frozen leadership line looks identical whether progress stopped or the benchmark was abandoned, and in August 2026 it was the second. The gap check watches the open-weight frontier against the proprietary one. And data validation enforces the source duties: an unsourced frontier record or an unsourced deciding value under a live recommendation is a hard failure, not a warning.

Two of those checks are deliberately noisy about their own blind spots. The health check reports orphaned mappings — a catalog entry pointed at an external row that has since been renamed or withdrawn — because the failure is otherwise silent: the model simply stops having a rating and nothing looks wrong. It also warns when an external snapshot has fallen more than thirty days behind the catalog.

Automation, and where it stops

Model announcements are picked up weekly from vendor feeds and filed as issues for review. LMArena thrones are synced daily, the Epoch index and the open-weight-versus-proprietary gap weekly, and the benchmark health report monthly, opening an issue only when there is a decision to make. None of it writes a benchmark value on its own; every score in the catalog was entered and sourced by a person.

Data, code and attribution

The catalog and the code that renders it are published in the public repository at github.com/FXND-DE/AI-Status. The code is MIT-licensed; the curated data in src/data/ is licensed CC BY 4.0, so it may be reused with attribution. That covers the selection, arrangement and verification of the figures — not the underlying vendor publications, and not company or model names. The site is vendor-neutral and carries no affiliate links. It is edited by Felix Neuland and published by R3ASON GmbH, Frankfurt am Main. Legal notice: Impressum, Datenschutz.