AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Beam

Reflection AI·

LLMsopen-weightcloud

Overview

Reflection AI's first model: a text-only sparse MoE with 501B total and 23B active parameters, pretrained on 23.8T tokens, built for coding, reasoning and agentic work. Waitlist API at launch; weights under Apache 2.0 announced for later in October 2026 together with a technical report. At Q4 the weights need roughly 330 GB, just above the Multi-GPU budget.

Capabilities and innovations

501B total / 23B active MoEText-onlyCoding, reasoning, agentic workApache 2.0 weights announcedFirst model from Reflection AI

Benchmarks

BenchmarkScoreSource
GPQA Diamond90.5%vendor
GPQA Diamond 90.5. Reflection's launch table, 2026-10-05; read via secondary coverage, reflection.ai was not reachable from the intake sandbox.
Humanity’s Last Exam (no tools)36.2%vendor
HLE 36.2 without tools. Reflection's launch table, 2026-10-05; read via secondary coverage, reflection.ai was not reachable from the intake sandbox.
SWE-bench Verified80.9%vendor
SWE-bench Verified 80.9. Reflection's launch table, 2026-10-05; read via secondary coverage, reflection.ai was not reachable from the intake sandbox.
Terminal-Bench 2.x80.1%vendor
Terminal-Bench 2.1 80.1, vendor harness. Reflection's launch table, 2026-10-05; read via secondary coverage, reflection.ai was not reachable from the intake sandbox.
DeepSWE v1.144.4%vendor
DeepSWE v1.1 44.4, explicitly v1.1. The same table lists GLM-5.2 at 44.0 and GLM-5.3 at 61.0, where Z.ai reports 46.2 and 66.9 for its own models; competitor rows are Reflection's measurements and are not taken over. Reflection's launch table, 2026-10-05; read via secondary coverage, reflection.ai was not reachable from the intake sandbox.
  • SWE-bench Pro: not reported by the vendor
  • MMLU-Pro: not reported by the vendor
  • MMMU-Pro: not reported by the vendor

Architecture and hardware

Parameters
501B total, 23B active per token (MoE)

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.