Beam
Reflection AI·
LLMsopen-weightcloud
Overview
Reflection AI's first model: a text-only sparse MoE with 501B total and 23B active parameters, pretrained on 23.8T tokens, built for coding, reasoning and agentic work. Waitlist API at launch; weights under Apache 2.0 announced for later in October 2026 together with a technical report. At Q4 the weights need roughly 330 GB, just above the Multi-GPU budget.
Capabilities and innovations
501B total / 23B active MoEText-onlyCoding, reasoning, agentic workApache 2.0 weights announcedFirst model from Reflection AI
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 90.5% | vendor GPQA Diamond 90.5. Reflection's launch table, 2026-10-05; read via secondary coverage, reflection.ai was not reachable from the intake sandbox. |
| Humanity’s Last Exam (no tools) | 36.2% | vendor HLE 36.2 without tools. Reflection's launch table, 2026-10-05; read via secondary coverage, reflection.ai was not reachable from the intake sandbox. |
| SWE-bench Verified | 80.9% | vendor SWE-bench Verified 80.9. Reflection's launch table, 2026-10-05; read via secondary coverage, reflection.ai was not reachable from the intake sandbox. |
| Terminal-Bench 2.x | 80.1% | vendor Terminal-Bench 2.1 80.1, vendor harness. Reflection's launch table, 2026-10-05; read via secondary coverage, reflection.ai was not reachable from the intake sandbox. |
| DeepSWE v1.1 | 44.4% | vendor DeepSWE v1.1 44.4, explicitly v1.1. The same table lists GLM-5.2 at 44.0 and GLM-5.3 at 61.0, where Z.ai reports 46.2 and 66.9 for its own models; competitor rows are Reflection's measurements and are not taken over. Reflection's launch table, 2026-10-05; read via secondary coverage, reflection.ai was not reachable from the intake sandbox. |
- SWE-bench Pro: not reported by the vendor
- MMLU-Pro: not reported by the vendor
- MMMU-Pro: not reported by the vendor
Architecture and hardware
- Parameters
- 501B total, 23B active per token (MoE)
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.