Macro Overview
How are thrones calculated?
The LLM, Coding and Image lines are step-function traces of the open-weight throne — each line sits in the current holder's provider lane and jumps vertically at every succession.
The LLM throne is capability-driven: it follows the highest-scoring open-weight model on the Epoch Capabilities Index (ECI), a benchmark-aggregate scale — not community votes. A transition renders dashed when the challenger's 95% CI overlaps the incumbent's (the lead is within noise) and solid when the lead is clear. Only models Epoch classifies as open-weight are eligible, so hosted-only tiers (e.g. Kimi K3, Qwen-Max) never hold this crown even at higher ECI.
The Coding and Image thrones remain arena-based (community Elo): post-2024 transitions come from LMArena WebDev/Code snapshots and the Artificial Analysis Text-to-Image Arena. Pre-Arena crowns are reconstructed editorially from benchmark consensus and release-impact, marked with reduced confidence (solid = high, dashed = medium, dotted = low).
For coding, the throne shifted from coding specialists (Code Llama, DeepSeek-Coder, Qwen-Coder) to general-purpose models (DeepSeek-V3 onwards) around December 2024 — reflecting the broader trend where SOTA generalists began outperforming task-specific architectures on SWE-Bench.
Below the timeline you'll find a second crown track By Hardware Class — editorial picks ranked by VRAM at Q4 quantization (Edge ≤24 GB, Workstation ≤80 GB, Multi-GPU ≤300 GB, Frontier no limit).
Video has no mature open-weight leaderboard on LMArena yet, so the pill stays disabled.
Hollow dashed dots on the timeline are crown holders that aren't in the local catalog yet (e.g. Vicuna, Tulu, WizardLM, Nemotron) — the line still passes through their provider lane.
Last sync: Sep 2, 2026·LMArena methodology
LMArena Champions
Externally verified via human preference votingNo open-weight throne holder yet — LMArena Text-to-Image has no ranked open-weight entries.
By Hardware Class
One model for everything- GPU
- Snapdragon X NPU
- Mac
- iPhone / iPad NPU
- GPU
- RTX 4090 · 24 GB
- Mac
- M5 · 32 GB
- GPU
- RTX PRO 6000 · 96 GB
- Mac
- M5 Max · 128 GB
- GPU
- 4× H100 · 320 GB
- Mac
- M5 Ultra · 512 GB
- GPU
- GB200 cluster
- Mac
- beyond one Mac
Benchmark Leader
Objective capability — independent of LMArena vote maturityNo benchmark-ranked entry — no academic image benchmark tracked.
How are thrones calculated?
The 👑 LMArena Champions follow weekly LMArena leaderboard snapshots — human preference voting on model outputs, filtered to open-weight models. A throne only changes hands when the new #1's 95% confidence interval clears the incumbent's, so near-ties don't flicker.
The Hardware Class crowns are editorial picks. We estimate the VRAM footprint of the whole checkpoint at Q4 quantization (+15% framework overhead) and assign one of five classes: Mobile / NPU (≤4 GB), Edge (≤24 GB), Workstation (≤80 GB), Multi-GPU (≤300 GB), Frontier (no limit). Total parameters, not active ones — also for MoE. Every weight has to be resident before the first token, so a 744B-A40B model needs ~430 GB no matter how few experts fire per token; sizing it by active parameters would file it under “runs on a 4090”, which is only true if you offload experts to system RAM or disk and accept an order of magnitude fewer tokens per second — at which point the API it was meant to replace is both faster and cheaper. A model therefore only wins a class it fits in unaided; sparsity is shown as an active figure, which predicts speed once the weights are loaded, not what can load them. Each class names reference hardware twice, once as a discrete GPU and once as an Apple-Silicon machine, because the two ladders diverge: a single M5 Ultra addresses up to 512 GB of unified memory and covers a class that otherwise takes four H100s. Unified memory is not all GPU-addressable though — macOS holds part of it back for the system, so treat the Mac figures as an upper bound.
A class is a budget, not a size bracket: every model that fits it competes, including everything a smaller class runs. A workstation runs what a 4090 runs, so ranking inside brackets let the dearer class recommend the worse model — workstation coding showed a 74 GB model at SWE-bench 77.6 while the edge card already offered a 21 GB one at 79.0. Where a larger budget buys nothing, the card says so instead of inventing a difference.
Within a class, the task pills decide what “best” means. Each task owns an ordered chain of signals and the first one with enough ranked candidates wins, so the label always names what the pick actually won on. Allround leads with the LMArena text rating — human preference over open-ended prompts is the one signal no single benchmark stands in for — then falls back to GPQA + MMLU-Pro, then GPQA + HLE, and only last to the legacy MMLU pair. That order is deliberate: current models report MMLU-Pro and HLE, while the models still carrying plain MMLU are mostly 2024/25 releases — asking for MMLU first ranked each class as it looked a year ago. Coding & Agents ranks on SWE-bench Pro + Terminal-Bench, falling back to SWE-bench Verified; Reasoning & Analysison GPQA + HLE (no-tools), falling back to GPQA. A model must report every benchmark of a step to be ranked by it, and scores are min-max normalized across the whole open-weight field so a value means the same thing in every cell. Every step with a full field nominates a winner, and the nominees are compared on the benchmarks they share — restricted to the ones the task ranks on, so a coding pick is never decided by a reasoning score. Until the arena snapshot covers more than the current champion, Allround falls back to the same academic pair as Reasoning in most classes; that overlap is the missing arena data, not a bug.
Three things are labelled rather than hidden. A step we rank ourselves needs three candidates to be a ranking at all — below that the pick is marked thin evidence (two) or only ranked model at this size (one). Where the top two are closer than 2% of the field's spread the pick reads too close to call — and there the more recent release is shown, because a board that recommends what to install today should default to the current model when the measurement is indifferent between them. Every pick carries its release month, so a class whose best measured model is a year old says so rather than reading as current. And an empty cell means no model of that size reports the task's signal at all — which is why coding starts at Edge and not below. The LMArena step is exempt from the three-candidate rule: that ranking was made against the entire open-weight field, not by us inside one class.
The Benchmark Leader crowns rank open-weight models by academic benchmarks — LLM by GPQA Diamond + HLE (no-tools), Coding by SWE-bench Pro + Terminal-Bench. Scores are min-max normalized across the open-weight field and averaged; a model must report all of a category's benchmarks to be ranked. Unlike the LMArena crowns this is vote-count independent, so a strong new release leads immediately rather than after its Arena votes mature. HLE is always the no-tools figure — with-tools scores are not comparable.
Video has no mature open-weight leaderboard on LMArena yet, so the 🎬 pill stays disabled.
Open-weight models without LMArena ratings still appear on the timeline but cannot hold an LMArena Champion crown. They remain eligible for task picks via the benchmark steps.
Last sync: Sep 2, 2026·LMArena methodology
Hy4-preview
2026-08-28Preview of Tencent's next-generation Hunyuan flagship: a 770B-parameter MoE activating 49B per token, with a 1M-token context window. Open-weight under Apache 2.0 with an FP8 checkpoint alongside BF16, tuned for coding agents, complex tool-use workflows and productivity tasks, and trained on domain data from Tencent's software, gaming and finance teams.
Key Capabilities
- 770B total / 49B active MoE
- 1M token context window
- Coding-agent and tool-use focus
- Apache 2.0 license
Innovations
- ★Tencent-internal software, gaming and finance domain data
- ★FP8 checkpoint shipped alongside BF16
Qwen3.8-Flash-Next
2026-08-26Open-weight multimodal MoE that Qwen ships as an early preview of the architecture behind Qwen4 — the same role Qwen3-Next played for Qwen3.5. A 125B main model plus a separate 51B N-gram embedding table, with 6B parameters activated per token. Four changes carry the release: a GDN + QSA hybrid attention stack (Gated DeltaNet compresses the history, Qwen Sparse Attention picks important context at micro-block granularity), a Gated Residual stream widened to four dynamically gated branches, the N-gram embedding table itself (looked up from local context, offloadable to host memory with asynchronous prefetching), and a Muon optimizer with a scaling law refitted for the new architecture. Native context is 262,144 tokens. Qwen reports roughly 1/9 the training cost of Qwen3.7-Plus. Because only 6B parameters are active and the embedding table can sit in system RAM, Unsloth's smallest quant runs on a 78 GB RAM / unified-memory machine without a discrete GPU. A production Qwen3.8-Flash for the QwenCloud API was announced at $0.16 / $0.47 per MTok but is not live yet, so no cloud entry is tracked; the license file ships with the weights and is not stated in the GitHub release.
Key Capabilities
- 125B main + 51B N-gram embedding parameters, 6B active per token
- 262K native context
- Natively multimodal (text and vision)
- GDN + QSA hybrid attention
- N-gram embedding table offloadable to host RAM
- Runs from system RAM without a discrete GPU (78 GB smallest quant)
Innovations
- ★Early open release of the Qwen4 architecture
- ★Qwen Sparse Attention (QSA) with micro-block context selection
- ★Gated Residual: 4-branch residual stream with dynamic gating
- ★N-gram Embedding with asynchronous host-memory prefetching
- ★Muon optimizer with refitted scaling law
Ornith-1.5-397B
2026-08-19MIT-licensed flagship of the Ornith-1.5 family from the research lab DeepReinforce. 397B total / 17B active MoE, post-trained from Qwen3.5-397B-A17B and keeping the base model's parameter layout unchanged — the first entry in this catalog to carry a base_model block, because it is a derivative of another lab's open weights rather than an in-house pretraining run. Ornith-1.5 turns the self-scaffolding of Ornith-1.0 into a closed self-improvement loop: the model generates its own tasks, scaffolds and rollouts and uses them as reinforcement-learning signal. Vendor-reported at launch: Terminal-Bench 2.1 86.1 (Harbor/Terminus-2, 128K context, averaged over 5 runs), SWE-bench Verified 86.0 (OpenHands harness, 256K context) and DeepSWE 56.0, up from 8.0 for Ornith-1.0. No independent reproduction published at intake. The MIT grant covers the weights Ornith publishes; the upstream Qwen3.5 terms still govern the lineage.
Key Capabilities
- 397B total / 17B active parameters (MoE)
- Agentic coding and terminal tasks
- Evaluated at 256K context (SWE-bench harness)
- GGUF, FP8 and NVFP4 builds
- MIT license without regional restrictions
Innovations
- ★Self-improvement loop: the model writes its own tasks, scaffolds and RL rollouts
- ★Post-trained derivative of open weights (Qwen3.5-397B-A17B) at unchanged parameter count
- ★DeepSWE jump from 8.0 (Ornith-1.0) to 56.0
Ornith-1.5-35B-A3B
2026-08-19Mid-size member of the Ornith-1.5 family: 35B total / 3B active MoE under MIT, trained with the same self-improvement loop as the 397B flagship and shipped with GGUF, MLX (including a 4-bit build) and NVFP4 checkpoints, so a quantised build fits a 24 GB consumer GPU. Vendor-reported at launch: Terminal-Bench 2.1 67.8 (Harbor/Terminus-2, 128K context, averaged over 5 runs) and SWE-bench Verified 79.0 (OpenHands harness, 256K context). Like the rest of the family it is a post-trained derivative of open weights, not an in-house pretraining run; the MIT grant covers Ornith's own weights, the upstream base terms govern the lineage.
Key Capabilities
- 35B total / 3B active parameters (MoE)
- Agentic coding and terminal tasks
- GGUF, MLX and NVFP4 builds
- Runs quantised on a single 24 GB GPU
- MIT license without regional restrictions
Innovations
- ★Self-improvement loop at 3B active parameters
- ★MLX 4-bit build for Apple Silicon at release
Ornith-1.5-9B
2026-08-19Smallest member of the Ornith-1.5 family: a dense 9B under MIT with a 262K native context and a quantised mobile build for iPhone and Android. GGUF checkpoints range from about 2.8 GB to 17.9 GB, so an 8 GB card runs a Q4/Q5 build and 24 GB runs BF16; on Apple Silicon the same builds run through Metal via MLX. Vendor-reported at launch: Terminal-Bench 2.1 47.0 (Harbor/Terminus-2, 128K context) and SWE-bench Verified 70.6 (OpenHands harness, 256K context). Ornith's size-comparison claims are size-adjusted — the 35B MoE of the same family still scores higher on both figures.
Key Capabilities
- 9B dense parameters
- 262K native context
- Agentic coding and terminal tasks
- GGUF and MLX builds from ~2.8 GB
- Quantised mobile build (iOS/Android)
- MIT license without regional restrictions
Innovations
- ★Self-improvement loop applied at 9B dense
- ★Mobile-quantised build shipped at release
Qwen3.8-27B
2026-08-14Dense 27B natively multimodal model under Apache 2.0 — the companion release to the Qwen3.8 open weights and the one that actually fits a single consumer GPU. Takes text, images and video through a 27-layer vision encoder, with a native 262,144-token context extensible to ~1M via YaRN and switchable thinking. Qwen positions it for agentic work: the launch table shows the largest gains over Qwen3.6-27B in agentic coding, computer use and vision-language tasks. At NVFP4 the weights are about 15.7 GB, a Q4_K_M GGUF roughly 16.4 GB, so 24 GB of VRAM is the comfortable target.
Key Capabilities
- 27B dense parameters
- Natively multimodal (text, image, video)
- 262K native context, ~1M via YaRN
- Switchable thinking mode
- Agentic coding and computer use
- Apache 2.0
Innovations
- ★Vision-language agent at 27B dense
- ★Integrated 27-layer vision encoder with image and video preprocessing
- ★Runs at NVFP4 in under 16 GB
GLM-5.3
2026-08-14Open-weight release of Z.ai's GLM-5.3 flagship: a 753B MoE activating ~40B parameters per token with a 1M-token context, architecturally identical to GLM-5.2 — the capability gain is post-training only. Weights landed on Hugging Face on 2026-08-28, fourteen days after launch, once Z.ai's safety review completed — but under a bespoke GLM-5.3 License instead of the MIT license GLM-5.2 shipped under: use, modification, distribution, sublicensing, sale, deployment and fine-tuning are all permitted, but a company with more than $10B aggregate revenue over any 12 consecutive months must pass a Z.ai security review before hosting the model commercially. Individual users and smaller companies are unaffected. The official checkpoint is FP8 (~743 GB of weights) and needs an 8-GPU H100/H200-class node; vLLM and SGLang serve it directly.
Key Capabilities
- 753B total / ~40B active MoE (base unchanged from GLM-5.2)
- 1M-token context
- Agentic / long-horizon coding
- Cybersecurity workflows
- GLM-5.3 License (custom, revenue-tiered)
Innovations
- ★Capability gain from post-training scaling alone, base model unchanged
- ★Revenue-gated open-weight license replacing MIT (security review above $10B revenue)
- ★Weights published 14 days after the API launch, after a safety review
DeepSeek-V4-Pro-0813
2026-08-13General-availability build of DeepSeek-V4-Pro. Identical 1.6T total / 49B active MoE architecture and size as the April preview — re-post-trained, with the gains concentrated in agentic and software-engineering tasks. 1M-token context, MIT license.
Key Capabilities
- 1.6T total / 49B active MoE
- 1M token context window
- Three reasoning effort modes (incl. Think Max)
- MIT license
Innovations
- ★Hybrid attention: CSA + HCA
- ★Manifold-Constrained Hyper-Connections
- ★Re-post-training of the April preview checkpoint
Qwen3.8-2.4T-A95B
2026-08-12Open-weight checkpoint of the Max-class flagship: 2.4T total / 95B active sparse MoE, published on Hugging Face and ModelScope alongside an FP8 build. It is the first Qwen-Max-class model with downloadable weights, but it is not the hosted Qwen3.8-Max: the checkpoint is text-only (no vision or video input) and always reasons — thinking mode cannot be disabled. Native context is 262,144 tokens, extensible to roughly 1M via YaRN. Shipped under the bespoke Qwen3.8-Max License rather than Apache 2.0. In FP8 the weights occupy about 2,325 GiB, so serving it takes a 16-GPU class deployment (NVIDIA documents four GB300 NVL72 trays).
Key Capabilities
- 2.4T total / 95B active parameters (sparse MoE)
- 262K native context, ~1M via YaRN
- Text-only input and output
- Always-on thinking (cannot be disabled)
- 128K max output tokens
- FP8 and BF16 checkpoints, vLLM / SGLang support
Innovations
- ★First Qwen-Max-class model released with open weights
- ★2.4T-parameter open-weight MoE
- ★Fine-grained FP8 block quantization (block size 128)
Nemotron 3.5 Lightning
2026-08-11NVIDIA's current open-weight release: a 30B / 3B-active hybrid model interleaving Mamba-2, MoE and select attention layers, tuned for throughput in agent loops rather than peak scores. Text-only with a 1M-token context, shipped under OpenMDW-1.1 together with training data and post-training recipes, in BF16, NVFP4 and community GGUF builds. A 4-bit quant runs on a single 24 GB consumer GPU. NVIDIA reports the largest gains over Nemotron 3 Nano on agentic evaluations; Artificial Analysis measures roughly 670 output tokens/s on pre-release infrastructure.
Key Capabilities
- 30B total / 3B active parameters (MoE)
- Hybrid Mamba-2 + MoE + attention layers
- 1M token context window
- Text-only reasoning
- BF16, NVFP4 and GGUF builds
- OpenMDW-1.1 with open training data and recipes
Innovations
- ★Mamba-2 / MoE hybrid at 3B active parameters
- ★NVFP4 checkpoint with near-BF16 benchmark parity
- ★Open training data, RL environments and post-training recipes
Muse Glimmer
2026-08-1030B dense agentic model distilled from Muse Spark 1.2 and released under Apache 2.0 — Meta's first open-weight release since the Llama 4 family. At 4-bit the checkpoint stays under 20 GB, so the full setup including KV cache and perception encoder fits a 24–32 GB memory envelope on a single consumer GPU or Mac. A separate perception encoder handles image input; block-level speculative decoding keeps latency inside a real agent loop.
Key Capabilities
- 30B dense parameters
- Multimodal input via separate perception encoder
- Agentic tool calling with planning and failure recovery
- Controllable reasoning effort levels
- Trained on 100+ languages
Innovations
- ★Distilled from Muse Spark 1.2
- ★Block-level speculative decoding for agent-loop latency
- ★4-bit K-Quant checkpoint under 20 GB
- ★Interleaved text/image via a dedicated perception encoder
LFM2.5-2.6B
2026-08-04On-device agentic model from Liquid AI, 2.69B parameters with a 131K-token context window and tool calling, trained to work inside real agent harnesses rather than to top knowledge benchmarks. Liquid measures 220 tok/s on an Apple M5 Max, 113 tok/s on an AMD Ryzen AI Max+ 395 and roughly 30 tok/s on a phone, all under 2.5GB of memory, and reports it beating the 4x larger Qwen3.5-9B on ToolSandbox (77.83 vs 76.44) and on every instruction-following benchmark. No MMLU or GPQA figures have been published, so the benchmark fields stay null.
Key Capabilities
- Runs under 2.5GB memory, ~30 tok/s on a phone
- 131K-token context window
- Tool calling for on-device agents
- Open weights
Innovations
- ★Agent-harness training at 2.6B scale
- ★Beats a 4x larger model on tool-use benchmarks
DeepSeek-V4-Flash-0731
2026-07-31General-availability build of DeepSeek-V4-Flash. Identical 284B total / 13B active MoE architecture and size as the April preview — re-post-trained, with the gains concentrated in agentic tool use. 1M-token context, MIT license. On 2026-08-21 DeepSeek added a vision variant on top of this checkpoint, DeepSeek-V4-Flash-Vision-Exp — API-only, without open weights.
Key Capabilities
- 284B total / 13B active MoE
- 1M token context window
- Three reasoning effort modes (incl. Think Max)
- Agentic tool use (Terminal-Bench, NL2Repo, Cybergym, Toolathlon)
- MIT license
Innovations
- ★Hybrid attention: CSA + HCA
- ★Manifold-Constrained Hyper-Connections
- ★Agent-focused re-post-training on the preview checkpoint
Kimi K3
2026-07-162.8T MoE (~50B active, 16 of 896 experts) native multimodal flagship with a 1M-token context window. Open weights released on Hugging Face 2026-07-27 under the bespoke Kimi K3 License — the largest open-weight model at release.
Key Capabilities
- 2.8T total / ~50B active MoE (896 experts, 16 routed)
- 1M token context window
- Natively multimodal (text, image, video)
- Vision-in-the-loop (screenshot inspect + code edit)
- Kimi K3 License (revenue-tiered)
Innovations
- ★Kimi Delta Attention
- ★2.8T Open-Weight MoE (largest open-weight at release)
- ★Vision-in-the-Loop Agent Feedback
- ★1M-Token Native Context
Soofi S 30B-A3B
2026-07-13Sovereign open-weight foundation model from the German SOOFI consortium (Fraunhofer IAIS/IIS, DFKI, TU Darmstadt, hessian.AI and others), coordinated by the KI Bundesverband and funded by the German Federal Ministry for Economic Affairs and Energy. Hybrid Mamba-Transformer MoE (~31.6B total, ~3.2B active), pretrained on ~27T tokens with up-weighted German and trained end-to-end on Deutsche Telekom infrastructure in Munich. Among fully-open models it reports the highest German and English aggregate scores (ahead of Apertus 70B, OLMo 3 32B, EuroLLM and Alia 40B), while still trailing broader open-weight leaders such as Qwen 3.5. Full transparency on data, training and evaluation. Vendor-reported code scores: HumanEval 73.8, MBPP 70.2, German-MBPP 84.2.
Key Capabilities
- German + English focus (sovereign EU model)
- Runs locally (~3.2B active parameters)
- Full training/data/eval transparency
- Trained end-to-end on EU infrastructure
Innovations
- ★Hybrid Mamba-2 / Transformer MoE (128+1 experts, 6 active per token)
- ★Strongest fully-open model on combined German+English benchmarks
- ★Sovereign AI: trained entirely in Germany on Deutsche Telekom compute
Hy3
2026-07-06Tencent's Hunyuan 3 flagship: a 295B-parameter MoE activating 21B per token (192 routed experts with top-8 routing plus an always-active shared expert), 256K context window and three selectable reasoning-effort modes. Full release under Apache 2.0 — unlike the April preview license, without territorial restrictions.
Key Capabilities
- 295B total / 21B active MoE
- 256K token context window
- Three reasoning-effort modes
- Apache 2.0 license
Innovations
- ★Dense-MoE hybrid with always-active shared expert
- ★Multi-Token-Prediction layers (3.8B)
GLM-5.2
2026-06-13753B Mixture-of-Experts (~40B active) open-weight model under MIT license. Introduces the IndexShare attention mechanism (~2.9x lower per-token FLOPs at 1M-token context) and a usable 1M-token context window. Ships with two thinking-effort levels; the highest tier is marketed as 'Max'. Available as open weights on Hugging Face and via the Z.ai API. Vendor-reported benchmarks (maximum thinking effort): GPQA Diamond 91.2, SWE-bench Pro 62.1, Terminal-Bench 2.1 81.0, HLE 40.5 (no tools).
Key Capabilities
- 753B total / ~40B active MoE
- 1M-token context
- Two thinking-effort levels (incl. 'Max')
- Agentic / long-horizon coding
- MIT license
Innovations
- ★IndexShare attention (~2.9x FLOP reduction at 1M-token context)
- ★Usable 1M-token context window
Nemotron 3 Ultra
2026-06-04Largest model of the Nemotron 3 family: 550B total / 55B active hybrid Mamba-Transformer MoE. Released with weights, training data and recipes under OpenMDW-1.1.
Key Capabilities
- 550B total / 55B active (MoE)
- Hybrid Mamba-Transformer
- 262K context (1M with NVFP4)
- OpenMDW-1.1 license
Innovations
- ★Hybrid Mamba-Transformer MoE
- ★Open training data and recipes (OpenMDW-1.1)
Gemma 4 12B
2026-06-03Open-weight 12B multimodal model from Google, released under Apache 2.0. Processes text, images, audio and video in a single encoder-free transformer with a 256K-token context window, and runs locally on a machine with 16GB of RAM or VRAM. On 2026-06-05 Google released Quantization-Aware Training (QAT) checkpoints across the Gemma 4 family (Q4_0 plus a mobile-specialized format), which Google reports cuts memory use by roughly 72%, letting the 12B run on consumer GPUs with little quality loss.
Key Capabilities
- Encoder-free unified multimodal (text, image, audio, video)
- 256K-token context window
- Runs on 16GB RAM/VRAM
- Apache 2.0 license
Innovations
- ★Single encoder-free multimodal transformer
- ★Quantization-Aware Training checkpoints (~72% less memory)
- ★Local multimodality at the 12B size
Command A+
2026-05-20218B sparse Mixture-of-Experts (~25B active) open-weight flagship under Apache 2.0, runnable on as few as 2x H100 GPUs. Enterprise- and RAG-focused with strong multilingual coverage. Available as open weights and via Cohere's API. Early independent benchmarks (Artificial Analysis): GPQA Diamond ~76, HLE ~11, MMMU-Pro 63, Intelligence Index ~37 — flagged as estimates, AA's full evaluation is still ongoing.
Key Capabilities
- 218B sparse MoE / ~25B active
- Apache 2.0 open weights
- Multilingual
- Enterprise / RAG focus
- Runs on 2x H100
Innovations
- ★Apache 2.0 open-weight enterprise flagship
Mistral Medium 3.5
2026-04-30128B dense flagship from Mistral with 256k context, multimodal (text + vision), open weights under a modified MIT license. Self-hostable on 4 GPUs at ~70GB VRAM with Q4 quantization. First Mistral flagship trained on the new 13,800-GPU Paris facility, merging Magistral (reasoning) and Devstral 2 (coding) into one model.
Key Capabilities
- 128B dense parameters
- 256k context window
- Multimodal (text + vision)
- Reasoning + coding unified
- Modified MIT license
Innovations
- ★First flagship trained on Mistral's 13,800-GPU Paris facility
- ★Unified Magistral (reasoning) + Devstral 2 (coding) in one model
- ★Modified MIT release at flagship scale
MiMo-V2.5-Pro
2026-04-27Open-weights flagship from Xiaomi. 1.02T total / 42B active hybrid MoE with interleaved Sliding Window + Global Attention (6:1, 128-token window). 1M context, MIT-licensed for commercial use and fine-tuning without additional authorization. Day-0 SGLang/vLLM support. Reports SWE-bench Pro 57.2 and ClawEval 64% Pass³ at ~70K tokens per trajectory.
Key Capabilities
- 1.02T total / 42B active (hybrid MoE)
- 1M context window
- Hybrid SWA + Global Attention (6:1, 128 window)
- FP8 E4M3 native precision
- MIT license (commercial use + fine-tuning)
Innovations
- ★Interleaved SWA + Global Attention (6:1 ratio)
- ★Lightweight MTP modules with dense FFNs
- ★Day-0 hardware adaptation (Trainium2, ROCm, Kunlun, T-Head, Enflame, Muxi)
- ★Largest fully open-sourced MoE model at release (1.02T total / 42B active, MIT license)
DeepSeek-V4-Pro
2026-04-24Preview release. 1.6T total / 49B active MoE with hybrid attention and manifold-constrained hyper-connections. Reports 27% of single-token inference FLOPs and 10% of KV cache vs DeepSeek-V3.2 in a 1M-token setting.
Key Capabilities
- 1.6T total / 49B active (MoE)
- 1M context window
- Three reasoning effort modes (incl. Think Max)
- MIT license
Innovations
- ★Hybrid attention: CSA + HCA
- ★Manifold-Constrained Hyper-Connections
- ★27% inference FLOPs and 10% KV cache vs V3.2 at 1M tokens
DeepSeek-V4-Flash
2026-04-24Preview release, superseded by DeepSeek-V4-Flash-0731 on 2026-07-31. Smaller 284B total / 13B active MoE sibling of V4-Pro, sharing hybrid attention and hyper-connection architecture.
Key Capabilities
- 284B total / 13B active (MoE)
- 1M context window
- Three reasoning effort modes (incl. Think Max)
- MIT license
Innovations
- ★Hybrid attention: CSA + HCA
- ★Manifold-Constrained Hyper-Connections
- ★Shared architecture with V4-Pro at smaller scale
Kimi K2.6
2026-04-201T MoE (32B active) open-weight successor to K2.5 with native multimodal vision (MoonViT 400M), 256K context, and agent swarm support for up to 300 sub-agents across 4,000 coordinated steps.
Key Capabilities
- 1T total / 32B active MoE (384 experts, 8 routed + 1 shared)
- 256K context window
- Natively multimodal (MoonViT 400M — text, image, video)
- Agent swarms (300 sub-agents, 4,000 coordinated steps)
- Long-horizon coding (Rust, Go, Python)
- Modified MIT license
Innovations
- ★MLA (Multi-head Latent Attention)
- ★300 Sub-Agent Swarm (up from 100 in K2.5)
- ★4,000 Coordinated Tool-Call Steps
- ★Native Video Input
GR00T N1.7
2026-04-17Open vision-language-action model for humanoid robots, released in early access under Apache 2.0 — the first fully commercially licensed model of the GR00T line. A vision-language backbone is paired with a diffusion transformer head that denoises continuous actions via flow matching. NVIDIA credits the generalization and language-following gains over N1.6 to 20,000 hours of EgoScale human egocentric video in pretraining. At 3B parameters it is small enough to run on local hardware, which is why it also appears in the Local tab.
Key Capabilities
- Vision-language-action robot control
- Flow-matching action transformer head
- Humanoid whole-task policies
- Apache 2.0 open weights
Innovations
- ★Open, commercially licensed humanoid VLA
- ★20K hours of EgoScale human video in pretraining
Qwen3.6-35B-A3B
2026-04-16Successor family to Qwen3.5 from the Qwen team, focused on agentic coding. 35B total / 3B active MoE with 262K context window and thinking preservation across conversation turns. Compatible with OpenClaw, Claude Code, and Cline. Apache 2.0 weights on HuggingFace.
Key Capabilities
- 35B total / 3B active MoE
- 262K context window
- Agentic coding focus
- Repository-level reasoning
- Thinking preservation across turns
- Apache 2.0
Innovations
- ★Thinking Preservation Across Conversation
- ★Repository-Level Reasoning
- ★OpenClaw/Claude Code/Cline Compatible
Gemma 4
2026-04-02Google's strongest open-weight model family. Apache 2.0 license. Available in multiple sizes including on-device variants (E2B, E4B) and a 31B dense model for self-hosted deployments. Frontier-competitive on reasoning and coding benchmarks.
Key Capabilities
- 31B dense flagship
- On-device variants (2B, 4B)
- Apache 2.0 license
- Frontier-competitive reasoning and coding
Innovations
- ★Multi-size open-weight family
- ★On-device optimized variants
- ★Apache 2.0 frontier model
GLM-5.1
2026-03-27744B parameter MoE model, 40B active per forward pass. MIT license. Trained on Huawei Ascend hardware.
Key Capabilities
- 744B total / 40B active MoE
- Trained on Huawei Ascend
- MIT license
Innovations
- ★Ascend-native training pipeline
- ★Efficient MoE inference
MiniMax-M2.7
2026-03-17Open-weight agentic MoE: 230B total / 10B active, 256 experts, ~200K context. Coding and tool-use focus.
Key Capabilities
- 230B total / 10B active (MoE)
- 256 experts
- ~200K context window
- Agentic & coding focus
Innovations
Mistral Small 4
2026-03-16119B MoE model (6B active) unifying reasoning, vision, and coding with configurable reasoning effort.
Key Capabilities
- 119B MoE (6B active)
- Configurable reasoning effort
- Multimodal (vision)
- Agentic coding
- 256k context
- Apache 2.0
Innovations
- ★Unified Magistral + Pixtral + Devstral
- ★128-Expert MoE Architecture
- ★40% Latency Reduction vs Small 3
- ★3x Throughput vs Small 3
Qwen3.5
2026-02-16Major architectural upgrade over Qwen3. 397B MoE (17B active) flagship with 1M context, natively multimodal, 8-19x higher decoding throughput. Covers 200+ languages.
Key Capabilities
- 397B total / 17B active MoE (flagship)
- 1M context window
- Natively multimodal
- Reasoning by default
- 200+ languages
- Apache 2.0
Innovations
- ★8-19x Throughput vs Qwen3-Max
- ★Native Multimodal All Sizes
- ★122B-A10B Runs on MacBook 64GB
GLM-5
2026-02-11744B MoE (40B active) open-weight model with DeepSeek Sparse Attention and strong agentic/frontend coding performance.
Key Capabilities
- 744B total / 40B active MoE
- 200K context
- Agentic workflows
- Frontend coding (98% build success)
- MIT license
Innovations
- ★DeepSeek Sparse Attention
- ★Slime Async RL Framework
- ★98% Frontend Build Success Rate
- ★Stealth-Launched as Pony Alpha on OpenRouter
Qwen3-Max-Thinking
2026-01-27Trillion-parameter reasoning model with adaptive tool use and test-time scaling. Highest-capability open model from Alibaba.
Key Capabilities
- Trillion-parameter scale
- Adaptive tool use
- Test-time compute scaling
- Deep chain-of-thought reasoning
Innovations
- ★Trillion-parameter open-weight release
- ★Test-time scaling for reasoning
- ★Adaptive tool-use integration
Kimi K2.5
2026-01-271T MoE (32B active) open-weight model with native multimodal vision and agent swarm support for up to 100 sub-agents.
Key Capabilities
- 1T total / 32B active MoE
- 262K context
- Natively multimodal (MoonViT 400M)
- Agent swarms (100 sub-agents)
- 1,500 parallel tool calls
- Modified MIT license
Innovations
- ★MoonViT Vision Encoder (400M params)
- ★100 Sub-Agent Swarm Support
- ★1,500 Parallel Tool Calls
- ★Top-Ranked Artificial Analysis & LMArena at Launch
EuroLLM-22B
2025-12-14Open multilingual model from the EU-funded EuroLLM / UTTER consortium, covering 35 languages including all 24 official EU languages. Dense decoder-only transformer (22B; a 9B variant released earlier) pretrained on 4T+ tokens with balanced European-language data. Positioned as a sovereign European base model for translation and multilingual understanding. Base and instruct variants on Hugging Face.
Key Capabilities
- 35 languages incl. all official EU languages
- Sovereign EU base model (translation, multilingual)
- Runs locally (22B dense)
- Base + instruct variants
Innovations
- ★Balanced coverage of all official EU languages
- ★EU-funded open consortium model (UTTER project)
Mistral 3 Family
2025-12-02Flagship Mistral Large 3 (675B MoE) and efficient Ministral 3 edge models. Largest open MoE model from Mistral.
Key Capabilities
- 675B MoE flagship + edge variants
- Enterprise and edge deployment
- Multi-language support
- Advanced function calling
Innovations
- ★Largest open MoE from Mistral
- ★Edge-optimized Ministral variants
- ★Improved MoE routing efficiency
DeepSeek-V3.2
2025-12-01Successor to DeepSeek V3. 671B MoE (37B active) with DeepSeek Sparse Attention for long-context efficiency and large-scale agentic tool-use pipeline.
Key Capabilities
- 671B total / 37B active MoE
- 164K context
- DeepSeek Sparse Attention (DSA)
- Agentic tool-use (1,800+ environments)
- Strong reasoning without thinking mode
- MIT license
Innovations
- ★DeepSeek Sparse Attention (DSA)
- ★Large-Scale Task Synthesis Pipeline
- ★1,800+ Agentic Environments
Flux 2.0
2025-11-25Next generation with improved photorealism and typography.
Key Capabilities
- Image reference
- Photorealism
- Typography
- Prompt understanding
Innovations
- ★Flux.2 Pro
- ★Flux.2 Dev
- ★Apache 2.0 License
- ★32B Open Source
DeepSeek-Terminus
2025-09-22Experimental model pushing the boundaries of open weights. Achieved strong reasoning and coding scores rivaling closed models.
Key Capabilities
- Advanced reasoning capabilities
- Strong coding performance
- 128K context window
- Open-weight frontier model
Innovations
- ★Next-generation reasoning architecture
- ★Improved reward modeling
- ★Enhanced self-verification
Apertus 70B
2025-09-02Fully open, transparent, multilingual model family (8B and 70B) from the Swiss AI Initiative (EPFL, ETH Zurich, CSCS). Decoder-only dense transformer pretrained on ~15T tokens with a staged web/code/math curriculum, supporting 1000+ languages, long context, and using only compliant, fully open training data. Weights, data and training recipe are released. Distributed via Hugging Face, Swisscom and the Public AI network.
Key Capabilities
- 1000+ languages (sovereign Swiss/EU model)
- Fully open weights, data and training recipe
- Long-context support
- 8B variant runs on consumer GPUs
Innovations
- ★One of the most transparent open models: compliant, fully open training data
- ★Massively multilingual coverage (1000+ languages)
- ★Trained on Swiss national supercomputer (CSCS Alps)
DeepSeek-V3.1
2025-08-21Incremental update with better instruction following and improved benchmark scores across the board.
Key Capabilities
- 671B total / 37B active parameters (MoE)
- Improved instruction following
- 128K context window
- Enhanced coding and reasoning
Innovations
- ★Refined RLHF alignment
- ★Improved multi-turn consistency
- ★Better long-context utilization
GPT-OSS
2025-08-05OpenAI's first open-weight release, marking a strategic shift. A highly capable open model from the company that pioneered closed frontier AI.
Key Capabilities
- Open-weight model
- Competitive with frontier open models
- Function calling support
- Multilingual
Innovations
- ★First open-weight release from OpenAI
- ★Strategic pivot toward open ecosystem
- ★Optimized for local deployment
Qwen3-Coder
2025-07-15Code-specialized open model optimized for software engineering. Achieved top-tier SWE-bench scores among open-weight models.
Key Capabilities
- Code-specialized architecture
- Top SWE-bench open-weight performance
- Multi-language code generation
- Repository-level understanding
Innovations
- ★Code-optimized training pipeline
- ★Repository-aware context handling
- ★Specialized coding reward model
Gemma 3n E2B
2025-06-26Phone-first open-weight model built on the MatFormer (Matryoshka Transformer) architecture, where the larger E4B model contains a fully functional E2B sub-model trained alongside it. Per-layer embeddings keep the resident footprint near the 2B effective parameters rather than the 5B raw ones, so it runs in about 2GB of memory on a handset. Handles text, image and audio input.
Key Capabilities
- Runs in ~2GB memory on a phone
- Text, image and audio input
- MatFormer elastic inference
- Fully offline via LiteRT / Google AI Edge
Innovations
- ★MatFormer nested sub-model (E2B inside E4B)
- ★Per-layer embeddings cut resident memory
Magistral Small
2025-06-10Mistral's first reasoning model, released as open weights at 24B under Apache 2.0. The larger Magistral Medium sibling announced the same day is API-only and has no published weights — it is tracked as a separate cloud entry.
Key Capabilities
- Reasoning traces in the response
- 128K context window
- Multilingual
- Apache 2.0 license
Innovations
- ★First reasoning model from Mistral
Qwen3 235B MoE
2025-04-29Hybrid thinking model supporting 119 languages. Seamlessly switches between fast responses and deep chain-of-thought reasoning.
Key Capabilities
- 235B total / 22B active parameters (MoE)
- Hybrid thinking (fast + deep reasoning)
- 119 languages supported
- MCP and tool-use support
Innovations
- ★Thinking mode toggle within single model
- ★Massive multilingual coverage (119 langs)
- ★Thinking budget control
- ★4-stage post-training pipeline
Qwen3
2025-04-29235B MoE flagship with hybrid thinking modes.
Key Capabilities
- 235B MoE
- Hybrid thinking
- 119 languages
- 128k context
Innovations
- ★Thinking/Non-thinking Modes
- ★MoE Architecture
- ★Agentic Coding
Llama 4 Scout
2025-04-05Open MoE multimodal model with 16 experts. 10M token context window via iRoPE. Efficient inference at 17B active parameters.
Key Capabilities
- 109B total / 17B active parameters (MoE)
- 16 experts, 1 active per token
- 10M token context window
- Native multimodal (text + image)
Innovations
- ★Interleaved early-fusion for multimodality
- ★MetaP for hyperparameter prediction
- ★Mixture of Experts at massive scale
- ★iRoPE architecture for long context
Llama 4 Maverick
2025-04-05Open MoE multimodal model with 128 experts. Frontier-level quality at efficient inference cost with 17B active parameters.
Key Capabilities
- 400B total / 17B active parameters (MoE)
- 128 experts, 1 active per token
- Frontier reasoning quality
- Native multimodal (text + image)
Innovations
- ★Interleaved early-fusion for multimodality
- ★MetaP for hyperparameter prediction
- ★Mixture of Experts at massive scale
- ★iRoPE architecture for long context
Llama 4 Behemoth
2025-04-052T MoE (288B active) teacher model. Announced alongside Scout/Maverick, still in training at launch. Tops benchmarks in STEM reasoning.
Key Capabilities
- 2T total / 288B active MoE
- 16 experts, 4 active
- STEM reasoning
- Natively multimodal
- In training at announcement
Innovations
- ★Largest Llama Model Ever
- ★Teacher Model Architecture
- ★288B Active Parameters
Command A
2025-03-13111B Dense Enterprise-Flagship (command-a-03-2025), 256k Kontext, 23 Sprachen, Fokus auf Tool-Use, agentische Workflows und Translation. Läuft auf wenigen GPUs, offene Gewichte. Direkter Vorgänger von Command A+. Vectara HHEM 9.3% Halluzination.
Key Capabilities
- 111B Dense
- Enterprise / Tool-Use / Agentic
- 23 Sprachen
- 256k Context
- Open weights
Innovations
- ★Effizientes Enterprise-Flagship
- ★Open weights
Mistral Small 3
2025-01-3024B open-weight model under Apache 2.0 with strong efficiency.
Key Capabilities
- 24B Parameters
- Apache 2.0
- On-device capable
Innovations
- ★Efficient 24B Scale
- ★Apache 2.0 License
- ★Edge Deployment
DeepSeek-R1
2025-01-20First open reasoning model using reinforcement learning. Matched OpenAI o1 on math and coding reasoning tasks without supervised fine-tuning for chain-of-thought.
Key Capabilities
- 671B total / 37B active parameters (MoE)
- 128K context window
- Chain-of-thought reasoning
- Distilled variants (1.5B to 70B)
Innovations
- ★Pure RL-based reasoning (no SFT for CoT)
- ★Group Relative Policy Optimization (GRPO)
- ★Self-verification via reasoning chains
- ★Knowledge distillation to smaller models
DeepSeek-V3
2024-12-26Cost-efficient SOTA model using FP8 mixed-precision training. Achieved GPT-4o-level performance at a fraction of the training cost ($5.5M).
Key Capabilities
- 671B total / 37B active parameters (MoE)
- 128K context window
- Multi-Token Prediction
- Strong math and coding
Innovations
- ★FP8 mixed-precision training
- ★Multi-Token Prediction (MTP)
- ★Auxiliary-loss-free load balancing
- ★DualPipe pipeline parallelism
Llama 3.3
2024-12-06Efficient 70B text model matching Llama 3.1 405B on key benchmarks.
Key Capabilities
- 70B Parameters
- 405B-level performance
- Cost efficient
Innovations
- ★Efficiency Breakthrough
- ★Distillation Techniques
QwQ-32B-Preview
2024-11-27Reasoning-focused model with chain-of-thought capabilities.
Key Capabilities
- 32B Parameters
- Chain-of-thought
- Math reasoning
- Self-reflection
Innovations
- ★Deliberative Reasoning
- ★Extended Thinking
- ★Self-Correction
Flux.1 Tools
2024-11-21Suite of editing tools including Fill, Depth, Canny, and Redux.
Key Capabilities
- Inpainting
- Outpainting
- Depth control
- Edge control
Innovations
- ★Flux.1 Fill
- ★Flux.1 Depth
- ★Flux.1 Canny
- ★Flux.1 Redux
Pixtral Large
2024-11-18124B multimodal model with frontier-class vision capabilities.
Key Capabilities
- 124B Parameters
- Frontier vision
- Multimodal
Innovations
- ★Large-scale Vision
- ★128k Context
- ★Multimodal Reasoning
Stable Diffusion 3.5 Medium
2024-10-29Compact 2.5B parameter variant optimized for consumer GPUs. Runs on hardware with 4-16 GB VRAM.
Key Capabilities
- 2.5B parameters
- Text-to-image
- Consumer GPU optimized (4-16 GB VRAM)
- Open weights (Stability Community License)
Innovations
- ★Efficient MMDiT at 2.5B Scale
- ★Consumer Hardware Target
Aya Expanse 32B
2024-10-2432B multilinguales Open-Weight-Modell aus der Aya-Forschung von Cohere Labs, 23 Sprachen, Schwerpunkt auf non-English-Evaluierungen (Global-MMLU, mArenaHard). Die englische Standard-Suite reportet der Vendor nicht, multilinguale Stärke und offene Gewichte sind der Fokus. Vectara HHEM 10.9% Halluzination.
Key Capabilities
- 32B Dense
- Multilingual (23 Sprachen)
- Open weights
- Aya-Forschungslinie
Innovations
- ★Multilingual-fokussiertes Open-Weight-Modell
- ★Data Arbitrage und multilinguales Preference-Training
Stable Diffusion 3.5 Large
2024-10-22Flagship 8.1B parameter open-weight image model with improved quality and prompt adherence. Available in standard and Turbo variants.
Key Capabilities
- 8.1B parameters
- Text-to-image
- Large + Large Turbo variants
- Open weights (Stability Community License)
Innovations
- ★MMDiT-X Architecture
- ★QK-Normalization
- ★Turbo Distillation Variant
Qwen2.5 72B
2024-09-19Rivaled GPT-4o on key benchmarks. Set a new standard for open-weight models with strong performance across math, coding, and reasoning.
Key Capabilities
- 0.5B to 72B parameter range
- 128K context window
- Structured output support
- Strong math and coding
Innovations
- ★18T token training corpus
- ★Improved synthetic data pipeline
- ★Better long-context handling
Qwen2.5
2024-09-19Flagship 72B model rivaling GPT-4o and Claude 3.5 on key benchmarks.
Key Capabilities
- 72B Parameters
- 128k context
- Structured output
- Tool use
Innovations
- ★Post-training Scaling
- ★Improved Math & Code
- ★Long-context Generation
Pixtral 12B
2024-09-17First multimodal model from Mistral with vision capabilities.
Key Capabilities
- Vision understanding
- 12B Parameters
- Open weights
Innovations
- ★Multimodal Vision
- ★Variable Image Resolution
- ★Native Vision Encoder
FLUX.1 [dev]
2024-08-01Open-weight text-to-image model from Black Forest Labs. Guidance-distilled 12B param model for non-commercial use.
Key Capabilities
- 12B parameters
- Text-to-image
- Guidance distilled
- Open weights (non-commercial)
Innovations
- ★Flow Matching Architecture
- ★Hybrid Transformer (MMDiT)
- ★Guidance Distillation
FLUX.1 [schnell]
2024-08-01Fastest FLUX variant, Apache 2.0 licensed. Optimized for local inference and rapid prototyping.
Key Capabilities
- 12B parameters
- Text-to-image
- Fastest FLUX variant
- Apache 2.0 license
Innovations
- ★Timestep Distillation (4 steps)
- ★Apache 2.0 Open Weight Image Gen
Mistral Large 2
2024-07-24123B parameter open model with strong coding capabilities. Competitive with GPT-4o and Claude 3.5 Sonnet on code generation tasks.
Key Capabilities
- 123B parameters
- 128K context window
- 80+ coding languages
- Function calling and JSON mode
Innovations
- ★Instruction-following improvements
- ★Reduced hallucination rate
- ★Improved multi-turn conversation
Llama 3.1 405B
2024-07-23First open-weight frontier-class model. The 405B variant matched GPT-4 on many benchmarks, proving open models could compete at the highest tier.
Key Capabilities
- 8B / 70B / 405B parameters
- 128K context window
- Tool use and function calling
- Multilingual (8 languages)
Innovations
- ★FP8 quantization for training
- ★128K context via progressive training
- ★Iterative DPO alignment
Llama 3.1
2024-07-23Massive 405B open-source model competing with frontier closed models.
Key Capabilities
- 405B Parameters
- 128k Context
- Tool use
- Multilingual
Innovations
- ★405B Open Weights
- ★128k Context
- ★Synthetic Data Scaling
Mistral Nemo
2024-07-1812B parameter model developed in collaboration with NVIDIA.
Key Capabilities
- 12B Parameters
- Apache 2.0
- 128k context
Innovations
- ★NVIDIA Collaboration
- ★Tekken Tokenizer
- ★Efficient 12B Scale
Stable Diffusion 3
2024-06-12Next-gen architecture with Multimodal Diffusion Transformer.
Key Capabilities
- MMDiT architecture
- Improved text rendering
- Better composition
Innovations
- ★Multimodal Diffusion Transformer
- ★Flow Matching
- ★Triple Text Encoders
Qwen2 72B
2024-06-07First non-Western top-tier open model. Demonstrated that frontier-class open-weight models could come from outside the US/EU ecosystem.
Key Capabilities
- 0.5B / 1.5B / 7B / 72B parameters
- 128K context window
- 29 languages supported
- Strong math and coding
Innovations
- ★YARN-based context extension
- ★Dual-chunk attention
- ★Multilingual data curation pipeline
Qwen2
2024-06-07Major upgrade with 72B flagship, competitive with leading Western models.
Key Capabilities
- 72B Parameters
- 128k context
- Multilingual (29 languages)
Innovations
- ★GQA Architecture
- ★Dual Chunk Attention
- ★YARN Scaling
DeepSeek-V2
2024-05-01Strong performance at a lower cost.
Key Capabilities
- MLA Architecture
- GPT-4 class
- Extremely cheap API
Innovations
- ★Multi-head Latent Attention (MLA)
- ★KV Cache Compression (93.3%)
- ★DeepSeekMoE Architecture
Command R+
2024-04-24104B Dense, RAG- und Tool-Use-optimiertes Open-Weight-Flagship (CC-BY-NC), 128k Kontext, citation-fähiges Retrieval. Starkes Grounding beim Zusammenfassen (Vectara HHEM 6.9% Halluzination für den 08-2024-Build), akademische Benchmarks dagegen schwach. Direkter Vorläufer der Command-A-Linie.
Key Capabilities
- 104B Dense
- RAG-optimiert
- Tool-Use
- Multilingual (10 Sprachen)
- 128k Context
Innovations
- ★Citation-fähiges RAG
- ★Open weights (CC-BY-NC)
Llama 3
2024-04-18Closed the quality gap with proprietary models significantly. The 70B variant rivaled GPT-3.5 Turbo on many tasks.
Key Capabilities
- 8B / 70B parameters
- 8K context window
- Tiktoken-based 128K vocabulary
- Strong code and reasoning
Innovations
- ★128K token vocabulary (4x Llama 2)
- ★15T token training corpus
- ★Improved post-training alignment
Grok-1
2024-03-17314B MoE (86B active) open-weight model released under Apache 2.0. Largest open-weight model at the time of release.
Key Capabilities
- 314B total / 86B active MoE
- 8 experts, 2 active
- Apache 2.0 license
- Largest open-weight model at release
Innovations
- ★Largest Open-Weight MoE at Launch
- ★Apache 2.0 Full Weight Release
Mistral Large
2024-02-26Flagship model with top-tier reasoning capabilities.
Key Capabilities
- 32k Context
- Multi-lingual
- Strong reasoning
Innovations
- ★32k Context Window
- ★Multi-lingual Excellence
- ★Function Calling
Qwen 1.5
2024-02-04Multi-size release from 0.5B to 110B with strong multilingual support.
Key Capabilities
- 0.5B–110B sizes
- Multilingual
- 32k context
Innovations
- ★Scalable Model Family
- ★RLHF Alignment
- ★Extended Context
DeepSeek-MoE
2024-01-01Mixture-of-Experts architecture.
Key Capabilities
- MoE efficiency
- High throughput
- Cost effective
Innovations
- ★MoE Efficiency
- ★Cost-effective Training
- ★High Throughput
Mixtral 8x7B
2023-12-11First mainstream open Mixture-of-Experts model. Used 8 expert networks with only 2 active per token, achieving 70B-quality at 13B inference cost.
Key Capabilities
- 46.7B total / 12.9B active parameters
- 8 expert networks, 2 active per token
- 32K context window
- Multilingual (EN, FR, IT, DE, ES)
Innovations
- ★Sparse Mixture of Experts (SMoE)
- ★Expert routing with top-2 gating
- ★Efficient inference via sparse activation
DeepSeek-Coder
2023-11-02Open-source code generation model.
Key Capabilities
- Code specialization
- Open source
- Various sizes
Innovations
- ★Code Specialization
- ★Multiple Model Sizes
- ★Open Source Focus
Mistral 7B
2023-09-27Efficiency breakthrough from Mistral AI. A 7B model that outperformed Llama 2 13B on all benchmarks using Grouped-Query Attention and Sliding Window Attention.
Key Capabilities
- 7B parameters
- Outperforms Llama 2 13B
- 8K context with sliding window
- Apache 2.0 license
Innovations
- ★Sliding Window Attention (SWA)
- ★Grouped-Query Attention (GQA)
- ★Rolling buffer KV cache
Code Llama
2023-08-24Code-specialized Llama model for code generation and understanding.
Key Capabilities
- Code generation
- Infilling
- Instruction following
Innovations
- ★Code Specialization
- ★Infilling Capability
- ★Long Context Fine-tuning
Qwen-7B
2023-08-03Alibaba's first open-source large language model.
Key Capabilities
- 7B Parameters
- Multilingual
- Code generation
Innovations
- ★Open-source Chinese LLM
- ★Multilingual Training
- ★Tool Use
Stable Diffusion XL
2023-07-26Major upgrade with native 1024x1024 resolution and improved generation.
Key Capabilities
- 1024x1024 native
- Two-stage pipeline
- Improved aesthetics
Innovations
- ★Higher Resolution
- ★Refiner Model
- ★Enhanced Prompt Understanding
Llama 2
2023-07-18First commercially licensed open LLM. Made open-weight models viable for business use and kickstarted the open-source LLM ecosystem.
Key Capabilities
- 7B / 13B / 70B parameters
- Commercial license
- 4K context window
- Chat-tuned variants (Llama 2-Chat)
Innovations
- ★Grouped-Query Attention (GQA) in 70B
- ★Ghost Attention (GAtt) for multi-turn
- ★RLHF with rejection sampling
LLaMA
2023-02-24First open-weight LLM from Meta, released to researchers. Proved that smaller open models could match much larger proprietary ones.
Key Capabilities
- 7B / 13B / 33B / 65B parameters
- Competitive with GPT-3 at smaller scale
- Research-only license
Innovations
- ★RMSNorm pre-normalization
- ★Rotary positional embeddings (RoPE)
- ★SwiGLU activation function
Stable Diffusion
2022-08-22Revolutionary open-source text-to-image model that democratized AI art.
Key Capabilities
- Text-to-Image
- Open source
- Local inference
- Community ecosystem
Innovations
- ★Latent Diffusion Model
- ★Open Source Democratization
- ★Diffusers Ecosystem