AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Muse Glimmer

Meta·

LLMsopen-weightcloud + localEdge / Consumer

Overview

30B dense agentic model distilled from Muse Spark 1.2 and released under Apache 2.0 — Meta's first open-weight release since the Llama 4 family. Served through the API as well as downloadable: at 4-bit the checkpoint stays under 20 GB, so the full setup including KV cache and perception encoder fits a 24–32 GB envelope on a single consumer GPU or Mac. A separate perception encoder handles image input; block-level speculative decoding keeps latency inside a real agent loop. 131K-token context.

Capabilities and innovations

30B dense parametersMultimodal input via separate perception encoderAgentic tool calling with planning and failure recoveryControllable reasoning effort levels131K-token contextApache 2.0Distilled from Muse Spark 1.2Block-level speculative decoding for agent-loop latency4-bit K-Quant checkpoint under 20 GBMeta's first open-weight release since Llama 4

Benchmarks

BenchmarkScoreSource
GPQA Diamond83.5%vendor
GPQA Diamond 83.5 from Meta's launch comparison table, mirrored from the local entry so both tabs read alike. Artificial Analysis independently measures 83.54 at high effort, i.e. the vendor figure holds.
Humanity’s Last Exam (no tools)21.96%third party
HLE 21.96% no-tools, measured by Artificial Analysis at high effort (accessed 2026-08-17). Fills the gap the local entry flagged as evaluation_pending: Meta shows HLE only in an image table with an unreachable methodology report.
SWE-bench Verified76%vendor
SWE-bench Verified 76.0 from Meta's launch table. Era-1 benchmark, recorded only because Meta reports it.
SWE-bench Pro51.2%vendor
SWE-bench Pro 51.2 from Meta's launch table. Artificial Analysis does not run SWE-bench Pro, so no independent cross-check exists.
Terminal-Bench 2.x51.7%vendor
Terminal-Bench 2.1 51.7 from Meta's launch table. Artificial Analysis independently measures 51.69 at high effort, i.e. the vendor figure holds.
MMMU-Pro74%vendor
MMMU-Pro 74 from Meta's launch table. Artificial Analysis independently measures 74.34 at high effort.
  • MMLU-Pro: independent evaluation pending — Not identifiable in Meta's launch table (published as an image); neither Vals AI nor Artificial Analysis reports MMLU-Pro for this model as of 2026-08-17.
  • Terminal-Bench 3.0: not reported by the vendor

API pricing

$0.35 per million input tokens, $1.5 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-08-17)

Architecture and hardware

Parameters
30B total (Dense)
Estimated VRAM at Q4
~18 GB, Edge / Consumer class

Links

More from Meta

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.