Muse Glimmer
Meta·
Overview
30B dense agentic model distilled from Muse Spark 1.2 and released under Apache 2.0 — Meta's first open-weight release since the Llama 4 family. Served through the API as well as downloadable: at 4-bit the checkpoint stays under 20 GB, so the full setup including KV cache and perception encoder fits a 24–32 GB envelope on a single consumer GPU or Mac. A separate perception encoder handles image input; block-level speculative decoding keeps latency inside a real agent loop. 131K-token context.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 83.5% | vendor GPQA Diamond 83.5 from Meta's launch comparison table, mirrored from the local entry so both tabs read alike. Artificial Analysis independently measures 83.54 at high effort, i.e. the vendor figure holds. |
| Humanity’s Last Exam (no tools) | 21.96% | third party HLE 21.96% no-tools, measured by Artificial Analysis at high effort (accessed 2026-08-17). Fills the gap the local entry flagged as evaluation_pending: Meta shows HLE only in an image table with an unreachable methodology report. |
| SWE-bench Verified | 76% | vendor SWE-bench Verified 76.0 from Meta's launch table. Era-1 benchmark, recorded only because Meta reports it. |
| SWE-bench Pro | 51.2% | vendor SWE-bench Pro 51.2 from Meta's launch table. Artificial Analysis does not run SWE-bench Pro, so no independent cross-check exists. |
| Terminal-Bench 2.x | 51.7% | vendor Terminal-Bench 2.1 51.7 from Meta's launch table. Artificial Analysis independently measures 51.69 at high effort, i.e. the vendor figure holds. |
| MMMU-Pro | 74% | vendor MMMU-Pro 74 from Meta's launch table. Artificial Analysis independently measures 74.34 at high effort. |
- MMLU-Pro: independent evaluation pending — Not identifiable in Meta's launch table (published as an image); neither Vals AI nor Artificial Analysis reports MMLU-Pro for this model as of 2026-08-17.
- Terminal-Bench 3.0: not reported by the vendor
API pricing
$0.35 per million input tokens, $1.5 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-08-17)
Architecture and hardware
- Parameters
- 30B total (Dense)
- Estimated VRAM at Q4
- ~18 GB, Edge / Consumer class
Links
More from Meta
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.