Muse Spark 1.1
Meta·
Overview
Second Meta Superintelligence Labs model and the launch vehicle for the paid Meta Model API. Multimodal reasoning model built for agentic work: takes text, image, video, audio and PDF input, returns text, with a 1M-token context window. Runs as main agent that plans and delegates or as a subagent, and generalises zero-shot to new tools, MCP servers and custom skills. Free for consumers in the Meta AI app (Thinking mode); developer preview US-only at launch. $1.25 / $4.25 per million input/output tokens.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 91.2% | independent GPQA Diamond 91,16% ± 2,10 (rank 17/129), measured by Vals AI. Meta reports no GPQA value. Artificial Analysis measures 89,8% in its own run. |
| Humanity’s Last Exam (no tools) | 45.1% | third party HLE 45,1% no-tools, measured by Artificial Analysis (xhigh, Intelligence Index v4.1 component). Used as primary value because it is the only independent measurement and deviates >3 points from Meta's self-report. |
| SWE-bench Pro | 61.5% | vendor Meta evaluation report, Figure 44: SWE-Bench Pro 61,5%. The report attributes the value to the Scale AI SWE-bench Pro leaderboard (731 tasks, mini-swe-agent harness) rather than an internal run. |
| MMLU-Pro | 88.7% | independent MMLU Pro 88,73% ± 0,32 (rank 13/128), measured by Vals AI. Meta reports no MMLU-Pro value. |
| Terminal-Bench 2.x | 76.2% | independent Primary source: official Terminal-Bench 2.1 leaderboard, verified entry mini-swe-agent + Muse Spark 1.1 at 76,2% ± 1,2% (2026-07-09). Meta itself cites this 76,2% for 1.1 in its Muse Code launch chart of 2026-08-05. |
- MMMU-Pro: not reported by the vendor
API pricing
$1.25 per million input tokens, $4.25 per million output tokens (USD, provider list price, checked 2026-08-06)
Links
More from Meta
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.