Muse Spark 1.2
Meta·
Overview
Coding-focused update of the Muse Spark frontier family, released alongside and co-trained with the Muse Code terminal agent. Third Meta Superintelligence Labs model in four months. Meta evaluates it inside its own Muse Code harness at xhigh reasoning effort. Same API pricing as Muse Spark 1.1 at $1.25 / $4.25 per million input/output tokens, with cached input at $0.15; a contributor tier cuts the rate by over 90% in exchange for permission to train on prompts and completions. Rate limits reach 3.000 requests and 4M tokens per minute per team.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 90.4% | third party GPQA Diamond 90,4%, measured by Artificial Analysis (xhigh, Intelligence Index v4.1 component). Meta reports no GPQA value for 1.2; no Vals AI run available. |
| Humanity’s Last Exam (no tools) | 43.9% | third party HLE 43,9% no-tools, measured by Artificial Analysis (xhigh). Same harness as the 1.1 value in this dataset, so the two are directly comparable. Meta reports no HLE value for 1.2. |
| Terminal-Bench 2.x | 82.9% | vendor Primary source: Meta's own launch chart, Terminal-Bench 2.1 82,9% for Muse Spark 1.2 in the Muse Code harness at xhigh effort (same chart lists the verified 76,2% for 1.1 with mini-swe-agent). No verified entry on the official Terminal-Bench leaderboard as of 2026-08-06; Artificial Analysis measures 80,1% in its own harness. Meta additionally reports DeepSWE v1.1 59,3% here; the Muse Spark 1.3 scorecard later restates the same model as 55.0, which is why deepswe_v1_1 is null on this entry. |
- SWE-bench Pro: independent evaluation pending — Meta reports DeepSWE v1.1 instead of SWE-bench Pro for 1.2; see the deepswe_v1_1 entry below for why that value is not recorded. The Scale AI SWE-bench Pro leaderboard, which supplied the 1.1 value, has no 1.2 entry as of 2026-08-06.
- MMLU-Pro: independent evaluation pending — Not reported by Meta. Vals AI has a Muse Spark 1.2 page but has so far only run three evaluations on it; MMLU Pro (present for 1.1) is not among them as of 2026-08-06.
- MMMU-Pro: not reported by the vendor
- DeepSWE v1.1: independent evaluation pending — Meta reports DeepSWE v1.1 instead of SWE-bench Pro, but the two Meta scorecards disagree about this model: the Muse Spark 1.2 launch scorecard (2026-08-05) gives 59.3, the Muse Spark 1.3 scorecard (2026-09-02) restates the same model on the same benchmark version as 55.0. Neither scorecard names the harness, and DeepSWE's own runs are fixed to mini-swe-agent while Meta reports its Terminal-Bench figures out of its own Muse Code harness — the most likely explanation, but not one Meta states. Launch coverage inherited the conflict: the 1.3 press round quoted a 16-point jump (from 59.3) while Meta's own chart implies 20.4 (from 55.0). Field left null until one of the two figures can be attributed to a named harness; entering either would date a Meta-internal restatement as progress. Datacurve's own board was not reachable from this environment.
API pricing
$1.25 per million input tokens, $4.25 per million output tokens (USD, provider list price, checked 2026-08-06)
Links
More from Meta
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.