Muse Spark 1.3
Meta·
Overview
Fourth Muse Spark release in five months, live at launch in Muse Code and the Meta Model API. Meta frames the update as an efficiency change as much as a capability one: for the same agentic work it reports about 20% fewer tool calls and about 25% fewer tokens, with fewer turns and less verbose output. 1M-token context, multimodal input; xhigh reasoning effort is the generally available tier, a max tier is in limited preview for Meta partners. Vendor-reported (Meta launch scorecard, September 2026): Terminal-Bench 2.1 88.8, DeepSWE v1.1 75.4, SWEAtlas CodeBase QnA 59.4, MRCR 256K-512K 98.5 and MRCR 512K-1M 98.1. API pricing unchanged from 1.1 and 1.2 at $1.25 / $4.25 per million input/output tokens with cached input at $0.15; the contributor tier (over 90% discount in exchange for training on prompts and completions) continues.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 94% | third party GPQA Diamond 94% at xhigh (max: also 94%), measured by Artificial Analysis. The article rounds to whole percent, so no decimal is recorded here - the value is deliberately less precise than the 1.2 entry's 90,4%, which came from AA's model page. Meta reports no GPQA value for 1.3. |
| Humanity’s Last Exam (no tools) | 47% | third party HLE 47% no-tools at xhigh (max: 49%), measured by Artificial Analysis, again rounded to whole percent. Caveat: the same table lists 1.2 at 45% while this dataset carries 43,9 from AA's 1.2 model page - more than rounding explains. AA now scores under Intelligence Index v4.1.1 (the 1.2 value was taken under v4.1), so the 1.2 column in that article is most likely a re-run. Flagged rather than reconciled; both entries keep the figure from the source they were taken from. |
| Terminal-Bench 2.x | 88.8% | vendor Terminal-Bench 2.1 88,8% from Meta's launch scorecard (Muse Spark 1.3 in the Muse Code harness at xhigh effort). Kept as primary so the 1.1 (76,2) - 1.2 (82,9) - 1.3 line stays inside one harness, exactly as the 1.2 entry does. The scorecard on the blog page is an image and not machine-readable from this environment; the figure was read from launch coverage quoting it (officechai, investing.com, 2026-09-02), which also reports GPT-5.6 Sol at the same 88,8. Not verified on the official Terminal-Bench leaderboard. |
- SWE-bench Pro: independent evaluation pending — Meta again reports DeepSWE v1.1 (75,4%) and SWEAtlas CodeBase QnA (59,4%) instead of SWE-bench Pro. Same situation as for 1.2: the Scale AI SWE-bench Pro leaderboard, which supplied the 1.1 value (61,5%), carries no 1.3 entry as of 2026-09-03.
- MMLU-Pro: independent evaluation pending — Not reported by Meta, and Vals AI has no Muse Spark 1.3 run as of 2026-09-03 (its newest Muse Spark page is 1.1). The 1.2 entry has the same gap.
- MMMU-Pro: not reported by the vendor
API pricing
$1.25 per million input tokens, $4.25 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-09-03)
Reliability
- tau3_banking
- 47 ()third party
Links
More from Meta
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.