AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Muse Spark 1.3

Meta·

LLMsflagshipcloud

Overview

Fourth Muse Spark release in five months, live at launch in Muse Code and the Meta Model API. Meta frames the update as an efficiency change as much as a capability one: for the same agentic work it reports about 20% fewer tool calls and about 25% fewer tokens, with fewer turns and less verbose output. 1M-token context, multimodal input; xhigh reasoning effort is the generally available tier, a max tier is in limited preview for Meta partners. Vendor-reported (Meta launch scorecard, September 2026): Terminal-Bench 2.1 88.8, DeepSWE v1.1 75.4, SWEAtlas CodeBase QnA 59.4, MRCR 256K-512K 98.5 and MRCR 512K-1M 98.1. API pricing unchanged from 1.1 and 1.2 at $1.25 / $4.25 per million input/output tokens with cached input at $0.15; the contributor tier (over 90% discount in exchange for training on prompts and completions) continues.

Capabilities and innovations

Agentic codingLong-horizon tool useMultimodal input1M-token contextReasoning effort control (xhigh; max in limited preview)Token and tool-call efficiency as a stated release goal, at unchanged price and contextFourth frontier release from Meta Superintelligence Labs in five months

Benchmarks

BenchmarkScoreSource
GPQA Diamond94%third party
GPQA Diamond 94% at xhigh (max: also 94%), measured by Artificial Analysis. The article rounds to whole percent, so no decimal is recorded here - the value is deliberately less precise than the 1.2 entry's 90,4%, which came from AA's model page. Meta reports no GPQA value for 1.3.
Humanity’s Last Exam (no tools)47%third party
HLE 47% no-tools at xhigh (max: 49%), measured by Artificial Analysis, again rounded to whole percent. Caveat: the same table lists 1.2 at 45% while this dataset carries 43,9 from AA's 1.2 model page - more than rounding explains. AA now scores under Intelligence Index v4.1.1 (the 1.2 value was taken under v4.1), so the 1.2 column in that article is most likely a re-run. Flagged rather than reconciled; both entries keep the figure from the source they were taken from.
Terminal-Bench 2.x88.8%vendor
Terminal-Bench 2.1 88,8% from Meta's launch scorecard (Muse Spark 1.3 in the Muse Code harness at xhigh effort). Kept as primary so the 1.1 (76,2) - 1.2 (82,9) - 1.3 line stays inside one harness, exactly as the 1.2 entry does. The scorecard on the blog page is an image and not machine-readable from this environment; the figure was read from launch coverage quoting it (officechai, investing.com, 2026-09-02), which also reports GPT-5.6 Sol at the same 88,8. Not verified on the official Terminal-Bench leaderboard.
  • SWE-bench Pro: independent evaluation pending — Meta again reports DeepSWE v1.1 (75,4%) and SWEAtlas CodeBase QnA (59,4%) instead of SWE-bench Pro. Same situation as for 1.2: the Scale AI SWE-bench Pro leaderboard, which supplied the 1.1 value (61,5%), carries no 1.3 entry as of 2026-09-03.
  • MMLU-Pro: independent evaluation pending — Not reported by Meta, and Vals AI has no Muse Spark 1.3 run as of 2026-09-03 (its newest Muse Spark page is 1.1). The 1.2 entry has the same gap.
  • MMMU-Pro: not reported by the vendor

API pricing

$1.25 per million input tokens, $4.25 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-09-03)

Reliability

tau3_banking
47 ()third party

Links

More from Meta

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.