Muse Spark
Muse Spark is capable in instruction following, factuality, long context, multimodal tasks, reasoning, and coding; and behind the leaders in agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Math or Multilingual.
Price
No current price is tracked for this model. See the rate card
Evidence
46results on40benchmarks
- 9 independently verified
- 19 aggregator
- 14 vendor-reported
- 4 cross-referenced
From 12 sources · latest Oct 8, 2026 · How verification works
Research
3 papers reference Muse SparkMuse Spark benchmark results
46 results on 40 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
14.0% behind the leader2 of 3 ranked benchmarks measured
- 75.52Oct 8, 2026
- IFBench75.92Oct 8, 2026aa_ifbench
14.7% behind the leader3 of 4 ranked benchmarks measured
- 66.30May 20, 2026
- AA-Omniscience · Accuracy49.60Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination15.81Oct 8, 2026omniscienceNonHallucination
14.9% behind the leader1 of 3 ranked benchmarks measured
- 78.00Oct 8, 2026
18.3% behind the leader2 of 6 ranked benchmarks measured
- CharXiv (reasoning)86.40Oct 7, 2026CharXiv-R
- MMMU-Pro80.52Oct 8, 2026aa_mmmu_pro
Show 3 more multimodal resultsHide 3 multimodal results
- 1305.40Aug 25, 2026
- 1295Jun 17, 2026
- 80.40Oct 7, 2026
19.3% behind the leader4 of 6 ranked benchmarks measured
- GPQA Diamond88.38Oct 8, 2026gpqa
- Humanity's Last Exam40.69Oct 8, 2026aa_hle
- ARC-AGI-242.50Oct 7, 2026ARC-AGI v2
- 11.33Oct 8, 2026
Show 3 more reasoning resultsHide 3 reasoning results
- GPQA Diamond89.50Oct 7, 2026GPQA
- 58.40Oct 7, 2026
- 58.00Jun 4, 2026
23.1% behind the leader6 of 10 ranked benchmarks measured
- 77.40Oct 7, 2026
- 1539.95May 22, 2026
- SciCode51.50Sep 4, 2026aa_scicode
- Terminal-Bench Hard45.45Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench 2.162.17Oct 8, 2026terminalbenchV21
- 52.40Oct 7, 2026
Show 1 more coding resultHide 1 coding result
- 55.00Oct 8, 2026
31.9% behind the leader3 of 7 ranked benchmarks measured
- 82.20Oct 8, 2026
- τ-Bench V3 · Banking19.59Aug 10, 2026tauBanking
- 25.05Oct 8, 2026
Show 1 more agentic resultHide 1 agentic result
- 16.21Jun 16, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 1489Oct 8, 2026
- frontiermath_tier_4_v114.60May 20, 2026frontiermath_tier_4
- AA Intelligence31.30Oct 8, 2026aa_intelligence_index
- 7.17Oct 8, 2026
- τ²-Bench Telecom (AA run)91.52Oct 8, 2026aa_tau2
- Artificial Analysis Coding Index58.62Sep 9, 2026aa_coding_index
Show 11 more resultsHide 11 results
- 28.69Aug 10, 2026
- CyberGym43.50Jul 16, 2026CyberGym (pass@1)
- DeepSearchQA (F1)74.80Oct 7, 2026DeepSearchQA
- 64.70Oct 7, 2026
- 50.30Jul 16, 2026
- 8.00Jul 16, 2026
- ScreenSpot-Pro (No tools)84.10Oct 7, 2026ScreenSpot Pro
- 71.30Oct 7, 2026
- 59.00Oct 7, 2026
- WMDP85.60Jul 16, 2026WMDP-Chem
- 33.00Oct 7, 2026
Muse Spark: common questions
Who makes Muse Spark?
Muse Spark is made by Meta.
When was Muse Spark released?
Muse Spark was released on Apr 8, 2026, according to Artificial Analysis.
What is Muse Spark good at?
Muse Spark is capable in instruction following, factuality, long context, multimodal tasks, reasoning, and coding; and behind the leaders in agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.
How many benchmarks has Muse Spark been tested on?
We track 46 results for Muse Spark on 40 benchmarks from 12 sources, 9 of them independently verified. The latest was recorded on Oct 8, 2026.
About this record
Where Muse Spark's numbers come from, and every name it appears under.
- Tracked since
- Apr 25, 2026
- Newest source mention
- Jul 13, 2026
Where the results come from
Verification: 46 scores · 9 independently verified · 19 aggregator-attributed · 4 vendor cross-reference · 14 vendor-reported. How these tiers are assigned
From 12 sources on 7 sites. Artificial Analysis supplies 19 of them; the 9 independently verified results come from 4 sites. Bars are coloured by trust tier.
- artificialanalysis.ai19
- api.llm-stats.com13
- ai.meta.com5
- labs.scale.com3
- datasets-server.huggingface.co2
- epoch.ai2
- lmarena.ai2