Each opens the head-to-head page: shared benchmarks, price and the receipt chain side by side.
Strong on Long Context and Factuality; capable on Reasoning and Coding +1 more; limited on Math and Instruction Following — at $1.25/M input across 3 providers. 1 paper references this model. Insufficient data on 3 of 10 axes.
| Provider | Model name | Input $/M | Output $/M | Effective | Current |
|---|---|---|---|---|---|
| direct ↗ | muse-spark-1-1 | $1.25 | $4.25 | 2026-09-02 | ● |
| meta ↗ | muse-spark-1.1 | $1.25 | $4.25 | 2026-09-02 | ● |
| meta-official ↗ | muse-spark-1.1 | $1.25 | $4.25 | 2026-09-01 | ● |
| Benchmark | Score | Settings | Source | Verification | Methodology | Measured | Reference |
|---|---|---|---|---|---|---|---|
aa_lcr | 81.33 | x-high | artificialanalysis.ai | AGGREGATOR | AGGR | 55m ago | view ↗ |
aa_omniscience | 28.08 | x-high | artificialanalysis.ai | AGGREGATOR | AGGR | 55m ago | view ↗ |
omniscienceNonHallucination | 50.02 | x-high | artificialanalysis.ai | AGGREGATOR | AGGR | 55m ago | view ↗ |
aa_intelligence_index | 53.20 | x-high | artificialanalysis.ai | AGGREGATOR | AGGR | 55m ago | view ↗ |
terminalbenchV21 | 77.90 | x-high | artificialanalysis.ai | AGGREGATOR | AGGR | 55m ago | view ↗ |
aa_coding_index | 71.34 | x-high | artificialanalysis.ai | AGGREGATOR | AGGR | 55m ago | view ↗ |
aa_hle | 46.20 | x-high | artificialanalysis.ai | AGGREGATOR | AGGR | 55m ago | view ↗ |
tauBanking | 31.75 | x-high | artificialanalysis.ai | AGGREGATOR | AGGR | 55m ago | view ↗ |
gpqa | 89.80 | x-high | artificialanalysis.ai | AGGREGATOR | AGGR | 55m ago | view ↗ |
| 15.14 | x-high | artificialanalysis.ai | AGGREGATOR | AGGR | 55m ago | view ↗ | |
aa_agentic_index | 39.74 | x-high | artificialanalysis.ai | AGGREGATOR | AGGR | 55m ago | view ↗ |
| 43.57 | x-high | artificialanalysis.ai | AGGREGATOR | AGGR | 55m ago | view ↗ | |
aa_scicode | 58.22 | x-high | artificialanalysis.ai | AGGREGATOR | AGGR | 55m ago | view ↗ |
omniscienceAccuracy | 52.05 | x-high | artificialanalysis.ai | AGGREGATOR | AGGR | 55m ago | view ↗ |
swe_bench_pro | 61.50 | — | labs.scale.comswe_bench_pro_public | VERIFIED | AGGR | 55m ago | view ↗ |
| 75.30 | — | labs.scale.commultichallenge | VERIFIED | OTHER | 56m ago | view ↗ | |
| 88.10 | — | labs.scale.commcp_atlas | VERIFIED | OTHER | 56m ago | view ↗ | |
| 1490 | — | lmarena.ai | VERIFIED | AGGR | 56m ago | view ↗ | |
BabyVision | 76.30 | — | api.llm-stats.comscores | VENDOR | OTHER | 58m ago | view ↗ |
| 53.00 | — | api.llm-stats.comscores | VENDOR | OTHER | 58m ago | view ↗ | |
| 57.20 | — | api.llm-stats.comscores | VENDOR | OTHER | 58m ago | view ↗ | |
| 80.80 | — | api.llm-stats.comscores | VENDOR | OTHER | 58m ago | view ↗ | |
Toolathlon | 75.60 | — | api.llm-stats.comscores | VENDOR | OTHER | 58m ago | view ↗ |
Job Bench | 54.70 | — | api.llm-stats.comscores | VENDOR | OTHER | 58m ago | view ↗ |
livebench_language@2026-06-25 | 74.34 | x-high | livebench.ai | VERIFIED | OTHER | 23h ago | view ↗ |
livebench_math@2026-06-25 | 87.14 | x-high | livebench.ai | VERIFIED | OTHER | 23h ago | view ↗ |
livebench_instruction_following@2026-06-25 | 69.64 | x-high | livebench.ai | VERIFIED | OTHER | 23h ago | view ↗ |
livebench_data_analysis@2026-06-25 | 72.55 | x-high | livebench.ai | VERIFIED | OTHER | 23h ago | view ↗ |
livebench_coding@2026-06-25 | 77.16 | x-high | livebench.ai | VERIFIED | OTHER | 23h ago | view ↗ |
livebench_agentic_coding@2026-06-25 | 58.54 | x-high | livebench.ai | VERIFIED | OTHER | 23h ago | view ↗ |
livebench_reasoning@2026-06-25 | 87.73 | x-high | livebench.ai | VERIFIED | OTHER | 23h ago | view ↗ |
| 57.79 | — | epoch.aibenchmarks.csv | VERIFIED | OTHER | 1d ago | view ↗ | |
| 1292.67 | — | datasets-server.huggingface.coleaderboard-dataset | VERIFIED | OTHER | 8d ago | view ↗ | |
CharXiv Reasoning (w/ tools) | 88.40 | — | ai.meta.commuse-spark-1-1-evaluation-report | VENDOR | OTHER | 43d ago | view ↗ |
HealthBench Professional | 59.30 | — | ai.meta.commuse-spark-1-1-evaluation-report | VENDOR | OTHER | 43d ago | view ↗ |
| 1381.00 | — | ai.meta.commuse-spark-1-1-evaluation-report | VENDOR | OTHER | 43d ago | view ↗ | |
DeepSearchQA | 84.90 | — | ai.meta.commuse-spark-1-1-evaluation-report | VENDOR | OTHER | 43d ago | view ↗ |
| 69.00 | — | ai.meta.commuse-spark-1-1-evaluation-report | VENDOR | OTHER | 43d ago | view ↗ | |
| 14.20 | — | ai.meta.commuse-spark-1-1-evaluation-report | VENDOR | OTHER | 43d ago | view ↗ | |
Toolathlon-Verified | 75.60 | — | ai.meta.commuse-spark-1-1-evaluation-report | VENDOR | OTHER | 43d ago | view ↗ |
| 54.10 | — | ai.meta.commuse-spark-1-1-evaluation-report | VENDOR | OTHER | 43d ago | view ↗ | |
| 62.10 | — | ai.meta.commuse-spark-1-1-evaluation-report | VENDOR | OTHER | 43d ago | view ↗ | |
| 1539.69 | — | datasets-server.huggingface.coleaderboard-dataset | VERIFIED | OTHER | 48d ago | view ↗ | |
| 23.40 | — | ai.meta.commuse-spark-1-1-evaluation-report | VENDOR | OTHER | 51d ago | view ↗ | |
| 4.80 | — | ai.meta.commuse-spark-1-1-evaluation-report | VENDOR | OTHER | 51d ago | view ↗ | |
| 77.00 | — | ai.meta.commuse-spark-1-1-evaluation-report | VENDOR | OTHER | 51d ago | view ↗ | |
CyberGym (pass@1) | 59.00 | — | ai.meta.commuse-spark-1-1-evaluation-report | VENDOR | OTHER | 51d ago | view ↗ |
WMDP-Chem | 87.00 | — | ai.meta.commuse-spark-1-1-evaluation-report | VENDOR | OTHER | 51d ago | view ↗ |
SWE-Bench Pro | 61.50 | — | api.llm-stats.comscores | VENDOR | OTHER | 58m ago | view ↗ |
| 80.00 | — | api.llm-stats.comscores | VENDOR | OTHER | 58m ago | view ↗ | |
Humanity's Last Exam | 62.10 | — | api.llm-stats.comscores | VENDOR | OTHER | 58m ago | view ↗ |
| 88.10 | — | api.llm-stats.comscores | VENDOR | OTHER | 58m ago | view ↗ | |
| 53.30 | — | ai.meta.commuse-spark-1-1-evaluation-report | VENDOR | OTHER | 43d ago | view ↗ | |
Humanity's Last Exam (no tools) | 52.20 | — | ai.meta.commuse-spark-1-1-evaluation-report | VENDOR | OTHER | 43d ago | view ↗ |