VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

AI benchmark leaderboards

Vector Wire publishes 449 benchmark leaderboards, each ranking published models with the source and verification tier shown on every score. The most measured is GPQA Diamond at 442 models; coverage across the set runs from 5 to 442. A board is published once at least five published models carry a comparable score on it.

449 boards·90 families
Most measuredRanked by models scored
#LeaderboardModels scoredLatest score
1GPQA Diamond4422026-08-24
2AA Intelligencecomposite4142026-08-22
3HLE3942026-08-24
4SciCode3782026-08-24
5AA-LCR3332026-08-23
6CritPt3192026-08-23
7Artificial Analysis Coding Indexcomposite3182026-08-22
8IFBench3172026-08-24
9AA-Omnisciencecomposite3162026-08-22
10OmniScience Accuracy3162026-08-22
11OmniScience Non-Hallucination3162026-08-22
12AA Agentic Indexcomposite2992026-08-22
13τ²-Bench Telecom (AA run)2932026-07-20
14GDPVal2912026-08-24
15Terminal-Bench Hard2872026-08-23
16MMLU-Procomposite1872026-08-24
17Terminal-Bench 2.11652026-08-24
18MMMU-Pro1622026-08-24
19TauBench V3 - Banking1542026-08-22
20GSM8K1492026-08-24
21SWE-bench Verified1482026-08-24
22MMLUcomposite1472026-08-24
23AIME 20251212026-08-24
24τ³-Bench1082026-08-23
Benchmark familiesBoards sharing a permalink stem — the same suite at different versions, subsets or context lengths

gsm8k2

deepswe2

docvqa2

Individual boards141 leaderboards with no sibling suite · A–Z

Counts are the number of published models carrying a comparable default-variant score. A board page ranks the top 100 of them. Internal and scaffold evaluations are excluded. How we choose what to publish.