VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

MMLU-Pro (HF Open LLM)

23 models tracked

MMLU-Pro as measured by the HF Open LLM Leaderboard v2 (5-shot direct). Values are RAW accuracy, recovered from the published normalization (raw = norm×0.9+10; baseline 10%, verified). Protocol match to standard MMLU-Pro unconfirmed → unpooled.

#ModelVendorBest scoreRunsLast seen
1Qwen 14B B100Alibaba51.812026-07-04
2Yi 1.5 34B 32K01.AI47.112026-07-04
3Yi 1.5 34B01.AI46.712026-07-04
4Yi 1.5 34B Chat 16K01.AI45.512026-07-04
5Yi 1.5 34B Chat01.AI45.212026-07-04
6Marco-o1AIDC-AI41.212026-07-04
7Yi 1.5 9B Chat 16K01.AI39.912026-07-04
8Yi 1.5 9B Chat01.AI39.812026-07-04
9Yi 1.5 9B01.AI39.212026-07-04
10Qwen2.5 7B Test NovelistAlibaba38.712026-07-04
11Yi 1.5 9B 32K01.AI37.612026-07-04
12Yi 9B 200K01.AI36.212026-07-04
13Yi 9B01.AI35.712026-07-04
14smartllama3.1-8B-001Meta34.912026-07-04
15Yi 1.5 6B Chat01.AI31.912026-07-04
16Yi 1.5 6B01.AI31.412026-07-04
17NuminaMath-7B-CoTAI-MO28.712026-07-04
18Qwen2.5 1.5B Continuous LearntAlibaba28.112026-07-04
19NuminaMath-7B-TIRAI-MO27.312026-07-04
20Llama 3 Instruct 8BMeta2612026-07-04
21Yi Coder 9B Chat01.AI24.312026-07-04
22Llama Squared 8BMeta23.712026-07-04
23Llama 3.1 8B SquarerootMeta17.512026-07-04

Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.