VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

HumanEval+

28 models tracked

HumanEval+ (EvalPlus rigorous test-case extension) — Well-known stricter-test-suite extension of HumanEval, distinct benchmark from base HumanEval.

#ModelVendorBest scoreRunsLast seen
1Phi 4 Reasoning PlusMicrosoft92.312026-08-23
2Granite 3.3 8B InstructIBM86.122026-08-23
3Granite 3.1 8B InstructIBM85.812026-06-11
4Granite 3.2 8B InstructIBM85.712026-06-11
5Granite 4.1 30BIBM85.412026-06-15
6Kimi K2 BaseMoonshot84.852026-06-13
7Phi 4Microsoft82.812026-08-23
8Seed Coder 8B InstructByteDance82.312026-06-05
9Llama 3.1 8B InstructMeta80.212026-06-11
10Granite 4.1 8BIBM79.912026-06-15
11DeepSeek-R1-Distill-Qwen-7BDeepSeek78.412026-06-11
12Granite 3.3 2B InstructIBM75.712026-06-11
13MiMo V2.5 Pro BaseXiaomi75.632026-06-05
14MiMo V2.5 ProXiaomi75.612026-08-23
15Granite 3.1 2B InstructIBM75.312026-06-11
16Granite 3.2 2B InstructIBM73.412026-06-11
17Step3.5 Flash BaseStepFun7222026-06-05
18MiMo V2.5 BaseXiaomi71.332026-06-05
19MiMo V2 Flash BaseXiaomi70.742026-06-13
20DeepSeek-V3.2-Exp-BaseDeepSeek67.742026-06-13
21Qwen3.5 2BAlibaba6612026-08-15
22DeepSeek-V3.1-BaseDeepSeek64.642026-06-13
23DeepSeek-R1-Distill-Llama-8BDeepSeek62.912026-06-11
24Granite 4.1 8B BaseIBM62.822026-06-15
25Granite 4.1 30B BaseIBM62.222026-06-15
26Granite 4.1 3B BaseIBM54.312026-06-15
27Qwen3.5 0.8BAlibaba46.812026-08-15
28ERNIE 4.5Baidu2512026-08-23

Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.