VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which model leads HiL-Bench?
Across 15 models scored on HiL-Bench, Opus 5 leads at 57, ahead of Claude Fable 5 at 56.3. The median tracked score is 35.3, and the field spans 4.3 to 57.

HiL-Bench

Across 15 models scored on HiL-Bench, Opus 5 leads at 57, ahead of Claude Fable 5 at 56.3. The median tracked score is 35.3, and the field spans 4.3 to 57.

15 models tracked
Data as of August 24, 2026
#ModelVendorBest scoreRunsLast seen
1Opus 5Anthropic5712026-08-24
2Claude Fable 5Anthropic56.312026-08-24
3GLM-5.2 Full Open SourceZ.ai43.712026-08-24
4Claude Opus 4.7Anthropic41.712026-08-24
5GPT-5.5OpenAI39.712026-08-24
6Claude Opus 4.6Anthropic38.312026-08-24
7Gemini 3.1 ProGoogle35.312026-08-24
8Claude Opus 4.8Anthropic35.312026-08-24
9GPT-5.6 SolOpenAI32.312026-08-24
10Gemini 3.5 FlashGoogle27.712026-08-24
11Grok 4.20SpaceXAI2012026-08-24
12Kimi K2.6Moonshot18.712026-08-24
13GPT-5.4OpenAI9.712026-08-24
14MiniMax M2.5MiniMax6.312026-08-24
15GPT-5.3 CodexOpenAI4.312026-08-24

Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.