VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which model leads DocVQA_test?
Across 13 models scored on DocVQA_test, Qwen3 VL 235B A22B Instruct leads at 97.1, ahead of Qwen3 VL 32B Instruct at 96.9. The median tracked score is 95.3, and the field spans 92.8 to 97.1.

DocVQA_test

Across 13 models scored on DocVQA_test, Qwen3 VL 235B A22B Instruct leads at 97.1, ahead of Qwen3 VL 32B Instruct at 96.9. The median tracked score is 95.3, and the field spans 92.8 to 97.1.

13 models tracked
Data as of August 25, 2026
#ModelVendorBest scoreRunsLast seen
1Qwen3 VL 235B A22B InstructAlibaba97.112026-08-25
2Qwen3 VL 32B InstructAlibaba96.912026-08-25
3Qwen2 VL 72B InstructAlibaba96.522026-08-25
4Qwen3 VL 235B A22B ReasoningAlibaba96.512026-08-25
5Qwen3 VL 8B InstructAlibaba96.112026-08-25
6Qwen3 VL 32B ReasoningAlibaba96.112026-08-25
7Qwen3 VL Thinking (8B)Alibaba95.312026-08-25
8Qwen3 VL 4B InstructAlibaba95.312026-08-25
9Claude 3.5 SonnetAnthropic95.212026-08-24
10Qwen3 VL 30B A3B InstructAlibaba9512026-08-25
11Qwen3 VL 30B A3B ReasoningAlibaba9512026-08-25
12Qwen3 VL 4B (Reasoning)Alibaba94.212026-08-25
13GPT-4oOpenAI92.812026-08-24

Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.