Pre-existing pooled aggregate of Finance Agent module/overall scores (already merged in DB) — DB rollup mixes 0.7-0.8 fractions with 57.9-64.4 percentages — scale-incompatible; flag.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.6 | Anthropic | 76.7 | 4 | 2026-08-23 |
| 2 | Claude Opus 4.7 | Anthropic | 71.5 | 5 | 2026-08-23 |
| 3 | Claude Sonnet 4.6 | Anthropic | 63.3 | 1 | 2026-08-23 |
| 4 | GPT-5.4 Pro | OpenAI | 61.5 | 2 | 2026-06-21 |
| 5 | GPT-5.5 | OpenAI | 60 | 2 | 2026-08-23 |
| 6 | Gemini 3.1 Pro | 59.7 | 3 | 2026-06-21 | |
| 7 | Gemini 3.5 Flash | 57.9 | 3 | 2026-08-23 | |
| 8 | GPT-5.4 | OpenAI | 57.2 | 3 | 2026-08-23 |
| 9 | Claude Opus 4.5 | Anthropic | 55.2 | 1 | 2026-07-29 |
| 10 | Claude Opus 4.8 | Anthropic | 53.9 | 2 | 2026-08-23 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.