VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

Every number traces back.

Provenance on every row
Loading the audit…
The life of a numberA live row — today's reasoning leader's top-weighted basket entry, traced end to end
Source → published rank · refetched on every loadNothing publishes unsourced
Tracing today's leader…
Below-threshold confidence never reaches this wire — it routes to review, not publication. Missing data prints an em-dash, never a guess.
The ladderFive provenance marks — the same ◆ every score wears across the terminal
IVIndependently verified

A third party ran the model.

Replicable from the harness's published prompts and scoring code. The strongest mark a score can wear.

HELM · LiveBench · SWE-Bench Verified · LMArena evals
AAAggregator attested

An aggregator surfaced the score.

Model card, leaderboard table, registry entry — a neutral party re-attesting a primary source.

Hugging Face · Papers with Code · model registries
VAVendor attributed

The maker's own number.

System card, model card, launch post. Authoritative on the model — not independently replicated.

OpenAI system cards · Anthropic model cards · vendor blogs
CRVendor cross-reference

A vendor re-states a rival's number.

Launch-post comparison tables, a blog quoting a competitor's eval. Kept, attributed — and never the headline.

Demoted below first-party attribution when the headline score is selected
SASource attributed

Extracted from prose, not re-run.

Research papers, conference talks, technical briefs. Surfaced with its source — the lowest rung, worn openly.

arXiv papers · conference talks · community evals
The pipelineFive stages, source poll → published board · runs without human hands
01

Ingest

200+ sources poll on schedule — leaderboards, system cards, audited harnesses, papers.

Sources register on first detection
02

Canonicalize

Aliases collapse onto one canonical model and one canonical benchmark; most-recent wins at the database layer.

One row of truth per (model, benchmark)
03

Score

Per-axis manifestos define the basket. Composites are reliability-weighted percentiles, shrunk toward the median for thin samples.

Two bases, toggleable everywhere
04

Verify

Every score carries its provenance mark; within-model disagreement surfaces as a band, never averaged away.

Low confidence → review, not publication
05

Publish

The boards render straight from the scored view — no editorial overlay between the data and the page.

Provenance + reference URL on every row
The basketsWhat each axis composites · weights ratified per quarter · retirements owned openly
Manifestos: agent-drafted, founder-ratified, revisited each quarter · deprecations listed with reasons in every manifesto← The frontier board
Which pages we publishSelection rules, not permutations

737 models carry a permalink on Vector Wire, which is 271,216 possible pairings. Our sitemap submits 1,770 of them — 0.65% — the 60 most-benchmarked models crossed pairwise, and nothing else.

  • ComparisonsSubmitted only for the 60 models carrying the most tracked benchmarks, crossed pairwise. A submitted comparison is withheld from indexing when the two models share fewer than 10 comparable benchmarks — below that there is not enough overlap to compare, and we would rather publish nothing than pad a table.
  • Benchmark boardsA leaderboard is published once at least 5 published models carry a comparable default-variant score on it. Internal and scaffold evaluations are excluded outright. Every board we publish is listed here.
  • Model pagesOne page per published model. No variants, no per-benchmark permutations, no per-vendor duplicates of the same model.

That ratio describes the URLs we submit for indexing. The comparison route itself answers any two tracked models, and a model page links the comparisons most relevant to that model whether or not the pair is submitted — a reader following a link is never sent to a dead end. What is deliberately bounded is what we ask a search engine to index.