Replicable from the harness's published prompts and scoring code. The strongest mark a score can wear.
Model card, leaderboard table, registry entry — a neutral party re-attesting a primary source.
System card, model card, launch post. Authoritative on the model — not independently replicated.
Launch-post comparison tables, a blog quoting a competitor's eval. Kept, attributed — and never the headline.
Research papers, conference talks, technical briefs. Surfaced with its source — the lowest rung, worn openly.
200+ sources poll on schedule — leaderboards, system cards, audited harnesses, papers.
Aliases collapse onto one canonical model and one canonical benchmark; most-recent wins at the database layer.
Per-axis manifestos define the basket. Composites are reliability-weighted percentiles, shrunk toward the median for thin samples.
Every score carries its provenance mark; within-model disagreement surfaces as a band, never averaged away.
The boards render straight from the scored view — no editorial overlay between the data and the page.
737 models carry a permalink on Vector Wire, which is 271,216 possible pairings. Our sitemap submits 1,770 of them — 0.65% — the 60 most-benchmarked models crossed pairwise, and nothing else.
That ratio describes the URLs we submit for indexing. The comparison route itself answers any two tracked models, and a model page links the comparisons most relevant to that model whether or not the pair is submitted — a reader following a link is never sent to a dead end. What is deliberately bounded is what we ask a search engine to index.