A reporting layer
over evaluation
infrastructure.

Evaluation Cards is a collection of reported model–benchmark results, organized under a five-level rollout hierarchy and four interpretive signals computed over the joined record.

Corpus snapshot · August 18, 2026
6,482
Models
Tracked across reporting sources
209,382
Reported results
(model, benchmark, metric) triples
44
Reporting organizations
Distinct evaluator initiatives in this corpus
901
Model developers
Distinct model-publishing organizations
81
Benchmark families
Top of the rollout hierarchy
889
Single benchmarks
1,174 slices · 1,242 metrics

Interpretive signals

Four signals computed over each (model, benchmark, metric-path) record and aggregated to the corpus level. Per-record instances appear on every model and benchmark page.

Benchmark families

All 81
Five-level rollout hierarchy

Every score resolves to an explicit path through this hierarchy, so aggregate claims drill down to the evidence supporting them.