A reporting layer
over evaluation
infrastructure.
Evaluation Cards is a collection of reported model–benchmark results, organized under a five-level rollout hierarchy and four interpretive signals computed over the joined record.
Interpretive signals
Four signals computed over each (model, benchmark, metric-path) record and aggregated to the corpus level. Per-record instances appear on every model and benchmark page.
of reported scores have a complete setup recorded. The rest cannot be independently re-run.
94% have at least one undocumented field. Most often missing: max tokens (93%), temperature (92%).
mean across 61,723 reported score triples.
Observed range: 7% to 93%.
of reported score triples have reports from more than one party.
95% third-party, 5% first-party of 61,723 unique triples.
of setup-eligible groups diverge across variants (343 of 862).
Cross-party divergence: 6%.
Benchmark families
All 81 →Mercor ACE
1 reported benchmark across this family.
AgentHarm
1 reported benchmark across this family.
Alpaca-EVAL-V1
4 reported benchmarks across this family.
Alpaca-EVAL-V2
1 reported benchmark across this family.
AlpacaEval
4 reported benchmarks across this family.
Mercor APEX
2 reported benchmarks across this family.
Every score resolves to an explicit path through this hierarchy, so aggregate claims drill down to the evidence supporting them.