01
One number, several worlds
A zero may mean the task is genuinely difficult. It may also mean the specification is incomplete, the environment is broken, the verifier recognizes only one privileged path, or the harness suppressed a capability the model otherwise has.
Those worlds are operationally different and statistically identical if all we publish is reward.
- Task validity: was the thing being asked coherent and feasible?
- Instrument validity: which model, harness, prompt assembly, and provider state produced the trace?
- Verifier validity: did the reward correspond to the task’s actual success condition?
- Temporal validity: would the same named system behave the same way next week?
02
Provenance is part of the result
A credible score should travel with its task revision, verifier, environment, harness, trajectory, and adjudication record. This is not administrative metadata. It is the evidence that tells us what kind of statement the score permits.
Epistemic correctness begins by refusing to promote an observation into a capability claim before its alternatives have been tested.
03
An evaluation-integrity lens
A task-quality audit can separate genuinely hard cases from specification-hard, verifier-hard, environment-hard, and construction-induced difficulty. Its purpose is not to rescue or attack one benchmark; it is to build the apparatus that lets any benchmark explain its own zeros.