Artifact walkthrough

Change the file. Check the evidence.

Observed in this synthetic run
CheckExitWhat it tells us
Unchanged referenced files0Match the saved baseline
After mutation · bundle only0Bundle still internally consistent
After mutation · with run root1Referenced score file has changed

The second pass is as important as the failure: the command’s scope determines what it can tell you.

01

Start with a claim small enough to check

The claim is that verification can detect a changed referenced file relative to a saved evidence bundle. We will not run a model, reproduce a benchmark score, or decide whether a reported answer is true.

This uses Eval Evidence’s public, deterministic synthetic demo. The demo starts with accuracy 0.75 and n = 4. These are fixture values, not research findings. The walkthrough changes the score file to accuracy 1.0 while preserving the saved bundle.

02

Install the exact source and run the example

The source is pinned to 6d4a25b5f288f9646f30e0d1c9f5923cc6c1ec8c, which identifies itself as 0.2.0rc1. It is a development candidate, not a claim of a stable final release. This receipt was produced with Python 3.14.7 and jsonschema 4.26.0.

Use Python 3.11–3.14 and curl. Download the two files below into a fresh working directory, then run these commands. Installation needs internet access; the demonstration itself makes no model calls or uploads. The script creates a new temporary workspace and leaves its synthetic files there for inspection.

python3 -m venv .venv
. .venv/bin/activate
curl -fL --max-time 60 -o eval-evidence.tar.gz "https://codeload.github.com/edward-lcl/eval-evidence/tar.gz/6d4a25b5f288f9646f30e0d1c9f5923cc6c1ec8c"
python -m pip install -r requirements.txt ./eval-evidence.tar.gz
python reproduce.py > receipt.json

03

Save the baseline before changing anything

The script runs the following command sequence inside its new workspace. Check reports on the current state; bundle saves a baseline; verify with --run-root compares that baseline against the selected local files. The original file check passed with exit code 0.

The optional missing debug file remains explicitly unavailable. A successful integrity check does not mean that every potentially useful piece of evidence was recorded.

python -m eval_evidence demo -o run
python -m eval_evidence check run
python -m eval_evidence bundle run -o evidence.json
python -m eval_evidence verify evidence.json --run-root run

04

Change one file, then ask two different questions

The script changes only outputs/scores.json from {accuracy: 0.75, n: 4} to {accuracy: 1.0, n: 4}. It leaves evidence.json unchanged and checks its SHA-256 before and after. Then it invokes verify twice.

Without --run-root, verification still returns exit 0: it checks the bundle’s schema and internal digest, not the current source files. With --run-root, it returns exit 1 and identifies both a file-size mismatch and a file-digest mismatch for outputs/scores.json. That expected failure is the result we wanted to reproduce.

python -m eval_evidence verify evidence.json
# exit 0; referenced_files_checked: false

python -m eval_evidence verify evidence.json --run-root run
# exit 1; referenced_files_checked: true
# Referenced file size mismatch for outputs/scores.json
# Referenced file digest mismatch for outputs/scores.json

05

Inspect the receipt, and keep the boundary

The linked receipt contains the captured JSON output and exit status for each step. Only temporary workspace paths were replaced with <workspace>; the inputs and saved bundle were not normalized or edited for the check. Your temporary path will differ.

This proves a byte difference against this saved baseline. It does not prove which score is correct, who changed the file, whether the verifier measures the intended task, or who created the bundle. An unsigned bundle can be altered and its digest recomputed. Keeping a trusted baseline is a separate responsibility.

That distinction is the point of the walkthrough: keep reported outcomes, provenance, retained bytes and trust claims separate enough that another person can check exactly what is being asserted.