tests/ proves the code works. This measures whether the model's verdicts
are right, which is a different question and needs human labels.
Read DESIGN.md before trusting any number out of here. The short
version: what ships today is a demonstration of five cases from a synthetic
manuscript, labelled by the same person who wrote the prompts. It shows the
harness runs. It measures nothing about real manuscripts.
python evals/runners/score_only.py \
--gold evals/gold/demo_v1.gold.json \
--results evals/gold/demo_v1.observed.json \
--case-dir demo_case \
--out evals/runsWrites eval.json (machine-readable) and EVAL.md (readable) into a run
directory. No model call, no network.
--case-dir <case> points the frozen-source check at that case's
refs_manifest.json. A gold not_retrieved case only tests retrieval for as
long as the DOI stays paywalled, so the runner compares each source's
expected_ref_status against what the manifest actually observed. A source
whose status changed invalidates the cases resting on it; they are excluded
from the judgement metrics, kept in the record, and listed in their own report
section.
Omit --case-dir and the check still runs — it reports every expected
source as not verified and says so on the report's face. It does not exclude
anything, because unverifiable is not the same as wrong. Earlier this flag
gated the check entirely, and this README's own documented invocation omitted
it, so the freeze check never ran and its silence was indistinguishable from a
pass.
Compare repeated runs:
python evals/runners/score_only.py --agreement evals/runs/<a> evals/runs/<b> evals/runs/<c>Reports the intersection (upper bound — only cases every run produced) and the union (lower bound — every case seen in any run, with the gaps charged to the model) side by side, and names which cases each run omitted. Runs from two different gold sets are refused outright rather than averaged.
python evals/runners/run_eval.py --gold evals/gold/demo_v1.gold.json --runs 3Prompts for confirmation with the call count, fetches sources over the network,
and defaults its case directory to a fresh path under $TMPDIR — outside this
repo, so claude -p cannot pick up the project's own skills and settings as
ambient context. It refuses to start under CI.
| Path | What it is |
|---|---|
DESIGN.md |
the protocol: unit, pairing, metrics, provenance, gold policy |
PROPOSAL.md |
draft text for the paired-benchmark issue |
align.py |
gold ↔ prediction matching, global score-sorted assignment |
tool_coverage.py |
shape-defensive reading of the tool's own coverage audit |
eligibility.py |
which cases may be scored, and if not, why |
metrics.py |
pure metric arithmetic, no I/O |
agreement.py |
repeated-run stability |
provenance.py |
what produced a run — model, prompt hash, commit |
scoring.py |
orchestration → the eval record |
eval_report.py + templates/ |
the Markdown artefact and its guardrails |
gold/ |
frozen gold sets and the transcribed demo run |
runners/ |
score_only.py (offline) · run_eval.py (live, paid) |
tests/ |
deterministic tests of the arithmetic — these do run in CI |
runs/ |
output; gitignored except for .gitkeep |
Validate against
schemas/eval_gold.schema.json; the tests
also check that every gold_anchor_phrases entry is a substring of its
decisive_passage, which catches authoring typos that would otherwise look
like model errors.
Two rules worth repeating:
uncheckedis not a legal gold verdict. It is a harness error, never a correct answer.- If labellers cannot agree, set
gold_verdict: nulland keep the case. Deleting it makes the set easier and inflates every score. The case is excluded from every denominator byeligibility.py, stays inper_case, and is listed under Excluded cases on the report.