Skip to content

Verdict harness: judge every rendered claim, over every state that renders one - #6

Merged
systemslibrarian merged 2 commits into
mainfrom
harness/split-point
Sep 30, 2026
Merged

systemslibrarian merged 2 commits into
mainfrom
harness/split-point

Conversation

@systemslibrarian

Copy link
Copy Markdown
Owner

The verdict-harness work, dormant on a local branch since 2026-09-22, routed onto main. Preserved first at 8f39f7cd on origin/verdict-harness, kept permanently at tag verdict-harness-preserved-2026-09-22.

What it does

  • Every rendered claim is judged, over the whole set of states that render one — not over the states a spec happened to visit.
  • Coverage is taken from the assertions that actually ran, recorded as they execute, rather than from what the source appears to contain.

Verification on the rebased branch

Gate Result
npm test 24/24 (3 files, with coverage)
npm run build clean
Playwright 23/23 — claims 9, flows 5, verdict-coverage 6, verdict-ledger 1, plus a11y

Mutation spot-check. Pinning all four computed data-status sites to "pass" (source diff confirmed non-empty: 4 lines):

Error: collusion-recovery data-status
expect(locator).toHaveAttribute(expected) failed

Restored, baseline back to 23/23, tree clean. collusion-recovery is the right marker to fail on — it is the one that says α became visible when both keys reached one server.

Rebase

main was already an ancestor of the branch, so nothing replayed and no conflict arose. No package.json or lockfile touched.

🤖 Generated with Claude Code

systemslibrarian and others added 2 commits September 21, 2026 18:41
Five of the six lane fixes apply here. Two gaps were confirmed by running the
mutation against the suite as it stood at ab679b9 and watching it pass.

Fix 1 — one helper asserts words, state and plate together.
e2e/verdict-assertions.ts adds expectVerdict(page, id, {text, status, tone}) and
expectClaim(page, id, {value, text}), and the coverage test now requires every
recorded mutation's assertedBy spec to call the helper on that marker, instead of
merely mentioning the id. This lab set className and dataset.status in separate
expressions, so pinning the class green while the words and the status still
followed the run was invisible: at ab679b9 that mutation passed 21/21. tone is
checked as the painted background, because .verdict switches class and
.mini-verdict switches on [data-status].

Fix 2 — measurements join the coverage loop on the same terms.
The page rendered eighteen-odd numbers and not one carried a marker, so the
coverage rule enforced itself over one family of a page that had two. Eight
data-claim markers now carry the measurement in data-value, are registered in
e2e/verdict-mutations.json under "claims", and fail the build both ways: a
rendered measurement with no record, and a record naming a measurement the page
no longer renders.

Fix 3 — the digit-plus-unit scan.
findStrayMeasurements walks the result regions for digit-plus-unit text and for
bare integers in stats cells that sit outside a marker, with its own injection
self-test so it cannot pass by being blind. Hex dumps are exempt as raw material
(claims.spec checks them byte for byte against the size markers), and labels and
explanatory prose are exempt the same way the verdict-word scan exempts them.
The one static number on the page that was really a claim — "AES-128-CTR → 2(λ +
1) bits" — is now measured from a real expandSeed() and marked.

Fix 4 — the oracle multiplied where the artifact iterates.
The old key-size test asserted root + 17 * domainBits + final === measured at the
one default domain: one per-level size multiplied by a level count, both
literals, where serializeKey WRITES one correction word per level. It could not
tell "sixteen words of seventeen bytes" from "one 272-byte block", which is
exactly what the parts list claims. The oracle now measures the per-level cost as
the slope between domains, derives λ and the root from it, counts the correction
words actually present in the dump by walking it in stride, and SUMS them.

Fix 5 — branch protection untouched. Confirmed read-only at the API
(build + verdict-coverage required, enforce_admins false, force-push disabled),
which is why this is a pull request.

Fix 6 — driveEveryState visits every option of every control.
Per control, not the cross-product. #alpha-range all 16 indices; the expansion
stepper all 5 levels down and back; #new-keys once; #size-domain all 4 domains;
#collusion-toggle both states, and again through a fetch for the tile it paints;
#tamper-toggle both states with its own fetch; #fetch-record throughout. Two
controls are not enumerable option sets and are recorded as such: #shelf-alpha is
a 65,536-value number input, so its two rendering classes stand in — the in-range
fetch and the rejection that renders no result — and the .key-inspector
disclosure is opened because both scans skip elements with no client rects, so a
closed disclosure hid its contents from the denominator entirely. No control on
this page was judged not to change what renders.

That widening is load-bearing, not tidiness: the formula's total was a hard
literal correct only at 65,536 records, and at ab679b9 that mutation also passed
21/21 because nothing ever moved #size-domain.

Baseline before any edit: 21 passed (15.0s). After: 23 passed (31.7s), plus
24 unit tests. Every mutation below was run with CI=1 so Playwright started its
own server on the pinned port 4698.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UH9YUUhdeWXJz141Zy8FxW
D6, in this lab. Fix 1 as it shipped at 521d03d did not satisfy Fix 1. The
rule was FILE-granular and read with a regex: a record naming
e2e/claims.spec.ts was satisfied by the string expectVerdict(page,
'record-match' appearing ANYWHERE in that file — inside a comment, inside a
call handed the page's own values, inside a call asserting a different state
in a different test. And the shipped tree already contained an instance:
tree-point's recorded mutation was killed by two raw assertions that never
reached the helper, so the expectedFlip its record describes was a flip no
run had ever demonstrated.

expectVerdict/expectClaim now append the (spec, test, marker) triple of every
call they EXECUTE to a run-scoped ledger under test-results/verdict-ledger/,
cleared once per run in e2e/global-setup.ts. e2e/verdict-ledger.spec.ts reads
it back and fails when a record's triple is not among them. Playwright runs
each spec in its own worker process, so a module-level Set cannot aggregate;
the sink is one append-only file per process, and the reader is its own
project with dependencies: ['claims'] rather than an ordinary test. That
project is what npm run test:verdicts runs now, so the required check reads
the ledger with the specs that write it — .github/ is untouched.

assertedBy stops being a filename. It is {spec, test, status, says}: the
triple, plus the part of the expectation the RECORD decides — the outcome the
killing assertion must require, and one fragment it must hand the helper
verbatim (a "re:" prefix is a RegExp source). That second half is the only
thing that can catch a call kept but made tautological, and it has to work
that way round: on an unmutated tree an expectation read off the page and one
decided in advance are the same values, so no amount of observing the run
separates them. Only something decided before the run can.

The plate is no longer an argument. tone: 'pass' | 'alarm' is gone and the
plate is derived from the asserted status through one table, because a reader
is never shown an alarm outcome on the success plate — so a test cannot
expect one either. That closes the residue: a tautologist who has read the
record and hands the declared fragment still cannot mask record-match's
mutation, because the plate follows the status it read rather than the paint
it read.

Two recorded mutations were killed by a precursor rather than by the helper,
and both are reordered so the helper is reached first:

- tree-point — reconstruction[alpha], the lit-bit count and the #tree-xor
  readback all failed on the direction-bit mutation before line 232's
  expectVerdict ran.
- key-size-formula — found by running the set: the formulaNumbers comparison
  caught the pinned-290 mutation at claims.spec.ts:209, before expectClaim.
  Through the helper it now fails on its own terms, "data-value 86 is not in
  17 + (17 × log₂ 16) + 1 = 290 bytes".

Escapes. Each run in an isolated tree archived from this commit, node_modules
symlinked, port moved to 4931, spec-only rewrites with the built bundle
IDENTICAL before and after (index-RVWjpVch.js 3a8b3a777965e297ba523fe90f68143b),
baseline 21 passed / exit 0 in the same tree:

- Commented out — verdict-ledger fails: "e2e/verdict-mutations.json says its
  mutation is killed by expectVerdict in e2e/claims.spec.ts › 'tampering
  violates integrity…', and that call did not execute this run. expectVerdict
  ran for record-match in: e2e/claims.spec.ts › honest PIR output equals the
  indexed shelf record (status 'pass')."
- Kept but tautological — expectVerdict fails in the test itself: "declares
  the words this assertion must require — 'RETRIEVED — AND WRONG' — and it
  was handed ['RETRIEVED — AND WRONG · 1 of 16 bytes differ from shelf[α]; no
  PIR authentication failure was raised']." With the declared fragment handed
  over and the mutation applied, it fails one layer down instead: "record-match
  is painted as pass while its outcome is 'alarm'", expected rgb(66, 29, 36),
  received rgb(15, 53, 43).
- Satisfied from an unrelated line in the same file — this lab's own failure
  mode, the killing assertion rewritten as text + data-status with the
  honest-fetch call left in place. Same ledger failure as the first, naming
  the test the surviving call actually ran in. At 521d03d this shipped 23/23.

All 16 recorded mutations re-run against this tree with CI=1: none survived,
and every failure is now inside e2e/verdict-assertions.ts — line 110 (text),
112 (data-status), 118 (plate), 154 (measured value), 161 (the sentence
states its own value).

README: the count is 47 (24 Vitest + 23 Playwright), the gate paragraph
describes the ledger, and README:17's "every number it renders carries a
data-claim marker" is corrected. #fetch-status renders "Both full-domain
folds completed in 1,019 ms" outside every marker; verdict-scan.ts scans
result regions and not status prose, so the rule is met as scoped and it was
the sentence that was false. The figure is left unmarked on purpose — it
measures the device, not the protocol.

Branch protection untouched, in either direction.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@systemslibrarian
systemslibrarian merged commit 340dbca into main Sep 30, 2026
8 checks passed
@systemslibrarian
systemslibrarian deleted the harness/split-point branch September 30, 2026 01:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant