Repository navigation
Verdict harness: judge every rendered claim, over every state that renders one - #6
Merged
Merged
Conversation
Five of the six lane fixes apply here. Two gaps were confirmed by running the mutation against the suite as it stood at ab679b9 and watching it pass. Fix 1 — one helper asserts words, state and plate together. e2e/verdict-assertions.ts adds expectVerdict(page, id, {text, status, tone}) and expectClaim(page, id, {value, text}), and the coverage test now requires every recorded mutation's assertedBy spec to call the helper on that marker, instead of merely mentioning the id. This lab set className and dataset.status in separate expressions, so pinning the class green while the words and the status still followed the run was invisible: at ab679b9 that mutation passed 21/21. tone is checked as the painted background, because .verdict switches class and .mini-verdict switches on [data-status]. Fix 2 — measurements join the coverage loop on the same terms. The page rendered eighteen-odd numbers and not one carried a marker, so the coverage rule enforced itself over one family of a page that had two. Eight data-claim markers now carry the measurement in data-value, are registered in e2e/verdict-mutations.json under "claims", and fail the build both ways: a rendered measurement with no record, and a record naming a measurement the page no longer renders. Fix 3 — the digit-plus-unit scan. findStrayMeasurements walks the result regions for digit-plus-unit text and for bare integers in stats cells that sit outside a marker, with its own injection self-test so it cannot pass by being blind. Hex dumps are exempt as raw material (claims.spec checks them byte for byte against the size markers), and labels and explanatory prose are exempt the same way the verdict-word scan exempts them. The one static number on the page that was really a claim — "AES-128-CTR → 2(λ + 1) bits" — is now measured from a real expandSeed() and marked. Fix 4 — the oracle multiplied where the artifact iterates. The old key-size test asserted root + 17 * domainBits + final === measured at the one default domain: one per-level size multiplied by a level count, both literals, where serializeKey WRITES one correction word per level. It could not tell "sixteen words of seventeen bytes" from "one 272-byte block", which is exactly what the parts list claims. The oracle now measures the per-level cost as the slope between domains, derives λ and the root from it, counts the correction words actually present in the dump by walking it in stride, and SUMS them. Fix 5 — branch protection untouched. Confirmed read-only at the API (build + verdict-coverage required, enforce_admins false, force-push disabled), which is why this is a pull request. Fix 6 — driveEveryState visits every option of every control. Per control, not the cross-product. #alpha-range all 16 indices; the expansion stepper all 5 levels down and back; #new-keys once; #size-domain all 4 domains; #collusion-toggle both states, and again through a fetch for the tile it paints; #tamper-toggle both states with its own fetch; #fetch-record throughout. Two controls are not enumerable option sets and are recorded as such: #shelf-alpha is a 65,536-value number input, so its two rendering classes stand in — the in-range fetch and the rejection that renders no result — and the .key-inspector disclosure is opened because both scans skip elements with no client rects, so a closed disclosure hid its contents from the denominator entirely. No control on this page was judged not to change what renders. That widening is load-bearing, not tidiness: the formula's total was a hard literal correct only at 65,536 records, and at ab679b9 that mutation also passed 21/21 because nothing ever moved #size-domain. Baseline before any edit: 21 passed (15.0s). After: 23 passed (31.7s), plus 24 unit tests. Every mutation below was run with CI=1 so Playwright started its own server on the pinned port 4698. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UH9YUUhdeWXJz141Zy8FxW
D6, in this lab. Fix 1 as it shipped at 521d03d did not satisfy Fix 1. The rule was FILE-granular and read with a regex: a record naming e2e/claims.spec.ts was satisfied by the string expectVerdict(page, 'record-match' appearing ANYWHERE in that file — inside a comment, inside a call handed the page's own values, inside a call asserting a different state in a different test. And the shipped tree already contained an instance: tree-point's recorded mutation was killed by two raw assertions that never reached the helper, so the expectedFlip its record describes was a flip no run had ever demonstrated. expectVerdict/expectClaim now append the (spec, test, marker) triple of every call they EXECUTE to a run-scoped ledger under test-results/verdict-ledger/, cleared once per run in e2e/global-setup.ts. e2e/verdict-ledger.spec.ts reads it back and fails when a record's triple is not among them. Playwright runs each spec in its own worker process, so a module-level Set cannot aggregate; the sink is one append-only file per process, and the reader is its own project with dependencies: ['claims'] rather than an ordinary test. That project is what npm run test:verdicts runs now, so the required check reads the ledger with the specs that write it — .github/ is untouched. assertedBy stops being a filename. It is {spec, test, status, says}: the triple, plus the part of the expectation the RECORD decides — the outcome the killing assertion must require, and one fragment it must hand the helper verbatim (a "re:" prefix is a RegExp source). That second half is the only thing that can catch a call kept but made tautological, and it has to work that way round: on an unmutated tree an expectation read off the page and one decided in advance are the same values, so no amount of observing the run separates them. Only something decided before the run can. The plate is no longer an argument. tone: 'pass' | 'alarm' is gone and the plate is derived from the asserted status through one table, because a reader is never shown an alarm outcome on the success plate — so a test cannot expect one either. That closes the residue: a tautologist who has read the record and hands the declared fragment still cannot mask record-match's mutation, because the plate follows the status it read rather than the paint it read. Two recorded mutations were killed by a precursor rather than by the helper, and both are reordered so the helper is reached first: - tree-point — reconstruction[alpha], the lit-bit count and the #tree-xor readback all failed on the direction-bit mutation before line 232's expectVerdict ran. - key-size-formula — found by running the set: the formulaNumbers comparison caught the pinned-290 mutation at claims.spec.ts:209, before expectClaim. Through the helper it now fails on its own terms, "data-value 86 is not in 17 + (17 × log₂ 16) + 1 = 290 bytes". Escapes. Each run in an isolated tree archived from this commit, node_modules symlinked, port moved to 4931, spec-only rewrites with the built bundle IDENTICAL before and after (index-RVWjpVch.js 3a8b3a777965e297ba523fe90f68143b), baseline 21 passed / exit 0 in the same tree: - Commented out — verdict-ledger fails: "e2e/verdict-mutations.json says its mutation is killed by expectVerdict in e2e/claims.spec.ts › 'tampering violates integrity…', and that call did not execute this run. expectVerdict ran for record-match in: e2e/claims.spec.ts › honest PIR output equals the indexed shelf record (status 'pass')." - Kept but tautological — expectVerdict fails in the test itself: "declares the words this assertion must require — 'RETRIEVED — AND WRONG' — and it was handed ['RETRIEVED — AND WRONG · 1 of 16 bytes differ from shelf[α]; no PIR authentication failure was raised']." With the declared fragment handed over and the mutation applied, it fails one layer down instead: "record-match is painted as pass while its outcome is 'alarm'", expected rgb(66, 29, 36), received rgb(15, 53, 43). - Satisfied from an unrelated line in the same file — this lab's own failure mode, the killing assertion rewritten as text + data-status with the honest-fetch call left in place. Same ledger failure as the first, naming the test the surviving call actually ran in. At 521d03d this shipped 23/23. All 16 recorded mutations re-run against this tree with CI=1: none survived, and every failure is now inside e2e/verdict-assertions.ts — line 110 (text), 112 (data-status), 118 (plate), 154 (measured value), 161 (the sentence states its own value). README: the count is 47 (24 Vitest + 23 Playwright), the gate paragraph describes the ledger, and README:17's "every number it renders carries a data-claim marker" is corrected. #fetch-status renders "Both full-domain folds completed in 1,019 ms" outside every marker; verdict-scan.ts scans result regions and not status prose, so the rule is met as scoped and it was the sentence that was false. The figure is left unmarked on purpose — it measures the device, not the protocol. Branch protection untouched, in either direction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The
verdict-harnesswork, dormant on a local branch since 2026-09-22, routed ontomain. Preserved first at8f39f7cdonorigin/verdict-harness, kept permanently at tagverdict-harness-preserved-2026-09-22.What it does
Verification on the rebased branch
npm testnpm run buildMutation spot-check. Pinning all four computed
data-statussites to"pass"(source diff confirmed non-empty: 4 lines):Restored, baseline back to 23/23, tree clean.
collusion-recoveryis the right marker to fail on — it is the one that says α became visible when both keys reached one server.Rebase
mainwas already an ancestor of the branch, so nothing replayed and no conflict arose. Nopackage.jsonor lockfile touched.🤖 Generated with Claude Code