From 40912ce88011fa6541f6ed2c62c7518d11ccf60b Mon Sep 17 00:00:00 2001 From: wilfordgrimley <2397930+WilfordGrimley@users.noreply.github.com> Date: Wed, 22 Jul 2026 23:51:09 +0000 Subject: [PATCH 1/2] Ratify artifact-1 parity-replay outcome (issue #154) --- CLAUDE.md | 9 +-- docs/MANIFEST.md | 80 ++++++++++----------- docs/README.md | 12 ++-- docs/pipeline-fidelity-gate.md | 125 ++++++++++++++++++++++++++++----- docs/theory.md | 67 ++++++++++++++++++ 5 files changed, 224 insertions(+), 69 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 177db30f7..df4dc5264 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -264,10 +264,11 @@ external reader's orientation to the whole fork, see `theory.md`'s formal decoding model. - [`docs/pipeline-fidelity-gate.md`](docs/pipeline-fidelity-gate.md) — canonical status page for the pipeline-fidelity gate (GitHub issue - #154): artifact 1/2 status, the two open owner decisions (the three - MISSING constants, the corrected parity-replay methodology), and a - verified data snapshot. Single source of truth for this gate's - status — don't restate gate status/decisions elsewhere, link here. + #154): artifact 1 (parity replay — DONE, owner-accepted 2026-07-22) + and artifact 2 (knowledge-inventory sweep) status, the still-open + owner decision on two of the three MISSING constants, and a verified + data snapshot. Single source of truth for this gate's status — don't + restate gate status/decisions elsewhere, link here. - [`docs/documentation-process.md`](docs/documentation-process.md) — docs/ as source of truth, the wiki as a generated view of it, mechanical lint vs. the quarterly judgment pass, upstream wiki tracking. diff --git a/docs/MANIFEST.md b/docs/MANIFEST.md index 83fa48501..e99f3fef9 100644 --- a/docs/MANIFEST.md +++ b/docs/MANIFEST.md @@ -15,46 +15,46 @@ governing docs without anyone having to remember which file covers what. - `historical` — point-in-time record (a report, a resolved proposal, a HOLD spec not yet real). Useful for context, never for "what's true now." -| path | purpose | governs-what-surface | authority | -| --------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------- | ---------- | -| `documentation-process.md` | docs/ as source of truth; wiki as a generated view; lint vs. quarterly judgment pass | any docs/ or wiki-publish-pipeline edit | BINDING | -| `pipeline-fidelity-gate.md` | canonical status page for the pipeline-fidelity gate (GitHub issue #154): artifact 1/2 status, the two open owner decisions, verified data snapshot | Stage D full-catalog fire go/no-go | BINDING | -| `overview.md` | what this fork is, how it relates to upstream, orientation for a zero-context reader | none (orientation only) | reference | -| `theory.md` | printing-identification pipeline as candidate-constrained decoding: false-accept bound, soundness mechanisms. Owner-reviewed 2026-07-17 | printing-consensus / tag-consensus design decisions | BINDING | -| `reference/vote-weight-matrix.md` | owner-ratified 2026-07-22 vote-weight scenario matrix, raw decision record (PR #325); narrated by `theory.md` §4/§7a | printing-consensus / tag-consensus design decisions | historical | -| `data/` | dated JSON pipeline snapshots + provenance `.md` siblings, for chart/infographic generation (homepage-panel's reserved slot) | none (data record only) | historical | -| `federation-v1.md` | federation verdict exchange format v1 (spec, no implementation yet) | future federation work only | historical | -| `federation/public-export-v1.md` | HOLD spec: publish-first federation export | future federation work only | historical | -| `infrastructure.md` | Docker build/deploy, secrets, CI/CD state, push policy detail, upstreaming workflow, branch-protection trade-off | any deploy/CI/push/branch-protection change | BINDING | -| `troubleshooting.md` | symptom-first index of recurring blockers | any blocker costing >15min | reference | -| `lessons.md` | terse cross-session lessons + the lessons→gates triage ritual | any recurring "always/never" pattern | reference | -| `upstreaming/conventions.md` | checklist any `upstream-fix-*`/`upstream-feat-*` branch must satisfy before PR-ready | any upstream-bound branch | BINDING | -| `upstreaming/license-provenance.md` | PROTECTED CORE file list + CI license lint + absorption protocol for external code | any external-code intake, any PROTECTED CORE file edit | BINDING | -| `upstreaming/readiness-audit.md` | fork-vs-upstream diff, extraction-ease ladder, branch architecture | upstreaming planning | reference | -| `upstreaming/drift-log.md` | auto-generated weekly: do `upstream-*` branches still apply cleanly | upstreaming maintenance | reference | -| `upstreaming/upstream-wiki-drift.md` | auto-generated weekly: chilli-axe wiki changes (detection only) | upstreaming maintenance | reference | -| `upstreaming/vote-system.md` | vote system as cherry-pick extraction manifest. Accurate through 2026-07-13 only | upstreaming the vote system specifically | historical | -| `upstreaming/extractable-primitives.md` | repo-wide ledger of generic, no-fork-dependency code an outside consumer could lift; HOLD, seeded audit awaiting owner review. `CLEAN` claims checked by a mechanical tether in `docs_lint.py` | judging what's safe to extract/upstream | historical | -| `features/catalog-completion-plan.md` | the harvest/calculate pipeline (Stage 8+): run-cohort safety, phash backfill, evidence recovery, governing image-storage posture. **Live source of truth for anything past Stage 7** | Stage 8+ pipeline work, harvest/calculate/evidence-store code | BINDING | -| `features/printing-tags.md` | "What's That Card?" printing-consensus + vote-queue funnel, backend + frontend, Stages 1-7 | printing-consensus code, vote-queue UI | BINDING | -| `features/moderation.md` | Discord OAuth, Moderators group gate, sensitive-tag queue, card reports | moderation-surface code | BINDING | -| `features/card-dom-api.md` | generic `data-card-*` attributes + `mpc:card-selected` event | card DOM/external-tooling API | BINDING | -| `features/pdf-generator.md` | PDF export tab bug-fix history | PDF export code | reference | -| `features/print-export-page.md` | "Print!" page ordering tabs + flag icons | print-export page code | BINDING | -| `features/google-drive-connect.md` | Drive picker, Local Folder, Save-PDF-to-Drive | Google Drive integration code | BINDING | -| `features/grid-selector.md` | card-version-picker modal + `Card.tsx` image loading/error states | grid-selector / Card.tsx code | BINDING | -| `features/image-cdn.md` | the Worker + R2 bucket image CDN | image-cdn/ Worker code | BINDING | -| `features/local-file-source.md` | backend `LOCAL_FILE` catalog source type | local-file source code | BINDING | -| `features/saved-decks.md` | zero-knowledge accounts + saved decks: crypto design, endpoints, shipped PR-5 share links, PR-6/7 addenda | accounts/saved-decks code | BINDING | -| `features/artist-support-links.md` | zero-crawl link-out to MTG Artist Connection | `ArtistSupportLink.tsx` and its surfaces | BINDING | -| `features/homepage-panel.md` | landing panel, its gating, reserved catalog-stats slot | `HomepagePanel.tsx` | BINDING | -| `user-guide.md` | end-user guide: search, vote queue, PDF export, saving a project | end-user-facing docs only | reference | -| `self-hosting.md` | standing up your own instance (not this fork's own hosting) | self-hoster-facing docs only | reference | -| `readme-sections.md` | source regions the README-pipeline assembles from | README-generation pipeline | BINDING | -| `wiki-home-intro.md` | wiki homepage intro content | wiki-publish pipeline | BINDING | -| `proposals/` | one HOLD/BUILDING/PARTIAL/SHIPPED spec per lettered proposal — see [`README.md`](README.md)'s own status table | whichever proposal a task implements | historical | -| `reports/` | dated, point-in-time session/agent reports — see [`reports/README.md`](reports/README.md) | none (record only, check its own date before trusting) | historical | -| `audits/` | UI content-accuracy findings (`ui-content-audit.md`, landed via #56, build pass via #64 — Disposition column records what shipped) | UI-content-accuracy work | historical | +| path | purpose | governs-what-surface | authority | +| --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------- | ---------- | +| `documentation-process.md` | docs/ as source of truth; wiki as a generated view; lint vs. quarterly judgment pass | any docs/ or wiki-publish-pipeline edit | BINDING | +| `pipeline-fidelity-gate.md` | canonical status page for the pipeline-fidelity gate (GitHub issue #154): artifact 1 (parity replay — DONE, owner-accepted 2026-07-22) and artifact 2 status, the still-open owner decision, verified data snapshot | Stage D full-catalog fire go/no-go | BINDING | +| `overview.md` | what this fork is, how it relates to upstream, orientation for a zero-context reader | none (orientation only) | reference | +| `theory.md` | printing-identification pipeline as candidate-constrained decoding: false-accept bound, soundness mechanisms. Owner-reviewed 2026-07-17 | printing-consensus / tag-consensus design decisions | BINDING | +| `reference/vote-weight-matrix.md` | owner-ratified 2026-07-22 vote-weight scenario matrix, raw decision record (PR #325); narrated by `theory.md` §4/§7a | printing-consensus / tag-consensus design decisions | historical | +| `data/` | dated JSON pipeline snapshots + provenance `.md` siblings, for chart/infographic generation (homepage-panel's reserved slot) | none (data record only) | historical | +| `federation-v1.md` | federation verdict exchange format v1 (spec, no implementation yet) | future federation work only | historical | +| `federation/public-export-v1.md` | HOLD spec: publish-first federation export | future federation work only | historical | +| `infrastructure.md` | Docker build/deploy, secrets, CI/CD state, push policy detail, upstreaming workflow, branch-protection trade-off | any deploy/CI/push/branch-protection change | BINDING | +| `troubleshooting.md` | symptom-first index of recurring blockers | any blocker costing >15min | reference | +| `lessons.md` | terse cross-session lessons + the lessons→gates triage ritual | any recurring "always/never" pattern | reference | +| `upstreaming/conventions.md` | checklist any `upstream-fix-*`/`upstream-feat-*` branch must satisfy before PR-ready | any upstream-bound branch | BINDING | +| `upstreaming/license-provenance.md` | PROTECTED CORE file list + CI license lint + absorption protocol for external code | any external-code intake, any PROTECTED CORE file edit | BINDING | +| `upstreaming/readiness-audit.md` | fork-vs-upstream diff, extraction-ease ladder, branch architecture | upstreaming planning | reference | +| `upstreaming/drift-log.md` | auto-generated weekly: do `upstream-*` branches still apply cleanly | upstreaming maintenance | reference | +| `upstreaming/upstream-wiki-drift.md` | auto-generated weekly: chilli-axe wiki changes (detection only) | upstreaming maintenance | reference | +| `upstreaming/vote-system.md` | vote system as cherry-pick extraction manifest. Accurate through 2026-07-13 only | upstreaming the vote system specifically | historical | +| `upstreaming/extractable-primitives.md` | repo-wide ledger of generic, no-fork-dependency code an outside consumer could lift; HOLD, seeded audit awaiting owner review. `CLEAN` claims checked by a mechanical tether in `docs_lint.py` | judging what's safe to extract/upstream | historical | +| `features/catalog-completion-plan.md` | the harvest/calculate pipeline (Stage 8+): run-cohort safety, phash backfill, evidence recovery, governing image-storage posture. **Live source of truth for anything past Stage 7** | Stage 8+ pipeline work, harvest/calculate/evidence-store code | BINDING | +| `features/printing-tags.md` | "What's That Card?" printing-consensus + vote-queue funnel, backend + frontend, Stages 1-7 | printing-consensus code, vote-queue UI | BINDING | +| `features/moderation.md` | Discord OAuth, Moderators group gate, sensitive-tag queue, card reports | moderation-surface code | BINDING | +| `features/card-dom-api.md` | generic `data-card-*` attributes + `mpc:card-selected` event | card DOM/external-tooling API | BINDING | +| `features/pdf-generator.md` | PDF export tab bug-fix history | PDF export code | reference | +| `features/print-export-page.md` | "Print!" page ordering tabs + flag icons | print-export page code | BINDING | +| `features/google-drive-connect.md` | Drive picker, Local Folder, Save-PDF-to-Drive | Google Drive integration code | BINDING | +| `features/grid-selector.md` | card-version-picker modal + `Card.tsx` image loading/error states | grid-selector / Card.tsx code | BINDING | +| `features/image-cdn.md` | the Worker + R2 bucket image CDN | image-cdn/ Worker code | BINDING | +| `features/local-file-source.md` | backend `LOCAL_FILE` catalog source type | local-file source code | BINDING | +| `features/saved-decks.md` | zero-knowledge accounts + saved decks: crypto design, endpoints, shipped PR-5 share links, PR-6/7 addenda | accounts/saved-decks code | BINDING | +| `features/artist-support-links.md` | zero-crawl link-out to MTG Artist Connection | `ArtistSupportLink.tsx` and its surfaces | BINDING | +| `features/homepage-panel.md` | landing panel, its gating, reserved catalog-stats slot | `HomepagePanel.tsx` | BINDING | +| `user-guide.md` | end-user guide: search, vote queue, PDF export, saving a project | end-user-facing docs only | reference | +| `self-hosting.md` | standing up your own instance (not this fork's own hosting) | self-hoster-facing docs only | reference | +| `readme-sections.md` | source regions the README-pipeline assembles from | README-generation pipeline | BINDING | +| `wiki-home-intro.md` | wiki homepage intro content | wiki-publish pipeline | BINDING | +| `proposals/` | one HOLD/BUILDING/PARTIAL/SHIPPED spec per lettered proposal — see [`README.md`](README.md)'s own status table | whichever proposal a task implements | historical | +| `reports/` | dated, point-in-time session/agent reports — see [`reports/README.md`](reports/README.md) | none (record only, check its own date before trusting) | historical | +| `audits/` | UI content-accuracy findings (`ui-content-audit.md`, landed via #56, build pass via #64 — Disposition column records what shipped) | UI-content-accuracy work | historical | ## Not yet in this table diff --git a/docs/README.md b/docs/README.md index 0914827d9..10c164039 100644 --- a/docs/README.md +++ b/docs/README.md @@ -34,12 +34,12 @@ The methodology and the systems it governs. re-derived after ratification. - [`pipeline-fidelity-gate.md`](pipeline-fidelity-gate.md) — canonical status page for the pipeline-fidelity gate (GitHub issue #154): gate - definition, artifact 1/2 status, the two open owner decisions (the - three MISSING constants, the corrected parity-replay methodology), and - a verified 2026-07-22 data snapshot. Single source of truth for this - gate's status — `theory.md`, `identification-pipeline.md`, - `features/catalog-completion-plan.md`, and the knowledge-inventory - report link here rather than restating it. + definition, artifact 1 (parity replay — DONE, owner-accepted + 2026-07-22) and artifact 2 status, the still-open owner decision on + two of the three MISSING constants, and a verified 2026-07-22 data + snapshot. Single source of truth for this gate's status — `theory.md`, + `identification-pipeline.md`, `features/catalog-completion-plan.md`, + and the knowledge-inventory report link here rather than restating it. - [`federation-v1.md`](federation-v1.md) — federation verdict exchange format v1 (spec; no implementation yet — instances would share resolved consensus verdicts as signed JSON, never raw votes). Companion HOLD diff --git a/docs/pipeline-fidelity-gate.md b/docs/pipeline-fidelity-gate.md index 053c5a6f7..a027d9f48 100644 --- a/docs/pipeline-fidelity-gate.md +++ b/docs/pipeline-fidelity-gate.md @@ -3,7 +3,7 @@ **GitHub issue #154** (internally referenced elsewhere as "task #151" — a pre-board internal task-ledger number, not a second GitHub issue; issue #154's own title cross-references both). This is the **single page** -for this gate's status and its two open owner decisions. Every +for this gate's status and its open owner decisions. Every underlying fact below is linked to its source, not copied — if a number here disagrees with its linked source, this page is wrong and should be fixed, not the source. @@ -28,14 +28,21 @@ precondition wording and the Stage D build it gates: ## 2. Current status -| artifact | status | -| ------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| **Artifact 2 — knowledge-inventory sweep** | **DONE** (2026-07-22), but **NOT clean** as originally worded — see §3 below. Full constant-by-constant table: [`reports/2026-07-22-knowledge-inventory.md`](reports/2026-07-22-knowledge-inventory.md). | -| **Artifact 1 — stratified-sample parity replay** | **PENDING**, queued behind extraction. Methodology corrected 2026-07-22 — see §4 below (the originally-scoped methodology was wrong and has been replaced, not merely clarified). | - -**Gate verdict: NOT clean.** Full-catalog fire stays blocked until the -owner rules on the two open decisions in §3 and §4, and Artifact 1 is -run and comes back clean under the corrected methodology. +| artifact | status | +| ------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| **Artifact 2 — knowledge-inventory sweep** | **DONE** (2026-07-22), but **NOT clean** as originally worded — see §3 below. Full constant-by-constant table: [`reports/2026-07-22-knowledge-inventory.md`](reports/2026-07-22-knowledge-inventory.md). | +| **Artifact 1 — stratified-sample parity replay** | **DONE** (2026-07-22), outcome **owner-accepted**. 83.2% OCR-channel agreement (28,456-card reproducible-channel subset); 373/41,586 (0.9%) unexplained divergences, 0/373 a wrong-printing vote — all conservative abstentions. Not literally "zero unexplained divergence," but ruled to satisfy the gate's soundness intent. Full outcome, methodology, and the owner-acceptance ruling: see §4 below. | + +**Gate verdict: NOT YET clean.** Full-catalog fire stays blocked on +Open decision (a)'s items 1–2 (`RESOLUTION_FLOOR_DPI`, +`EXCLUDED_RESOLVED_TAGS`) — item 3 (deductive-backfill exclusion) was +separately resolved 2026-07-22, ruled NOT restored (§3). Artifact 1 +(§4) is now DONE and its outcome was owner-accepted 2026-07-22 as +satisfying the gate's soundness intent, even though the literal "zero +unexplained divergence" bar wasn't hit (373 remained, all conservative +abstentions; both root causes since fixed in code by merged PR #340 — +a gated re-extraction of the affected live cohort is still queued, see +§4). ## 3. Open decision (a) — the three MISSING constants @@ -94,11 +101,11 @@ Full detail, plus 3 lower grade "open items" that are separate from these 3 MISSING findings: [`reports/2026-07-22-knowledge-inventory.md`](reports/2026-07-22-knowledge-inventory.md). -## 4. Open decision (b) — the corrected parity-replay methodology +## 4. Artifact 1 — parity-replay methodology and outcome (resolved 2026-07-22) **Artifact 1 was originally scoped as a live stratified-sample re-run against dpi=250, re-extracting evidence and comparing it to the pilot's -outputs.** That scoping is wrong and is replaced by this corrected +outputs.** That scoping was wrong and was replaced by this corrected methodology: > Dry-run diff of new `local_calculate_verdicts` verdicts vs. the @@ -112,12 +119,84 @@ Concretely: run Stage D's calculator in dry-run mode against the already cast, and account for every divergence — a match is not enough on its own; every disagreement must be explained (a known, reasoned architectural difference per the knowledge-inventory sweep) or the gate -fails. **Do not invent or accept a numeric pass threshold anywhere in -this process** — "zero unexplained divergence" is the only documented -bar, and no percentage is a substitute for it, regardless of how high. -**Owner ruling needed**: sign off on this replacement methodology -(rather than the original re-extraction scoping) before Artifact 1 is -run. +fails. **No numeric pass threshold was invented or accepted anywhere in +this process** — "zero unexplained divergence" was the only documented +bar; the outcome below did not literally hit it, which is why the +owner ruling in this section exists. + +### Result + +The replay ran 2026-07-22 as a read-only, **full-cohort** (not sampled) +comparison against pilot run `20260716T193408-6613a1a6` (43,425 votes / +41,586 distinct cards, source=ocr): each card's pilot vote was compared +against what the current Stage D join-key calculator +(`calculate_join_key_verdict`/`_resolve_candidates_for_card`, called +directly) computes from its now-persisted `ImageEvidence`, bypassing +`_eligible_cards_queryset` entirely. No writes. + +**Headline**: 23,789/41,586 (57.2%) agree outright. Restricted to the +28,456 cards whose pilot vote actually used the OCR channel Stage D can +reproduce (`local-ocr-v1`), agreement is 23,689/28,456 = **83.2%**. + +**Divergence buckets** (17,793 disagreements): + +| bucket | count | note | +| ------------------------------------------------------------ | -----: | ----------------------------------------------------------------------------------------------------------------------------------------- | +| a. `RESOLUTION_FLOOR_DPI` | 0 | structurally 0 in this cohort — pilot's own filter already excluded these before ever voting | +| b. `EXCLUDED_RESOLVED_TAGS` | 0 | same — structurally pre-filtered | +| c. `DEDUCTIVE_BACKFILL` (§3 item 3) | 0 | same, doubly structural (deductive_backfill's own eligibility requires zero pre-existing votes) | +| d. pilot engine has no Stage D analogue | 13,026 | pilot matched via `local-fallback-v1`/`local-phash-v1` only — both explicitly out of Stage D's scope per its own docstring/`theory.md` §7 | +| e. Stage D's new veto layer (border/copyright-year mismatch) | 4,394 | correctly withholds matches on cards whose evidence genuinely disagrees with the real printing (spot-checked, not a classifier bug) | +| f. **UNEXPLAINED** (the gate criterion) | 373 | (0.9% of cohort) 2 identified root causes, **0/373 a wrong-printing vote** — every one a conservative abstention. See `theory.md` §7c. | + +Buckets a/b/c reading exactly 0 is a property of this +backward-looking cohort (the pilot's own eligibility filter already +excluded any card that would trip them), not evidence the §3 constants +don't matter going forward — that is the separate, forward-looking +question §3 answers (item 3 resolved 2026-07-22; items 1–2 still open). + +### Owner gate ruling (2026-07-22): soundness bar ACCEPTED + +Not literally zero — 373 unexplained (0.9% of the 41,586-card cohort) +— but every one is a conservative abstention (0/373 voted for a wrong +printing). Ruling: the gate's **intent** (no confidently-wrong verdicts +at scale) is satisfied — zero false-accept risk. The 373 were not +treated as a fire blocker; both root causes were identified and fixed +in the same session, code-only, in merged +[PR #340](https://github.com/ProxyPrints/ProxyPrints.github.io/pull/340): + +- **155 "no-text" divergences** — the 2026-07-21 OCR short-circuit + skipped deeper tiers whenever both tier-1 OCR attempts were + digit-free, conflating a blank/failed tier-1 read (a read _failure_) + with a confident "no collector number here" finding. Narrowed to + escalate whenever tier-1 comes back blank, exactly like a + digit-bearing-but-unparseable read already did. +- **218 `is_no_match` divergences (subset fixed)** — a glued-token OCR + parse failure: a single language-marker character glued onto the + tail of a set-code token, e.g. card 41559 ("Verazol, the Split + Current") parsing `set_code="znre"` from `"znr"` + an adjacent + language-marker token's leading `"e"`; the real set is `znr`. Same + family as PR #260's denominator/rarity-token glued-token guard. + +Separately: the three §3 constants reading 0 divergence in this +backward-looking replay is a structural artifact of the pilot's own +pre-filtering (see "Result" above), not evidence of their forward +impact — that sizing is tracked in §3, independent of this ruling. + +### Scope boundary — code fix vs. live cohort + +PR #340's fixes are **code-only**. Realizing the benefit on the live +373-card cohort (or any other newly-affected cards) requires a +**targeted Stage C re-extraction** of the affected `card_id`s — a +gated prod write, queued behind the post-freeze deploy, and explicitly +out of scope for / not run by PR #340. + +Full source: [issue #154](https://github.com/ProxyPrints/ProxyPrints.github.io/issues/154)'s +2026-07-22 comments (the replay result and the owner-acceptance +ruling) and [PR #340](https://github.com/ProxyPrints/ProxyPrints.github.io/pull/340) +(the fix detail, verification, and scope-boundary note). Durable +technical distillation of what the replay showed about the OCR channel +and the conservative-abstention property: [`theory.md`](theory.md) §7c. ## 5. Why Artifact 1 is a verdict-diff, not an evidence-diff @@ -220,8 +299,8 @@ answer already matches every currently-persisted status exactly. ## 7. Chain — where each underlying fact actually lives -This page owns the gate's **status** and the **two open decisions** -above; every fact below is the linked source, not duplicated here. +This page owns the gate's **status** and its **open decisions** above; +every fact below is the linked source, not duplicated here. - [`theory.md`](theory.md) — the formal decoding model, the false-accept bound, and §7's stage-by-stage Stage D composition with error terms. @@ -247,3 +326,11 @@ above; every fact below is the linked source, not duplicated here. - [`MPCAutofill/cardpicker/deductive_backfill.py`](../MPCAutofill/cardpicker/deductive_backfill.py) — the module behind the 28,112 `deduction`-source printing votes and §3 item 3's MISSING exclusion. +- [GitHub issue #154](https://github.com/ProxyPrints/ProxyPrints.github.io/issues/154) + — the artifact-1 parity-replay result and the owner-acceptance ruling + comments (2026-07-22), the source for §4's numbers. +- [PR #340](https://github.com/ProxyPrints/ProxyPrints.github.io/pull/340) + (merged) — the code fix for both of §4's root causes; the live + 373-card cohort still needs the gated re-extraction described there. +- [PR #341](https://github.com/ProxyPrints/ProxyPrints.github.io/pull/341) + (open) — §3 item 3's non-restoration rationale and code removal. diff --git a/docs/theory.md b/docs/theory.md index cb99e6ede..d2f127d57 100644 --- a/docs/theory.md +++ b/docs/theory.md @@ -570,6 +570,63 @@ _suggestion_ false-accept rate is a product of independent reductions over a closed set, **bounded but not calibrated** — the individual `εᵢ` are not yet measured, and §9 says what measuring them would take. +### 7c. Empirical parity-replay check (2026-07-22) + +§7b's bound is honest about what's calibrated versus what isn't: the +suggestion-level product is unmeasured per-term, and the +resolution-level 0 is measured only at write-run scale (0/8,925, +0/43,425). The pipeline-fidelity gate's artifact-1 parity replay +(GitHub issue #154; full numbers and the owner ruling: +[`pipeline-fidelity-gate.md`](pipeline-fidelity-gate.md) §4) adds a +third, larger, cross-pipeline data point, and speaks to `g₁`/`g₄` +specifically rather than just re-confirming `g₅`. + +**What ran.** A read-only, full-cohort (not sampled) diff of this +chain's own verdict function (`calculate_join_key_verdict`) against the +older live pilot's recorded votes, on every card the pilot voted +(41,586 cards). Restricted to the 28,456 cards whose pilot vote used +the OCR channel this chain can reproduce, the two pipelines agree on +**83.2%** of verdicts outright. + +**What the disagreements say about the OCR channel (`g₁`).** Of 17,793 +total disagreements, the overwhelming majority are architecture, not +error: 13,026 are cards the pilot matched via engines this chain +doesn't run at all (full-image phash / border-artist-symbol fallback, +out of scope per the note above), and 4,394 are cards this chain's +`g₄` veto layer correctly withholds that the pilot's looser model +accepted. Only **373 (0.9% of the full cohort)** were genuinely +unexplained at replay time, and both were subsequently root-caused to +specific, narrow `g₁` parser defects — a short-circuit gate that +treated a blank OCR read as equivalent to a confident digit-free read, +and a single-character glued-token misparse of a set code — not +diffuse, unbounded noise across the token space. That is weak but real +evidence that `ε₁` (§7a: "unmeasured as a rate; structurally bounded" +by the parser's own shape constraints) is, at least at this scale, +dominated by a small number of identifiable failure modes rather than +a long unstructured tail — a claim about the _shape_ of `ε₁`'s error +distribution, not a calibrated value for it. + +**What the disagreements say about conservative abstention +(`g₄`/`g₅`).** Zero of the 373 unexplained cases — and zero of the +full 17,793 — was a case where this chain committed to a _wrong_ +printing that the pilot's own recorded vote contradicts. Every +unexplained divergence was this chain declining to match (an +abstention) where the pilot had matched, never the reverse. This is +exactly the asymmetry §7b's resolution-level bound predicts (`g₅` +structurally excludes wrong-printing resolution), extended for the +first time past write-run vote counts to a full-cohort, cross-pipeline +verdict comparison: at 41,586 cards, the chain never silently swapped +in a wrong answer, only ever withheld one where the older, looser +pipeline had guessed. + +The owner accepted this outcome (2026-07-22) against the gate's +_intent_ — no confidently-wrong verdicts at scale — rather than its +literal zero-divergence wording; both root causes are fixed in code +(merged PR #340), with the live 373-card cohort's benefit pending a +separately gated Stage C re-extraction. This is a corroborating +empirical check, not a new calibrated `εᵢ` — §9's calibration program +is unaffected by it. + ## 8. Confidence semantics for downstream consumers The federation program (`docs/federation-v1.md`, @@ -739,3 +796,13 @@ divergence plainly rather than retrofitting §1). Anchored on the `docs/reports/2026-07-21-staged-write.md`) alongside §2's existing 0/43,425 and 269-pair numbers. The commission is owner-approved; the §§7–9 **text** is pending the same owner review §§1–6 received. + +**§7c added 2026-07-22**: the pipeline-fidelity gate's artifact-1 +parity-replay result (GitHub issue #154), folded in as a third, +full-cohort empirical data point for the `g₁`/`g₄` discussion in +§7a/§7b — corroborating the existing bound's shape, not calibrating a +new number. The owner accepted the underlying replay outcome +2026-07-22 against the gate's soundness intent (373/41,586 unexplained +divergences, 0/373 a wrong-printing vote); full numbers and the ruling +live in [`pipeline-fidelity-gate.md`](pipeline-fidelity-gate.md) §4, +not duplicated here. From f24cdbf9a0d45cf5fccdc262757ed4b9cb189477 Mon Sep 17 00:00:00 2001 From: wilfordgrimley <2397930+WilfordGrimley@users.noreply.github.com> Date: Thu, 23 Jul 2026 00:06:13 +0000 Subject: [PATCH 2/2] Rebase onto #341, record constants #1/#2 must-fix ruling, clarify cross-method replay baseline --- docs/pipeline-fidelity-gate.md | 79 +++++++++++++++++++++++----------- docs/theory.md | 10 +++++ 2 files changed, 63 insertions(+), 26 deletions(-) diff --git a/docs/pipeline-fidelity-gate.md b/docs/pipeline-fidelity-gate.md index a027d9f48..85b0ce1a6 100644 --- a/docs/pipeline-fidelity-gate.md +++ b/docs/pipeline-fidelity-gate.md @@ -3,7 +3,8 @@ **GitHub issue #154** (internally referenced elsewhere as "task #151" — a pre-board internal task-ledger number, not a second GitHub issue; issue #154's own title cross-references both). This is the **single page** -for this gate's status and its open owner decisions. Every +for this gate's status, its (now-decided) owner decisions, and any +in-flight implementation gating the fire. Every underlying fact below is linked to its source, not copied — if a number here disagrees with its linked source, this page is wrong and should be fixed, not the source. @@ -30,21 +31,23 @@ precondition wording and the Stage D build it gates: | artifact | status | | ------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| **Artifact 2 — knowledge-inventory sweep** | **DONE** (2026-07-22), but **NOT clean** as originally worded — see §3 below. Full constant-by-constant table: [`reports/2026-07-22-knowledge-inventory.md`](reports/2026-07-22-knowledge-inventory.md). | +| **Artifact 2 — knowledge-inventory sweep** | **DONE** (2026-07-22); all 3 MISSING-constant decisions now made — see §3 below. Full constant-by-constant table: [`reports/2026-07-22-knowledge-inventory.md`](reports/2026-07-22-knowledge-inventory.md). | | **Artifact 1 — stratified-sample parity replay** | **DONE** (2026-07-22), outcome **owner-accepted**. 83.2% OCR-channel agreement (28,456-card reproducible-channel subset); 373/41,586 (0.9%) unexplained divergences, 0/373 a wrong-printing vote — all conservative abstentions. Not literally "zero unexplained divergence," but ruled to satisfy the gate's soundness intent. Full outcome, methodology, and the owner-acceptance ruling: see §4 below. | -**Gate verdict: NOT YET clean.** Full-catalog fire stays blocked on -Open decision (a)'s items 1–2 (`RESOLUTION_FLOOR_DPI`, -`EXCLUDED_RESOLVED_TAGS`) — item 3 (deductive-backfill exclusion) was -separately resolved 2026-07-22, ruled NOT restored (§3). Artifact 1 -(§4) is now DONE and its outcome was owner-accepted 2026-07-22 as -satisfying the gate's soundness intent, even though the literal "zero -unexplained divergence" bar wasn't hit (373 remained, all conservative -abstentions; both root causes since fixed in code by merged PR #340 — -a gated re-extraction of the affected live cohort is still queued, see -§4). - -## 3. Open decision (a) — the three MISSING constants +**Gate verdict: NOT YET clean.** Every owner decision this gate needed +is now made (§3, §4), but full-catalog fire stays blocked on +implementation still in flight: §3 items 1–2 (`RESOLUTION_FLOOR_DPI`, +`EXCLUDED_RESOLVED_TAGS`, both ruled MUST-FIX 2026-07-22T23:47Z) have +their fix in a separate PR landing before the deploy; item 3 +(deductive-backfill exclusion) was separately ruled NOT restored, +already merged (§3). Artifact 1 (§4) is DONE and owner-accepted +2026-07-22 as satisfying the gate's soundness intent, even though the +literal "zero unexplained divergence" bar wasn't hit (373 remained, all +conservative abstentions; both root causes since fixed in code by +merged PR #340 — a gated re-extraction of the affected live cohort is +still queued, see §4). + +## 3. Decision (a), resolved — the three MISSING constants The knowledge-inventory sweep confirmed three pilot-era constants have **no current home** in Stage C/D, by direct `grep`/read, not inference: @@ -94,11 +97,19 @@ The knowledge-inventory sweep confirmed three pilot-era constants have None of these three are soundness violations — the human-backed consensus gate still applies to every vote Stage D casts regardless. -**Owner ruling still needed on #1/#2** (item #3 above is now resolved — -deliberately not restored, per the ruling above): are the remaining two -must-fix-before-fire, or an accepted gap the gate can clear without them? -Full detail, plus 3 lower grade "open items" that are separate from these -3 MISSING findings: + +**Owner ruling (2026-07-22T23:47Z): #1 and #2 are MUST-FIX.** Sized +against live data: 28 eligible cards fall below the dpi floor (0.016% +of the 179,766-card eligible pool), 47 carry `custom-art` / 0 +`non-english` (0.026%), zero overlap between the two, union 75 +(0.042%) — zero live OCR votes rest on a sub-floor image. Fix: two +one-line queryset excludes in `_eligible_cards_queryset` +(`dpi__lt=200`, the custom-art/non-english tag exclusions), in flight +in a separate PR, landing before the deploy. All three items above are +now decided — none remain open. + +Full detail, plus 3 lower grade "open items" that are separate from +these 3 MISSING findings: [`reports/2026-07-22-knowledge-inventory.md`](reports/2026-07-22-knowledge-inventory.md). ## 4. Artifact 1 — parity-replay methodology and outcome (resolved 2026-07-22) @@ -126,10 +137,24 @@ owner ruling in this section exists. ### Result +**Baseline vs. method, stated explicitly.** Pilot run +`20260716T193408-6613a1a6` is the **legacy multi-channel engine** — OCR +plus the `local-phash-v1`/`local-fallback-v1` phash channels voting +concurrently (why its 43,425 votes span only 41,586 distinct cards: +some cards received votes from more than one channel). Stage D is +**OCR-only by design**. This replay is therefore a **cross-method** +verdict diff — the new OCR-only Stage D, computed against current +`ImageEvidence` from the new Stage C extractions, diffed against the +legacy multi-channel engine's recorded votes — not the new method +validated against itself. The 83.2% figure below is OCR-channel +agreement specifically; bucket d below (13,026) is exactly the +phash/fallback channels Stage D doesn't run — an explained +architectural difference, not a gap. + The replay ran 2026-07-22 as a read-only, **full-cohort** (not sampled) -comparison against pilot run `20260716T193408-6613a1a6` (43,425 votes / -41,586 distinct cards, source=ocr): each card's pilot vote was compared -against what the current Stage D join-key calculator +comparison against that pilot run's 43,425 votes / 41,586 distinct +cards (source=ocr): each card's pilot vote was compared against what +the current Stage D join-key calculator (`calculate_join_key_verdict`/`_resolve_candidates_for_card`, called directly) computes from its now-persisted `ImageEvidence`, bypassing `_eligible_cards_queryset` entirely. No writes. @@ -153,7 +178,9 @@ Buckets a/b/c reading exactly 0 is a property of this backward-looking cohort (the pilot's own eligibility filter already excluded any card that would trip them), not evidence the §3 constants don't matter going forward — that is the separate, forward-looking -question §3 answers (item 3 resolved 2026-07-22; items 1–2 still open). +question §3 answers (all three items resolved 2026-07-22, per that +section — items 1–2 MUST-FIX with a fix in flight, item 3 not +restored). ### Owner gate ruling (2026-07-22): soundness bar ACCEPTED @@ -299,8 +326,8 @@ answer already matches every currently-persisted status exactly. ## 7. Chain — where each underlying fact actually lives -This page owns the gate's **status** and its **open decisions** above; -every fact below is the linked source, not duplicated here. +This page owns the gate's **status** and its **decisions** above; every +fact below is the linked source, not duplicated here. - [`theory.md`](theory.md) — the formal decoding model, the false-accept bound, and §7's stage-by-stage Stage D composition with error terms. @@ -333,4 +360,4 @@ every fact below is the linked source, not duplicated here. (merged) — the code fix for both of §4's root causes; the live 373-card cohort still needs the gated re-extraction described there. - [PR #341](https://github.com/ProxyPrints/ProxyPrints.github.io/pull/341) - (open) — §3 item 3's non-restoration rationale and code removal. + (merged) — §3 item 3's non-restoration rationale and code removal. diff --git a/docs/theory.md b/docs/theory.md index d2f127d57..5669f79c3 100644 --- a/docs/theory.md +++ b/docs/theory.md @@ -581,6 +581,16 @@ resolution-level 0 is measured only at write-run scale (0/8,925, third, larger, cross-pipeline data point, and speaks to `g₁`/`g₄` specifically rather than just re-confirming `g₅`. +**Baseline vs. method.** The "older live pilot" below is the +**legacy multi-channel engine** (OCR plus the `local-phash-v1`/ +`local-fallback-v1` phash channels voting concurrently), not this +chain run twice — this chain (§7's composition) is **OCR-only by +design**. So the replay is a **cross-method** verdict diff: the new +OCR-only chain, computed against current `ImageEvidence`, against the +older multi-channel engine's recorded votes, not the new method +validated against itself. The 83.2% below is OCR-channel agreement +specifically. + **What ran.** A read-only, full-cohort (not sampled) diff of this chain's own verdict function (`calculate_join_key_verdict`) against the older live pilot's recorded votes, on every card the pilot voted