@@ -688,17 +688,22 @@ Old/Modern Border, Future Frame) 400 on tap; `Full Art`/`Borderless`/
688688
689689## Stage 8: local (zero-API-cost) printing-identification backfill pilot
690690
691- ** Status: environment proposal delivered, not yet built** (2026-07-15,
692- ` worktree-local-printing-id-pilot ` ) - this section records the plan shape,
693- not a shipped feature; see ` journal/2026-07-15-local-printing-id-pilot.md `
694- (machine-local, not committed) for the full investigation.
691+ ** Status: built and pilot-run against live production** (2026-07-15,
692+ ` worktree-local-printing-id-pilot ` , PR #22 ). Code + tests merged; a real
693+ ` --limit 300 --engine both --nice ` invocation ran against the live DB and
694+ its full results are summarized below. ** Full-catalog run explicitly NOT
695+ executed** - see "Real pilot run results" below for why (a ~ 13-day
696+ single-process projection). See
697+ ` journal/2026-07-15-local-printing-id-pilot.md ` (machine-local, not
698+ committed) for the complete data dump this summary is drawn from.
695699
696700Sibling to Stage 6's deductive backfill, same non-negotiable principle
697- (a deduction is always a vote, never a direct resolve - the human-backed
701+ (a vote is always just a vote, never a direct resolve - the human-backed
698702gate in ` vote_consensus.resolve_weighted_consensus ` still applies), but
699703sourced from actually looking at the card image instead of pure logical
700704deduction from existing structured data - two independent local (no paid
701- API calls) engines:
705+ API calls) pass-1 engines, plus a pass-2 fallback for cards pass 1 can't
706+ reach at all:
702707
703708- ** L1, OCR** : Tesseract on a cropped, preprocessed collector-line region
704709 (bottom-left corner, grayscale/upscaled/thresholded). Parses candidate
@@ -708,39 +713,101 @@ API calls) engines:
708713 trusting the OCR output itself. Never writes ` is_no_match ` .
709714- ** L2, perceptual hash** : art-region phash comparison against each
710715 name-candidate's Scryfall art crop, voting only when there's a clear
711- single best match (distance threshold + margin over the second-best).
712- ` CanonicalCard.image_hash ` (` models.py ` ) already exists as a
713- ` BigIntegerField ` for exactly this - added when ` import_canonical_card_ data ` first shipped ("CanonicalCard population fix" above) but never
714- computed in production (` --skip-image-hash ` was used for the real
715- import; confirmed live, 113,224/113,224 rows still at the placeholder
716- ` 0 ` ) - so this pilot is the first thing to actually populate it, lazily,
717- only for candidates it needs.
718- - Both engines vote under their own ` anonymous_id `
719- (` local-ocr-v1 ` /` local-phash-v1 ` ), same weight/gate treatment as any
720- other AI-sourced vote - when both vote on the same card and agree, both
721- votes stand as independent evidence; on disagreement, ** neither** is
722- written (logged instead - the disagreement set is the interesting
723- output of a pilot like this, not noise to discard).
724-
725- ** Environment** : the OCR engine needs the ` tesseract-ocr ` system binary
726- (not pip-installable - ` pytesseract ` is a thin subprocess wrapper around
727- it), which the django container's image doesn't currently have. Two
728- options, no Dockerfile change made without explicit sign-off given this is
729- pilot-only tooling: bake ` tesseract-ocr ` into ` docker/django/Dockerfile ` 's
730- existing apt-get line (matches how every other ` manage.py ` command in this
731- repo is invoked, at the cost of a production image rebuild/restart and a
732- permanent size increase for a one-off tool), or run from a host-side venv
733- pointed at the already-` 127.0.0.1 ` -exposed Postgres/Elasticsearch ports
734- (zero image/container change, trivially reversible, matches this
735- project's precedent of running one-off scripts against the live DB from
736- outside Docker). Command itself (` manage.py local_identify_printing_tags ` )
737- is unaffected either way - only the invocation environment differs.
738-
739- ** Pilot discipline** : ` --limit 300 ` default, explicit hold before any
740- full-catalog run: this is genuinely new signal (visual inspection, not
741- pure logical deduction) and needs a human spot-check of yield/accuracy
742- before scaling up, unlike Stage 6's deduction which was provably exact by
743- construction.
716+ single best match (distance threshold + margin over the second-best,
717+ recalibrated from real production data - see the journal for the
718+ calibration history). ` CanonicalCard.image_hash ` (` models.py ` ) already
719+ exists as a ` BigIntegerField ` for exactly this - added when
720+ ` import_canonical_card_data ` first shipped ("CanonicalCard population
721+ fix" above) but never computed in production (` --skip-image-hash ` was
722+ used for the real import) - so this pilot is the first thing to
723+ actually populate it, lazily, only for candidates it needs. Capped at
724+ 12 candidates per name (basic lands/staples can have hundreds - see
725+ "Real pilot run results" for how often this cap fires).
726+ - ** Pass 2, fallback** (` local_fallback.py ` , ` local-fallback-v1 ` ): fires
727+ only when pass 1 (either engine) produced no accepted vote for a card -
728+ the old-border-frame case (no collector line printed on the card face
729+ at all, just an "Illus. ` <artist> ` " credit). Evidence-combination model
730+ across border-color sample, artist-name OCR fuzzy match, and set-symbol
731+ phash (found unreliable in practice, kept but effectively disabled via
732+ a strict threshold - see ` local_fallback.py ` 's module docstring for the
733+ full negative finding): a vote is cast only when the intersection of
734+ every sub-check that produced a reading narrows to exactly one
735+ candidate.
736+ - Border-color sampling and frame-style classification (OCR-collector-
737+ line-present vs. Illus.-anchor-present) run for ** every** processed
738+ card regardless of printing-vote success, casting standalone
739+ attribute-chip votes (Black/White/Silver Border, Borderless, Old/Modern
740+ Border) - and, when a printing vote ** is** confirmed for that card this
741+ run, preferring ground truth from that printing's own
742+ ` CanonicalPrintingMetadata ` (Scryfall ` border_color ` /` frame ` ) over the
743+ heuristic estimate. The same heuristic reading also feeds a
744+ ** consistency check** : if a card's observed frame class contradicts its
745+ matched printing's real frame value, the printing vote itself is
746+ withheld (kept as a frame-vote-only outcome) rather than trusting an
747+ art/OCR match that likely landed on the wrong printing.
748+ - All engines vote under ` VoteSource.OCR ` (the 2026-07-15 split of the
749+ old single ` VoteSource.AI ` value into ` DEDUCTION ` /` OCR ` - see
750+ ` models.py ` 's ` VoteSource ` docstring; same weight/gate treatment as
751+ before, individual technique still distinguishable via ` anonymous_id ` )
752+ - when OCR and phash both vote on the same card and agree, both votes
753+ stand as independent evidence; on disagreement, ** neither** is written
754+ (logged instead - see the journal's disagreement examples).
755+
756+ ** Environment** : resolved via a host-side venv pointed at the
757+ already-` 127.0.0.1 ` -exposed Postgres/Elasticsearch ports (zero
758+ Docker/container change) - ` tesseract-ocr ` installed via host apt,
759+ ` pytesseract ` /` ImageHash ` /` Pillow ` via the venv's ` requirements.txt `
760+ install. No Dockerfile change made. (A future full-catalog run, if one
761+ ever happens, should revisit baking ` tesseract-ocr ` into
762+ ` docker/django/Dockerfile ` instead, per the original tradeoff writeup.)
763+
764+ ### Real pilot run results (2026-07-15, ` --limit 300 --engine both --nice ` )
765+
766+ ** 32m4.6s wall-clock, 19m36s user + 4m37s sys CPU (≈76% avg utilization of
767+ one core on this 2-CPU box), exit 0.** ` --nice ` confirmed actually
768+ throttling (process niceness observed alternating 5↔19 during the run).
769+
770+ | Engine | Attempted | Votes written | Yield |
771+ | --------- | --------- | ------------- | ----- |
772+ | OCR | 300 | 77 | 25.7% |
773+ | Phash | 300 | 13 | 4.3% |
774+ | Fallback | 210 | 4 | 1.9% |
775+ | ** Total** | — | ** 94** | — |
776+
777+ ** Gate check: 0/94 affected cards resolved** - the human-backed gate held
778+ perfectly at this scale, same result as Stage 6's 0/28,112.
779+
780+ Largest skip bucket by far: OCR's "parsed-but-no-match" at 176/300
781+ (58.7%) - a syntactically valid collector line that didn't match any of
782+ the card's own candidates. Not investigated further in this pilot (out of
783+ scope), but the single most promising lead for improving yield before any
784+ larger run - see the journal for the plausible-causes breakdown.
785+
786+ Attribute votes: border ` {black: 280, borderless: 17, white: 3} ` (91 from
787+ ground truth, 209 from the pixel heuristic); frame ` {modern: 258, old: 14} ` , 28 abstains (91 from ground truth, 181 from the OCR/Illus.-anchor
788+ heuristic); ** 6 frame-mismatches** (printing vote withheld by the
789+ consistency check) - see the journal for all 10 sampled examples and the
790+ per-case reasoning.
791+
792+ ** Full-catalog projection: ~ 171,800 eligible cards remain (of 179,002 raw
793+ eligible pool) → naive linear projection ≈ 306 hours ≈ 12.8 days of
794+ continuous single-process runtime.** This is the key number for any
795+ future decision to scale up - not attempted in this pilot, and not
796+ practical as a single uninterrupted process. Before attempting it:
797+ parallelizing across multiple processes/pk-range partitions, and (more
798+ urgently) switching from the current one-giant-` bulk_create ` -at-the-end
799+ write pattern to periodic batch flushing (matching
800+ ` deductive_backfill.py ` 's existing ` batch_size ` /` flush() ` precedent) so a
801+ multi-day run's progress survives a crash/restart/deploy instead of
802+ losing everything accumulated since the last completed run - both raised
803+ but not implemented in this pilot, out of its locked scope.
804+
805+ 5-vote spot check, 20-vote random admin-link sample, 3 disagreement
806+ examples, and the filename tag-gap census (1,097 unresolved cards with an
807+ unmatchable ` expansion_hint ` ) are all in the journal, not duplicated here.
808+
809+ ** Pilot discipline honored** : ` --limit 300 ` , no full-catalog run attempted
810+ per the original hold.
744811
745812## Key files
746813
@@ -779,9 +846,11 @@ construction.
779846 via iterative screenshot review, not built against a real design system -
780847 owner has flagged that this needs a proper pass with the ` /dataviz ` skill
781848 in the future rather than further ad hoc CSS tuning.
782- - ` CanonicalCard.image_hash ` is bootstrapped to ` 0 ` for every row
783- (` --skip-image-hash ` ); real perceptual-hash-based matching isn't
784- implemented yet.
849+ - ` CanonicalCard.image_hash ` was bootstrapped to ` 0 ` for every row
850+ (` --skip-image-hash ` ) at import time; Stage 8's phash engine is the
851+ first thing to actually populate it, lazily and only for rows it
852+ needs - most of the table (any candidate no pilot run has hashed yet)
853+ is still at the placeholder ` 0 ` .
785854- Client-side (Orama) search has no Stage 3 parity — see above.
786855- Upstreaming this feature is deprioritized — see
787856 [[ ../infrastructure.md]] 's Upstreaming section.
@@ -805,3 +874,27 @@ construction.
805874 deductive printing-tag backfill) reflect three concurrently-developed
806875 branches sharing this one doc file, numbered in landing order to avoid
807876 collisions.
877+ - ** Future work: anonymous_id trust scoring via honeypot questions**
878+ (2026-07-15, raised during Stage 8's pilot run). Idea: periodically
879+ serve a voter a card whose printing is already known with very high
880+ confidence — ideally an already-` RESOLVED ` card (real human-backed
881+ consensus), Stage 6's D1 tier as a fallback pool (0 false positives
882+ across 27,424 live cards, but still AI-derived, not independently
883+ human-verified, so using it as "trusted" ground truth to police other
884+ submissions has a circularity worth being honest about) — without
885+ telling the voter it's a check, and score their ` anonymous_id ` based on
886+ whether they answer correctly. Deprioritize/downweight low-scoring
887+ anonymous_ids to make data poisoning more costly. Same crowdsourcing
888+ pattern as reCAPTCHA/Mechanical-Turk gold-standard questions. Known
889+ limitation before this is worth building: ` anonymous_id ` is a
890+ client-generated, trivially rotatable value
891+ (` frontend/src/common/anonymousId.ts ` ) with no persisted identity —
892+ a trust score raises the cost of poisoning (a fresh ID needed per
893+ abuse attempt) but doesn't stop a determined actor, so it's a speed
894+ bump, not a hard Sybil defense. Also a genuinely new subsystem, not a
895+ small addition: a honeypot-injection point in ` question_feed.py `
896+ (nothing currently interrupts the three-tier ranked union with a
897+ planted question), somewhere to persist per-` anonymous_id ` trust state
898+ (no such model exists today), and a way to feed that score back into
899+ ` vote_consensus ` 's per-source weighting — worth its own design pass
900+ rather than bolting onto an existing stage.
0 commit comments