Skip to content

Commit 1b439ed

Browse files
Document Stage 8 pilot run results and trust-scoring future work
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016i9S7LQsCL3FGaih3ZTRBJ
1 parent ff7bedd commit 1b439ed

1 file changed

Lines changed: 135 additions & 42 deletions

File tree

‎docs/features/printing-tags.md‎

Lines changed: 135 additions & 42 deletions
Original file line numberDiff line numberDiff line change
@@ -688,17 +688,22 @@ Old/Modern Border, Future Frame) 400 on tap; `Full Art`/`Borderless`/
688688

689689
## Stage 8: local (zero-API-cost) printing-identification backfill pilot
690690

691-
**Status: environment proposal delivered, not yet built** (2026-07-15,
692-
`worktree-local-printing-id-pilot`) - this section records the plan shape,
693-
not a shipped feature; see `journal/2026-07-15-local-printing-id-pilot.md`
694-
(machine-local, not committed) for the full investigation.
691+
**Status: built and pilot-run against live production** (2026-07-15,
692+
`worktree-local-printing-id-pilot`, PR #22). Code + tests merged; a real
693+
`--limit 300 --engine both --nice` invocation ran against the live DB and
694+
its full results are summarized below. **Full-catalog run explicitly NOT
695+
executed** - see "Real pilot run results" below for why (a ~13-day
696+
single-process projection). See
697+
`journal/2026-07-15-local-printing-id-pilot.md` (machine-local, not
698+
committed) for the complete data dump this summary is drawn from.
695699

696700
Sibling to Stage 6's deductive backfill, same non-negotiable principle
697-
(a deduction is always a vote, never a direct resolve - the human-backed
701+
(a vote is always just a vote, never a direct resolve - the human-backed
698702
gate in `vote_consensus.resolve_weighted_consensus` still applies), but
699703
sourced from actually looking at the card image instead of pure logical
700704
deduction from existing structured data - two independent local (no paid
701-
API calls) engines:
705+
API calls) pass-1 engines, plus a pass-2 fallback for cards pass 1 can't
706+
reach at all:
702707

703708
- **L1, OCR**: Tesseract on a cropped, preprocessed collector-line region
704709
(bottom-left corner, grayscale/upscaled/thresholded). Parses candidate
@@ -708,39 +713,101 @@ API calls) engines:
708713
trusting the OCR output itself. Never writes `is_no_match`.
709714
- **L2, perceptual hash**: art-region phash comparison against each
710715
name-candidate's Scryfall art crop, voting only when there's a clear
711-
single best match (distance threshold + margin over the second-best).
712-
`CanonicalCard.image_hash` (`models.py`) already exists as a
713-
`BigIntegerField` for exactly this - added when `import_canonical_card_ data` first shipped ("CanonicalCard population fix" above) but never
714-
computed in production (`--skip-image-hash` was used for the real
715-
import; confirmed live, 113,224/113,224 rows still at the placeholder
716-
`0`) - so this pilot is the first thing to actually populate it, lazily,
717-
only for candidates it needs.
718-
- Both engines vote under their own `anonymous_id`
719-
(`local-ocr-v1`/`local-phash-v1`), same weight/gate treatment as any
720-
other AI-sourced vote - when both vote on the same card and agree, both
721-
votes stand as independent evidence; on disagreement, **neither** is
722-
written (logged instead - the disagreement set is the interesting
723-
output of a pilot like this, not noise to discard).
724-
725-
**Environment**: the OCR engine needs the `tesseract-ocr` system binary
726-
(not pip-installable - `pytesseract` is a thin subprocess wrapper around
727-
it), which the django container's image doesn't currently have. Two
728-
options, no Dockerfile change made without explicit sign-off given this is
729-
pilot-only tooling: bake `tesseract-ocr` into `docker/django/Dockerfile`'s
730-
existing apt-get line (matches how every other `manage.py` command in this
731-
repo is invoked, at the cost of a production image rebuild/restart and a
732-
permanent size increase for a one-off tool), or run from a host-side venv
733-
pointed at the already-`127.0.0.1`-exposed Postgres/Elasticsearch ports
734-
(zero image/container change, trivially reversible, matches this
735-
project's precedent of running one-off scripts against the live DB from
736-
outside Docker). Command itself (`manage.py local_identify_printing_tags`)
737-
is unaffected either way - only the invocation environment differs.
738-
739-
**Pilot discipline**: `--limit 300` default, explicit hold before any
740-
full-catalog run: this is genuinely new signal (visual inspection, not
741-
pure logical deduction) and needs a human spot-check of yield/accuracy
742-
before scaling up, unlike Stage 6's deduction which was provably exact by
743-
construction.
716+
single best match (distance threshold + margin over the second-best,
717+
recalibrated from real production data - see the journal for the
718+
calibration history). `CanonicalCard.image_hash` (`models.py`) already
719+
exists as a `BigIntegerField` for exactly this - added when
720+
`import_canonical_card_data` first shipped ("CanonicalCard population
721+
fix" above) but never computed in production (`--skip-image-hash` was
722+
used for the real import) - so this pilot is the first thing to
723+
actually populate it, lazily, only for candidates it needs. Capped at
724+
12 candidates per name (basic lands/staples can have hundreds - see
725+
"Real pilot run results" for how often this cap fires).
726+
- **Pass 2, fallback** (`local_fallback.py`, `local-fallback-v1`): fires
727+
only when pass 1 (either engine) produced no accepted vote for a card -
728+
the old-border-frame case (no collector line printed on the card face
729+
at all, just an "Illus. `<artist>`" credit). Evidence-combination model
730+
across border-color sample, artist-name OCR fuzzy match, and set-symbol
731+
phash (found unreliable in practice, kept but effectively disabled via
732+
a strict threshold - see `local_fallback.py`'s module docstring for the
733+
full negative finding): a vote is cast only when the intersection of
734+
every sub-check that produced a reading narrows to exactly one
735+
candidate.
736+
- Border-color sampling and frame-style classification (OCR-collector-
737+
line-present vs. Illus.-anchor-present) run for **every** processed
738+
card regardless of printing-vote success, casting standalone
739+
attribute-chip votes (Black/White/Silver Border, Borderless, Old/Modern
740+
Border) - and, when a printing vote **is** confirmed for that card this
741+
run, preferring ground truth from that printing's own
742+
`CanonicalPrintingMetadata` (Scryfall `border_color`/`frame`) over the
743+
heuristic estimate. The same heuristic reading also feeds a
744+
**consistency check**: if a card's observed frame class contradicts its
745+
matched printing's real frame value, the printing vote itself is
746+
withheld (kept as a frame-vote-only outcome) rather than trusting an
747+
art/OCR match that likely landed on the wrong printing.
748+
- All engines vote under `VoteSource.OCR` (the 2026-07-15 split of the
749+
old single `VoteSource.AI` value into `DEDUCTION`/`OCR` - see
750+
`models.py`'s `VoteSource` docstring; same weight/gate treatment as
751+
before, individual technique still distinguishable via `anonymous_id`)
752+
- when OCR and phash both vote on the same card and agree, both votes
753+
stand as independent evidence; on disagreement, **neither** is written
754+
(logged instead - see the journal's disagreement examples).
755+
756+
**Environment**: resolved via a host-side venv pointed at the
757+
already-`127.0.0.1`-exposed Postgres/Elasticsearch ports (zero
758+
Docker/container change) - `tesseract-ocr` installed via host apt,
759+
`pytesseract`/`ImageHash`/`Pillow` via the venv's `requirements.txt`
760+
install. No Dockerfile change made. (A future full-catalog run, if one
761+
ever happens, should revisit baking `tesseract-ocr` into
762+
`docker/django/Dockerfile` instead, per the original tradeoff writeup.)
763+
764+
### Real pilot run results (2026-07-15, `--limit 300 --engine both --nice`)
765+
766+
**32m4.6s wall-clock, 19m36s user + 4m37s sys CPU (≈76% avg utilization of
767+
one core on this 2-CPU box), exit 0.** `--nice` confirmed actually
768+
throttling (process niceness observed alternating 5↔19 during the run).
769+
770+
| Engine | Attempted | Votes written | Yield |
771+
| --------- | --------- | ------------- | ----- |
772+
| OCR | 300 | 77 | 25.7% |
773+
| Phash | 300 | 13 | 4.3% |
774+
| Fallback | 210 | 4 | 1.9% |
775+
| **Total** | — | **94** | — |
776+
777+
**Gate check: 0/94 affected cards resolved** - the human-backed gate held
778+
perfectly at this scale, same result as Stage 6's 0/28,112.
779+
780+
Largest skip bucket by far: OCR's "parsed-but-no-match" at 176/300
781+
(58.7%) - a syntactically valid collector line that didn't match any of
782+
the card's own candidates. Not investigated further in this pilot (out of
783+
scope), but the single most promising lead for improving yield before any
784+
larger run - see the journal for the plausible-causes breakdown.
785+
786+
Attribute votes: border `{black: 280, borderless: 17, white: 3}` (91 from
787+
ground truth, 209 from the pixel heuristic); frame `{modern: 258, old: 14}`, 28 abstains (91 from ground truth, 181 from the OCR/Illus.-anchor
788+
heuristic); **6 frame-mismatches** (printing vote withheld by the
789+
consistency check) - see the journal for all 10 sampled examples and the
790+
per-case reasoning.
791+
792+
**Full-catalog projection: ~171,800 eligible cards remain (of 179,002 raw
793+
eligible pool) → naive linear projection ≈ 306 hours ≈ 12.8 days of
794+
continuous single-process runtime.** This is the key number for any
795+
future decision to scale up - not attempted in this pilot, and not
796+
practical as a single uninterrupted process. Before attempting it:
797+
parallelizing across multiple processes/pk-range partitions, and (more
798+
urgently) switching from the current one-giant-`bulk_create`-at-the-end
799+
write pattern to periodic batch flushing (matching
800+
`deductive_backfill.py`'s existing `batch_size`/`flush()` precedent) so a
801+
multi-day run's progress survives a crash/restart/deploy instead of
802+
losing everything accumulated since the last completed run - both raised
803+
but not implemented in this pilot, out of its locked scope.
804+
805+
5-vote spot check, 20-vote random admin-link sample, 3 disagreement
806+
examples, and the filename tag-gap census (1,097 unresolved cards with an
807+
unmatchable `expansion_hint`) are all in the journal, not duplicated here.
808+
809+
**Pilot discipline honored**: `--limit 300`, no full-catalog run attempted
810+
per the original hold.
744811

745812
## Key files
746813

@@ -779,9 +846,11 @@ construction.
779846
via iterative screenshot review, not built against a real design system -
780847
owner has flagged that this needs a proper pass with the `/dataviz` skill
781848
in the future rather than further ad hoc CSS tuning.
782-
- `CanonicalCard.image_hash` is bootstrapped to `0` for every row
783-
(`--skip-image-hash`); real perceptual-hash-based matching isn't
784-
implemented yet.
849+
- `CanonicalCard.image_hash` was bootstrapped to `0` for every row
850+
(`--skip-image-hash`) at import time; Stage 8's phash engine is the
851+
first thing to actually populate it, lazily and only for rows it
852+
needs - most of the table (any candidate no pilot run has hashed yet)
853+
is still at the placeholder `0`.
785854
- Client-side (Orama) search has no Stage 3 parity — see above.
786855
- Upstreaming this feature is deprioritized — see
787856
[[../infrastructure.md]]'s Upstreaming section.
@@ -805,3 +874,27 @@ construction.
805874
deductive printing-tag backfill) reflect three concurrently-developed
806875
branches sharing this one doc file, numbered in landing order to avoid
807876
collisions.
877+
- **Future work: anonymous_id trust scoring via honeypot questions**
878+
(2026-07-15, raised during Stage 8's pilot run). Idea: periodically
879+
serve a voter a card whose printing is already known with very high
880+
confidence — ideally an already-`RESOLVED` card (real human-backed
881+
consensus), Stage 6's D1 tier as a fallback pool (0 false positives
882+
across 27,424 live cards, but still AI-derived, not independently
883+
human-verified, so using it as "trusted" ground truth to police other
884+
submissions has a circularity worth being honest about) — without
885+
telling the voter it's a check, and score their `anonymous_id` based on
886+
whether they answer correctly. Deprioritize/downweight low-scoring
887+
anonymous_ids to make data poisoning more costly. Same crowdsourcing
888+
pattern as reCAPTCHA/Mechanical-Turk gold-standard questions. Known
889+
limitation before this is worth building: `anonymous_id` is a
890+
client-generated, trivially rotatable value
891+
(`frontend/src/common/anonymousId.ts`) with no persisted identity —
892+
a trust score raises the cost of poisoning (a fresh ID needed per
893+
abuse attempt) but doesn't stop a determined actor, so it's a speed
894+
bump, not a hard Sybil defense. Also a genuinely new subsystem, not a
895+
small addition: a honeypot-injection point in `question_feed.py`
896+
(nothing currently interrupts the three-tier ranked union with a
897+
planted question), somewhere to persist per-`anonymous_id` trust state
898+
(no such model exists today), and a way to feed that score back into
899+
`vote_consensus`'s per-source weighting — worth its own design pass
900+
rather than bolting onto an existing stage.

0 commit comments

Comments
 (0)