A personal analytics tool for the mobile game 三国谋定天下 (演武). The core recommendation pipeline is: game screenshots → OCR extraction → per-battle JSON → a deterministic offline model builder → a single generated artifact → a client-side React app that recommends heroes/skills and builds LLM prompts. Recommendation remains fully client-side. Isolated, write-only Cloudflare Pages Functions collect anonymous draft-choice telemetry and optional community battle reports without participating in scoring or static page reads. Scheduled GitHub workflows export only the relevant D1 table into runner-temporary storage, publish deterministic static artifacts and aggregate checkpoints, and purge only rows covered by a successfully published checkpoint. Raw telemetry, submission IDs, and telemetry timestamps are never committed. Accepted community battle files intentionally retain the exact contributor name, a normalized upload timestamp, and the selected season so suspicious submission patterns can be reviewed later; transport submission IDs remain D1-only.
Game rules: see GAME_RULE.md.
web/public/game-data/database.jsonholds the catalog plus the imported 飞将吕布 hero/skill rankings, complete strong/championship builds, matchup matrix, and analysis guide.- Copy game screenshots into
data/images/. make extract— OCR the images intodata/battles/*.json, then rebuildweb/src/recommendation_data.json.make sync-yanwu-corpus— download, checksum-verify, and normalize the pinned external Yanwu release when the Git-ignored cache is absent or stale.make build-recommendation— synchronize the pinned external corpus if needed, then rebuild fromdata/battles/, accepted reports indata/web-upload/, and the external normalized corpus.make evaluate-recommendation— run the deterministic grouped stable-hash evaluation and write the ignoredresults_recommendation_evaluation.json; this never changes production weights orweb/src/recommendation_data.json.make import-web-battles EXPORT=/path/to/web_battle_submissions.sql— revalidate and import one bounded D1 export, update the static leaderboard, and rebuild the recommendation artifact in one full batch.make build-telemetry EXPORT=/path/to/round_telemetry.sql— validate the current D1 table export, fold rows newer than the committed cursor, and rebuild the public aggregate artifact plusdata/telemetry_state.json.make import-yanwu— validate the local seven-sheet三谋演武-飞将吕布.xlsxwithout writing;make import-yanwu APPLY=1atomically updates the derived guide data indatabase.json.make web— start the React dev server (http://localhost:3000).
The recommender is an opponent-aware paired model trained offline and scored in the browser:
- Offline builder (
data/build_recommendation_data.py): validates manualdata/battles/*.json, accepteddata/web-upload/*.json, and the verified normalized external Yanwu release (failing clearly on unknown/invalid winners rather than counting both teams as losses), then trains a single regularized logistic / Bradley-Terry model. Each complete battle is one paired observation —features(team1) − features(team2)with the winner as the label. Features are hero presence, non-default skill presence, supported hero pairs, assigned hero-skill, and supported within-hero skill pairs; sparse interactions are filtered by a support floor and shrunk by L2. After fitting, the builder applies a deterministic, bounded, always-subtractive popularity penalty to atomicH/Sitems whose support rate is low relative to season-aware exposure. Catalog heroes and standalone skills below the fitting floor use a zero fitted baseline, so extremely rare or unused old items receive explicit negative weights without fitting unstable one- or two-battle coefficients. Newly introduced items receive grace while few battles have occurred since introduction. Hero signatures and explicit shadow skills are not synthesized at zero support, although observed non-default transfers remain eligible.HP/HS/SPinteractions remain governed by their support floors and L2, and rawmodel.supportremains literal evidence. Unknown-season battles still train the logistic model but are excluded from both observed and availability counts in the season-dependent popularity adjustment, so an untrusted corpus cannot manufacture thousands of exposures. The adjustment adds no artifact maps or client-side scoring logic. Catalog introduction seasons are required positive integers; a trusted known-season battle that predates one of its items fails validation. The builder emitsweb/src/recommendation_data.json(schema/catalog metadata, clean battle counts, model weights + per-feature support/evidence, smoothed hero/skill analytics, and a lightweight grouped stable-hash backtest). That check keeps capture/upload sessions intact, starts external reports from stable report identities, merges exact and one-skill-different matchup clusters, and assigns whole groups with the fixed seedsanmou-grouped-holdout-v2. Season, chronology, winner, and outcome do not determine split membership. The build is fail-closed — if any battle file is invalid or unreadable it aborts before writing, so a corrupt capture can never partially overwrite the artifact — and byte-reproducible: no wall-clock or prior-output fields, so re-running on the same corpus yields a byte-identical file. A deterministiccorpus_versioncontent hash identifies the runtime training inputs. Trusted localBattle.seasonremains model metadata for catalog consistency and known-season popularity exposure; imported Yanwu battles always use null. - No runtime opponent. The user never enters an opponent. A team's score is
its relative roster strength (
w · features(team)) against the learned metagame — not an opponent-specific win probability. The opponent term is a shared constant across a user's options and is dropped. - Client engine (
web/src/services/recommendationEngine.ts, backed byrecommendationModel.ts): offered-set picks rank options by marginal roster-strength improvement over the current pool + evidence. The two-support- skill pick is chosen as a joint pair (each skill's presence + the best feasible hero routing + the within-hero skill-pair bonus when both land on one hero), not two independent top-1 picks. The Team Builder uses an evidence-only policy. A hero must independently clear the atomic hero (H) gate, and every relationship inside a pair/trio must independently clear the hero-pair (HP) gate. Each gate uses the fitted model's support floor (currently 5 battles for atomic hero/skill features and 8 for pair features). Positive, zero, and negative fitted weights remain eligible and affect ranking; missing or under-supported features still fail, so one well-supported hero cannot rescue an unobserved partner. If a qualified group matches two or three members of a known team inweb/public/game-data/database.json, its formation and canonical hero slots are preserved; guide data never bypasses the model gates. A skill must independently clear both its atomic skill (S) and hero-skill (HS) gates. Supported within-hero skill-pair (SP) evidence, including negative weights, ranks qualified choices without vetoing the pairing. Owned guide skills keep their canonical slots only after passing the same gates. Unsupported heroes and skills stay in the warehouse for manual placement instead of being forced into a complete 9-hero/18-skill result. The deterministic search runs in a client Web Worker, with a yielding main-thread fallback and an in-memory result cache, so it adds no Cloudflare Function usage and keeps the loading UI responsive. Players can then drag, tap, or use the keyboard to rearrange its three teams. The UI shows each team's live 评分 and compact positive evidence (武将配合 / 武将与战法 / 战法搭配, each with 加分 and reference battle counts); displayed evidence keeps a +0.1 visibility floor so tiny accepted gains are not rendered as “+0.0”, and there is no aggregate 总评分.
make evaluate-recommendation runs the full evaluation-only harness in
data/evaluate_recommendation_model.py. Inputs retain three reported source
categories:
data/battles/→uploaded_by_medata/web-upload/→uploaded_by_others- pinned normalized Yanwu release →
external_yanwu
Protocol version 2 is deliberately season-independent. Leakage groups keep capture/upload sessions together using a 30-minute inactivity window (web uploads are partitioned first by exact contributor identity). Each external Yanwu report starts from its immutable report identity rather than making the release one giant group. Exact and one-skill-different matchup clusters are then merged with those initial groups. Winner and outcome are excluded from matchup identity, and season is never read while grouping or splitting.
The locked test was selected once from the pre-Yanwu corpus: 20% of its whole
leakage groups by the fixed seed
sanmou-grouped-holdout-v2:pre-yanwu-locked-test. Its source-qualified battle
identities and original group IDs are persisted in
data/evaluation/locked-pre-yanwu-test.json, so later manual captures or web
uploads cannot enter, displace, or rename the locked population. Any new or
Yanwu group that touches a locked-test session or exact/near-duplicate matchup
is removed. The remaining whole pre-Yanwu and eligible Yanwu groups are divided
into training and development with the independent fixed seed
sanmou-grouped-holdout-v2:development (20% development). The test is not used
for configuration selection.
Training/development groups tune logistic regularization C, single/pair
support floors, the SP within-hero skill-pair ablation, and the popularity
penalty (gamma and tau). Season-recency weighting and season-trend variants
were removed rather than replaced with another temporal assumption. Selected
and current production configurations are refit on training plus development,
then scored once on the locked test. The report includes split source/outcome
balance, accuracy, log loss, Brier score, feature coverage, source breakdowns,
and deterministic 95% percentile confidence intervals that resample whole
locked-test leakage groups. Intervals are omitted below five groups and marked
exploratory below twenty.
The same report includes the controlled Yanwu comparison. A baseline production configuration is trained on all non-test pre-Yanwu groups; a candidate with the identical configuration adds all eligible Yanwu groups; both score the exact same locked pre-Yanwu rows. It reports sample/group counts, coverage, paired metric deltas and uncertainty, plus source-level results where group evidence permits. The report labels the result inconclusive and makes no improvement claim unless the paired 95% intervals support better accuracy, Brier, and log loss together.
The harness atomically rewrites only the ignored
results_recommendation_evaluation.json. Candidate settings are recommendations
for review: they are never fed back into the builder, and no production weight
or support-threshold change happens automatically.
The /contribute page is intentionally a small no-auth experiment. A player can
copy a catalog-backed DeepSeek OCR prompt and paste its JSON to prefill every
recognized catalog value in the confirmation form; missing or unrecognized
values remain editable, and final submission still requires strict validation.
The player can also skip OCR and enter both teams manually. The prompt asks
DeepSeek to recognize each hero's first/signature skill before reverse-mapping
the hero, and supplies rough normalized portrait and landscape positions rather
than device-specific pixel crops. The player reviews every hero, skill, winner,
and the two teams' scores from the current model before submission.
The optional public contributor name is stored in a one-year cookie, remains
editable even after an anonymous submission, and may be empty; printable
Unicode is preserved exactly. A separate contribution-season cookie defaults to
the highest numeric season in database.json. It does not read or modify the
homepage setup season.
POST /api/battles repeats validation against
web/public/game-data/database.json and writes through the existing
TELEMETRY_DB D1 binding to web_battle_submissions. The endpoint is
write-only, requires JSON and an explicit uploader string (the empty string is
anonymous), and rejects browser requests from a different origin. It has no
login: direct clients can still call it, but the live D1 queue is atomically
capped at 500 reports. Once full, new uploads receive HTTP 429 until the daily
job drains it; retries of an existing submission ID remain idempotent. The
separate /contributors page fetches the generated static
web/public/game-data/web_upload_data.json; neither it nor the homepage reads
D1 or a Function.
web/public/_routes.json limits Pages Function execution to /api/*, so
Function quota exhaustion does not route static pages or assets through a
Worker.
The daily update-web-battles.yml workflow:
- applies the idempotent D1 table migration and exports only
web_battle_submissionsinto runner-temporary storage; - processes at most 500 rows in ascending AUTOINCREMENT order and revalidates each report;
- allows at most two occurrences of one semantic fingerprint, where ordered hero/skill positions, winning lineup, and exact uploader are significant but swapping the two team sides is not; the selected season is deliberately not part of the fingerprint, so changing it cannot bypass duplicate detection;
- commits accepted reports with
uploader_name, normalizeduploaded_at, andseasonmoderation metadata, together with the aggregate checkpoint, static leaderboard, and a full one-shot recommendation rebuild; and - deletes D1 rows only through the high-water mark read back from that successful commit.
Malformed and third-or-later duplicate reports advance the checkpoint as aggregate rejections and receive no leaderboard credit. A transport retry with the same UUID is idempotent and is separate from semantic duplicate handling. Both data-publishing workflows share one concurrency group so their generated data commits cannot race each other.
data/external/yanwu-release.json pins one immutable release asset from
CharlesWang505/yanwu-battle-reports,
including its byte size, SHA-256, source count, schema, attribution, and
CC BY 4.0 licence.
Adopting a later release is a reviewed manifest update; scheduled jobs never
follow a mutable “latest” URL. Any s16 text in the immutable release tag,
filename, or URL is treated as an opaque asset identifier. The raw report's
season label is ignored, is absent from manifest model metadata, and every
normalized Yanwu battle has "season": null.
make sync-yanwu-corpus verifies and normalizes the release into
.cache/yanwu/. The raw and normalized files are regenerable and Git-ignored.
A warm cache makes no network request: sync checksum-verifies the pinned raw
asset, deterministically regenerates the expected normalized value, and accepts
the normalized cache only when it matches. A cold cache downloads to a
temporary file, verifies the manifest before publication, and normalizes
atomically. Unknown catalog names or a checksum/schema/count mismatch fail
closed.
make build-recommendation depends on this sync step, so it never silently
falls back to a local-only model.
The daily web-battle workflow restores .cache/yanwu/ through GitHub Actions
using a key derived from the manifest, normalizer, and game catalog. If GitHub
evicts the cache, the next job reconstructs it once from the immutable release.
Model fitting consumes only the verified local normalized file and remains
deterministic and offline.
Before accepting traffic, apply
web/migrations/0003_web_battle_submissions.sql to the same D1 database bound
as TELEMETRY_DB. The scheduled workflow also applies it idempotently, but the
deployed Function needs the table immediately.
cd web
pnpm dlx wrangler@4.112.0 d1 execute "$CLOUDFLARE_D1_DATABASE_NAME" \
--remote \
--file=migrations/0003_web_battle_submissions.sql \
--yesimage_extraction/— OCR skill extraction (PaddleOCR).skill_extraction_system.pyis the engine;batch_extract_battles.pyruns it overdata/images/and writesdata/battles/*.json.test_image_extraction.pyvalidates against golden image fixtures inimage_extraction/fixtures/(~69 MB, intentionally committed).study-battle-report/ocr_battle_log.py— a separate OCR script for battle-log screenshots. It deliberately duplicates some OCR/db/fuzzy-match logic fromimage_extractionbecause the two live in different workspaces; do not merge them unless they start changing in lockstep.data/build_recommendation_data.py— the deterministic offline model builder: validates all three battle sources and emitsweb/src/recommendation_data.json(the single artifact the web app reads).data/test_build_recommendation_data.pycovers validation/feature-extraction/training and the lightweight grouped stable-hash backtest. Manual and web observations share a fail-closed maximum-two semantic duplicate policy.data/evaluate_recommendation_model.py,data/recommendation_evaluation.py, anddata/evaluation/locked-pre-yanwu-test.json— the deterministic full grouped-holdout experiment harness, its checked-in locked-test identities, and shared stable-hash split, session grouping, near-duplicate, metric, and cluster-bootstrap helpers. Its ignored JSON report is evaluation-only.data/import_web_battles.py— validates a boundedweb_battle_submissionsD1 export, advances the aggregate checkpoint over accepted and rejected rows, writes accepted reports plus contributor/time/ season moderation metadata todata/web-upload/, renders the static leaderboard, and drives a complete recommendation rebuild.data/yanwu_corpus.py,data/sync_yanwu_corpus.py, anddata/external/yanwu-release.json— pin, verify, normalize, and cache the external CC BY 4.0 release without committing its large data artifacts.data/web_upload_state.json— generated aggregate checkpoint containing the D1 cursor, cumulative accepted/rejected totals, public contributor totals, and versioned duplicate-fingerprint counts. It contains no raw battle payloads, submission UUIDs, or per-row timestamps.data/build_telemetry_data.py— the deterministic telemetry builder. It fails closed when the D1 export or schema cannot be verified; individual malformed, catalog-mismatched, or impossible events are quarantined and exposed only as an aggregateinvalid_event_count. Recommendation scores, recommendation positions, and model-version labels are client-reported, indicative telemetry: they are checked for bounded shape and internal consistency but are not replayed against historical recommendation models. Valid rows are reduced atomically toweb/public/game-data/telemetry_data.json. The cumulative schema-v5 artifact covers all ten rounds and adds offer/pick, round, position, score-margin, and model-disagreement aggregates plus a deterministic online conditional-choice model. The model remains unavailable until explicit event/estimated-session/disagreement/evaluation evidence gates and a quality gate pass. The raw export remains outside the repository. During each incremental build, it validates and advancesdata/telemetry_state.json, which contains only cumulative counters, a fixed-size anonymous session estimate, resumable model state, and the last processed D1 row ID. Schema v5 is rendered solely from that checkpoint, so old raw rows can be deleted without reducing public totals. Optimizer features, optimizer deltas, and model-quality statistics are persisted only in groups supported by at least ten new events, so a small batch's pool/offer/choice correlations or probability vector are not committed. Cumulative recommendation-model labels are capped at 32 entries by folding low-support historical labels into anotherbucket without changing the event total.data/telemetry_retention.py— validates the D1 AUTOINCREMENT migration, sequence/cursor safety, and aggregate Wrangler results, then prepares one bounded 14-day purge. It also appends an aggregate-only publication report (newly validated and cumulative validated event counts, derived from committed checkpoints) to the job summary. It never reads or prints row-level telemetry.data/telemetry_state.json— generated, aggregate-only telemetry checkpoint committed atomically with the public telemetry artifact. It contains no raw event records, event/session identifiers, or timestamps. Its public-style offer/pick counters can include small totals, while correlated model features and evaluation deltas are support-gated.web/— React (Vite) + MUI; recommendation and leaderboard reads are client-side/static, with isolated write-only Pages Functions for anonymous telemetry and battle submissions. TypeScript-enabled (type-check withpnpm typecheck, backed by the Go-nativetypescript@7). Notable modules:src/services/recommendationEngine.ts— offered-set/support/formation recommendations + analytics, scored against the artifact.src/services/recommendationModel.ts— pure paired-model primitives (feature extraction + scoring), kept in lockstep with the Python builder.src/services/promptGenerator.ts— builds the LLM prompts (uses model weights + analytics).src/context/GameContext.tsx— global game state (useReducer); getdispatchviauseGame().src/utils/{clipboard,rankings,storage,usePinyin*}— shared utilities.src/types/— hand-written domain types (domain.ts,recommendation.ts,game.ts) fordatabase.json/recommendation_data.jsonand the game state/reducer.src/data.ts— the central typed boundary that imports and casts the bundled JSON once.
agent/— local TypeScript HTTP/CLI runtime for model-backed experiments. It uses an OpenAI-compatible Responses API provider boundary, is never required by the public static site, and hosts one parent LangGraph recommendation workflow. Its internal hero, formation/row, and skill subgraphs fill only null values, preserve every already-filled value, and then run a read-only team review. Already-complete inputs route directly to that review. Its loopback HTTP server exposes a typed team-recommendation endpoint to explicitly allowed browser origins while keeping generic chat outside browser CORS. See agent/README.md for the graph nodes and thepnpm recommendfixture run.data/import_yanwu_workbook.py— strict, deterministic seven-sheet workbook importer. It defaults to a no-write dry run and requires--applyto updateweb/public/game-data/database.json; the source workbook itself stays untracked.web/public/game-data/database.json— catalog and guide data. Hero/skill rankings, known builds, championship references, matchup relationships, and analysis are attributed in the guide metadata to 飞将吕布; contact details from the workbook are never published.web/public/game-data/telemetry_data.json— generated, aggregate-only player-choice analytics and gated preference-model artifact; updated weekly by GitHub Actions.web/public/game-data/web_upload_data.json— generated static upload totals and contributor leaderboard; updated daily with accepted battle imports.web/src/recommendation_data.json— generated bybuild_recommendation_data.py; don't hand-edit.autojs/— AutoJS (Android) scripts that capture the screenshots. Device-specific.
make extract— OCR all images indata/images/, then rebuild the recommendation artifact.make sync-yanwu-corpus— populate or validate the Git-ignored pinned external corpus cache; a valid warm cache makes no network request.make build-recommendation— synchronize the pinned release if needed and regenerateweb/src/recommendation_data.jsonfrom manual, accepted web-upload, and external battles.make evaluate-recommendation— run the grouped stable-hash model evaluation and write ignoredresults_recommendation_evaluation.json; it does not update the production recommendation artifact.make test— image-extraction Python tests (pytest image_extraction/, parallel). ~40s (loads PaddleOCR).make test-data— the offline data-builder Python suites, including the incremental-checkpoint tests (fast, no PaddleOCR).make test-web-battles— the web-battle importer plus recommendation-builder suites.make test-telemetry— telemetry-builder and incremental-checkpoint Python tests (fast, stdlib-compatible).make web— start the Vite dev server (port 3000).- Web unit tests:
cd web && pnpm test(Vitest). Type-check:cd web && pnpm typecheck(Go-nativetsc). E2e:cd web && pnpm test:e2e(Playwright). Build:cd web && pnpm build. - Local agent:
cd agent && pnpm start. Token-free checks:pnpm typecheck && pnpm test && pnpm build. Explicit live model check:pnpm smoke. Explicit combined LangGraph hero + formation + skill check:pnpm recommend fixtures/partial-teams.json. - Python runs under uv (Python 3.12):
uv run python <script>.make syncinstalls deps.
web/src/recommendation_data.json is generated; never hand-edit it. It contains:
schema/catalog— model + database metadata (incl. hero→default-skill map and acatalog_versioncontent hash).battle_counts— clean total / team1 / team2 wins, invalid count, and a deterministiccorpus_versioncontent hash over runtime training inputs. Trusted localBattle.seasonremains part of that hash because it can affect catalog checks and popularity adjustment; Yanwu contributes only null (no build timestamp — the artifact is byte-reproducible).model— the paired logistic weights keyed by feature id, plus per-featuresupport(evidence). Feature ids are pipe-joined, with pairs sorted for order-independence:H|hero,S|skill,HP|a|b,HS|hero|skill,SP|hero|s1|s2. AtomicH/Sweights include the bounded, always-subtractive low-popularity adjustment. Its observed and exposure counts exclude unknown-season rows, although those rows still train the logistic fit. Below-floor catalog heroes and standalone skills may therefore have penalty-only weights; they are not added to logistic fitting.HP/HS/SPremain support-floor/L2-only. Rawmodel.supportis literal observed evidence, not penalty-adjusted. Zero-support entries are omitted from that map to avoid repeating names—the client already interprets missing support as0. No separate penalty/exposure maps are serialized or needed by client scoring. Build the same ids in TS viaweb/src/services/recommendationModel.ts; never re-derive them inline. JS[a,b].sort()equals Pythonsorted()for these CJK (BMP) names — the invariant the keying relies on.analytics— smoothed per-hero/skill win rates + usage.backtest— the lightweight grouped stable-hash check for the current production configuration, including accuracy, log loss, Brier, cluster-aware uncertainty, source/outcome split balance, source breakdowns, and a separateevaluation_versionfor evaluation-only source/session metadata beyond the runtimecorpus_version. Capture/upload sessions and exact/near-duplicate matchups stay together; season and outcome do not affect membership. Hyperparameter and controlled-corpus comparisons live only in the full evaluator's ignored result file.
- Recommendation and leaderboard reads are static/client-side only;
src/services/api.tsis an in-memory scoring shim, not HTTP.web/functions/api/telemetry/rounds.jsandweb/functions/api/battles.jsare isolated write-only Cloudflare Pages endpoints. - When changing recommendation/prompt logic, protect it with the behavior-focused
unit tests in
web/src/services/__tests__/(paired feature extraction, model scoring, global optimisation, deterministic output, no runtime opponent). - Regenerable/scratch dirs are gitignored:
extracted_results/,tmp_crops/,test-results/, plusstudy-battle-report/battles/*/OCR artifacts.
This README is the canonical project doc for humans and coding agents. Claude Code
loads it via CLAUDE.md; Codex and Rovo Dev via AGENTS.md. Directory-scoped agent
notes live in web/AGENTS.md and image_extraction/.agent.md.