A bookstore recommender that makes its evaluation inspectable. Browse typographic book covers, compare training-derived recommendation signals, and read the evidence behind model and off-policy estimates.
Open the bookstore · Source · Measured results · Evaluation report
These are demonstration recommendations on a public dataset; no real reader's identity is present and no recommendation is personalized to a real person.
The pinned public CSV -> parquet -> NumPy/SciPy exact scoring backend evaluated 10,000 books and 5,000 readers using Global source-order holdout, full catalog. The source has no rating timestamps, so row order is a temporal proxy. These results apply to the recorded eligible reader sample and its evaluated catalog; they are not a date-verified backtest.
| Model | NDCG at ten with interval | Recall at two hundred with interval |
|---|---|---|
| Popularity | 0.0406 [0.0373, 0.0441] | 0.1298 [0.1249, 0.1346] |
| Item cosine | 0.0541 [0.0507, 0.0579] | 0.2585 [0.2517, 0.2654] |
| Implicit ALS | 0.0546 [0.0513, 0.0579] | 0.2607 [0.2543, 0.2672] |
| Fixed blend | 0.0538 [0.0504, 0.0576] | 0.2707 [0.2641, 0.2774] |
| Content TF-IDF | 0.0411 [0.0381, 0.0442] | 0.1399 [0.1348, 0.1451] |
| LambdaMART | 0.0440 [0.0415, 0.0466] | 0.2154 [0.2097, 0.2214] |
| Two-tower neural | 0.0019 [0.0014, 0.0024] | 0.0162 [0.0146, 0.0181] |
| Session cosine | 0.0900 [0.0855, 0.0947] | 0.2900 [0.2830, 0.2967] |
| Elman recurrent | 0.0232 [0.0212, 0.0252] | 0.1041 [0.0995, 0.1086] |
Observed leader: session_cosine. Corrected selection: session_cosine. The full report includes paired comparisons, shortcut diagnostics and limitations.
uv sync --frozen
uv run pytest
uv run python scripts/render_reports.py --check
cd web
npm ci
npm run devFor durable recommendations and interaction logging, start the API in a separate terminal from the repository root:
uv run uvicorn packages.api.main:app --host 127.0.0.1 --port 8000 --no-access-logThe default database is local SQLite. Configure DATABASE_URL, a persistent
SESSION_HASH_KEY, ADMIN_TOKEN, and CORS_ORIGINS for the intended environment.
Administrative writes are unavailable until their token is set. See
serving and logging for endpoint and security
behavior. GitHub Pages hosts the frontend.
The live API uses a Worker and durable D1 storage on
the selected free Sites host. Public request checks and a separate D1 read
verified persisted interactions. It serves bounded SSE responses with
EventSource reconnection and Last-Event-ID replay because the host buffers long
streaming responses. See deployment verification.
The Worker/D1 deployment replaces a paid managed-service dependency. PostgreSQL
remains optional for the Python reference and local ANN benchmark. Browser-only
state does not count as a server exposure log.
| Module | Recorded status | Evidence or limit |
|---|---|---|
| Popularity / cosine / implicit ALS / blend | measured | Train-only scorers with exact full-catalog offline evaluation and persisted factors. |
| LambdaMART learned ranker | measured | Learned over a retrieval candidate union using an earlier source-order window. Evaluated against all baselines over the complete unseen catalog. |
| pgvector / approximate nearest neighbors | measured_local_postgres | Exact SQL and HNSW measured on the frozen two-tower embeddings in local PostgreSQL, with the same training-seen filters. This is separate from the hosted D1 serving path. |
| Content TF-IDF path | measured | Normalized title, author and tag TF-IDF profiles score all books including training-cold items. Uses undated metadata; no historical metadata availability claim. |
| OPE simulator | measured | Local IPS, SNIPS, DM, DR with 200 seeds, two sample sizes and oracle/misspecified reward models. |
| Real logged-bandit OPE | measured_six_file_bts_benchmark | All six random/BTS samples. Official campaign beta priors define the BTS target through beta draws ranked into three positions. Random logs after a shared cutoff provide OPE; BTS logs after the same cutoff provide an empirical on-policy benchmark with Wilson intervals. Logged BTS propensities vary within item/slot, so the frozen-prior approximation is not proven identical to the deployed per-impression policy. Differences include target mismatch and sampling error, not pure estimator error. Row bootstrap is conditional on the fitted reward model and Monte Carlo target, ignores repeated-reader dependence, and is not joint-slate inference. |
| Verified temporal features | unavailable | Source ratings lack timestamps. Source-order train-only features are tested; metadata is a snapshot. |
| Two-tower neural retrieval | measured | Independent user-ID and item-ID embeddings32 -> learned projection16 -> ReLU -> L2 normalization |
| Recurrent session model | measured | 16-dimensional Elman tanh recurrent encoder with learned input/output item embeddings |
The measured models and additional experiments are listed above. The earlier
bounded release is retained under results/history/; its scores must not be
compared directly with the expanded population. A score blend and a learned
ranker are identified separately. See scope decisions for
differences from the originating brief and the recorded deployment choices.
| Skill | Inspectable evidence |
|---|---|
| Ranking evaluation and comparison | Protocol, results, paired tests and shortcut diagnostics |
| Learned and baseline models | Model cards, recorded training artifacts and candidate-feature contract |
| Counterfactual policy evaluation | Six-file OBD benchmark, simulator truth and pinned source crosscheck |
| Serving and durable interactions | Serving contract, independent-connection and subprocess tests |
| Reproducibility and honest reporting | Strict document renderer, executed notebooks, population labels |
| Release verification | CI workflow, lint, strict types, branch coverage, browser tests and semantic palette checks |
uv run python scripts/crosscheck_obp.py verifies the local OPE point estimates
against pinned upstream arithmetic methods. This is an explicitly scoped source
reference check; it does not certify the full OBP package on this Python version.
uv run python scripts/build_evidence.py downloads the declared sources and
rebuilds the core manifest and browser artifacts. Then run
uv run python scripts/build_neural_evidence.py,
uv run python scripts/render_reports.py,
uv run python scripts/build_notebooks.py and the checks above. This is a model
and evidence rebuild, not a live database reset. Review the
evaluation protocol, runbook,
architecture, and contribution guide.
Original code is MIT licensed. Goodbooks-10k and derived data retain CC BY-SA 4.0 with attribution to Zygmunt Zając. Open Bandit sample and reference materials retain their upstream notices. No remote book-cover images are used.