Skip to content

Latest commit

 

History

26 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

stacks

A bookstore recommender that makes its evaluation inspectable. Browse typographic book covers, compare training-derived recommendation signals, and read the evidence behind model and off-policy estimates.

Open the bookstore · Source · Measured results · Evaluation report

These are demonstration recommendations on a public dataset; no real reader's identity is present and no recommendation is personalized to a real person.

What was measured

The pinned public CSV -> parquet -> NumPy/SciPy exact scoring backend evaluated 10,000 books and 5,000 readers using Global source-order holdout, full catalog. The source has no rating timestamps, so row order is a temporal proxy. These results apply to the recorded eligible reader sample and its evaluated catalog; they are not a date-verified backtest.

Model NDCG at ten with interval Recall at two hundred with interval
Popularity 0.0406 [0.0373, 0.0441] 0.1298 [0.1249, 0.1346]
Item cosine 0.0541 [0.0507, 0.0579] 0.2585 [0.2517, 0.2654]
Implicit ALS 0.0546 [0.0513, 0.0579] 0.2607 [0.2543, 0.2672]
Fixed blend 0.0538 [0.0504, 0.0576] 0.2707 [0.2641, 0.2774]
Content TF-IDF 0.0411 [0.0381, 0.0442] 0.1399 [0.1348, 0.1451]
LambdaMART 0.0440 [0.0415, 0.0466] 0.2154 [0.2097, 0.2214]
Two-tower neural 0.0019 [0.0014, 0.0024] 0.0162 [0.0146, 0.0181]
Session cosine 0.0900 [0.0855, 0.0947] 0.2900 [0.2830, 0.2967]
Elman recurrent 0.0232 [0.0212, 0.0252] 0.1041 [0.0995, 0.1086]

Observed leader: session_cosine. Corrected selection: session_cosine. The full report includes paired comparisons, shortcut diagnostics and limitations.

Run locally

uv sync --frozen
uv run pytest
uv run python scripts/render_reports.py --check
cd web
npm ci
npm run dev

For durable recommendations and interaction logging, start the API in a separate terminal from the repository root:

uv run uvicorn packages.api.main:app --host 127.0.0.1 --port 8000 --no-access-log

The default database is local SQLite. Configure DATABASE_URL, a persistent SESSION_HASH_KEY, ADMIN_TOKEN, and CORS_ORIGINS for the intended environment. Administrative writes are unavailable until their token is set. See serving and logging for endpoint and security behavior. GitHub Pages hosts the frontend. The live API uses a Worker and durable D1 storage on the selected free Sites host. Public request checks and a separate D1 read verified persisted interactions. It serves bounded SSE responses with EventSource reconnection and Last-Event-ID replay because the host buffers long streaming responses. See deployment verification. The Worker/D1 deployment replaces a paid managed-service dependency. PostgreSQL remains optional for the Python reference and local ANN benchmark. Browser-only state does not count as a server exposure log.

Implemented scope

Module Recorded status Evidence or limit
Popularity / cosine / implicit ALS / blend measured Train-only scorers with exact full-catalog offline evaluation and persisted factors.
LambdaMART learned ranker measured Learned over a retrieval candidate union using an earlier source-order window. Evaluated against all baselines over the complete unseen catalog.
pgvector / approximate nearest neighbors measured_local_postgres Exact SQL and HNSW measured on the frozen two-tower embeddings in local PostgreSQL, with the same training-seen filters. This is separate from the hosted D1 serving path.
Content TF-IDF path measured Normalized title, author and tag TF-IDF profiles score all books including training-cold items. Uses undated metadata; no historical metadata availability claim.
OPE simulator measured Local IPS, SNIPS, DM, DR with 200 seeds, two sample sizes and oracle/misspecified reward models.
Real logged-bandit OPE measured_six_file_bts_benchmark All six random/BTS samples. Official campaign beta priors define the BTS target through beta draws ranked into three positions. Random logs after a shared cutoff provide OPE; BTS logs after the same cutoff provide an empirical on-policy benchmark with Wilson intervals. Logged BTS propensities vary within item/slot, so the frozen-prior approximation is not proven identical to the deployed per-impression policy. Differences include target mismatch and sampling error, not pure estimator error. Row bootstrap is conditional on the fitted reward model and Monte Carlo target, ignores repeated-reader dependence, and is not joint-slate inference.
Verified temporal features unavailable Source ratings lack timestamps. Source-order train-only features are tested; metadata is a snapshot.
Two-tower neural retrieval measured Independent user-ID and item-ID embeddings32 -> learned projection16 -> ReLU -> L2 normalization
Recurrent session model measured 16-dimensional Elman tanh recurrent encoder with learned input/output item embeddings

The measured models and additional experiments are listed above. The earlier bounded release is retained under results/history/; its scores must not be compared directly with the expanded population. A score blend and a learned ranker are identified separately. See scope decisions for differences from the originating brief and the recorded deployment choices.

Skills demonstrated

Skill Inspectable evidence
Ranking evaluation and comparison Protocol, results, paired tests and shortcut diagnostics
Learned and baseline models Model cards, recorded training artifacts and candidate-feature contract
Counterfactual policy evaluation Six-file OBD benchmark, simulator truth and pinned source crosscheck
Serving and durable interactions Serving contract, independent-connection and subprocess tests
Reproducibility and honest reporting Strict document renderer, executed notebooks, population labels
Release verification CI workflow, lint, strict types, branch coverage, browser tests and semantic palette checks

Reproduce and change it

uv run python scripts/crosscheck_obp.py verifies the local OPE point estimates against pinned upstream arithmetic methods. This is an explicitly scoped source reference check; it does not certify the full OBP package on this Python version.

uv run python scripts/build_evidence.py downloads the declared sources and rebuilds the core manifest and browser artifacts. Then run uv run python scripts/build_neural_evidence.py, uv run python scripts/render_reports.py, uv run python scripts/build_notebooks.py and the checks above. This is a model and evidence rebuild, not a live database reset. Review the evaluation protocol, runbook, architecture, and contribution guide.

Original code is MIT licensed. Goodbooks-10k and derived data retain CC BY-SA 4.0 with attribution to Zygmunt Zając. Open Bandit sample and reference materials retain their upstream notices. No remote book-cover images are used.

About

A bookstore recommender that measures its own evaluation: transparent recommendations, reproducible metrics, and off-policy estimates.

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages