Chess-World-Model is a benchmark for testing whether sequence models can maintain an exact latent state over action histories. The task is not to play chess. Given a legal sequence of UCI moves, a model must reconstruct the full chess state after every prefix of the game.
The benchmark is built around real games from the public Lichess open database and an out-of-distribution random-uniform split generated by uniformly sampling legal moves. This separates in-distribution generalisation on human games from rule-consistent state tracking under legal but non-human trajectories.
Each input prefix is aligned with the complete chess state reached after that prefix: piece placement, side to move, castling rights, en passant state, and move counters.
|
Exact state tracking Predicts 75 categorical state labels at every timestep, including all 64 board squares and FEN-style auxiliary variables. |
Real and random play Uses held-out human games plus uniformly random legal self-play to expose shortcuts hidden by familiar chess positions. |
Matched model interface Trains Transformer, SLiCE, Mamba-3, and Gated DeltaNet models under a shared move-to-state prediction protocol. |
.
|-- main.py # training entrypoint
|-- training.py # training loop, validation, checkpointing
|-- data/ # PGN processing, JSONL loading, random test generation
|-- evaluation/ # checkpoint evaluation and metrics
|-- models/ # Transformer, SLiCE, Mamba, Gated DeltaNet
|-- experiment_configs/ # named JSON presets
|-- assets/ # paper PDF and dataset figure assets
`-- pyproject.toml # uv project metadata
This project uses Python 3.12 and uv.
uv syncThe base environment supports the Transformer and SLiCE presets. Gated DeltaNet and Mamba use optional dependency groups described below.
For local smoke tests or runs without Weights & Biases logging, pass:
--wandb_mode disabledProcessed chess data is intentionally not included. Download PGN exports directly from the Lichess open database:
https://database.lichess.org/
Standard rated game files follow this pattern:
https://database.lichess.org/standard/lichess_db_standard_rated_YYYY-MM.pgn.zst
The paper uses March 2025 Lichess games as the training/validation source and April 2025 games for the held-out real-game test condition. The commands below use the same months as examples.
mkdir -p data/raw
curl -L \
-o data/raw/lichess_db_standard_rated_2025-03.pgn.zst \
https://database.lichess.org/standard/lichess_db_standard_rated_2025-03.pgn.zst
zstd -d \
-o data/raw/lichess_db_standard_rated_2025-03.pgn \
data/raw/lichess_db_standard_rated_2025-03.pgn.zstLichess files are large. Full monthly standard-rated exports are tens of GB compressed and much larger after decompression. The database page also provides torrent downloads, which are often more reliable for full months.
The processor replays each PGN game with python-chess, packs each UCI move into a categorical move token, and writes sharded JSONL examples containing aligned moves and states.
mkdir -p data/processed/lichess_2025_03
export CHESS_WORLD_MODEL_ID_KEY="$(openssl rand -hex 32)"
uv run python -m data.process \
--pgn data/raw/lichess_db_standard_rated_2025-03.pgn \
--out data/processed/lichess_2025_03 \
--max-games 10000000 \
--min-fullmoves 10 \
--shard-size 1000000This writes up to 10M retained games, matching the benchmark scale used in the paper. Omit or change --max-games for a different corpus size.
CHESS_WORLD_MODEL_ID_KEY is used only to derive stable opaque example IDs from source game URLs. Keep the same value if you want deterministic IDs and train/validation splits across repeated preprocessing runs.
For a quick local smoke test:
mkdir -p data/processed_smoke
uv run python -m data.process \
--pgn data/raw/lichess_db_standard_rated_2025-03.pgn \
--out data/processed_smoke \
--max-games 1000 \
--min-fullmoves 10Create a local held-out real-game test set from another Lichess month:
mkdir -p data/test/lichess_april_2025_10k
curl -L \
-o data/raw/lichess_db_standard_rated_2025-04.pgn.zst \
https://database.lichess.org/standard/lichess_db_standard_rated_2025-04.pgn.zst
zstd -d \
-o data/raw/lichess_db_standard_rated_2025-04.pgn \
data/raw/lichess_db_standard_rated_2025-04.pgn.zst
uv run python -m data.process \
--pgn data/raw/lichess_db_standard_rated_2025-04.pgn \
--out data/test/lichess_april_2025_10k \
--max-games 10000 \
--min-fullmoves 10Generate a random-uniform legal-play test set:
mkdir -p data/test/random_uniform_10k
uv run python -m data.generate_random_uniform_test_set \
--out data/test/random_uniform_10k \
--games 10000 \
--seed 0Processed JSONL files, decompressed PGNs, checkpoints, logs, and evaluation outputs are ignored by git.
Run a small Transformer smoke test:
uv run python -m main \
--data_path data/processed_smoke \
--experiment_config transformer_d128_defaults_simple \
--epochs 1 \
--val_every 0 \
--save_every 0 \
--wandb_mode disabledRun a preset on the processed training corpus:
uv run python -m main \
--data_path data/processed/lichess_2025_03 \
--experiment_config mamba_d384_defaults_simple \
--ckpt_path checkpoints/mamba_d384 \
--wandb_run_name mamba_d384 \
--wandb_mode disabled--experiment_config accepts either a JSON file path or a preset name under experiment_configs/, with or without the .json suffix. CLI flags override values from the selected config.
| Family | Preset names | Extra dependency group |
|---|---|---|
| Transformer | transformer_d{128,256,384,512}_defaults_simple |
base install |
| SLiCE | slice_d{128,256,384,512}_defaults_simple |
base install |
| Gated DeltaNet | gated_deltanet_d{128,256,384,512}_defaults_simple |
fla |
| Mamba-3 | mamba_d{128,256,384,512}_defaults_simple |
mamba |
The size ladder corresponds to approximately 3M, 8M, 18M, and 38M parameters in the paper configurations.
Optional dependency setup
Gated DeltaNet uses flash-linear-attention:
uv sync --group flaMamba runs use the optional Mamba dependency group:
MAMBA_FORCE_BUILD=TRUE \
uv sync --group mamba --refresh-package mamba-ssm --reinstall-package mamba-ssmThe repo pins mamba-ssm to a Git revision because the tested PyPI release did not expose mamba_ssm.modules.mamba3. Verify the install before launching mamba-3 runs:
uv run python -c "from mamba_ssm.modules.mamba3 import Mamba3; print(Mamba3.__name__)"If that import still fails, rebuild without using cached wheels:
MAMBA_FORCE_BUILD=TRUE \
uv pip install --python .venv/bin/python \
--no-build-isolation \
--no-deps \
--force-reinstall \
--no-cache-dir \
"mamba-ssm @ git+https://github.com/state-spaces/mamba@b267be48e9e71a3a37310ade04b058625409da2d"Training supports optional per-step scheduling:
lr_scheduler: none | cosine | linear
lr_warmup_steps: linear warmup steps
lr_decay_steps: scheduler horizon in optimiser steps
lr_min_ratio: final LR ratio relative to base LR
Example:
uv run python -m main \
--data_path data/processed/lichess_2025_03 \
--experiment_config transformer_d256_defaults_simple \
--batch_size 128 \
--lr 3e-4 \
--lr_scheduler cosine \
--lr_warmup_steps 5000 \
--lr_decay_steps 620000 \
--lr_min_ratio 0.1 \
--wandb_mode disabledEvaluate one or more checkpoints against local JSONL test sets:
uv run python -m evaluation.evaluate_on_test_sets \
--checkpoint checkpoints/mamba_d384/latest.pt \
--dataset lichess_april_2025_10k=data/test/lichess_april_2025_10k \
--dataset random_uniform_10k=data/test/random_uniform_10kThe evaluator accepts checkpoint files or checkpoint directories containing latest.pt. Results are written under test_eval/ by default, with one JSON file per checkpoint and a combined summary.tsv.
Useful evaluation flags:
--cpu force CPU evaluation
--output_root choose the output directory
--max_batches cap batches per dataset for a quick check
--batch_size override the checkpoint batch size
The accompanying preprint is included as assets/Chess_World_Model.pdf.
@misc{walker2026chessworldmodel,
title = {Chess-World-Model: A 10M-Game Benchmark for Exact State Tracking from Chess Move Sequences},
author = {Walker, Benjamin and Lyons, Terry},
year = {2026}
}This code is released under the MIT licence. Lichess database exports are distributed separately by Lichess under CC0.
