A battle AI for Pokémon Showdown gen 9 random battles — singles and doubles — learned from public ladder replays. No hand-coded strategy: it learns a win predictor V(s) and a policy prior π(a|s) from replays, then plays by one-ply expectimax with a damage-formula transition model.
Train one on your own machine and battle it.
Every turn, for every legal action it can take:
EV(action) = V(state after my action + opponent's most likely response)
- V(s) — LightGBM win predictor (AUC 0.716) over a perspective-symmetric feature vector: species, stats, HP, boosts, status, item/ability semantic tags, hazards, weather, terrain, tera — plus a momentum block (speed control, both-side kill-clocks, priority presence, switch counters). Trained on ~200k decision points from public ladder replays.
- Opponent model — the opponent's response is scored with a gen-9 damage formula (stats, boosts, STAB, type matrix, weather/terrain, accuracy, effective speed) over the moves their species could be running, weighted by the π prior — so protect-heavy or switch-happy foes discount naturally.
- π(a|s) — two-stage policy clone: a move-vs-switch classifier (89% acc) plus move/switch rankers (41% top-1 / 78% top-3 over 310 moves) learned from what humans actually did in similar states.
- Terastallization — a first-class candidate whenever it's available, with tera typing handled in both offense and defense.
- Tempo heuristics — the search layer prices what a 1-ply value function can't: entry hazards on switch-in evals, a last-mon urgency penalty (no more setup-and-die), a repetition tax that breaks recovery-stall loops (roost ×10), and a switch-commit threshold so pivots must clear a real EV margin instead of bleeding tempo.
| judge | MAE | corr. with outcome | notes |
|---|---|---|---|
| V(s) | 0.26 | 0.92 | on 40 held-out mid-game states |
| Jev (TypeSafe System One) | 0.48 | 0.23 | semantic judge, same states |
The Jev comparison (scripts/jev_benchmark.py) was the fun experiment: a
calibrated decision model with full semantic knowledge of Pokémon loses to
the replay-learned V(s) at battle-state judgment — mostly because it ignores
fainted actives and compresses toward 0.5. Its edge shows only where V lacks
move-level knowledge. Scaffold included, API costs ~1¢ per 40 states.
Playing strength: ~50% win rate vs poke-env's SimpleHeuristics baseline in random singles (10-game batches, zero decision errors). Doubles pipeline is complete but younger — same architecture, less mature.
You need a Pokémon Showdown server. The bot was developed against a locally
patched server (pokemon-showdown fork with offline account registration), but
any server the bot can log into works.
git clone <this repo> && cd showdown-brain
uv sync
# 1. start a showdown server (see https://github.com/smogon/pokemon-showdown)
# and register a bot account on it
# 2. point the bot at the server + give it credentials
export SHOWDOWN_BOT_USER=mybot SHOWDOWN_BOT_PASS=hunter2
# (or put {"username": ..., "password": ...} in data/bot-account.json — never committed)
# 3. run it — lobby-visible, accepts singles AND doubles, forever
uv run python scripts/serve.pyThen open your server in a browser, find the bot's username, and challenge it to [Gen 9] Random Battle (or Random Doubles, once you flip the format).
Out of the box it loads the pretrained models/v_model.txt + models/pi_model.txt
shipped in this repo.
uv sync
uv run python -m sb.harvest --limit 4000 --rating-min 1400 # ladder replays -> data/raw/
uv run python -m sb.parser # -> data/rows_v2.jsonl (full-fidelity states)
uv run python -m sb.parser_doubles # -> data/rows_v2_doubles.jsonl
uv run python scripts/train_v_v2.py # V(s) over the full state -> models/v2_model.txt
uv run python scripts/train_pi2.py # 2-stage policy prior -> data/pi2_*
uv run python scripts/gauntlet.py --games 60 # plays its past selves + LEARNS from itIt learns from every game it plays: finished battles are dumped to
data/learned_games/, folded into rows_personal.jsonl, and
models/v2_model.txt rebuilds on public+personal rows between games
(sb/learn.py). See ARCHITECTURE.md for the whole loop.
Splits are by game id, never by row (rows from one game correlate). LightGBM is capped at 4 threads by default because uncapped training schedulers will eat your machine.
Watch it learn by playing itself:
uv run python scripts/selfplay.py # bot vs poke-env SimpleHeuristicsPlayersrc/sb/harvest.py— replay archive search + log downloadersrc/sb/parser.py— showdown protocol → battle state →(state, action, outcome)rows (singles)src/sb/parser_doubles.py— same for random doublessrc/sb/features.py— state dict → numeric vector;assets/game data resolutionsrc/sb/transitions.py— gen-9 damage formula + non-damage move effects, in dict-spacesrc/sb/showdown_state.py— poke-env live battle → parser-shaped state dictsrc/sb/agent.py— VsAgent: one-ply expectimax playersrc/sb/pi2.py— 2-stage policy prior inferencescripts/— train_v, train_pi, train_pi2, selfplay, humanassets/— game data (species, moves, type matrix, semantic tags, randbats sets)models/— pretrained V / π (git-tracked, ~2 MB total)data/— harvested replays + rows (gitignored, can be GBs)
Game data in assets/ comes from the community-maintained
pokemon-data pipeline
(Smogon sets + TypeSafe semantic tags). Override the location with
POKEMON_DATA=/path/to/dir.
- V(s) AUC ~0.64 — mid-game EVs live in a narrow band, so decisions cluster. More (and higher-rated) replays is the main lever.
- One-ply search: no move-sequencing awareness (setup → attack patterns are emergent from V, not planned).
- Opponent model assumes a π-weighted damaging response; status/protect play is under-modeled.
- Doubles support is young — protect/redirect positioning is not modeled.
MIT (see LICENSE). Replay data is downloaded from the public Showdown replay archive — respect their rate limits (the harvester sleeps between requests).