TOD is a decision API from Dasein Labs / parseclab (docs). A request is one image, some text, and one or more questions, each with labelled choices; the answer is a probability distribution over each question's choices.
This repo connects TOD to Papers, Please (Story mode, Days 1-3) with a loop where every in-game decision is TOD's: which element to click or drag, where to drop it, what each paper is, what the passport says, and whether the entrant is approved or denied. The code around TOD only sees the screen, turns it into labelled options, asks, and executes the answer.
Claim demonstrated: a general decision model, given only what is on screen plus the written
rules, can play a rules-heavy document-inspection game through Day 3. In the reference run
20261004_115735 it reached Day 3's end: 13 entrants, 10 handled correctly, 3
citations, and the no-documents entrant interrogated through the game's inspect-mode sequence.
Contents: Non-negotiables · Architecture · Extraction · Set-of-Mark and option text · Question design · Input conventions · Guardrails · Ground truth · Results · Limitations · How to verify TOD made the decision · Teach TOD your own game · Setup and running · Dev log
These rules define the experiment. A change that breaks one of them is out of scope.
- TOD decides. Every click, drag, drop target, paper identity, reading and verdict is the top answer to a TOD question. No other model picks inputs.
- Options only come from the screen. Every choice is a box found by extraction on the current frame (or a fixed screen region that is visibly there, such as the empty counter).
- One image per request. Each TOD request carries exactly one frame.
- No scripted sequences. There is no hard-coded order of clicks. The rules are written down as text; TOD reads them every tick.
- No fixed coordinates as actions. Coordinates are only used to draw options (static layout fixtures, derived drop zones). Which one is used is TOD's pick.
- Nothing leaked into option text. An option says what the element is and where it is (OCR text, label, TOD's own earlier identity answer). It never says "this is the right one".
- Ground truth is never shown to TOD. The game-memory reader (
gt.py) is used for scoring and for the harness's stop condition only.
flowchart LR
G["Game window"] -->|grab frame| F["Raw frame"]
F --> X["Extraction<br/>YOLO + Grounding DINO + OCR<br/>+ static layout + sheet segmentation"]
F --> R1["TOD request 1<br/>state questions<br/>raw frame"]
X --> R1b["TOD request 1b<br/>paper identities, readings, verdict<br/>raw frame"]
X --> S["Set-of-Mark frame<br/>numbered boxes"]
R1 --> B["Situation block<br/>+ dynamic block"]
R1b --> B
S --> R2["TOD request 2<br/>action, source, target<br/>SoM frame"]
B --> R2
R2 --> C["Input conventions<br/>+ guardrails"]
C -->|click, drag or wait| G
C --> L["runs/ts/tick_NNNN.json"]
Request 1 runs in parallel with extraction. Request 1b needs the extracted boxes (it asks about each paper and about OCR'd candidate strings). Request 2 needs both. Each request carries one image. Every answer, probability, option text and executed input is logged per tick.
sequenceDiagram
participant L as loop.py
participant T as TOD
L->>T: request 1, raw frame: screen, person at window, paper on counter, passport open, tray open, passport under each stamp, inspect mode, no documents
T-->>L: probabilities per question
L->>T: request 1b, raw frame: identity of each paper, which OCR date is EXP., which token is the ISS. city, which VALID ON date, verdict
T-->>L: probabilities per question
Note over L: situation block and WHAT APPLIES NOW, built only from TOD answers
L->>T: request 2, SoM frame plus rules, situation and last 15 actions: action, source element, target element
T-->>L: e.g. click element 3 with p 0.96
flowchart TD
A["TOD readings<br/>country, EXP. date, ISS. city, entry ticket"] --> V
M["Rule text for today<br/>from the manual"] --> V
V{"TOD verdict question<br/>approved / denied / cannot_decide_yet"}
V -->|cannot_decide_yet| N["No stamp offered<br/>dynamic block names the missing reading"]
V -->|approved or denied| O["Only the stamp named by the verdict is drawn<br/>the other is struck: hidden_by tod_verdict"]
O --> P{"TOD picks the stamp?"}
P -->|no| Q["Other action executed"]
P -->|yes| GATE{"Press gate<br/>TOD's own strip answer: passport under that stamp?"}
GATE -->|yes| X["Click executed"]
GATE -->|no| RF["Refused, no input<br/>refusal shown next tick"]
The rule is applied by TOD, not by code: the verdict question states today's rule and lists TOD's own earlier readings; code only displays them and gates the press on TOD's own strip answer.
src/tod_papers/extract.py (+ layout.py, extract_server.py for a GPU box). Output: a list
of boxes, each with a kind, a label or OCR text, and a position; describe(box) turns a box into
option text. Details and comparison tables: docs/extraction.md.
| piece | what it produces | why it exists |
|---|---|---|
YOLO icon detector (OmniParser v2 icon_detect) |
class-agnostic UI element boxes | finds buttons, stamps, the horn, the tray tab without a game-specific model |
| Grounding DINO (open vocabulary) + CLIP labels | a label per untexted box ("yellow lever", "rubber stamp", "speaker/horn", "passport booklet", "folded page corner", ...) | a box without text or label gave TOD nothing to match; with labels the idle-booth pick went from ~0.2 (bulletin) to 0.71-0.75 (horn) |
| OCR (PP-OCRv5 via RapidOCR, line crops upscaled 3x nearest-neighbour) | text lines with boxes | the game's pixel font; OCR text goes into option text and into reading candidates |
Static layout finder (layout.py) |
fixtures that are always at the same native-pixel place: stamp tray tab, stamp bar, counter, inspect button, empty counter, microphone, day tiles | YOLO misses some of these on some frames; they are drawn as options only when the screen state shows them, and TOD chooses them like any other box |
| Derived drop zones | stamp landing strips (from the detected stamp heads), the entrant (largest person box), the put-away spot, the rulebook slot | drop targets must be the places that matter in the game; a single "empty space" target failed |
| Sheet segmentation | one box per paper when papers overlap (passport vs entry ticket vs flyer), each with its own OCR text | an identity question per paper needs one paper per box |
Local extraction on an RTX 4070 takes 0.5-0.8 s on a quiet machine and 2-5 s with the game running; the same server on a cloud L4 is about 1.1 s per tick round trip (docs/remote_extraction.md).
som.py draws a numbered marker on every box offered this tick; boxes ruled out (no effect
twice, or hidden by a guardrail) are struck through. The numbers are the choice labels of the
request-2 source and target questions, and each number's criteria text is the box's
description. Real option texts from run 20261004_115735, tick 22:
1 object - shutter lever at the top-right corner of the booth window -- drag it down to open
or close the window shutter (middle-left) - drag
3 object - green APPROVED stamp (knob and body) on the open tray -- click it to stamp the
passport lying in the strip beneath it (middle-right) - click (press to stamp)
7 drop target - stamp landing strip (under the APPROVED stamp head)
Paper boxes carry TOD's own identity answer from request 1b, e.g. passport 0.86 -- <position>.
The option says what the element is; it does not say whether to use it.
All questions are in manual.py (text) and loop.py (when they are asked).
State facts (request 1, raw frame). screen (menu, bulletin, booth_idle,
documents_on_desk, stamp_tray_open, inspect_mode, day_end, ...), person_at_window,
document_on_counter_shelf, document_open_on_desk, stamp_tray_open,
passport_under_denied / passport_under_approved, inspect_mode_on,
no_documents_presented, photo_matches_person, rulebook_page. Yes/no questions are TOD
"noul" questions (probability of yes).
Per-paper identity (request 1b). One question per paper box, with that box's OCR text in
the question: passport / rulebook / bulletin / entry_ticket / transcript / flyer / citation /
other. Tick 22: doc0 = passport (p 0.895).
Readings as choices over OCR'd candidates (request 1b). TOD is not asked "is it expired?"; it is asked which of the strings OCR found is the one that matters:
exp_date: "OCR found these dates on the documents on the desk. On the open passport data
page, which one is printed after 'EXP.'?"
D1 = EXP. 1925.10.19 D2 = EXP. 1984.02.17 none
-> D2 (p 0.955)
ticket_date: "OCR read these 'VALID ON' lines on the desk. On the ENTRY TICKET ..., which one
is printed after 'VALID ON'?" D1 = VALID ON 1982.11.25 none -> D1 (p 0.994)
issuing_country: ARSTOTZKA / KOLECHIA / IMPOR / ANTEGRIA / OBRISTAN / REPUBLIA /
UNITED FEDERATION / unreadable -> OBRISTAN (p 0.835)
This moved offline expiry accuracy from 3/30 to 28/30 and city from 26/30 to 29/30.
Verdict (request 1b). A choice approved / denied / cannot_decide_yet. The question text is
today's rule plus TOD's own earlier readings of this entrant:
TODAY'S RULE: Day 3 (1982.11.25): ... APPROVED if the passport is valid: it is not expired (its
EXP. date is after 1982.11.25) AND its ISS. city is in the rulebook list for the passport's
country ... a foreigner also needs an ENTRY TICKET VALID ON 1982.11.25 ...
YOUR OWN EARLIER READINGS of this entrant's papers:
- issuing country: OBRISTAN (your answer at tick 20, p=0.83)
- EXP. date: 1984.02.17 (your answer at tick 20, p=0.95)
- ISS. city: 'Lorndaz' (your answer at tick 20, p=0.94)
- entry ticket: dated_today (your answer at tick 20, p=0.99)
-> approved (p 0.768)
Action (request 2, SoM frame). action (click / drag / wait), source (numbered elements +
wait), target (numbered drop targets). The text is the ~2.4k-character rules manual
(BOOTH_MANUAL), the situation block (TOD's answers, each labelled "your reading"), a short
"WHAT APPLIES NOW" dynamic block derived from those answers, and the last 15 actions with their
observed effect. Tick 22's dynamic block:
WHAT APPLIES NOW (from your own answers above):
- Your verdict (this tick): APPROVED p=0.77. The passport lies under APPROVED: press the
APPROVED stamp once.
TOD's answer: action = click (0.89), source = 3 (0.96) -> the APPROVED stamp was pressed.
Some elements are press-only (stamps, buttons, the horn, page corners) and some are drag-only
(papers, the tray tab, the lever). This is the game's input model, not a decision: if TOD picks
drag on a stamp, the press is executed as a click; if it picks click on a paper, the drag
goes to TOD's chosen target. Each coercion is logged in input_convention /
convention_mismatch. The element and the target are always TOD's.
Guardrails remove options or refuse an input. They never choose one. Each keys on TOD's own answers:
| guardrail | keyed on |
|---|---|
only the verdict's stamp drawn (stamp_hidden) |
TOD's verdict answer |
| press gate (refused press) | TOD's passport_under_<side> answer |
| horn hidden while someone is at the window | TOD's person_at_window |
| stowed flyers / citations struck | TOD's identity answer for that paper |
| stuck detection: an element with no visible effect twice is struck for 6 ticks | measured pixel change after the input |
| repeat-drag ban, tray open/close limit | the action history |
| delete / trash veto on menu screens (only remaining regex) | OCR text of the option |
stop flags --stall-stop, --pick-stop, --refuse-stop, --stop-on-screen |
repeated state / picks / refusals |
src/tod_papers/gt.py reads the game's memory (Unity IL2CPP; offsets in
docs/ground_truth.md): screen, day, clock, entrant name, the correct
verdict, the verdict given, error classes, processed count, citations, savings. It is written to
each tick JSON under gt and used by tools/report.py for scoring and by the harness to stop at
a given day's night screen. It never appears in any TOD request text.
Scored by ground truth. "given" counts entrants that received a stamp (or, for the no-documents entrant, were interrogated). Days 1-2 rows are from runs that continued into later days.
| day | run | entrants reached | given | correct | notes |
|---|---|---|---|---|---|
| 1 | 20261003_092642 |
8 | 7 | 7 | --pause-think |
| 1 | 20261003_182519 |
4 | 3 | 3 | no pause |
| 2 | 20261003_182519 |
7 | 6 | 6 | no pause |
| 2 | 20261003_092642 |
7 | 5 | 4 | photo look-alike approved |
| 3 | 20261004_115735 |
13 | 12 + interrogation | 10 | reached Day 3 night; 3 citations |
Reference run 20261004_115735: 254 ticks, 684 TOD calls, $0.31 in TOD calls, median
extract 1.09 s, request 2 1.47 s, --pause-think. The three citations: an OCR misread city
(Passport/IssuingCity), a look-alike photo approved (Passport/Face), an OCR misread expiry
date. The no-documents entrant: rulebook to desk (t132) -> Basic Rules (t133) -> inspect button
(t134) -> rule line "Entrant must have a passport" (t135) -> empty counter (t136) -> microphone
(t138) -> "Where is your passport?" (t139); he left without a stamp or a citation.
One unbroken take from cold start through Day 3 night, no --pause-think, with the fast scorer box
(deploy/gcp_fast.sh, --fast) and A100 extraction (docs/remote_extraction.md):
| days | 1-3 in one loop run |
| entrants seen / processed | 22 / 19 |
| correct | 17 / 19 (2 citations, both Day 2 approvals that should have been denials) |
| no-documents entrant | Jorji interrogated on Day 3 |
| ticks / TOD requests / decisions | 274 / 741 / 4,294 |
| cost | $1.29 at $0.30 per 1k decisions (docs/tally.md) |
| wall-clock | 16.5 min; median tick 3.7 s |
The demo video was recorded in one take with the OBS pipeline in
docs/recording.md (tools/obs_setup.py, tools/obs_director.py,
tools/produce_demo.py; live tally overlay from docs/tally.md, viewer layout in
docs/viewer.md). The older capture-and-zoom pipeline is kept in
docs/recording_cap_legacy.md. The MP4 itself is not in this repo
(recordings and runs/ are gitignored).
python tools/report.py runs/<run> prints this table for any run.
- Photo check. Denying on the photo needs a confident
different; look-alikes get approved. - OCR misreads. The pixel font occasionally yields
Eist Grestinor1943.11.25; TOD then reads the wrong string correctly. - Clock vs. tick time. A tick takes 6-10 s (three TOD calls plus extraction); the game day is ~4 real minutes. Without help, entrants after 18:00 are unpaid and Day 2 can end in debt.
--pause-think. The harness can suspend the game process while TOD thinks (no input, no menu; the frame TOD sees is from before the suspend). The reference Day 3 run used it. It is a timing aid, off by default, and visible as stutter in recordings.- Step N (no documents) re-runs for ~8 ticks after the interrogation starts.
- Days 4+ are not attempted.
Every tick writes runs/<ts>/tick_NNNN.json, plus raw_NNNN.png (the unmarked frame of
requests 1/1b) with --save-raw and, in current code, tick_NNNN.png (the numbered SoM frame of
request 2). Read, for any tick:
| field | what it shows |
|---|---|
state_text |
the full request-2 text TOD received (rules, situation, dynamic block, history) |
descriptions |
the option text for every numbered element |
answers |
TOD's raw probabilities for action, source, target, screen |
tod_pick |
the argmax that was used; compare with answers |
state |
every request-1 / 1b answer with value, p, and the full probs (incl. verdict, doc0.., exp_date, ticket_date) |
stamp_hidden, excluded |
which options were struck and why (keyed on TOD answers or no-effect) |
input_convention, convention_mismatch |
any click/drag coercion |
executed, changed |
the input actually sent and whether the screen changed |
tod_request_id, tod_ms |
the TOD request id and latency for request 2 |
gt |
ground truth for scoring; compare it with state_text to confirm it never appears there |
Check: tod_pick equals the argmax of answers.source / answers.target; executed names the
same element; no gt value (entrant name, correct) occurs in state_text. tools/viewer.py
shows the same fields side by side with the frame, live or for a past run
(docs/viewer.md).
The loop is game-specific only in four files. To adapt it:
- Capture and input (
io_win.py). Find the window, grab the client area, send clicks and interpolated drags. Reuse as is for any Windows game. - Extraction (
extract.py). Run it on a few saved frames and look at the SoM image. Every element a player would use must have a box and a description. Add Grounding DINO prompts for untexted objects in your game; add OCR upscaling if the font is pixel art. - Fixtures (
layout.py). For elements that never move but are hard to detect, add a static box in native coordinates, drawn only when the screen state (a TOD answer) says it is visible. Add derived drop zones for the places where drops matter. - Rules (
manual.py,BOOTH_MANUAL). Write the rules a new player would need, as rules, not as a sequence of steps. Keep it short (ours is ~2.4k characters); long scripts made answers worse and eventually hit the request size limit. - State questions (
manual.py, request 1). One question per fact that changes what to do: screen kind, who is present, where the key objects are. Ask them on the raw frame, separately from the action question. - Readings (request 1b). For anything to be read off a document, give TOD the OCR'd
candidate strings as choices plus
none, rather than a yes/no about the conclusion. - Decision questions. If the game has a judgement (approve/deny), make it a TOD question whose text contains the rule and TOD's own earlier readings; offer only the inputs that carry out that judgement.
- Dynamic block. Per tick, derive one or two sentences from TOD's answers ("your verdict is X; the passport is under X") and put them right before the history.
- Evaluation. Get ground truth from somewhere other than the screen (memory, logs, a save
file) and keep it out of every request. Score with it; dry-run fixes offline on saved frames
(
loop.py --frames ...) before going live.
Windows 11, Python 3.13, an NVIDIA GPU (or the remote extractor), Papers, Please on Steam
(AppID 239030) in a 2280x1280 window. A TOD API key in .env as TOD_API_KEY=... (gitignored).
:: extraction stack (GPU)
py -3.13 -m venv .venv-extract
.venv-extract\Scripts\python -m pip install torch torchvision --index-url https://download.pytorch.org/whl/cu126
.venv-extract\Scripts\python -m pip install -r requirements-extract.txt
.venv-extract\Scripts\python deploy\fetch_models.py
:: loop venv, layered on .venv-extract (see the header of requirements-loop.txt)
py -3.13 -m venv .venv-loop
.venv-loop\Scripts\python -m pip install --no-deps -r requirements-loop.txt
:: ground-truth reader (evaluation only)
py -3.13 -m venv .venv-gt
.venv-gt\Scripts\pip install -r requirements-gt.txtExtraction runs in-process by default. To use a GPU server instead:
python -m tod_papers.extract_server --host 0.0.0.0 --port 8765 :: on the GPU box (or deploy/)
set TOD_EXTRACT_URL=http://127.0.0.1:8765 :: in the loop's shelldeploy/ has a Dockerfile, a Modal app and a GCP L4 spot script
(docs/remote_extraction.md).
Run:
.venv-loop\Scripts\python tools\demo_run.py :: fresh story, Day 1 -> Day 3 night
.venv-loop\Scripts\python tools\demo_run.py --from-day 3 :: keep saves, TOD picks the day tile
.venv-loop\Scripts\python -m tod_papers.loop --max-ticks 400 --save-raw --stall-stop 12
.venv-loop\Scripts\python -m tod_papers.loop --frames runs\<ts>\raw_0022.png --day 3 :: offline, no input
.venv-loop\Scripts\python tools\report.py runs\<ts> :: per-entrant table, timings
.venv-loop\Scripts\python tools\viewer.py --run runs\<ts> :: frame + TOD answers, live or replayMore: docs/demo.md (demo run), docs/viewer.md (viewer), docs/loop.md (tick, flags, decision path), docs/game.md (game rules and UI), docs/io.md, docs/extraction.md, docs/ground_truth.md, docs/devlog.md.
Apache-2.0, see LICENSE. Papers, Please is (c) 3909 LLC and is not included; you
need your own copy. The repo ships only small UI crops used for template matching
(layout_assets.npz, rulebook_corner.png). Contributions: CONTRIBUTING.md.