Track: research · Level: core · Effort: ~16h, GPU dependent · Phase: 7 · Depends on: FIX-7-1, FIX-3-2, FIX-0-5, FIX-0-6, FIX-0-7
Source: T20 findings Proposal A and Part VI item 3 (#389)
Why this matters
This is the issue that turns None into numbers, and the only one allowed to. Everything else
in the programme refuses to guess precisely so that this study can decide.
Steps
- Fit thresholds on shadow-mode records (FIX-3-2) plus the existing T20 matrices, on the
training seed split only.
- Sweep the loss quantile
q alongside the thresholds.
- Validate the resulting regime assignments against Waterbirds ground-truth masks. That is what
SpuriousBench masks are for. The rule must not consume them at run time, or it stops being
deployable.
- Evaluate on the held-out seed split against the pre-registered endpoint.
- Use the corrected statistics from
benchmarks/stats.py (FIX-0-5), and report faithfulness
only against the coverage baseline (FIX-0-6) and the noise floor (FIX-0-7).
- Ship the result as a versioned threshold profile, loadable per FIX-1-4.
Done when
Either a calibrated, held-out-validated threshold profile exists, or the kill criterion fired
and that is written down.
Risk: GPU availability is the single-point dependency of the whole plan. Shadow mode
mitigates it by making collection a by-product of runs that were happening anyway. If GPU time
does not materialise before CW12, this becomes a post-cohort study with the data already
collected, and that is an acceptable outcome.
How to take this: comment "taking this" and wait to be assigned. Branch fix-7-2-calibrate-thresholds from upstream/main, and put Closes #<this issue number> in your PR. Full workflow: the Cohort Handbook (pinned in Discord).
Plan: plans/FIX-PLAN-T20-2026-08-23.md, section 9.
Track: research · Level: core · Effort: ~16h, GPU dependent · Phase: 7 · Depends on: FIX-7-1, FIX-3-2, FIX-0-5, FIX-0-6, FIX-0-7
Source: T20 findings Proposal A and Part VI item 3 (#389)
Why this matters
This is the issue that turns
Noneinto numbers, and the only one allowed to. Everything elsein the programme refuses to guess precisely so that this study can decide.
Steps
training seed split only.
qalongside the thresholds.SpuriousBench masks are for. The rule must not consume them at run time, or it stops being
deployable.
benchmarks/stats.py(FIX-0-5), and report faithfulnessonly against the coverage baseline (FIX-0-6) and the noise floor (FIX-0-7).
Done when
Either a calibrated, held-out-validated threshold profile exists, or the kill criterion fired
and that is written down.
Risk: GPU availability is the single-point dependency of the whole plan. Shadow mode
mitigates it by making collection a by-product of runs that were happening anyway. If GPU time
does not materialise before CW12, this becomes a post-cohort study with the data already
collected, and that is an acceptable outcome.
How to take this: comment "taking this" and wait to be assigned. Branch
fix-7-2-calibrate-thresholdsfromupstream/main, and putCloses #<this issue number>in your PR. Full workflow: the Cohort Handbook (pinned in Discord).Plan:
plans/FIX-PLAN-T20-2026-08-23.md, section 9.