Every claim reported to the project owner that was later refuted, with what refuted it. Kept because seven have now occurred, and because the pattern they share is reporting a number faster than verifying it.
None of these reached a paper, a preprint, or the run machine's confirmatory record. All were caught internally. That is the system working, but the frequency is the problem, not the catching.
The structural fix is scripts/status.py. Reported numbers are generated from
artifacts, not typed. Three of these involved a number that was transcribed or transferred
by hand.
R7 is the one to read. It is the only entry where the refuted claim had already been built on — an entire preregistered study existed to explain an artifact — and the only one where the error survived an explicit defence (A4.8) before a direct measurement killed it. Its lesson is narrower than "verify before reporting": when a result is surprising, the next move is one more measurement, not one more argument.
- Reported: that on pairs the instrument places at
|Δθ| < 0.16the model answers "which do you prefer?" at p = 0.977, that stated confidence tracks measured preference at only ρ = 0.093, and — A4.8, defending it — that position bias accounts for 5% of the readout's variance, so saturation was not an artifact of the elicitation's geometry. - Refuted by: decomposing the readout instead of arguing about it. Every pair is asked
in both option orders, so
logit p = d + β·sseparates algebraically — no model, no estimate. On gemma the position term is 81% of the readout, not 5%. Once removed, the preference component correlates with the instrument at ρ = 0.500, not 0.093. Preference is graded; the readout was measuring slot position. - Cause: two compounding errors in A4.8, both committed while defending A4.7. It reasoned
about the DV's position term using Pass A's β — a different elicitation answering a
different question — and inferred
var(d)frommedian|readout|, which assumes the conclusion it was testing. The framing mismatch was A3.12's own finding, and A4.8 committed it while arguing that a framing-related result was robust. - Scope of the damage. A4.7 was not a stray number: it generated
preregistration_probe.mdin full, and that study — activation collection on three models, probes at every layer — was designed to explain something that was not happening. It is cancelled by A4.9's gate, before any activation was collected. The cost was a preregistration and a day. - What caught it. Not a sharper argument — one more check on data already in hand. Every confident piece of reasoning offered in defence of A4.7 was wrong; the decomposition took minutes and needed no new run.
- Corrected once more by A4.10. A4.9 stated the 81% without qualification. Across all three models the position share is 3.1× / 1.9× / 1.2× (gemma / llama / qwen-1.5b), so 81% is gemma's, not a property of instruction-tuned models. The general claim is the range.
R6 — hidden-state cache sized at 5.9 GB
- Reported: as the recommended cache scope ("post only"), computed at 10,000 post passes.
- Refuted by: arithmetic. The same message added two conditions (
self-recounted,chose-provisional), taking post from 10,000 to 14,000 — so "post only" was the full 8.23 GB, not 5.9. With the structure control now added it is 16,000 post passes and 9.4 GB. - Cause: a downstream number not recomputed after an upstream change, in the same message that made the change.
- Third instance of this specific failure, after the R2 magnitude transfer and the
z-statistic that divided by one variance component. Distinct from the estimator errors
in R3/R4 — this class is arithmetic hygiene, and the structural fix is that any figure
derived from a pass count must be regenerated by
scripts/status.pyrather than carried in prose.
- Reported: in conversation, as an argument that a position-capture mixture was small enough not to need fitting.
- Refuted by: the bound was derived against a null of 0.5. That is the same wrong null that produced R3, so the bound is void rather than merely imprecise.
- Does not need redoing. Independently of the bound, lambda is ruled out on a sign argument: dC/dlambda is negative at every gap, and its magnitude is largest at wide gaps (-0.95 at gap 4.5 versus -0.25 at gap 0.3), where the residual is already positive. A grid search over (lambda, beta) against the observed curve puts the optimum at lambda = 0.000 exactly. A mixture weight cannot produce a residual that changes sign with gap.
- Consequences: none downstream. Recorded because the log's value is being complete, and because this is the third instance of the same error class.
posterior SD understated ~1.8x)"
- Reported: in conversation, and committed in
1d600aewithsrc/analysis/template_dependence.py. - Refuted by: a proper over-dispersion test. Within-cell success counts against the binomial expectation implied by the fitted probabilities give dispersion 1.070, 95% CI [0.907, 1.258] — consistent with independence.
- Cause: the ICC treated templates as raters over cells, which conflates genuine
between-cell variation in
p— already captured byθ − α— with excess dependence. Cells differ inpby design, so that ICC is large under perfect independence. - Consequences: no design effect exists. No cell-level random effect is added. §4.4's disjoint-split independence is intact. The separation ratio is not reduced by 4.39 → 2.5; under the corrected model it rises to 4.77. Methods-paper candidate #4 is withdrawn.
- This one was ours, not the reviewer's, and the review built a priority on it.
- Recorded in preregistration.md A2.3 W2.
- Reported: in conversation as a paradigm-level finding, and used as the basis of
a plan to abandon Stage 0 for a methods paper. Committed in
c86968a,48b1fe1. - Refuted by: fitting the order term the model was missing.
β = +1.367on Gemma, which makes the correct null for order-reversal consistency atx = 0equal to2s(1−s) = 0.324, not 0.5. Observed consistency tracks a content-plus-position model to within ±0.05. - Cause: the estimator could not distinguish "no content signal" from "signal plus unmodelled additive bias." It was the second. Compounding it, the pipeline stratified on gaps computed from the uncompressed θ, so the x-axis was wrong too.
- Consequences: Stage 0's hypothesis is live again. The strategic plan built on this claim is void. The diagnostic is redefined as discriminability (A2.6).
- Recorded in preregistration.md A2.3 W1.
- Reported: in the first draft of
docs/pass_c_assumption_audit.md, claiming it would invalidate most of thechosecondition. - Refuted by: measuring it directly on a Pass C-style choice prompt: 0.8365, degraded relative to 0.995–1.000 for the other templates but above the 0.5 floor.
- Cause: the 0.38 figure was measured in the anchor-comparison context and transferred to Pass C without checking that it applied.
- Consequences: a T2 tidy-up rather than a threat to the condition. The mechanism is real; the magnitude was not.
- Corrected in place in the audit, with the correction stated rather than the number quietly changed.
- Reported: consistency of 0.07–0.32 at every gap, verdict "window closed, instrument failing."
- Refuted by: internal incoherence — the anchor run had shown 0.98 consistency at the widest gap. Investigation found 39% of readouts below the mass floor.
- Cause: D3 switched option labels from letters to digits, but templates t2 and t3 still read "the single letter {label_a} or {label_b}", rendering as "the single letter 1 or 2". Models answered in prose.
- Consequences: that run's numbers discarded. It also exposed a worse bug —
stimulus file contents were not in any config hash, so fixing the templates
invalidated no cache and artifacts would have been silently reused against
different prompts. Fixed in
806579a.
R1 and R2 were transcription or transfer errors. R3 and R4 were estimator errors —
a statistic computed correctly but measuring something other than the intended
quantity, in both cases because a nuisance component was left unmodelled (β) or
misattributed (between-cell variance read as dependence).
The transcription class is addressed by generated status reporting. The estimator class is addressed by preregistered fit-quality diagnostics: the excess-consistency slope (A2.1) and the model-versus-empirical reliability gap (A2.2) both exist because of R3 and R4, and both are currently flagging that the order model remains misspecified.