|
| 1 | +# Amendment 1 to the v4 analysis plan (32ab699) |
| 2 | + |
| 3 | +Timestamped by its own commit, while the batch is still running. |
| 4 | +`32ab699` is left untouched; where this amendment contradicts it, the |
| 5 | +amendment governs. External audit found the plan not yet fully |
| 6 | +specified — several choices could still have been made after seeing |
| 7 | +outcomes. Each is closed here. |
| 8 | + |
| 9 | +Batch state at this amendment: Cell A complete, Cell B in progress |
| 10 | +(records exist on disk, unopened beyond what is disclosed below). |
| 11 | + |
| 12 | +## 1. Status reclassified — this is not preregistration |
| 13 | + |
| 14 | +`32ab699` and its commit message called itself a preregistration made |
| 15 | +"before outcomes exist to peek at". **That claim was false when made**: |
| 16 | +at commit time Cell A's five treatment and three control records and |
| 17 | +two Cell B records existed on disk, and log tails during health checks |
| 18 | +had exposed endpoint lines ("absent at window end: True") of at most |
| 19 | +two runs. Git verifies the timestamp; nothing can verify the files were |
| 20 | +never opened. |
| 21 | + |
| 22 | +Correct designation: **a mid-collection, partially outcome-blinded |
| 23 | +prospective analysis plan**, with the endpoint exposure above. Valuable |
| 24 | +because the analysis choices predate reading any resurrection count, |
| 25 | +trajectory, reset delta, or late-recurrence outcome — not equivalent to |
| 26 | +registration before collection. |
| 27 | + |
| 28 | +Auditor transparency, carried into the record: the audit that prompted |
| 29 | +this amendment inspected v4 schema/provenance fields and one validity |
| 30 | +flag, not substantive outcomes. |
| 31 | + |
| 32 | +## 2. Event intervals and horizons, exactly |
| 33 | + |
| 34 | +Resurrection witnesses are interval-censored. For every witness at the |
| 35 | +author: |
| 36 | + |
| 37 | +- **Event interval** = (previous author observation `done`, current |
| 38 | + author observation `done`], both in per-node observation time — never |
| 39 | + the shared row timestamp `t`. |
| 40 | +- **Evidence within horizon H**: interval upper bound ≤ H. |
| 41 | +- **Definitely late (threshold L)**: interval lower bound > L. |
| 42 | +- **Ambiguous**: interval straddles H or L — counted in neither bin, |
| 43 | + reported separately, never dropped. |
| 44 | +- **Horizon truncation** (Q2's 3-lifetime reanalysis): a truncated |
| 45 | + series keeps, per node, exactly the observations with `done` ≤ H; |
| 46 | + the analyser then runs unchanged on that subset. |
| 47 | + |
| 48 | +Lifetimes are multiples of `configured_insert_ttl` (96 s → H = 288 s |
| 49 | +for Q2; L = 192 s and 480 s for Q3). |
| 50 | + |
| 51 | +## 3. No formal hypothesis tests for v4 |
| 52 | + |
| 53 | +The conditional "if any formal test is run" left a post-outcome choice |
| 54 | +open. Closed: **no formal hypothesis tests will be run on v4 data.** No |
| 55 | +Fisher, no Holm family, no p-values. Reporting is: raw per-run |
| 56 | +outcomes, Wilson 95% intervals on proportions, and one exploratory |
| 57 | +effect estimate (Q2's risk difference). Anything beyond that in a |
| 58 | +future write-up is post-hoc and must be labelled so. |
| 59 | + |
| 60 | +## 4. Q1 primary estimate is v4-only |
| 61 | + |
| 62 | +v3 was acquired by script `d1366561…`, v4 by `c899a021…`, with |
| 63 | +substantial changes between them; reanalysis aligns derived fields, not |
| 64 | +acquisition behaviour. Therefore: |
| 65 | + |
| 66 | +- **Primary Q1 estimate: the five v4 Cell A treatment runs.** |
| 67 | +- v3's three runs: reported separately as historical replication. |
| 68 | +- A pooled 8-run figure is secondary at most, and only after a |
| 69 | + documented acquisition-path equivalence audit (a reviewed diff of |
| 70 | + `d1366561…` → `c899a021…` showing no change to probing, sampling |
| 71 | + cadence, gating, or recording of the raw series) — matching binary |
| 72 | + and cell dict alone is insufficient. |
| 73 | + |
| 74 | +## 5. Remaining specifications closed |
| 75 | + |
| 76 | +- **Q2 comparator**: v4 Cell A only (the five primary runs), truncated |
| 77 | + to the 3-lifetime horizon by the rule in §2. v3 is not a comparator. |
| 78 | +- **Interval methods**: proportions get Wilson 95% intervals; the Q2 |
| 79 | + risk difference gets a **95% Newcombe hybrid-score interval without |
| 80 | + continuity correction**. |
| 81 | +- **Q3 intervals**: Wilson 95% intervals ARE reported at n=3 (0/3 |
| 82 | + leaves an upper bound near 56%, and that width is the message). |
| 83 | + "No interval worth reporting" in 32ab699 is withdrawn. |
| 84 | +- **Missingness policy**: planned denominators are fixed by the |
| 85 | + committed driver — treatments 5 (A) / 3 (B) / 3 (C), controls 3 (A) / |
| 86 | + 2 (B). Every report states planned n, valid n, invalid filenames with |
| 87 | + their failure reasons, and best/worst-case bounds for each binary |
| 88 | + outcome (all-invalid-runs-positive / all-negative). **Invalid or |
| 89 | + timed-out runs are not re-run for v4** — selective repetition after |
| 90 | + seeing which runs failed is outcome-dependent sampling. |
| 91 | +- **Cell key**: the full record `cell` dict plus window, arm, topology |
| 92 | + (chain; `directed` flag), `sample_gap_s`, |
| 93 | + `propagation_check_t_requested`, and the acquisition-script hash. |
| 94 | + Batch (v3/v4) is a stratum, never silently pooled. |
| 95 | +- **Controls endpoint**: predeclared as `absent_at_end` by cell with a |
| 96 | + Wilson 95% interval, reported under the standing label "resurrection |
| 97 | + not observable (3 samples)". |
| 98 | +- **Immutability, corrected**: "source records stay immutable" |
| 99 | + conflicted with `--reanalyse`, which rewrites derived fields in |
| 100 | + place. The immutable object is the **raw series plus acquisition |
| 101 | + metadata**; the extract stage hashes that portion of each record |
| 102 | + separately and re-derives all statistics from the series itself, so |
| 103 | + v4 aggregate results never depend on in-record `analysis` fields. |
| 104 | + `--reanalyse` remains a convenience for humans reading single |
| 105 | + records; it is not an input to the pipeline. No `--reanalyse` runs on |
| 106 | + v4 records before extraction is complete. |
0 commit comments