|
| 1 | +# Complete Testing Suite — Fix & Finish Plan |
| 2 | + |
| 3 | +Drafted 2026-09-01 (Fable seat), amended same day for the standardized-only |
| 4 | +measurement policy. Execute from this file in a fresh Opus session. |
| 5 | + |
| 6 | +## Measurement policy (Alan, 2026-09-01 — supersedes prior scope) |
| 7 | + |
| 8 | +All pipeline comparison and ALL shared or published reporting uses |
| 9 | +standardized, publicly recognized benchmarks scored by their official |
| 10 | +evaluators — currently SWE-bench Verified on the pinned official harness. |
| 11 | +Homegrown corpora and judge panels (the bounce-protocol document benchmark, |
| 12 | +its 3-judge protocol, blind-judge calibration) are RETIRED from measurement |
| 13 | +and from every shared surface. Existing internal results are archived in |
| 14 | +place and never published. Rationale: results on a benchmark nobody outside |
| 15 | +this repo has seen are not comparable and not worth sharing; common tests |
| 16 | +with common baselines are. |
| 17 | + |
| 18 | +Boundary: the hermetic regression suite (`tests/run-all.sh`, 42 suites) is |
| 19 | +engineering QA that gates harness correctness — it is not a benchmark, its |
| 20 | +results are not comparison data, and it stays. |
| 21 | + |
| 22 | +| Surface | Status under policy | |
| 23 | +|---|---| |
| 24 | +| S1 Code — SWE-bench Verified battery (`benchmarks/code/`) | The measurement surface. Partial: B 5/5 (repair arm inert), C 4/5. A, D, solos unrun | |
| 25 | +| S2 Documents — bounce-protocol suite (`benchmarks/`) | RETIRED. Batch b1 complete on disk; archive as internal evidence, no further spend, never on the shared site | |
| 26 | +| S3 Regression — `tests/run-all.sh` | QA gate, 42/42 green. Not reported as benchmark data | |
| 27 | + |
| 28 | +## Issues ledger |
| 29 | + |
| 30 | +| # | Issue | Status | |
| 31 | +|---|---|---| |
| 32 | +| 1 | GLM/Kimi bill reasoning against `max_tokens`; capped critics returned empty content | FIXED `735722f` | |
| 33 | +| 2 | kimi-seat test could not simulate a missing key with a real key on disk | FIXED `5cf451c` | |
| 34 | +| 3 | Codex refuses all writes on Windows despite `--sandbox workspace-write` | OPEN — Phase 0.1 | |
| 35 | +| 4 | Conditions A and D never run on code; solos never run | OPEN — Phase 1 | |
| 36 | +| 5 | B scored with inert repair arm: 5/5 is really Fable-solo | OPEN — re-run after 0.1 | |
| 37 | +| 6 | GLM/Kimi have no agent loop — solo cells need a single-shot harness | OPEN — Phase 1.5 | |
| 38 | +| 7 | Judge `position_biased` verdicts discard t1/t7 cells (doc suite) | CLOSED-RETIRED — surface withdrawn; no re-judging spend | |
| 39 | +| 8 | `sanitize-leak` on t2 (doc suite) | CLOSED-RETIRED — same | |
| 40 | +| 9 | Codex-judge self-preference confound (doc suite) | CLOSED-RETIRED — same | |
| 41 | +| 10 | Two orchestrators wrote one status file (b1 watchdog stamped the SWE status) | OPEN — Phase 0.2 | |
| 42 | +| 11 | 5-task subset: one task = 20 points; B/C gap is one task | OPEN — Phase 4 decides scale | |
| 43 | +| 12 | HF Hub unauthenticated-rate-limit warnings during evaluation | OPEN — minor, Phase 0.3 | |
| 44 | +| 13 | Evaluator leaves 5 images per run | OPEN — hygiene, Phase 0.3, default OFF; never delete other projects' images | |
| 45 | + |
| 46 | +## Phase 0 — Unblock the harness |
| 47 | + |
| 48 | +**0.1 Codex writable workspace (the critical fix).** |
| 49 | +Codex 0.144.5 on Windows degrades `workspace-write` to read-only. Fix |
| 50 | +sequence, stop at the first that passes: |
| 51 | +1. Probe `-s danger-full-access` with the existing 1-file throwaway-repo test. |
| 52 | +2. If refused, probe `--dangerously-bypass-approvals-and-sandbox`. |
| 53 | +3. If neither, route codex through WSL against the same workspace path. |
| 54 | + |
| 55 | +Guardrails: elevated access is acceptable ONLY because benchmark workspaces |
| 56 | +are disposable clones under `benchmarks/results/code/runs/`. Gate behind |
| 57 | +`CODE_BENCH_CODEX_SANDBOX` (default stays `workspace-write`); record the mode |
| 58 | +in `run-manifest.json` — treatment-relevant fact. |
| 59 | +Exit: driver-path probe edits a file; mode recorded in manifest. |
| 60 | + |
| 61 | +**0.2 Status-file single-writer.** One writer per status file; observers get |
| 62 | +their own files; every status line carries `writer=`. |
| 63 | +Exit: tagged lines present; no cross-suite writes. |
| 64 | + |
| 65 | +**0.3 Small hygiene.** `HF_TOKEN` via the `.env.local` loader (never echo). |
| 66 | +Image-prune stays default OFF. |
| 67 | + |
| 68 | +## Phase 1 — Complete the code matrix (SWE-bench Verified, frozen 5-task subset) |
| 69 | + |
| 70 | +Order preserves pairing: never spend Fable dispatches on a condition whose |
| 71 | +comparator cannot run. |
| 72 | + |
| 73 | +| Cell set | Dispatches | Est. cost | Precondition | |
| 74 | +|---|---|---|---| |
| 75 | +| 1.1 A (Fable solo), 5 cells | 5 Fable | ~$5 | none | |
| 76 | +| 1.2 D (self-bounce), 5 cells | 10 Fable | ~$15-20 | none | |
| 77 | +| 1.3 B re-run (real repair), 5 cells | 5 Fable + 5 Codex | ~$5 + plan compute | 0.1 | |
| 78 | +| 1.4 Codex solo, 5 cells | 5 Codex | plan compute | 0.1 | |
| 79 | +| 1.5 GLM solo + Kimi solo, single-shot tier | 10 API calls | cents | new harness | |
| 80 | + |
| 81 | +1.5 harness: issue text + `git grep`-selected file context in one prompt → |
| 82 | +unified diff → `git apply --check` gate → prediction. Label the tier |
| 83 | +"single-shot" everywhere — never unlabeled beside agentic rows. |
| 84 | +All cells scored by the official Docker evaluator; caps A=1, D=2 on |
| 85 | +`--max-claude-dispatches`. |
| 86 | +Exit: every matrix row measured or explicitly blocked; zero infrastructure |
| 87 | +failures; prediction files validate 5/5 unique frozen IDs. |
| 88 | + |
| 89 | +## Phase 2 — Retire the homegrown document benchmark |
| 90 | + |
| 91 | +No model spend. Archive-only: |
| 92 | +1. Leave batch b1 results and `reports/b1.md` in place as internal evidence; |
| 93 | + they are never published, linked, or summarized on any shared surface. |
| 94 | +2. Add a retirement note to `benchmarks/README.md` (doc-suite root): retired |
| 95 | + from measurement 2026-09-01 per standardized-only policy; direct readers |
| 96 | + to `benchmarks/code/` for the active benchmark. |
| 97 | +3. Cancel outstanding doc-suite work: t1/t2/t7 re-judging, sanitizer fix for |
| 98 | + judging, blind-judge calibration baselines. Do not delete any code or |
| 99 | + results — retire, don't destroy. |
| 100 | +Exit: retirement note committed; no doc-suite job scheduled anywhere. |
| 101 | + |
| 102 | +## Phase 3 — Results site: standardized benchmarks only |
| 103 | + |
| 104 | +Update the existing artifact (same URL). Two sections: |
| 105 | +1. **Leaderboard** — SWE-bench Verified frozen-subset matrix, all Phase 1 |
| 106 | + rows, coverage labels, per-task dots. |
| 107 | +2. **Methodology & integrity** — evaluator pin + gold canary 1/1, dispatch |
| 108 | + counts, per-condition cost, harness commit, and the standing caveat: |
| 109 | + frozen 5-task probe, not comparable to published full-500 scores. |
| 110 | +Remove nothing that is already standardized; add no homegrown-benchmark |
| 111 | +content. One aggregator script (`benchmarks/site/aggregate.sh`) builds a |
| 112 | +single JSON from evaluator reports + run logs; the page renders only that. |
| 113 | +Exit: site rebuilt from aggregator output alone; every number traceable to a |
| 114 | +file on disk; zero references to the retired suite. |
| 115 | + |
| 116 | +## Phase 4 — Scale gate (go/no-go recommendation, never autonomous) |
| 117 | + |
| 118 | +Present with dollar estimates, run nothing: |
| 119 | +- Expand the SWE-bench Verified subset (25-50 tasks) if any pipeline-vs-solo |
| 120 | + gap from Phase 1 is worth confirming. |
| 121 | +- Candidate additional suites — standardized public benchmarks only, each |
| 122 | + with an official pinned harness (e.g. SWE-bench Lite, Terminal-Bench, |
| 123 | + Aider Polyglot, LiveCodeBench). No internal corpus is ever proposed. |
| 124 | +- Note for the document pipeline: it currently has NO standardized public |
| 125 | + benchmark. Until one exists and is adopted at this gate, document-pipeline |
| 126 | + quality claims stay unmeasured rather than internally measured. |
| 127 | + |
| 128 | +## Budget & sequencing |
| 129 | + |
| 130 | +Phase 0 is hours, two codex probes. Phase 1 ≈ 20 Fable dispatches (~$25-30), |
| 131 | +10 codex cells inside the daily guard cap, GLM/Kimi in cents. Phase 2 is a |
| 132 | +docs commit. Phase 3 after Phase 1. Phase 4 is a decision. Throughout: |
| 133 | +`.env.local`, results, workspaces, trajectories stay uncommitted; no key |
| 134 | +values in logs; one writer per status file; evidence never deleted. |
0 commit comments