Skip to content

Commit 06aa65d

Browse files
authored
Merge pull request #57 from alanshurafa/codex/code-benchmark-battery
Land the SWE-bench code benchmark battery on master
2 parents d68c310 + 05d52f9 commit 06aa65d

35 files changed

Lines changed: 3721 additions & 31 deletions

‎.gitignore‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -59,3 +59,7 @@ tests/.run-ledger/
5959

6060
# Machine-local operational notes (private paths, spend, secrets workflow)
6161
.planning/local/
62+
63+
# Python bytecode caches
64+
__pycache__/
65+
*.pyc
Lines changed: 67 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,67 @@
1+
# Code Benchmark Battery — Implementation Plan
2+
3+
Date: 2026-08-31
4+
5+
## Goal
6+
7+
Measure whether Co-Evolution workflows produce repository patches that pass
8+
standard deterministic evaluators, while keeping subscription usage bounded and
9+
making every live dispatch auditable before it runs.
10+
11+
## Conditions
12+
13+
- **A — Fable solo:** one Fable coding-agent dispatch.
14+
- **B — cross-vendor bounce:** Fable implements; Codex reviews and repairs.
15+
- **C — Fable-led panel:** Fable implements; Codex, GLM, and Kimi critique;
16+
Fable performs the final repair.
17+
- **D — Fable self-bounce:** Fable implements and then reviews/repairs its own
18+
patch. This controls for extra passes and compute.
19+
20+
Declared provider dispatches are a lower bound: one coding-agent dispatch may
21+
contain multiple internal model turns. The live runner must enforce a declared
22+
Claude-dispatch cap and must never infer a percentage of a Max subscription from
23+
tokens because Anthropic does not publish a fixed weekly token denominator.
24+
25+
## Delivery sequence
26+
27+
1. Add frozen suite, condition, external-source, and subset manifests.
28+
2. Add an offline compute estimator and fail-closed prediction validator.
29+
3. Add a pinned SWE-bench installer, metadata fetcher, and official evaluator
30+
bridge under an ignored cache/results tree.
31+
4. Add hermetic tests and wire them into the repository aggregate gate.
32+
5. Prepare the five-task canary and validate official gold scoring when Docker
33+
is available.
34+
6. Add live patch-generation drivers only after the offline harness is green.
35+
Begin with one task and conditions A/B/C, capped at four declared Fable
36+
dispatches. Measure the actual Settings > Usage change before expanding.
37+
38+
## Safety and scope
39+
40+
- No full SWE-bench run in this phase.
41+
- No live model calls during setup or hermetic verification.
42+
- No gold patches, hidden tests, or oracle solutions are exposed to agents.
43+
- Every condition starts from the same clean instance and produces a standard
44+
prediction record for the official evaluator.
45+
- Raw datasets, images, virtual environments, trajectories, and results stay
46+
below `benchmarks/results/code/` and remain uncommitted.
47+
48+
## Acceptance criteria
49+
50+
- `bash benchmarks/code/code-bench.sh check` passes.
51+
- The estimator reports exact declared dispatch counts and refuses a cap breach
52+
with exit 75.
53+
- Prediction validation rejects unknown instances, duplicates, empty patches,
54+
and malformed JSONL.
55+
- The pinned metadata fetch contains only public task inputs, never gold data.
56+
- The official SWE-bench gold canary passes once Docker is running.
57+
- The repository aggregate test gate includes the new hermetic suite.
58+
59+
## Execution result
60+
61+
- Pinned SWE-bench source and CLI installed under the ignored results cache.
62+
- Windows compatibility patch forces LF for the Linux `eval.sh` file only.
63+
- Docker Desktop 29.6.2 official gold canary `sympy__sympy-20916` completed
64+
and resolved 1/1 with zero infrastructure, ambiguous, or evaluator errors.
65+
- Hermetic code-benchmark suite passed 14/14; repository aggregate passed
66+
42/42 after isolating the intentionally installed private Kimi key.
67+
- No live Fable, Codex, GLM, or Kimi benchmark calls were made during setup.

‎benchmarks/COMPLETE-SUITE-PLAN.md‎

Lines changed: 134 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,134 @@
1+
# Complete Testing Suite — Fix & Finish Plan
2+
3+
Drafted 2026-09-01 (Fable seat), amended same day for the standardized-only
4+
measurement policy. Execute from this file in a fresh Opus session.
5+
6+
## Measurement policy (Alan, 2026-09-01 — supersedes prior scope)
7+
8+
All pipeline comparison and ALL shared or published reporting uses
9+
standardized, publicly recognized benchmarks scored by their official
10+
evaluators — currently SWE-bench Verified on the pinned official harness.
11+
Homegrown corpora and judge panels (the bounce-protocol document benchmark,
12+
its 3-judge protocol, blind-judge calibration) are RETIRED from measurement
13+
and from every shared surface. Existing internal results are archived in
14+
place and never published. Rationale: results on a benchmark nobody outside
15+
this repo has seen are not comparable and not worth sharing; common tests
16+
with common baselines are.
17+
18+
Boundary: the hermetic regression suite (`tests/run-all.sh`, 42 suites) is
19+
engineering QA that gates harness correctness — it is not a benchmark, its
20+
results are not comparison data, and it stays.
21+
22+
| Surface | Status under policy |
23+
|---|---|
24+
| S1 Code — SWE-bench Verified battery (`benchmarks/code/`) | The measurement surface. Partial: B 5/5 (repair arm inert), C 4/5. A, D, solos unrun |
25+
| S2 Documents — bounce-protocol suite (`benchmarks/`) | RETIRED. Batch b1 complete on disk; archive as internal evidence, no further spend, never on the shared site |
26+
| S3 Regression — `tests/run-all.sh` | QA gate, 42/42 green. Not reported as benchmark data |
27+
28+
## Issues ledger
29+
30+
| # | Issue | Status |
31+
|---|---|---|
32+
| 1 | GLM/Kimi bill reasoning against `max_tokens`; capped critics returned empty content | FIXED `735722f` |
33+
| 2 | kimi-seat test could not simulate a missing key with a real key on disk | FIXED `5cf451c` |
34+
| 3 | Codex refuses all writes on Windows despite `--sandbox workspace-write` | OPEN — Phase 0.1 |
35+
| 4 | Conditions A and D never run on code; solos never run | OPEN — Phase 1 |
36+
| 5 | B scored with inert repair arm: 5/5 is really Fable-solo | OPEN — re-run after 0.1 |
37+
| 6 | GLM/Kimi have no agent loop — solo cells need a single-shot harness | OPEN — Phase 1.5 |
38+
| 7 | Judge `position_biased` verdicts discard t1/t7 cells (doc suite) | CLOSED-RETIRED — surface withdrawn; no re-judging spend |
39+
| 8 | `sanitize-leak` on t2 (doc suite) | CLOSED-RETIRED — same |
40+
| 9 | Codex-judge self-preference confound (doc suite) | CLOSED-RETIRED — same |
41+
| 10 | Two orchestrators wrote one status file (b1 watchdog stamped the SWE status) | OPEN — Phase 0.2 |
42+
| 11 | 5-task subset: one task = 20 points; B/C gap is one task | OPEN — Phase 4 decides scale |
43+
| 12 | HF Hub unauthenticated-rate-limit warnings during evaluation | OPEN — minor, Phase 0.3 |
44+
| 13 | Evaluator leaves 5 images per run | OPEN — hygiene, Phase 0.3, default OFF; never delete other projects' images |
45+
46+
## Phase 0 — Unblock the harness
47+
48+
**0.1 Codex writable workspace (the critical fix).**
49+
Codex 0.144.5 on Windows degrades `workspace-write` to read-only. Fix
50+
sequence, stop at the first that passes:
51+
1. Probe `-s danger-full-access` with the existing 1-file throwaway-repo test.
52+
2. If refused, probe `--dangerously-bypass-approvals-and-sandbox`.
53+
3. If neither, route codex through WSL against the same workspace path.
54+
55+
Guardrails: elevated access is acceptable ONLY because benchmark workspaces
56+
are disposable clones under `benchmarks/results/code/runs/`. Gate behind
57+
`CODE_BENCH_CODEX_SANDBOX` (default stays `workspace-write`); record the mode
58+
in `run-manifest.json` — treatment-relevant fact.
59+
Exit: driver-path probe edits a file; mode recorded in manifest.
60+
61+
**0.2 Status-file single-writer.** One writer per status file; observers get
62+
their own files; every status line carries `writer=`.
63+
Exit: tagged lines present; no cross-suite writes.
64+
65+
**0.3 Small hygiene.** `HF_TOKEN` via the `.env.local` loader (never echo).
66+
Image-prune stays default OFF.
67+
68+
## Phase 1 — Complete the code matrix (SWE-bench Verified, frozen 5-task subset)
69+
70+
Order preserves pairing: never spend Fable dispatches on a condition whose
71+
comparator cannot run.
72+
73+
| Cell set | Dispatches | Est. cost | Precondition |
74+
|---|---|---|---|
75+
| 1.1 A (Fable solo), 5 cells | 5 Fable | ~$5 | none |
76+
| 1.2 D (self-bounce), 5 cells | 10 Fable | ~$15-20 | none |
77+
| 1.3 B re-run (real repair), 5 cells | 5 Fable + 5 Codex | ~$5 + plan compute | 0.1 |
78+
| 1.4 Codex solo, 5 cells | 5 Codex | plan compute | 0.1 |
79+
| 1.5 GLM solo + Kimi solo, single-shot tier | 10 API calls | cents | new harness |
80+
81+
1.5 harness: issue text + `git grep`-selected file context in one prompt →
82+
unified diff → `git apply --check` gate → prediction. Label the tier
83+
"single-shot" everywhere — never unlabeled beside agentic rows.
84+
All cells scored by the official Docker evaluator; caps A=1, D=2 on
85+
`--max-claude-dispatches`.
86+
Exit: every matrix row measured or explicitly blocked; zero infrastructure
87+
failures; prediction files validate 5/5 unique frozen IDs.
88+
89+
## Phase 2 — Retire the homegrown document benchmark
90+
91+
No model spend. Archive-only:
92+
1. Leave batch b1 results and `reports/b1.md` in place as internal evidence;
93+
they are never published, linked, or summarized on any shared surface.
94+
2. Add a retirement note to `benchmarks/README.md` (doc-suite root): retired
95+
from measurement 2026-09-01 per standardized-only policy; direct readers
96+
to `benchmarks/code/` for the active benchmark.
97+
3. Cancel outstanding doc-suite work: t1/t2/t7 re-judging, sanitizer fix for
98+
judging, blind-judge calibration baselines. Do not delete any code or
99+
results — retire, don't destroy.
100+
Exit: retirement note committed; no doc-suite job scheduled anywhere.
101+
102+
## Phase 3 — Results site: standardized benchmarks only
103+
104+
Update the existing artifact (same URL). Two sections:
105+
1. **Leaderboard** — SWE-bench Verified frozen-subset matrix, all Phase 1
106+
rows, coverage labels, per-task dots.
107+
2. **Methodology & integrity** — evaluator pin + gold canary 1/1, dispatch
108+
counts, per-condition cost, harness commit, and the standing caveat:
109+
frozen 5-task probe, not comparable to published full-500 scores.
110+
Remove nothing that is already standardized; add no homegrown-benchmark
111+
content. One aggregator script (`benchmarks/site/aggregate.sh`) builds a
112+
single JSON from evaluator reports + run logs; the page renders only that.
113+
Exit: site rebuilt from aggregator output alone; every number traceable to a
114+
file on disk; zero references to the retired suite.
115+
116+
## Phase 4 — Scale gate (go/no-go recommendation, never autonomous)
117+
118+
Present with dollar estimates, run nothing:
119+
- Expand the SWE-bench Verified subset (25-50 tasks) if any pipeline-vs-solo
120+
gap from Phase 1 is worth confirming.
121+
- Candidate additional suites — standardized public benchmarks only, each
122+
with an official pinned harness (e.g. SWE-bench Lite, Terminal-Bench,
123+
Aider Polyglot, LiveCodeBench). No internal corpus is ever proposed.
124+
- Note for the document pipeline: it currently has NO standardized public
125+
benchmark. Until one exists and is adopted at this gate, document-pipeline
126+
quality claims stay unmeasured rather than internally measured.
127+
128+
## Budget & sequencing
129+
130+
Phase 0 is hours, two codex probes. Phase 1 ≈ 20 Fable dispatches (~$25-30),
131+
10 codex cells inside the daily guard cap, GLM/Kimi in cents. Phase 2 is a
132+
docs commit. Phase 3 after Phase 1. Phase 4 is a decision. Throughout:
133+
`.env.local`, results, workspaces, trajectories stay uncommitted; no key
134+
values in logs; one writer per status file; evidence never deleted.

‎benchmarks/README.md‎

Lines changed: 28 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,33 @@
11
# Co-Evolution Benchmark Suite
22

3+
## Retired from measurement, 2026-09-01
4+
5+
The plan-composition benchmark described below no longer measures anything.
6+
Measurement and every shared report now use standardized, publicly recognized
7+
benchmarks scored by their official evaluators — currently SWE-bench Verified
8+
on the pinned official harness in [`benchmarks/code/`](code/README.md). Go
9+
there for the active benchmark.
10+
11+
A result on a corpus and a judge panel that exist only in this repository
12+
cannot be compared against anything anyone else has run, so it is not worth
13+
publishing. The batch b1 results and `reports/b1.md` stay on disk as internal
14+
evidence of what was built; they are never published, linked, or summarized on
15+
a shared surface. The outstanding work on this suite — re-judging the
16+
position-biased cells, the sanitizer fix for judging, and blind-judge
17+
calibration baselines — is cancelled rather than deferred.
18+
19+
Nothing here is deleted. The runbook below still describes what the scripts do
20+
if you need to read or re-derive an archived batch. Do not schedule new
21+
batches, and do not add this suite's numbers to any published page.
22+
23+
A future benchmark may be added only if it is a standardized public suite with
24+
an official pinned evaluator.
25+
26+
---
27+
28+
The rest of this document is the archived runbook for the retired
29+
plan-composition benchmark.
30+
331
Batch runbook for comparing plan-composition conditions (solo Fable, Codex
432
bounce, panel critique, self-bounce control) on identical planning tasks,
533
scored by three blind automated judges. See `PREREGISTRATION.md` for the

0 commit comments

Comments
 (0)