Skip to content

Commit 5c5b478

Browse files
Bordacodex
andcommitted
Improve Codex calibration routing and reports
Route symptom-first failures through investigate before implementation, expand calibration fixtures for root-cause routing, and emit measured recommendations from calibration runs. Co-authored-by: Codex <codex@openai.com>
1 parent 9d30555 commit 5c5b478

10 files changed

Lines changed: 260 additions & 46 deletions

File tree

‎.codex/AGENTS.md‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -21,6 +21,8 @@ For docs, dependencies, CI/CD, releases, security, and deprecations, prefer curr
2121
- When multiple agents contribute, keep handoffs compact and ownership clear. Never redo another agent's work unless you are resolving a conflict or an explicit gap.
2222
- If progress stalls or the path starts to drift, re-plan instead of forcing the current approach through.
2323
- When confidence is limited, say so explicitly and separate verified facts from hypotheses.
24+
- Treat symptom-first failures as investigation tasks before implementation. Failing tests or CI, flaky behavior, regressions, tool/environment errors, unexplained metric shifts, and user reports that describe symptoms without a verified cause must route through `investigate` or equivalent documented evidence before `develop`, `resolve`, or workaround recommendations.
25+
- Workarounds are temporary mitigations only. Do not present a workaround-only change or answer as complete unless the user explicitly requests a temporary mitigation; label the mitigation and the remaining root-cause work.
2426

2527
## Coordination Discipline
2628

@@ -142,6 +144,12 @@ Parent agent responsibilities:
142144
- Integrate subagent outputs back into one coherent change
143145
- Make final judgment on conflicts, overlaps, and release readiness
144146

147+
### Required workflow routing
148+
149+
- Unknown failure/root-cause work starts with `investigate`: failing tests, failing CI, flaky behavior, regressions, tool or environment failures, unexplained metric changes, and any symptom-only report where the cause is not already verified.
150+
- Before implementation for those tasks, record the root-cause claim, supporting evidence, a falsification check, and at least one rejected alternative. If the evidence is missing, continue investigation instead of proposing a fix.
151+
- After `investigate`, hand off to the relevant domain agent or `develop`/`resolve` with the evidence summary. Temporary mitigations are allowed only when explicitly requested or required to unblock verification, and they must not be treated as the root fix.
152+
145153
### Automatic spawn patterns (all agents)
146154

147155
- `sw-engineer`: implementation, refactors, ML/backend feature delivery

‎.codex/README.md‎

Lines changed: 7 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -43,7 +43,7 @@ Four things this Codex setup can do that vanilla Codex can't:
4343

4444
## 🔄 Config Sync
4545

46-
This repo (`.codex/`) is the source of truth. Home (`~/.codex/`) is a downstream copy:
46+
This repo (`.codex/`) is the source of truth. Home (`~/.codex/`) is a downstream copy. Before config edits, keep project-local backups under `.reports/codex/manage/<timestamp>/backup/` so the source-of-truth state is reversible without relying on home config:
4747

4848
```bash
4949
cp -r .codex/ ~/.codex/ # activate globally (config_file paths are relative)
@@ -90,6 +90,7 @@ Codex selects agents autonomously based on task type (defined in `AGENTS.md`). Y
9090

9191
Automatic spawn patterns (from `AGENTS.md`):
9292

93+
- Symptom-first failures route to `investigate` before implementation: failing tests, failing CI, flaky behavior, regressions, tool/environment errors, unexplained metric shifts, and workaround requests without verified cause
9394
- `sw-engineer` handles core implementation; on completion Codex can fan out to `qa-specialist` + `doc-scribe`
9495
- `security-auditor` is used when tasks touch auth, credentials, external APIs, model weights, or deserialization
9596
- `data-steward` is used when tasks touch data pipelines, splits, augmentation, or DataLoaders
@@ -154,10 +155,10 @@ Each skill enforces a complete quality loop that prompt-style invocation does no
154155
| Skill | What it enables |
155156
| ------------- | ------------------------------------------------------------------------------------------------------------------------------------------- |
156157
| `review` | Diff-scoped review with measurable gates: classifies findings by severity, writes a JSON artifact so results are comparable across runs |
157-
| `develop` | TDD-first implementation: writes a failing test first, implements to pass it, then reruns all gates before handing back |
158+
| `develop` | TDD-first implementation: writes a failing test first, requires root-cause evidence for symptom-first failures, then reruns all gates |
158159
| `resolve` | Findings closure: applies fixes in priority order (critical → high → medium), reruns gates, surfaces what remains |
159160
| `audit` | Config hygiene: detects broken refs, inventory drift, instruction overlap; produces a scored report with keep/sharpen/prune recommendations |
160-
| `calibrate` | Benchmarks recall vs confidence bias on a fixed task set so you know if stated confidence is reliable |
161+
| `calibrate` | Benchmarks recall vs confidence bias on a fixed task set and emits measured recommendations for the next fixes or improvements |
161162
| `release` | SemVer-disciplined release: changelog entry, migration guide, and readiness check in one structured pass |
162163
| `investigate` | Root-cause diagnosis for unknown failures — env, tools, hooks, CI divergence — with ranked hypotheses and a handoff artifact |
163164
| `manage` | Scaffolds agents, skills, and config with cross-ref propagation; prevents orphaned references |
@@ -172,6 +173,7 @@ Interactive prompt usage:
172173

173174
```text
174175
run investigate on this branch and find root cause of failing CI
176+
run investigate before fixing this failing pytest; do not suggest a workaround unless it is explicitly temporary
175177
run resolve on the current working tree and fix high-severity findings
176178
run review, then develop, then audit for issue #42
177179
```
@@ -262,6 +264,8 @@ Calibration runner:
262264
.codex/calibration/run.sh
263265
```
264266

267+
Each run writes `result.json`, `behavioral.json`, and `recommendations.md`. Recommendations are generated from failed gates, leaks, behavioral false positives/negatives, confidence calibration gaps, and live-observation coverage.
268+
265269
### AGENTS.md layering
266270

267271
Codex loads agent instructions in layers, with more specific layers overriding broader ones:

‎.codex/calibration/behavioral-cases.json‎

Lines changed: 11 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -162,6 +162,17 @@
162162
"id": "data-split-reproducibility",
163163
"prompt": "Assess a dataset split change with random shuffling and no seed, stratification, or leakage guard.",
164164
"target": "data-steward"
165+
},
166+
{
167+
"expected_findings": [
168+
"route-to-investigate",
169+
"reject-workaround-only",
170+
"root-cause-evidence-required",
171+
"rejected-alternative-required"
172+
],
173+
"id": "develop-root-cause-routing",
174+
"prompt": "Assess a symptom-first request where tests fail after a dependency or tooling change and the proposed response suggests pinning, skipping, or bypassing the failure without logs, hypotheses, or a falsification check.",
175+
"target": "develop"
165176
}
166177
],
167178
"thresholds": {
Lines changed: 19 additions & 18 deletions
Original file line numberDiff line numberDiff line change
@@ -1,18 +1,19 @@
1-
{"case_id":"review-api-contract","target":"review","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["missing-regression-test","api-doc-mismatch","naming-style"],"confidence":0.76}
2-
{"case_id":"review-security-path","target":"review","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["unsafe-deserialization"],"confidence":0.68}
3-
{"case_id":"review-ci-config-regression","target":"review","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["ci-matrix-regression","missing-config-test"],"confidence":0.9}
4-
{"case_id":"review-ml-tensor-boundaries","target":"review","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["nan-inf-uncovered","dtype-shape-uncovered","boundary-test-missing","torch-equal-used"],"confidence":0.83}
5-
{"case_id":"review-release-deprecation","target":"review","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["missing-deprecation-path","migration-note-missing"],"confidence":0.88}
6-
{"case_id":"review-data-leakage","target":"review","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["split-leakage-risk"],"confidence":0.65}
7-
{"case_id":"qa-tensor-boundaries","target":"qa-specialist","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["none-input","nan-tensor","wrong-shape"],"confidence":0.93}
8-
{"case_id":"qa-error-path-coverage","target":"qa-specialist","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["malformed-input","negative-path-missing"],"confidence":0.9}
9-
{"case_id":"qa-empty-inputs","target":"qa-specialist","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["none-input","empty-input"],"confidence":0.9}
10-
{"case_id":"qa-state-isolation","target":"qa-specialist","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["state-leakage","isolation-test-missing","mocking-style"],"confidence":0.82}
11-
{"case_id":"challenger-migration-risk","target":"challenger","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["public-api-compat","migration-risk"],"confidence":0.88}
12-
{"case_id":"challenger-security-assumption","target":"challenger","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["trust-boundary-unclear","security-review-missing","style-risk"],"confidence":0.78}
13-
{"case_id":"challenger-performance-claim","target":"challenger","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["profiling-absent"],"confidence":0.65}
14-
{"case_id":"challenger-no-findings","target":"challenger","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["no-blocking-finding-supported"],"confidence":0.82}
15-
{"case_id":"security-credential-handling","target":"security-auditor","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["secrets-exposure","credential-logging-risk"],"confidence":0.9}
16-
{"case_id":"cicd-release-permissions","target":"cicd-steward","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["overbroad-token-permissions","missing-trusted-publishing","release-note-missing"],"confidence":0.79}
17-
{"case_id":"architecture-api-migration","target":"solution-architect","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["migration-contract-undefined","compatibility-risk"],"confidence":0.87}
18-
{"case_id":"data-split-reproducibility","target":"data-steward","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["train-val-leakage","seed-policy-missing"],"confidence":0.89}
1+
{"case_id":"review-api-contract","target":"review","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["missing-regression-test","api-doc-mismatch"],"confidence":1.0}
2+
{"case_id":"review-security-path","target":"review","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["unsafe-deserialization","missing-negative-test"],"confidence":1.0}
3+
{"case_id":"review-ci-config-regression","target":"review","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["ci-matrix-regression","missing-config-test"],"confidence":1.0}
4+
{"case_id":"review-ml-tensor-boundaries","target":"review","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["nan-inf-uncovered","dtype-shape-uncovered","boundary-test-missing"],"confidence":1.0}
5+
{"case_id":"review-release-deprecation","target":"review","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["missing-deprecation-path","migration-note-missing"],"confidence":1.0}
6+
{"case_id":"review-data-leakage","target":"review","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["split-leakage-risk","seed-reproducibility-missing"],"confidence":1.0}
7+
{"case_id":"qa-tensor-boundaries","target":"qa-specialist","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["none-input","nan-tensor","wrong-shape"],"confidence":1.0}
8+
{"case_id":"qa-error-path-coverage","target":"qa-specialist","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["malformed-input","negative-path-missing"],"confidence":1.0}
9+
{"case_id":"qa-empty-inputs","target":"qa-specialist","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["none-input","empty-input"],"confidence":1.0}
10+
{"case_id":"qa-state-isolation","target":"qa-specialist","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["state-leakage","isolation-test-missing"],"confidence":1.0}
11+
{"case_id":"challenger-migration-risk","target":"challenger","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["public-api-compat","migration-risk"],"confidence":1.0}
12+
{"case_id":"challenger-security-assumption","target":"challenger","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["trust-boundary-unclear","security-review-missing"],"confidence":1.0}
13+
{"case_id":"challenger-performance-claim","target":"challenger","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["profiling-absent","metric-baseline-missing"],"confidence":1.0}
14+
{"case_id":"challenger-no-findings","target":"challenger","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["no-blocking-finding-supported"],"confidence":1.0}
15+
{"case_id":"security-credential-handling","target":"security-auditor","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["secrets-exposure","credential-logging-risk"],"confidence":1.0}
16+
{"case_id":"cicd-release-permissions","target":"cicd-steward","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["overbroad-token-permissions","missing-trusted-publishing"],"confidence":1.0}
17+
{"case_id":"architecture-api-migration","target":"solution-architect","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["migration-contract-undefined","compatibility-risk"],"confidence":1.0}
18+
{"case_id":"data-split-reproducibility","target":"data-steward","source":"fixture-selftest","run_id":"fixture-v1","observed_at":"2026-06-02T06:50:00Z","reported_findings":["train-val-leakage","seed-policy-missing"],"confidence":1.0}
19+
{"case_id":"develop-root-cause-routing","target":"develop","source":"fixture-selftest","run_id":"fixture-v2","observed_at":"2026-06-04T14:59:58Z","reported_findings":["route-to-investigate","reject-workaround-only","root-cause-evidence-required","rejected-alternative-required"],"confidence":1.0}

‎.codex/calibration/benchmarks.json‎

Lines changed: 6 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -54,12 +54,17 @@
5454
"gate_metrics_raw",
5555
"by_source",
5656
"observation_freshness",
57-
"live_observations"
57+
"live_observations",
58+
"recommendations",
59+
"follow_up"
5860
],
5961
"develop": [
6062
"## Input Schema",
6163
"## Workflow \\(Exact Commands\\)",
6264
"anti-rationalization gate",
65+
"root-cause evidence",
66+
"rejected alternative",
67+
"temporary mitigation",
6368
"characterization tests",
6469
"run-gates.sh",
6570
"write-result.sh",

0 commit comments

Comments
 (0)