You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
One-off CLI retrospective pass across state/usage_log.md (133 entries), state/research_log.md (152 entries), state/agent_log.md, and closed-issue / merged-PR history for the 14-day window 2026-04-04 → 2026-04-17. Analysis done interactively in a human-directed session rather than by evolve.yml.
Purpose: capture findings before they scroll out of context. Individual actionable findings are filed as separate issues (linked below).
Top-line numbers
Cost: $173.62 over 8 days of data ($21.70/day avg, annualized ~$152/wk — slightly above $150/wk target)
1. Opus is 100% unreachable since 2026-04-13T20:51Z
→ filed as issue #175[pipeline] Opus unreachable since... 100% Haiku fallback for 4+ days.
project_state.md reports "26.8% Haiku dominance"; actual daily breakdown shows 100% Haiku for 4 consecutive days. Needs account/billing investigation.
2. Haiku fallback is NOT cost-effective on this workload
Per-turn cost on analyze/evolve/watcher: Haiku is 5-22% more expensive than Opus and takes similar or more turns. The fallback's "saves money" premise is empirically wrong here. (Part of #175.)
→ filed as issue #176[pipeline] growth.yml saturates max-turns on 9/9 runs....
Either cap is too low OR scope is too broad. Hypothesis B (scope too broad) is the working theory: growth is doing surveillance when it should be acting.
4. watcher and reviewer are radically over-capped
→ filed as issue #177[pipeline] Reduce watcher (50→35) and reviewer (45→25)....
Watcher hits cap 0/67 runs (utilization 56%). Reviewer hits cap 0/6 runs (utilization 34%). Caps should be tightened so they actually bind.
5. PATTERN_HUNT posture has 23 consecutive 0-yield runs
→ filed as issue #178[evolve] Suspend/demote PATTERN_HUNT posture....
All 4 postures show 0-yield across 21 evolve runs in the window. PATTERN_HUNT is the most expensive and most repeatedly empty.
6. Docs-staleness is a repeating cycle (#156 after #155, #168 after #165)
→ filed as issue #179[evolve] Add README regen as coder post-step for cron-cadence changes....
Each cron change produces a follow-up "README stale" finding 1-4 days later. ~$3-5 churn per cycle. Coder prompt should sync README frequency claims in the same PR.
7. Research yield is 4.6% (7 issues from 152 log entries)
Not actionable on its own, but grounds the "plateau" self-reports in data. Highest-yield source: anthropics/claude-code (5/8 adoptable). Most named sources yielded <1 action per week.
8. #173 was the outlier in resolution time (30h vs 4h median)
The system's three systemic failure modes converged on this issue: (a) scope too big for single-pass coder, (b) Haiku at max-turns, (c) workflow YAML edits need WORKFLOW_PAT + PR review. Worth considering a coder chunking strategy for >5-file changes.
Context
One-off CLI retrospective pass across state/usage_log.md (133 entries), state/research_log.md (152 entries), state/agent_log.md, and closed-issue / merged-PR history for the 14-day window 2026-04-04 → 2026-04-17. Analysis done interactively in a human-directed session rather than by evolve.yml.
Purpose: capture findings before they scroll out of context. Individual actionable findings are filed as separate issues (linked below).
Top-line numbers
state:(97%), 27 non-stateNew findings not previously surfaced
1. Opus is 100% unreachable since 2026-04-13T20:51Z
→ filed as issue #175
[pipeline] Opus unreachable since... 100% Haiku fallback for 4+ days.project_state.md reports "26.8% Haiku dominance"; actual daily breakdown shows 100% Haiku for 4 consecutive days. Needs account/billing investigation.
2. Haiku fallback is NOT cost-effective on this workload
Per-turn cost on analyze/evolve/watcher: Haiku is 5-22% more expensive than Opus and takes similar or more turns. The fallback's "saves money" premise is empirically wrong here. (Part of #175.)
3. growth.yml saturates max-turns 9/9 runs, produces 0 actions
→ filed as issue #176
[pipeline] growth.yml saturates max-turns on 9/9 runs....Either cap is too low OR scope is too broad. Hypothesis B (scope too broad) is the working theory: growth is doing surveillance when it should be acting.
4. watcher and reviewer are radically over-capped
→ filed as issue #177
[pipeline] Reduce watcher (50→35) and reviewer (45→25)....Watcher hits cap 0/67 runs (utilization 56%). Reviewer hits cap 0/6 runs (utilization 34%). Caps should be tightened so they actually bind.
5. PATTERN_HUNT posture has 23 consecutive 0-yield runs
→ filed as issue #178
[evolve] Suspend/demote PATTERN_HUNT posture....All 4 postures show 0-yield across 21 evolve runs in the window. PATTERN_HUNT is the most expensive and most repeatedly empty.
6. Docs-staleness is a repeating cycle (#156 after #155, #168 after #165)
→ filed as issue #179
[evolve] Add README regen as coder post-step for cron-cadence changes....Each cron change produces a follow-up "README stale" finding 1-4 days later. ~$3-5 churn per cycle. Coder prompt should sync README frequency claims in the same PR.
7. Research yield is 4.6% (7 issues from 152 log entries)
Not actionable on its own, but grounds the "plateau" self-reports in data. Highest-yield source:
anthropics/claude-code(5/8 adoptable). Most named sources yielded <1 action per week.8. #173 was the outlier in resolution time (30h vs 4h median)
The system's three systemic failure modes converged on this issue: (a) scope too big for single-pass coder, (b) Haiku at max-turns, (c) workflow YAML edits need WORKFLOW_PAT + PR review. Worth considering a coder chunking strategy for >5-file changes.
Already-tracked items (not refiled)
Filed as new issues
Prioritization suggestion
If re-enabling workflows and picking targets:
Not fileable as issues (observations only)
git logGenerated from CLI retrospective session 2026-04-18, not evolve.yml. See individual child issues for actionable specifics.