Skip to content

[evolve] 14-day retrospective (Apr 4–17) — findings from CLI session 2026-04-18 #180

Description

@verkyyi

Context

One-off CLI retrospective pass across state/usage_log.md (133 entries), state/research_log.md (152 entries), state/agent_log.md, and closed-issue / merged-PR history for the 14-day window 2026-04-04 → 2026-04-17. Analysis done interactively in a human-directed session rather than by evolve.yml.

Purpose: capture findings before they scroll out of context. Individual actionable findings are filed as separate issues (linked below).

Top-line numbers

  • Cost: $173.62 over 8 days of data ($21.70/day avg, annualized ~$152/wk — slightly above $150/wk target)
  • Commits: 889 total, 862 state: (97%), 27 non-state
  • Issues closed: 13 (bot-authored fix PRs + human direct pushes)
  • PRs merged: 13 (11 bot, 2 human) — 85% self-heal ratio
  • Median bot-PR merge time: 1.6 minutes

New findings not previously surfaced

1. Opus is 100% unreachable since 2026-04-13T20:51Z

→ filed as issue #175 [pipeline] Opus unreachable since... 100% Haiku fallback for 4+ days.

project_state.md reports "26.8% Haiku dominance"; actual daily breakdown shows 100% Haiku for 4 consecutive days. Needs account/billing investigation.

2. Haiku fallback is NOT cost-effective on this workload

Per-turn cost on analyze/evolve/watcher: Haiku is 5-22% more expensive than Opus and takes similar or more turns. The fallback's "saves money" premise is empirically wrong here. (Part of #175.)

3. growth.yml saturates max-turns 9/9 runs, produces 0 actions

→ filed as issue #176 [pipeline] growth.yml saturates max-turns on 9/9 runs....

Either cap is too low OR scope is too broad. Hypothesis B (scope too broad) is the working theory: growth is doing surveillance when it should be acting.

4. watcher and reviewer are radically over-capped

→ filed as issue #177 [pipeline] Reduce watcher (50→35) and reviewer (45→25)....

Watcher hits cap 0/67 runs (utilization 56%). Reviewer hits cap 0/6 runs (utilization 34%). Caps should be tightened so they actually bind.

5. PATTERN_HUNT posture has 23 consecutive 0-yield runs

→ filed as issue #178 [evolve] Suspend/demote PATTERN_HUNT posture....

All 4 postures show 0-yield across 21 evolve runs in the window. PATTERN_HUNT is the most expensive and most repeatedly empty.

6. Docs-staleness is a repeating cycle (#156 after #155, #168 after #165)

→ filed as issue #179 [evolve] Add README regen as coder post-step for cron-cadence changes....

Each cron change produces a follow-up "README stale" finding 1-4 days later. ~$3-5 churn per cycle. Coder prompt should sync README frequency claims in the same PR.

7. Research yield is 4.6% (7 issues from 152 log entries)

Not actionable on its own, but grounds the "plateau" self-reports in data. Highest-yield source: anthropics/claude-code (5/8 adoptable). Most named sources yielded <1 action per week.

8. #173 was the outlier in resolution time (30h vs 4h median)

The system's three systemic failure modes converged on this issue: (a) scope too big for single-pass coder, (b) Haiku at max-turns, (c) workflow YAML edits need WORKFLOW_PAT + PR review. Worth considering a coder chunking strategy for >5-file changes.

Already-tracked items (not refiled)

Filed as new issues

Prioritization suggestion

If re-enabling workflows and picking targets:

  1. [pipeline] Opus unreachable since 2026-04-13T20:51Z — 100% Haiku fallback for 4+ days #175 — distorts every cost assumption, affects all other findings
  2. [pipeline] growth.yml saturates max-turns on 9/9 runs, produces 0 actions — cap/scope mismatch #176 — pure waste stream, no action produced
  3. [evolve] Suspend/demote PATTERN_HUNT posture — 23 consecutive 0-yield runs, plateau confirmed empirically #178 — free savings, zero risk
  4. [pipeline] Reduce watcher (50→35) and reviewer (45→25) max-turns — over-capped, never saturate #177 — tail variance reduction, zero risk
  5. [evolve] Add README regen as coder post-step for cron-cadence changes — eliminates docs-staleness loop #179 — eliminates a recurring cost loop

Not fileable as issues (observations only)

  • State commit churn is high (97% of commits are state:) but not broken — just noise for humans doing git log
  • Hourly distribution clusters correctly at cron-trigger times — no anomaly
  • Failed-run rate is low (~5/130 runs) and all categorized

Generated from CLI retrospective session 2026-04-18, not evolve.yml. See individual child issues for actionable specifics.

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions