If you have a darwin-skill installation with a populated results.tsv, this guide walks through a side-by-side migration. You do not have to abandon Darwin to try evolve-skill — the two can run on separate branches against the same skills.
- Your existing
test-prompts.jsonfiles become the training set inevolve-skill. They have been seen during prior optimization, so they cannot be used as holdouts. - Your
results.tsvhistory is compatible with the new schema except for three new columns:sd,function_overlap,confidence. Back-fill these as blanks; new rows populate them. - The
darwin-skillbranch convention (auto-optimize/YYYYMMDD-HHMM) can be renamed toevolve/YYYYMMDD-HHMMor run alongside without conflict.
Since your old test-prompts.json is compromised (seen during prior optimization), you need to write new holdout prompts — ones that the existing optimized skill has never been evaluated against.
Rule of thumb: a holdout prompt should be a task you'd actually want the skill to do, drawn from a scenario different enough from the training prompts that memorizing the training prompts wouldn't help.
Your current darwin-skill baseline scores were produced without measurement stability checks. They are point estimates of uncertain reliability.
Before continuing to optimize, run Phase 2 (three-run baseline) on each skill you plan to keep working on. If Gate 1 fails, do not continue optimizing that skill until you tighten its anchors — past "improvements" on that skill may have been noise-driven.
Before your next optimization round, ask an independent sub-agent:
"Read this SKILL.md and list the 5 core functions this skill provides, one per line, each as a short noun phrase."
Save the output to <skill_dir>/functions-pre.txt. This is the reference list for Gate 3.
Your Darwin runs probably averaged +8 to +12 points per skill across a full run. evolve-skill will likely report smaller per-round gains because it rejects sub-noise-floor changes that Darwin would have kept.
This is not regression. The real improvement — measured on the holdout set — is what matters, and it should be more consistent with evolve-skill.
Expect Gate 1 to fail on approximately 10–30% of skills on the first attempt. This is an honest signal about rubric calibration, not a problem with the skill.
For each Gate 1 failure, the worst-SD dimension in calibrate.py's output is the place to tighten the rubric anchor. After tightening, retry the three-run baseline.
Every experiment, including rejected ones, lives on a branch. After a few runs you will have many evolve/YYYYMMDD-HHMM/<skill>/exp-N branches. This is intentional — they are training data.
You can prune old branches by date (e.g. "drop branches older than 90 days") but do not prune them eagerly. A year of preserved experiments across all your skills is surprisingly useful when you are debugging why a pattern stopped working.
- Install side-by-side. Put
evolve-skill/in.claude/skills/alongsidedarwin-skill/. - Pick one skill that you've already optimized with Darwin and feel confident about. Its current score is your starting point.
- Write two holdout prompts for that skill. Store them in its
test-prompts.jsonunder theholdoutkey. - Run a three-run baseline (
calibrate.py) on that skill. - Check Gate 1. If SD ≤ 2.0, proceed. If not, open
anchored-rubric.mdand tighten the worst-SD dimension's anchors; retry. - Run one Phase 3 round. Observe which gates fire.
- Expand to more skills only after you are comfortable with the pace and gate behavior on the first one.
results.tsvfrom Darwin can be appended toevolve-skill'sresults.tsvwith the new columns blank. Sort by timestamp.playbook.mdis a new concept and starts empty. Do not back-fill it from Darwin's log — patterns need observedΔ ≥ 5evidence to qualify.- Result card templates from Darwin are visually different but conceptually identical. You can keep using Darwin's cards if you prefer them.
Yes, but not on the same skill at the same time. Pick one optimizer per skill per branch. They will fight for the baseline otherwise.
A reasonable split: use Darwin on experimental skills you're still shaping (fast iteration, loose rubric), and evolve-skill on production skills where drift and overfitting matter.