Skip to content

Latest commit

 

History

History
executable file
·
73 lines (42 loc) · 4.7 KB

File metadata and controls

executable file
·
73 lines (42 loc) · 4.7 KB

Migrating from darwin-skill

If you have a darwin-skill installation with a populated results.tsv, this guide walks through a side-by-side migration. You do not have to abandon Darwin to try evolve-skill — the two can run on separate branches against the same skills.

What You Can Keep

  • Your existing test-prompts.json files become the training set in evolve-skill. They have been seen during prior optimization, so they cannot be used as holdouts.
  • Your results.tsv history is compatible with the new schema except for three new columns: sd, function_overlap, confidence. Back-fill these as blanks; new rows populate them.
  • The darwin-skill branch convention (auto-optimize/YYYYMMDD-HHMM) can be renamed to evolve/YYYYMMDD-HHMM or run alongside without conflict.

What You Need to Add

1. Holdout prompts

Since your old test-prompts.json is compromised (seen during prior optimization), you need to write new holdout prompts — ones that the existing optimized skill has never been evaluated against.

Rule of thumb: a holdout prompt should be a task you'd actually want the skill to do, drawn from a scenario different enough from the training prompts that memorizing the training prompts wouldn't help.

2. Re-calibrated baselines

Your current darwin-skill baseline scores were produced without measurement stability checks. They are point estimates of uncertain reliability.

Before continuing to optimize, run Phase 2 (three-run baseline) on each skill you plan to keep working on. If Gate 1 fails, do not continue optimizing that skill until you tighten its anchors — past "improvements" on that skill may have been noise-driven.

3. Function lists for each skill

Before your next optimization round, ask an independent sub-agent:

"Read this SKILL.md and list the 5 core functions this skill provides, one per line, each as a short noun phrase."

Save the output to <skill_dir>/functions-pre.txt. This is the reference list for Gate 3.

What You Should Expect to Change

Slower per-skill progress

Your Darwin runs probably averaged +8 to +12 points per skill across a full run. evolve-skill will likely report smaller per-round gains because it rejects sub-noise-floor changes that Darwin would have kept.

This is not regression. The real improvement — measured on the holdout set — is what matters, and it should be more consistent with evolve-skill.

Some skills refusing optimization

Expect Gate 1 to fail on approximately 10–30% of skills on the first attempt. This is an honest signal about rubric calibration, not a problem with the skill.

For each Gate 1 failure, the worst-SD dimension in calibrate.py's output is the place to tighten the rubric anchor. After tightening, retry the three-run baseline.

More branches in your repo

Every experiment, including rejected ones, lives on a branch. After a few runs you will have many evolve/YYYYMMDD-HHMM/<skill>/exp-N branches. This is intentional — they are training data.

You can prune old branches by date (e.g. "drop branches older than 90 days") but do not prune them eagerly. A year of preserved experiments across all your skills is surprisingly useful when you are debugging why a pattern stopped working.

A Recommended Migration Sequence

  1. Install side-by-side. Put evolve-skill/ in .claude/skills/ alongside darwin-skill/.
  2. Pick one skill that you've already optimized with Darwin and feel confident about. Its current score is your starting point.
  3. Write two holdout prompts for that skill. Store them in its test-prompts.json under the holdout key.
  4. Run a three-run baseline (calibrate.py) on that skill.
  5. Check Gate 1. If SD ≤ 2.0, proceed. If not, open anchored-rubric.md and tighten the worst-SD dimension's anchors; retry.
  6. Run one Phase 3 round. Observe which gates fire.
  7. Expand to more skills only after you are comfortable with the pace and gate behavior on the first one.

Compatibility Notes

  • results.tsv from Darwin can be appended to evolve-skill's results.tsv with the new columns blank. Sort by timestamp.
  • playbook.md is a new concept and starts empty. Do not back-fill it from Darwin's log — patterns need observed Δ ≥ 5 evidence to qualify.
  • Result card templates from Darwin are visually different but conceptually identical. You can keep using Darwin's cards if you prefer them.

Can I Just Run Both?

Yes, but not on the same skill at the same time. Pick one optimizer per skill per branch. They will fight for the baseline otherwise.

A reasonable split: use Darwin on experimental skills you're still shaping (fast iteration, loose rubric), and evolve-skill on production skills where drift and overfitting matter.