Problem (capability gap)
Skill value is currently asserted via description and popularity, not measured. There is no baseline/repeated-run evidence of a skill's effect on quality, cost, or trigger accuracy.
Acceptance direction
- An evaluation harness with: skill + baseline digests, assertions frozen before execution, model/harness/repository/config identity, isolation mode, real token/time/tool/error metrics (or explicit unknown), repeated runs + aggregation, held-out near-miss trigger tests, and separate human feedback.
- Kept as a separate suite from the governance eval; reuse governance-eval primitives only where semantics match.
Dependencies
Depends on stable task identities and artifact boundaries - backlog #77 and #78.
Scope guards / non-goals
- Must not absorb or rewrite the existing governance eval.
- Popularity, curation, "official" labels, or author benchmarks are not effectiveness evidence.
Surfaced by the 2026-06-19 external skill/workflow-practices research. Recording phase only. Tracked as backlog #79.
Problem (capability gap)
Skill value is currently asserted via description and popularity, not measured. There is no baseline/repeated-run evidence of a skill's effect on quality, cost, or trigger accuracy.
Acceptance direction
Dependencies
Depends on stable task identities and artifact boundaries - backlog #77 and #78.
Scope guards / non-goals
Surfaced by the 2026-06-19 external skill/workflow-practices research. Recording phase only. Tracked as backlog #79.