Skip to content

[Feature] Skill effectiveness evaluation harness #254

Description

@KbWen

Problem (capability gap)

Skill value is currently asserted via description and popularity, not measured. There is no baseline/repeated-run evidence of a skill's effect on quality, cost, or trigger accuracy.

Acceptance direction

  • An evaluation harness with: skill + baseline digests, assertions frozen before execution, model/harness/repository/config identity, isolation mode, real token/time/tool/error metrics (or explicit unknown), repeated runs + aggregation, held-out near-miss trigger tests, and separate human feedback.
  • Kept as a separate suite from the governance eval; reuse governance-eval primitives only where semantics match.

Dependencies

Depends on stable task identities and artifact boundaries - backlog #77 and #78.

Scope guards / non-goals

  • Must not absorb or rewrite the existing governance eval.
  • Popularity, curation, "official" labels, or author benchmarks are not effectiveness evidence.

Surfaced by the 2026-06-19 external skill/workflow-practices research. Recording phase only. Tracked as backlog #79.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestskillSkill system (new skill, skill fix, conflict)

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions