You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
- Noise/distractor cases (inject irrelevant context, skills, or tools — does the agent stay focused?)
236
237
237
238
**Boundary cases** — Inputs at the exact boundary:
238
239
@@ -243,7 +244,7 @@ Target ratio: roughly 40% positive, 40% negative, 20% boundary for each eval dim
243
244
244
245
## Phase 6: Select Graders and Define Logic
245
246
246
-
For each task, define the specific grading approach:
247
+
For each task, define the specific grading approach. **Prefer execution-based validation over pattern matching** — actually run the generated code or invoke the generated artifact and verify runtime behavior. A source file that contains the right patterns but doesn't wire them up correctly will pass a regex check but fail execution.
Copy file name to clipboardExpand all lines: plugin/skills/define-evals/references/eval-guide.md
+16-2Lines changed: 16 additions & 2 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -30,11 +30,11 @@ Each task should have a known-good solution that proves the task is solvable and
30
30
31
31
### 3. Grade Outcomes, Not Paths
32
32
33
-
Don't check exact tool call sequences. Agents can find creative but valid solutions that rigid path-checking would fail unfairly.
33
+
Don't check exact tool call sequences. Agents can find creative but valid solutions that rigid path-checking would fail unfairly. Prefer **execution-based validation** over pattern matching — actually run the generated code and verify runtime behavior rather than checking whether the source contains certain strings or patterns. A file that contains `checkpointer = MemorySaver()` but never wires it up will pass a pattern match but fail execution.
34
34
35
35
### 4. Build Balanced Problem Sets
36
36
37
-
Test both where a behavior SHOULD and SHOULD NOT occur. One-sided evals create one-sided optimization. For example, a search agent needs evals for queries requiring search AND queries it should answer from existing knowledge.
37
+
Test both where a behavior SHOULD and SHOULD NOT occur. One-sided evals create one-sided optimization. For example, a search agent needs evals for queries requiring search AND queries it should answer from existing knowledge. Include **noise/distractor cases**: inject irrelevant context, skills, or tools alongside the relevant ones and verify the agent stays focused. An agent that succeeds in a clean environment but fails when given distractors has a routing or attention problem.
38
38
39
39
### 5. Build In Partial Credit
40
40
@@ -173,6 +173,20 @@ Run the agent without the component under test (control), then with it (treatmen
173
173
174
174
Pass/fail metrics alone are insufficient for iterating on evals. Full trajectory visibility — what the agent read, wrote, invoked, and in what order — is required to diagnose *why* a task failed. Without observability, you know something broke but not whether the failure was in retrieval, routing, generation, or grading. Design eval harnesses to capture full interaction traces, not just final outcomes.
175
175
176
+
## Eval Architecture Patterns
177
+
178
+
### Separate Tasks from Treatments
179
+
180
+
Decouple *what the agent does* (the task) from *what context or skills the agent receives* (the treatment). A task defines the scenario, expected behavior, and validation logic. A treatment defines the skills, documentation, and configuration provided to the agent. When tasks and treatments are independent, any treatment can be applied to any task, enabling combinatorial testing: does adding skill X improve performance on tasks A, B, C? Does removing context Y cause regressions? This separation is what makes baseline comparison (control vs. treatment) practical at scale.
181
+
182
+
### Declarative Task Metadata
183
+
184
+
Define task properties (difficulty, category, timeout, target artifacts, validation scripts) in a structured config file rather than embedding them in test code. This makes tasks scannable, filterable, and composable without reading implementation details.
185
+
186
+
### Check Functions with Mandatory Verdicts
187
+
188
+
Design grader functions that *must* call `passed()` or `failed()` — not calling either is an error. This eliminates the common antipattern of returning ambiguous values or silently passing when a check was never actually run. Each check should be independently reportable with a descriptive name.
189
+
176
190
## Anti-Patterns to Avoid
177
191
178
192
1. **Vibe-based development**: No evals at all, shipping on intuition.
Copy file name to clipboardExpand all lines: plugin/skills/define-evals/references/sources.md
+1Lines changed: 1 addition & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -35,6 +35,7 @@ All sources were fetched and verified on April 22, 2026.
35
35
## Skill Evaluation
36
36
37
37
-[Evaluating Skills](https://www.langchain.com/blog/evaluating-skills) — Robert Xu, LangChain, March 2026. Skill invocation as first-class metric. Bug-fixing as superior eval paradigm (constrained tasks easier to grade). ~12 skill ceiling for reliable disambiguation. Baseline comparison methodology (control vs. treatment). Full trajectory observability required for iteration. Claude Code with skills: 82% task completion vs. 9% without.
0 commit comments