Skip to content

Commit 6e45df6

Browse files
Add eval architecture patterns from LangChain skills-benchmarks
Co-Authored-By: Claude Code <noreply@anthropic.com> Signed-off-by: Savitha Raghunathan <saveetha13@gmail.com>
1 parent 0622d3a commit 6e45df6

3 files changed

Lines changed: 19 additions & 3 deletions

File tree

‎plugin/skills/define-evals/SKILL.md‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -233,6 +233,7 @@ For each task, define:
233233
- Inputs that look similar but shouldn't trigger the behavior
234234
- Boundary cases just outside the expected scope
235235
- Adversarial inputs (prompt injection, out-of-scope requests)
236+
- Noise/distractor cases (inject irrelevant context, skills, or tools — does the agent stay focused?)
236237

237238
**Boundary cases** — Inputs at the exact boundary:
238239

@@ -243,7 +244,7 @@ Target ratio: roughly 40% positive, 40% negative, 20% boundary for each eval dim
243244

244245
## Phase 6: Select Graders and Define Logic
245246

246-
For each task, define the specific grading approach:
247+
For each task, define the specific grading approach. **Prefer execution-based validation over pattern matching** — actually run the generated code or invoke the generated artifact and verify runtime behavior. A source file that contains the right patterns but doesn't wire them up correctly will pass a regex check but fail execution.
247248

248249
### Code-Based Graders
249250

‎plugin/skills/define-evals/references/eval-guide.md‎

Lines changed: 16 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -30,11 +30,11 @@ Each task should have a known-good solution that proves the task is solvable and
3030
3131
### 3. Grade Outcomes, Not Paths
3232
33-
Don't check exact tool call sequences. Agents can find creative but valid solutions that rigid path-checking would fail unfairly.
33+
Don't check exact tool call sequences. Agents can find creative but valid solutions that rigid path-checking would fail unfairly. Prefer **execution-based validation** over pattern matching — actually run the generated code and verify runtime behavior rather than checking whether the source contains certain strings or patterns. A file that contains `checkpointer = MemorySaver()` but never wires it up will pass a pattern match but fail execution.
3434

3535
### 4. Build Balanced Problem Sets
3636

37-
Test both where a behavior SHOULD and SHOULD NOT occur. One-sided evals create one-sided optimization. For example, a search agent needs evals for queries requiring search AND queries it should answer from existing knowledge.
37+
Test both where a behavior SHOULD and SHOULD NOT occur. One-sided evals create one-sided optimization. For example, a search agent needs evals for queries requiring search AND queries it should answer from existing knowledge. Include **noise/distractor cases**: inject irrelevant context, skills, or tools alongside the relevant ones and verify the agent stays focused. An agent that succeeds in a clean environment but fails when given distractors has a routing or attention problem.
3838

3939
### 5. Build In Partial Credit
4040

@@ -173,6 +173,20 @@ Run the agent without the component under test (control), then with it (treatmen
173173

174174
Pass/fail metrics alone are insufficient for iterating on evals. Full trajectory visibility — what the agent read, wrote, invoked, and in what order — is required to diagnose *why* a task failed. Without observability, you know something broke but not whether the failure was in retrieval, routing, generation, or grading. Design eval harnesses to capture full interaction traces, not just final outcomes.
175175

176+
## Eval Architecture Patterns
177+
178+
### Separate Tasks from Treatments
179+
180+
Decouple *what the agent does* (the task) from *what context or skills the agent receives* (the treatment). A task defines the scenario, expected behavior, and validation logic. A treatment defines the skills, documentation, and configuration provided to the agent. When tasks and treatments are independent, any treatment can be applied to any task, enabling combinatorial testing: does adding skill X improve performance on tasks A, B, C? Does removing context Y cause regressions? This separation is what makes baseline comparison (control vs. treatment) practical at scale.
181+
182+
### Declarative Task Metadata
183+
184+
Define task properties (difficulty, category, timeout, target artifacts, validation scripts) in a structured config file rather than embedding them in test code. This makes tasks scannable, filterable, and composable without reading implementation details.
185+
186+
### Check Functions with Mandatory Verdicts
187+
188+
Design grader functions that *must* call `passed()` or `failed()` — not calling either is an error. This eliminates the common antipattern of returning ambiguous values or silently passing when a check was never actually run. Each check should be independently reportable with a descriptive name.
189+
176190
## Anti-Patterns to Avoid
177191

178192
1. **Vibe-based development**: No evals at all, shipping on intuition.

‎plugin/skills/define-evals/references/sources.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -35,6 +35,7 @@ All sources were fetched and verified on April 22, 2026.
3535
## Skill Evaluation
3636

3737
- [Evaluating Skills](https://www.langchain.com/blog/evaluating-skills) — Robert Xu, LangChain, March 2026. Skill invocation as first-class metric. Bug-fixing as superior eval paradigm (constrained tasks easier to grade). ~12 skill ceiling for reliable disambiguation. Baseline comparison methodology (control vs. treatment). Full trajectory observability required for iteration. Claude Code with skills: 82% task completion vs. 9% without.
38+
- [skills-benchmarks](https://github.com/langchain-ai/skills-benchmarks) — LangChain, 2026. Reference implementation: task/treatment separation, TOML-based task metadata, Docker isolation per trial, execution-based validation (run code, don't pattern match), noise/distractor skill testing, section-level skill A/B testing via XML tags.
3839

3940
## Secondary Sources (Search-Verified)
4041

0 commit comments

Comments
 (0)