Skip to content

Add adaptive harness receipt eval - #11

Merged
uncfreak1255-code merged 2 commits into
mainfrom
codex/adaptive-harness-candidate
Jul 15, 2026
Merged

Add adaptive harness receipt eval#11
uncfreak1255-code merged 2 commits into
mainfrom
codex/adaptive-harness-candidate

Conversation

@uncfreak1255-code

Copy link
Copy Markdown
Owner

Summary

  • add an eval-only five-field adaptive-harness receipt for the DF-03, DF-08, and fuzzy-task failure shapes
  • add fail-closed custom-eval, overlay, strict-case, provenance, and frozen-score-floor support to the benchmark path
  • update cold-start continuity so future agents recover the rejected candidate and evidence hold instead of repeating this task

Decision

The focused three-sample run passed with +0.4250 weighted delta, 0.9750 candidate score, zero boundary violations, and all DF-03/DF-08 candidate samples strict-pass. The retained sealed run scored 0.9281, below the frozen v0.2.0 floor of 0.9673, so the behavior candidate is rejected and the shipped skill remains unchanged.

Proof

  • npm test — pass
  • npm run benchmark:adaptive-harness — pass; focused receipt at results/2026-07-15T00-11-39-008Z-gpt-5.5/
  • npm run benchmark:sealed — intentional failed decision gate; relative comparison passed, absolute 0.9281 < 0.9673
  • npm run benchmark:trajectories — pass, 6/6
  • npm run smoke:cold-start — pass
  • Codex autoreview — clean, no accepted/actionable findings

No global install, hook, permanent agent, skill-text change, deploy, or runtime mutation is included.

@uncfreak1255-code
uncfreak1255-code merged commit d5ac167 into main Jul 15, 2026
2 checks passed
@uncfreak1255-code
uncfreak1255-code deleted the codex/adaptive-harness-candidate branch July 15, 2026 00:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant