Flag unresolved Layer-0 grades, fix silent excepts, gitleaks allowlist - #11
Merged
Merged
Conversation
…s allowlist - ceis.grade_turn: a sub-threshold Layer 0 classification accepted because no judge is available is now marked unresolved=True (with a warning) instead of passing as a definitive detection_layer=0 grade; grade_corpus and grade_corpus_multi_trial count these and expose unresolved_turns in grading_metadata, warning when nonzero. - grader._compute_rubric_version: rubric 'unknown' is logged with its cause instead of a bare 'except Exception: pass' (ruff S110). - dashboard lift table: bind the loop's baseline as a default argument so the helper cannot close over a later row's value (ruff B023). - .gitleaks.toml: allowlist results/ and seeds_generated/ artifact paths and the judge_api_request_id field shape (provider message IDs, not credentials) so history scans stop flooding 10k+ false positives while code and config remain fully scanned. Tests cover unresolved marking for sub-threshold/no-judge, definitive Layer 0, judge-resolved turns, and the corpus-level unresolved count. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01142EwVwkXFRGnYu7YgHSYL
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Polish items from the readiness audit (P2-4, P2-6 top hits, P2-7):
Unresolved Layer-0 grades are flagged, never silently definitive (P2-4).
grade_turnaccepted a sub-threshold (< 0.8 confidence) pattern classification as the final grade whenever no judge was configured — indistinguishable in the output from a confident grade, despite the documented 30-40% pattern FN rate. Such turns now carryunresolved=Truewith a warning, and both corpus graders expose anunresolved_turnscount ingrading_metadata(warning when nonzero).Two real ruff S/B hits (P2-6).
grader._compute_rubric_versionswallowed every exception and silently reported rubric "unknown" (S110) — it now logs the cause. The dashboard lift-table helper closed over the loop'sbasevariable (B023) — the baseline is now bound as a default argument.gitleaks allowlist (P2-7). History scans flooded 10,519
generic-api-keyfindings, all fromjudge_api_request_idfields (provider message IDs) and model prose inside committed eval artifacts..gitleaks.tomlallowlistsresults/andseeds_generated/paths plus the request-id field shape; code and config remain fully scanned.Tests
tests/test_ceis_unresolved.py: sub-threshold/no-judge is unresolved, definitive Layer 0 and judge-resolved turns are not, and the corpus-levelunresolved_turnscount surfaces ingrading_metadata. 1,178 tests pass; ruff clean.🤖 Generated with Claude Code
https://claude.ai/code/session_01142EwVwkXFRGnYu7YgHSYL