Skip to content

Commit 0715857

Browse files
feat(mt#4970): Exclude log-only-family records from the injected fire count
## Summary `injectedFiresSinceLastReview` is the count the review thresholds key off, and its whole purpose (mt#3197) is to exclude detections the operator never saw. It computed that from `suppressionReasons` alone. That is one of two ways a fire fails to reach the operator. The other is **per-FAMILY log-only gating inside a LIVE detector**: `turn-end-untaken-action-scan.ts` declares `LOG_ONLY_FAMILIES`, whose matches reach the calibration record and never `additionalContext`. Those records carry no suppression reason — nothing suppressed them; they were never eligible — and are otherwise **byte-identical** to an injecting family's. So `untaken-action`'s 2026-09-04 window reported **120 injected fires where 23 reached the agent**. Every FP rate a reviewer computed over that denominator was 5x off, and the threshold that declared the log review-due fired on a population 81% invisible to the agent it was measuring. ## Key changes - **Writer declares, sweep reads.** `toCalibrationMatches` marks each match whose family is log-only; the sweep excludes it via `isLogOnlyFamilyRecord` and reports it under `logOnlyFamilySinceLastReview`. A per-detector family table on the *reading* side is the drift hazard mt#4465 recorded for `judgedText` — two places would have to agree about `LOG_ONLY_FAMILIES` forever, and only one is edited when an arm flips. - **`every`, not `some`.** A record carrying a log-only match *alongside* an injecting one DID reach the agent. 8 of the live window's 23 injected records are that shape; a `some` test would have swung the count wrong the other way. - **`allWithheld` closes the cliff this exclusion opens.** See below. - **Counted, not dropped.** The log-only arm's fire volume is exactly the evidence a decision to promote it would rest on. ### The cliff, and why it needed closing in the same PR mt#4049 shipped `allSuppressed` days ago because keying cadence off the injected count inverts for a detector suppressing *all* its fires — the count pins at 0 and no count-bearing leg can fire. Adding this exclusion opens the identical cliff through a **third column**: an all-log-only detector has `injectedFires === 0` **and** `suppressedSinceLastReview === 0`, so `atCountThreshold` is false, `allSuppressed` is false, and the log falls out of review entirely. `allWithheld` (`injected === 0 && suppressed + logOnly >= FIRES_THRESHOLD`) **replaces** `allSuppressed` in the routing and equals it wherever no log-only family exists — every detector but `untaken-action` today, asserted as a test. `evaluatedOnly` is deliberately **outside** the union: those records never matched, so nothing is there to classify and their absence from review is correct rather than a gap. ## Testing Execution evidence: ``` $ bun test --preload ./tests/setup.ts ./src/domain/calibration/calibration-sweep.log-only.test.ts 13 pass 0 fail 32 expect() calls ``` **AT2** (`a mixed record counts as injected, not log-only`), **AT3** (`a detector with no log-only families reports an unchanged count`, which also asserts `allWithheld === allSuppressed`), **AT4** (`false when the marker is ABSENT`) and **AT5** (`an all-log-only log at volume stays routable and keeps its records`) all pass in that run, named by number. `calibration-sweep.ts` and the untaken-action hook are both widely imported, so the full related-test closure was run: ``` $ bun scripts/run-related-tests.ts src/domain/calibration/calibration-sweep.ts \ .minsky/hooks/turn-end-untaken-action-scan.ts \ src/domain/calibration/calibration-sweep.log-only.test.ts scripts/verify-log-only-exclusion.ts 3631 pass 0 fail Ran 3631 tests across 93 files. [27.41s] ``` **AT1** — the live corpus, via the shipped predicate: ``` $ bun scripts/verify-log-only-exclusion.ts --log <state-dir>/untaken-action-calibration.jsonl window: first 190 of 217 records as-is (pre-declaration) injected= 120 logOnly= 0 suppressed= 70 evalOnly= 0 with declaration applied injected= 23 logOnly= 97 suppressed= 70 evalOnly= 0 OK: 97 records move from injected to log-only once the writer's declaration is present (120 -> 23); the columns balance, and the as-is pass is unchanged at 120 because an absent marker is not a declaration (AT4). ``` Guard canaries: `[PASS] turn-end-untaken-action-scan (registry, expects=warn)` — 74 total, 0 failed. Typecheck: 0 errors across 8 projects. Lint: 0 errors, 0 warnings. Negative control — removing `!isLogOnlyFamilyRecord(r)` from the injected filter, restoring the pre-fix count exactly: ``` (fail) AT1: log-only records are counted under their own name, not as injected (fail) AT2 end-to-end: a mixed record counts as injected, not log-only (fail) a SUPPRESSED log-only record counts as suppressed, not log-only (fail) AT5: an all-log-only log at volume stays routable and keeps its records (fail) suppressed and log-only volume COMBINE to clear the bar 8 pass 5 fail ``` and the live artifact fails diagnostically rather than silently: ``` with declaration applied injected= 120 logOnly= 97 FAIL: 0 records left the injected column but 97 arrived in the log-only column. The two must balance — a record may not be dropped from both or counted in both. ``` Reverted; the passing runs above are on the shipped tree. ## Deploy verification `src/domain/calibration/calibration-sweep.ts` and its two test files return **true** from `isDeploySurfaceFile` (predicate run over the full changed set before writing this — the other 8 files return false, so no `[no-deploy-impact]` tag is claimed). Post-merge I will run `deployment_wait-for-latest` for `minsky-mcp` with `notBefore` set to the merge timestamp and `expectCommitSha` set to the merge SHA, and read `buildIdentity` rather than the status alone. ## Spec-decision reconciliation Two calls, both recorded in the spec rather than left for review: **SC3 vs AT4 — a genuine conflict, resolved in AT4's favour.** SC3 asks that the 2026-09-04 window report 23. AT4 requires a pre-declaration record to count as injected so historical FP rates do not shift under a reader. Every record in that window predates the declaration, so both cannot hold of the same run. The sweep treats an absent marker as "not declared"; reclassifying 97 already-acked records would move an FP rate under anyone who had computed one. Practical impact is nil — the watermark was advanced to 215 on 2026-09-04, so that window is acked and will not be swept again. SC3's figure is instead produced by the verification artifact above. Read SC3 as *"a window whose records carry the declaration reports 23."* **SC5's "prefer generalizing over a parallel gate."** `allSuppressed` keeps its exact prior meaning and name; the generalization is `allWithheld`, which replaces it in the routing rather than sitting beside it. Widening `allSuppressed`'s own predicate was rejected: it would make `allSuppressed: true` for a log that suppressed nothing, which is a false statement to its six existing readers (including `scripts/verify-all-suppressed-routing.ts` from mt#4049) rather than a cosmetic naming issue. ## Notes for review - Consumer enumeration was done at plan time and typecheck confirmed it: the only two sites needing changes were the fixture builders in `calibration-sweep.review-due.test.ts` and `calibration-review-cadence-detector.test.ts`, both of which now DERIVE the new fields the way they already derive `allSuppressed`, so no existing fixture changes behavior. - `logOnly` is added to `CONSUMED_MATCH_KEYS` so it cannot also ride in `detectorFields` — mem#827 records reviewers reading that sub-object as if it were the record, and this field is read by the counting path. - **mt#4318 does not subsume this** and vice-versa: that task surfaces a whole DETECTOR's `injection_enabled` posture; this is a per-FAMILY gate inside a detector that is live. `untaken-action` would read `injection_enabled: true` under mt#4318's fix and the 97 would still be miscounted. Both specs already reconcile each other; the skill note added here points at mt#4318 for the sibling case. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01TPNyYTreM4F7PRbvh4xA5b Co-Authored-By: minsky-ai[bot] <minsky-ai[bot]@users.noreply.github.com>
2 parents be66b76 + c32b418 commit 0715857

15 files changed

Lines changed: 822 additions & 19 deletions

‎.claude/hooks/calibration-review-cadence-detector.ts‎

Lines changed: 13 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -358,6 +358,19 @@ export function formatCadenceWarning(due: ReviewDueLog[]): string {
358358
`Judge the suppressions, not a false-positive rate: there are no fires to rate.`
359359
);
360360
}
361+
// mt#4970: the same shape one column over, and it needs its own sentence for
362+
// the same reason the leg above does. Saying "suppressed N" here would be
363+
// false — nothing was suppressed; these matches were never eligible to
364+
// inject — and the number to judge is the log-only volume, not a
365+
// suppression count that is 0 on this leg by construction.
366+
if (d.reason === "all-withheld") {
367+
return (
368+
` - ${d.name}: matched ${d.logOnlyFamilySinceLastReview} of ` +
369+
`${d.firesSinceLastReview} detection(s) in LOG-ONLY families and injected NONE — the ` +
370+
`detector is running and the operator has seen nothing. Is this arm too broad? ` +
371+
`Judge the matches, not a false-positive rate: there are no fires to rate.`
372+
);
373+
}
361374
const reasonLabel =
362375
d.reason === "past-threshold"
363376
? "past review threshold (fires + diversity)"

‎.claude/hooks/turn-end-untaken-action-scan.ts‎

Lines changed: 37 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -1124,11 +1124,45 @@ export const STRANDED_TASK_FAMILY = "stranded-task-state";
11241124
* A family is added here by a DELIBERATE edit, so promoting an arm to injecting
11251125
* is a one-line removal that shows up in review rather than an emergent effect.
11261126
*/
1127-
const LOG_ONLY_FAMILIES: ReadonlySet<string> = new Set([
1127+
export const LOG_ONLY_FAMILIES: ReadonlySet<string> = new Set([
11281128
STRANDED_TASK_FAMILY,
11291129
PRESENT_PROGRESSIVE_FAMILY,
11301130
]);
11311131

1132+
/**
1133+
* Project matches into the calibration record, DECLARING which of them could
1134+
* never have injected (mt#4970).
1135+
*
1136+
* The sweep's `injectedFiresSinceLastReview` excludes what the operator never
1137+
* saw, and it derived that from `suppressionReasons` alone. A log-only family
1138+
* is the other way a fire fails to reach the operator, and it leaves no
1139+
* suppression reason — nothing suppressed it; it was never eligible. So the
1140+
* sweep counted 120 injected fires in a window where 23 reached the agent.
1141+
*
1142+
* The fact is declared HERE, by the writer, rather than inferred there. A
1143+
* per-detector family table in the sweep is the drift hazard mt#4465 recorded
1144+
* for `judgedText`: two places would have to agree about
1145+
* {@link LOG_ONLY_FAMILIES} forever, and only one of them is edited when an arm
1146+
* flips. Marking the match means promoting an arm to injecting stays the
1147+
* one-line removal from that set which this module already documents.
1148+
*
1149+
* `logOnly` is written ONLY when true — never `false`. Absent means "this
1150+
* writer does not declare the fact", which the sweep must treat as the status
1151+
* quo (counted as injected), exactly as `deferralOverlap`'s docblock in
1152+
* `calibration-sweep.ts` requires for the same reason: a projection that
1153+
* defaulted the field would manufacture a measurement for every record written
1154+
* before this change.
1155+
*/
1156+
function toCalibrationMatches(
1157+
matches: readonly { family: string; matchedPhrase: string }[]
1158+
): Array<{ family: string; phrase: string; logOnly?: true }> {
1159+
return matches.map((m) => ({
1160+
family: m.family,
1161+
phrase: m.matchedPhrase,
1162+
...(LOG_ONLY_FAMILIES.has(m.family) ? { logOnly: true as const } : {}),
1163+
}));
1164+
}
1165+
11321166
/**
11331167
* Calls whose RESULT carries a task's current status.
11341168
*
@@ -1511,7 +1545,7 @@ export function run(
15111545
timestamp: new Date().toISOString(),
15121546
session_id: input.session_id,
15131547
stop_hook_active: input.stop_hook_active === true,
1514-
matches: newMatches.map((m) => ({ family: m.family, phrase: m.matchedPhrase })),
1548+
matches: toCalibrationMatches(newMatches),
15151549
final_message_tail: finalMessage.slice(-TAIL_WINDOW_CHARS),
15161550
deferralOverlap: detectDeferralPhrases(finalMessage).length > 0,
15171551
suppressionReasons: [reason],
@@ -1644,7 +1678,7 @@ export function run(
16441678
timestamp: new Date().toISOString(),
16451679
session_id: input.session_id,
16461680
stop_hook_active: input.stop_hook_active === true,
1647-
matches: newMatches.map((m) => ({ family: m.family, phrase: m.matchedPhrase })),
1681+
matches: toCalibrationMatches(newMatches),
16481682
final_message_tail: finalMessage.slice(-TAIL_WINDOW_CHARS),
16491683
// Retained (mt#3620) so the overlap rate stays measurable across the
16501684
// direction change — but it no longer suppresses, so it is reported

‎.claude/skills/calibration-review/SKILL.md‎

Lines changed: 23 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -240,11 +240,33 @@ rather than recording "cannot classify" from the empty array — the
240240
say `classifiable`. See §"Cannot classify" is a claim about the corpus.
241241

242242
**An `all-suppressed` log is the exception: it DOES carry its records (mt#4049).**
243-
That gate is `atCountThreshold || allSuppressed`, because routing a log for
243+
That gate is `atCountThreshold || allWithheld`, because routing a log for
244244
review and then handing the reviewer an empty array would be a warning with no
245245
evidence attached. So read `newRecords` directly for this leg rather than
246246
reaching for the raw JSONL.
247247

248+
**`allWithheld` generalizes `allSuppressed` (mt#4970), and the two are equal for
249+
every log without a log-only family.** It fires when the injected count is 0 and
250+
the volume the operator never saw — SUPPRESSED plus LOG-ONLY-family — clears the
251+
same bar. Both columns mean "the record matched and nobody saw it", which is what
252+
keeps "is this gate too broad?" answerable. `evaluatedOnlySinceLastReview` is
253+
deliberately outside the union: those records never matched, so there is nothing
254+
to classify and their absence from review is correct.
255+
256+
**Bound the FP rate by `injectedFiresSinceLastReview`, and read
257+
`logOnlyFamilySinceLastReview` beside it (mt#4970).** A per-FAMILY log-only arm
258+
inside a LIVE detector produces records byte-identical to injecting ones — the
259+
gating lives in the hook's own family set, not in the record — so before mt#4970
260+
`untaken-action`'s 2026-09-04 window reported **120 injected where 23 reached the
261+
agent**, and any FP rate computed over that denominator was 5x off. The two
262+
counts now sit side by side: divide by the injected one, and treat the log-only
263+
one as what it is — the fire volume a decision to promote that arm would rest on,
264+
not warnings anyone received.
265+
266+
Note this is the per-FAMILY case. The per-DETECTOR sibling — a whole detector
267+
quieted via `injection_enabled` — is still unsurfaced and owned by mt#4318; a
268+
log can be affected by one, the other, or both.
269+
248270
### Step 1b — Coverage-receipt check (mt#2554)
249271

250272
Independent of the past-threshold gate above (a DEAD detector fires rarely, so it

‎.minsky/hooks/calibration-review-cadence-detector.test.ts‎

Lines changed: 27 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -123,6 +123,17 @@ function makeResult(
123123
((overrides.injectedFiresSinceLastReview ??
124124
merged.firesSinceLastReview - merged.suppressedSinceLastReview) === 0 &&
125125
merged.suppressedSinceLastReview >= FIRES_THRESHOLD),
126+
// mt#4970: defaults to 0 — a fixture that names no log-only family has none.
127+
logOnlyFamilySinceLastReview: overrides.logOnlyFamilySinceLastReview ?? 0,
128+
// mt#4970: DERIVED, same reasoning as the three above. With the 0 default it
129+
// reduces to `allSuppressed` for every existing fixture, so no prior case
130+
// changes behavior.
131+
allWithheld:
132+
overrides.allWithheld ??
133+
((overrides.injectedFiresSinceLastReview ??
134+
merged.firesSinceLastReview - merged.suppressedSinceLastReview) === 0 &&
135+
merged.suppressedSinceLastReview + (overrides.logOnlyFamilySinceLastReview ?? 0) >=
136+
FIRES_THRESHOLD),
126137
};
127138
}
128139

@@ -217,6 +228,7 @@ describe("shouldReWarn", () => {
217228
firesSinceLastReview: 43,
218229
injectedFiresSinceLastReview: 43,
219230
suppressedSinceLastReview: 0,
231+
logOnlyFamilySinceLastReview: 0,
220232
totalFires: 43,
221233
distinctPhrases: 31,
222234
reason: "past-threshold",
@@ -271,6 +283,7 @@ describe("shouldReWarn — policy-coverage kind (mt#2659)", () => {
271283
firesSinceLastReview: 1457,
272284
injectedFiresSinceLastReview: 1457,
273285
suppressedSinceLastReview: 0,
286+
logOnlyFamilySinceLastReview: 0,
274287
totalFires: 1457,
275288
distinctPhrases: 5,
276289
reason: "past-threshold",
@@ -342,6 +355,7 @@ describe("suppression-aware review-due legs (mt#3197, PR #2300 R1)", () => {
342355
makeResult(entry, {
343356
firesSinceLastReview: 12,
344357
suppressedSinceLastReview: 11,
358+
logOnlyFamilySinceLastReview: 0,
345359
totalFires: 40,
346360
}),
347361
];
@@ -394,6 +408,7 @@ describe("all-suppressed review-due leg (mt#4049)", () => {
394408
makeResult(entry, {
395409
firesSinceLastReview: 12,
396410
suppressedSinceLastReview: 12,
411+
logOnlyFamilySinceLastReview: 0,
397412
totalFires: 12,
398413
}),
399414
];
@@ -415,6 +430,7 @@ describe("all-suppressed review-due leg (mt#4049)", () => {
415430
makeResult(entry, {
416431
firesSinceLastReview: 12,
417432
suppressedSinceLastReview: 12,
433+
logOnlyFamilySinceLastReview: 0,
418434
totalFires: 40,
419435
}),
420436
];
@@ -436,6 +452,7 @@ describe("all-suppressed review-due leg (mt#4049)", () => {
436452
makeResult(entry, {
437453
firesSinceLastReview: 12,
438454
suppressedSinceLastReview: 11,
455+
logOnlyFamilySinceLastReview: 0,
439456
totalFires: 12,
440457
}),
441458
];
@@ -489,6 +506,7 @@ describe("all-suppressed review-due leg (mt#4049)", () => {
489506
makeResult(entry, {
490507
firesSinceLastReview: 12,
491508
suppressedSinceLastReview: 12,
509+
logOnlyFamilySinceLastReview: 0,
492510
totalFires: 12,
493511
distinctPhrases: 1,
494512
}),
@@ -503,6 +521,7 @@ describe("all-suppressed review-due leg (mt#4049)", () => {
503521
makeResult(entry, {
504522
firesSinceLastReview: 12,
505523
suppressedSinceLastReview: 12,
524+
logOnlyFamilySinceLastReview: 0,
506525
totalFires: 12,
507526
watermarkCount: 99,
508527
}),
@@ -522,6 +541,7 @@ describe("formatCadenceWarning — the all-suppressed line (mt#4049)", () => {
522541
firesSinceLastReview: 15,
523542
injectedFiresSinceLastReview: 0,
524543
suppressedSinceLastReview: 13,
544+
logOnlyFamilySinceLastReview: 0,
525545
totalFires: 15,
526546
distinctPhrases: 13,
527547
reason: "all-suppressed",
@@ -553,6 +573,7 @@ describe("formatCadenceWarning", () => {
553573
// mt#3197: no suppression outcome recorded -> every fire is injected.
554574
injectedFiresSinceLastReview: 43,
555575
suppressedSinceLastReview: 0,
576+
logOnlyFamilySinceLastReview: 0,
556577
totalFires: 43,
557578
distinctPhrases: 31,
558579
reason: "past-threshold",
@@ -565,6 +586,7 @@ describe("formatCadenceWarning", () => {
565586
firesSinceLastReview: 8,
566587
injectedFiresSinceLastReview: 8,
567588
suppressedSinceLastReview: 0,
589+
logOnlyFamilySinceLastReview: 0,
568590
totalFires: 20,
569591
distinctPhrases: 3,
570592
reason: "time-stale",
@@ -599,6 +621,7 @@ describe("formatCadenceWarning", () => {
599621
firesSinceLastReview: 0,
600622
injectedFiresSinceLastReview: 0,
601623
suppressedSinceLastReview: 0,
624+
logOnlyFamilySinceLastReview: 0,
602625
totalFires: 121,
603626
distinctPhrases: 0,
604627
reason: "watermark-stranded",
@@ -628,6 +651,7 @@ describe("formatCadenceWarning", () => {
628651
firesSinceLastReview: 1,
629652
injectedFiresSinceLastReview: 1,
630653
suppressedSinceLastReview: 0,
654+
logOnlyFamilySinceLastReview: 0,
631655
totalFires: 1,
632656
distinctPhrases: 1,
633657
reason: "never-reviewed",
@@ -665,6 +689,7 @@ describe("formatCadenceWarning", () => {
665689
firesSinceLastReview: 999,
666690
injectedFiresSinceLastReview: 999,
667691
suppressedSinceLastReview: 999,
692+
logOnlyFamilySinceLastReview: 0,
668693
totalFires: 9999,
669694
distinctPhrases: 999,
670695
reason: "never-fired",
@@ -794,6 +819,7 @@ describe("selectPendingAskLogs", () => {
794819
firesSinceLastReview: 20,
795820
injectedFiresSinceLastReview: 20,
796821
suppressedSinceLastReview: 0,
822+
logOnlyFamilySinceLastReview: 0,
797823
totalFires: 1477,
798824
distinctPhrases: 5,
799825
reason: "past-threshold",
@@ -854,6 +880,7 @@ describe("formatPendingAskLines", () => {
854880
firesSinceLastReview: 20,
855881
injectedFiresSinceLastReview: 20,
856882
suppressedSinceLastReview: 0,
883+
logOnlyFamilySinceLastReview: 0,
857884
totalFires: 1477,
858885
distinctPhrases: 5,
859886
reason: "past-threshold",

‎.minsky/hooks/calibration-review-cadence-detector.ts‎

Lines changed: 13 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -355,6 +355,19 @@ export function formatCadenceWarning(due: ReviewDueLog[]): string {
355355
`Judge the suppressions, not a false-positive rate: there are no fires to rate.`
356356
);
357357
}
358+
// mt#4970: the same shape one column over, and it needs its own sentence for
359+
// the same reason the leg above does. Saying "suppressed N" here would be
360+
// false — nothing was suppressed; these matches were never eligible to
361+
// inject — and the number to judge is the log-only volume, not a
362+
// suppression count that is 0 on this leg by construction.
363+
if (d.reason === "all-withheld") {
364+
return (
365+
` - ${d.name}: matched ${d.logOnlyFamilySinceLastReview} of ` +
366+
`${d.firesSinceLastReview} detection(s) in LOG-ONLY families and injected NONE — the ` +
367+
`detector is running and the operator has seen nothing. Is this arm too broad? ` +
368+
`Judge the matches, not a false-positive rate: there are no fires to rate.`
369+
);
370+
}
358371
const reasonLabel =
359372
d.reason === "past-threshold"
360373
? "past review threshold (fires + diversity)"

‎.minsky/hooks/turn-end-untaken-action-scan.ts‎

Lines changed: 37 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -1121,11 +1121,45 @@ export const STRANDED_TASK_FAMILY = "stranded-task-state";
11211121
* A family is added here by a DELIBERATE edit, so promoting an arm to injecting
11221122
* is a one-line removal that shows up in review rather than an emergent effect.
11231123
*/
1124-
const LOG_ONLY_FAMILIES: ReadonlySet<string> = new Set([
1124+
export const LOG_ONLY_FAMILIES: ReadonlySet<string> = new Set([
11251125
STRANDED_TASK_FAMILY,
11261126
PRESENT_PROGRESSIVE_FAMILY,
11271127
]);
11281128

1129+
/**
1130+
* Project matches into the calibration record, DECLARING which of them could
1131+
* never have injected (mt#4970).
1132+
*
1133+
* The sweep's `injectedFiresSinceLastReview` excludes what the operator never
1134+
* saw, and it derived that from `suppressionReasons` alone. A log-only family
1135+
* is the other way a fire fails to reach the operator, and it leaves no
1136+
* suppression reason — nothing suppressed it; it was never eligible. So the
1137+
* sweep counted 120 injected fires in a window where 23 reached the agent.
1138+
*
1139+
* The fact is declared HERE, by the writer, rather than inferred there. A
1140+
* per-detector family table in the sweep is the drift hazard mt#4465 recorded
1141+
* for `judgedText`: two places would have to agree about
1142+
* {@link LOG_ONLY_FAMILIES} forever, and only one of them is edited when an arm
1143+
* flips. Marking the match means promoting an arm to injecting stays the
1144+
* one-line removal from that set which this module already documents.
1145+
*
1146+
* `logOnly` is written ONLY when true — never `false`. Absent means "this
1147+
* writer does not declare the fact", which the sweep must treat as the status
1148+
* quo (counted as injected), exactly as `deferralOverlap`'s docblock in
1149+
* `calibration-sweep.ts` requires for the same reason: a projection that
1150+
* defaulted the field would manufacture a measurement for every record written
1151+
* before this change.
1152+
*/
1153+
function toCalibrationMatches(
1154+
matches: readonly { family: string; matchedPhrase: string }[]
1155+
): Array<{ family: string; phrase: string; logOnly?: true }> {
1156+
return matches.map((m) => ({
1157+
family: m.family,
1158+
phrase: m.matchedPhrase,
1159+
...(LOG_ONLY_FAMILIES.has(m.family) ? { logOnly: true as const } : {}),
1160+
}));
1161+
}
1162+
11291163
/**
11301164
* Calls whose RESULT carries a task's current status.
11311165
*
@@ -1508,7 +1542,7 @@ export function run(
15081542
timestamp: new Date().toISOString(),
15091543
session_id: input.session_id,
15101544
stop_hook_active: input.stop_hook_active === true,
1511-
matches: newMatches.map((m) => ({ family: m.family, phrase: m.matchedPhrase })),
1545+
matches: toCalibrationMatches(newMatches),
15121546
final_message_tail: finalMessage.slice(-TAIL_WINDOW_CHARS),
15131547
deferralOverlap: detectDeferralPhrases(finalMessage).length > 0,
15141548
suppressionReasons: [reason],
@@ -1641,7 +1675,7 @@ export function run(
16411675
timestamp: new Date().toISOString(),
16421676
session_id: input.session_id,
16431677
stop_hook_active: input.stop_hook_active === true,
1644-
matches: newMatches.map((m) => ({ family: m.family, phrase: m.matchedPhrase })),
1678+
matches: toCalibrationMatches(newMatches),
16451679
final_message_tail: finalMessage.slice(-TAIL_WINDOW_CHARS),
16461680
// Retained (mt#3620) so the overlap rate stays measurable across the
16471681
// direction change — but it no longer suppresses, so it is reported

‎.minsky/skills/calibration-review/SKILL.md‎

Lines changed: 23 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -238,11 +238,33 @@ rather than recording "cannot classify" from the empty array — the
238238
say `classifiable`. See §"Cannot classify" is a claim about the corpus.
239239

240240
**An `all-suppressed` log is the exception: it DOES carry its records (mt#4049).**
241-
That gate is `atCountThreshold || allSuppressed`, because routing a log for
241+
That gate is `atCountThreshold || allWithheld`, because routing a log for
242242
review and then handing the reviewer an empty array would be a warning with no
243243
evidence attached. So read `newRecords` directly for this leg rather than
244244
reaching for the raw JSONL.
245245

246+
**`allWithheld` generalizes `allSuppressed` (mt#4970), and the two are equal for
247+
every log without a log-only family.** It fires when the injected count is 0 and
248+
the volume the operator never saw — SUPPRESSED plus LOG-ONLY-family — clears the
249+
same bar. Both columns mean "the record matched and nobody saw it", which is what
250+
keeps "is this gate too broad?" answerable. `evaluatedOnlySinceLastReview` is
251+
deliberately outside the union: those records never matched, so there is nothing
252+
to classify and their absence from review is correct.
253+
254+
**Bound the FP rate by `injectedFiresSinceLastReview`, and read
255+
`logOnlyFamilySinceLastReview` beside it (mt#4970).** A per-FAMILY log-only arm
256+
inside a LIVE detector produces records byte-identical to injecting ones — the
257+
gating lives in the hook's own family set, not in the record — so before mt#4970
258+
`untaken-action`'s 2026-09-04 window reported **120 injected where 23 reached the
259+
agent**, and any FP rate computed over that denominator was 5x off. The two
260+
counts now sit side by side: divide by the injected one, and treat the log-only
261+
one as what it is — the fire volume a decision to promote that arm would rest on,
262+
not warnings anyone received.
263+
264+
Note this is the per-FAMILY case. The per-DETECTOR sibling — a whole detector
265+
quieted via `injection_enabled` — is still unsurfaced and owned by mt#4318; a
266+
log can be affected by one, the other, or both.
267+
246268
### Step 1b — Coverage-receipt check (mt#2554)
247269

248270
Independent of the past-threshold gate above (a DEAD detector fires rarely, so it

0 commit comments

Comments
 (0)