Skip to content

Commit 39a9a46

Browse files
committed
fix: bound the effective threshold, not one of its two handles (#134)
Verification found that bounding `tolerance` had closed one handle of a two-handled lever and left the other free. The decision is `actual - golden > tol`, so the gate passes anything at or below `golden + tolerance` — that sum is the effective threshold, and the two fields enter it identically. A golden of 2.0 with an honest tolerance of 0.02 admitted every physically possible error rate, and the tolerance this same issue had just put on the passing line read a reassuring 0.0200, because the number that moved was not the one being rendered. Reproduced against the real committed baseline with one field edited. The bound is now on the sum (<= 0.5, against a committed maximum of 0.1749). It is the only bound invariant to trading one handle for the other; a per-field bound always leaves its partner as a substitute. `golden`'s own constant is deleted rather than kept. With the sum bounded it can never be the guard that fires, and mutation testing put it at zero failed assertions both before and after. A constant no test can distinguish from its absence is not a guard. Its stated justification was also wrong: "CER can exceed 1.0 via insertions" describes a measurement, which ERROR_RATE_MAX governs. A truncated comparison also still reported success. Closing the empty case closed cardinality 0, not 1-of-N: both sides derive from the same baseline.json, so deleting entries shrinks the baseline, the work list and the measured set together and the gate prints "all 1 corpora within tolerance" over a release that verified one twelfth of what it should have — by pure deletion, with no unusual value anywhere. The expected set is pinned in baseline-meta.json and, for the repo's own baseline, is now an INVARIANT of regression-gate.sh: a missing, unreadable or malformed meta is fatal there, an overridden BESTASR_BASELINE only warns. The compare stage still tolerates an absent anchor with a printed NOTE, because a stdin filter cannot know whether it was handed a whole sweep — the invariant belongs where the answer is knowable. That conditional also upgrades the #48 model-artifact pin, which was skipped SILENTLY when the same file was missing. A satisfied anchor now says so. Its only previous evidence was the absence of the warning, which asks a reader to already know the warning exists — the argument this issue makes for rendering the tolerance on passing lines, applied to the guard it just added. `error_rate: false` was the one boolean position still able to flip a verdict, and nothing tested it: float(false) is 0.0, a boolean laundered into a PERFECT SCORE rather than a bad one. Bounding golden at 0.5 had quietly made the existing boolean test vacuous, which is the general hazard — tightening one bound can hollow out another guard's only test. OverflowError escaped the numeric converter; a bare 400-digit integer parses to an arbitrary-precision int and float() raises neither TypeError nor ValueError. Non-zero exit, so never a bypass, but it reached the log as a stack trace where the verdict belongs. Rejected values are capped at 120 characters and corpus lists at 12 names, the anchor is shape-checked before use, it names the file it actually read rather than the default path, and the sum-bound message renders at 17 significant digits so a rejected value cannot print as equal to the threshold it exceeded. A second verify round found the anchor could still be switched off by omitting a key, and that three guards had no test able to tell them from their absence. The gate now REFUSES an unanchored run on the repo's own baseline (missing, unreadable, or corpora absent/null/not a unique non-empty string array); an overridden BESTASR_BASELINE warns instead. The compare stage keeps tolerating an absent anchor with a NOTE, because a stdin filter cannot know whether it was handed a whole sweep — the invariant belongs where the answer is knowable, and both halves now have a test so the permissive half is no longer the whole specification. Every guard is non-vacuous, measured by deleting each independently: sum bound 9, completeness 7, boolean 6, isfinite 5, TOLERANCE_MAX 3, OverflowError 3, ERROR_RATE_MAX 2, echo cap 2, and the gate's own anchor wiring 1 — a one-token typo in that shell heredoc used to revert the whole completeness guard with a green suite, because nothing ran regression-gate.sh. Those counts come from scripts/mutation-check.sh, committed alongside them. A count the reader cannot re-run is the same defect as a threshold the reader cannot see, one level out. The PR body has been regenerated from this tree. An earlier version of this commit message claimed the body had already been corrected when only the CHANGELOG had been; that claim was false when written and is the reason this message was amended rather than followed by another commit. 474 tests / 88 suites green (+29 over main). Refs #134
1 parent 7c84d28 commit 39a9a46

6 files changed

Lines changed: 826 additions & 32 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 94 additions & 11 deletions
Original file line numberDiff line numberDiff line change
@@ -125,15 +125,23 @@ All notable changes to bestASR are documented here. The format follows
125125
- **An unbounded `tolerance`.** This path needs no unusual number at all,
126126
which is what made it hard to see: `tolerance: 1e308` rendered a 94-point
127127
regression as an ordinary pass line, indistinguishable from a real one.
128-
`tolerance` now has the tightest bound of the three fields — the committed
129-
baseline uses 0.02, and 0.5 already accepts a 50-point jump.
128+
`tolerance` is now capped at **0.25**, the tightest per-field bound of the
129+
three, against a committed baseline that uses 0.02 throughout. Stated
130+
plainly, since the point of this change is that the prose must not flatter
131+
the gate: both that ceiling and the comparison are *inclusive*, so a
132+
regression of exactly 0.25 against a golden of 0 does pass. The bound
133+
refuses the absurd values; the tolerance now rendered on every passing line
134+
is what exposes the merely-too-lenient ones.
130135
- **A garbage numeric field** raised `ValueError` mid-run. A traceback is not
131-
a verdict; it now fails as a named gate error like every other.
136+
a verdict; numeric fields now fail as a named gate error. (Scoped honestly:
137+
this covers the *numeric* converter. Malformed document shapes — a
138+
top-level array, a missing `corpus` or `language` key — still surface as
139+
tracebacks. They all exit non-zero, so none is fail-open, but the class is
140+
not uniformly handled and a shape guard after `json.load` is the follow-up.)
132141

133-
Bounds are deliberately **asymmetric**, because the danger is: `golden` and
134-
`tolerance` are hazardous when *large* (both make the comparison vacuous),
135-
while `error_rate` is hazardous only when *non-finite* — a large measurement
136-
correctly fails, so its ceiling is a garbage filter, not a security boundary.
142+
`error_rate` is bounded differently on purpose: it is hazardous only when
143+
*non-finite*, because a large measurement correctly **fails**. Its ceiling is
144+
a garbage filter, not a security boundary.
137145

138146
A cross-model verification round found four more ways the same verdict could
139147
come out wrong, each fixed here:
@@ -155,10 +163,85 @@ All notable changes to bestASR are documented here. The format follows
155163
`float(true)` is 1.0 and a bare `true` passed every bound as an ordinary
156164
number.
157165

158-
`TOLERANCE_MAX` also tightened from 0.5 to **0.25** on the same review. No
159-
fixed ceiling can separate *generous* from *disabled* on its own, which is why
160-
the bound handles the absurd values and the newly-rendered tolerance handles
161-
the merely-too-lenient ones.
166+
A second verification round found that bounding `tolerance` had closed one
167+
handle of a two-handled lever, and named the other:
168+
169+
- **An inflated `golden` bought exactly the slack a wide `tolerance` was
170+
refused.** The decision is `actual - golden > tol`, so the gate passes
171+
anything at or below **`golden + tolerance`** — that sum is the effective
172+
threshold, and the two fields enter it identically. A `golden` of 2.0 with
173+
an honest `tolerance` of 0.02 admitted every physically possible error
174+
rate, and the tolerance newly rendered on the pass line read a reassuring
175+
`0.0200`, because the number that had moved was not the one being shown.
176+
Reproduced against the real committed baseline with a single field edited.
177+
The bound is now on **the sum** (`≤ 0.5`, against a committed maximum of
178+
0.1749) — the only bound invariant to trading one handle for the other.
179+
`golden`'s separate ceiling was **deleted**: with the sum bounded it could
180+
never be the guard that fired, and mutation testing put it at 0 of 34
181+
assertions. A constant no test can distinguish from its absence is not a
182+
guard.
183+
- **A truncated comparison still reported success.** Closing the empty case
184+
closed cardinality 0, not 1-of-N. Both sides of the comparison derive from
185+
the same `baseline.json` — the worklist stage reads it and the assembler
186+
iterates it — so deleting entries shrinks the baseline, the work list and
187+
the measured set together, and the gate prints `all 1 corpora within
188+
tolerance` over a release that verified one twelfth of what it should
189+
have. By pure deletion, with no unusual value anywhere. The expected corpus
190+
set is now pinned in `benchmarks/baseline-meta.json` and supplied to the
191+
compare stage by the gate, so truncating the baseline requires a second,
192+
deliberate edit to the file that carries the seeding provenance. A run with
193+
no anchor says so on stdout rather than passing quietly — a completeness
194+
check that can be skipped by omitting a key is the shape of the hole it
195+
closes.
196+
- **`OverflowError` escaped the numeric converter.** JSON has no integer
197+
width limit, so a bare 400-digit literal parses to an arbitrary-precision
198+
`int` and `float()` raises — neither `TypeError` nor `ValueError`. Exits
199+
non-zero, so never a bypass, but it reached the log as a stack trace where
200+
the verdict belongs. Universal, not version-specific. Rejected values are
201+
now also truncated at 120 characters, so a hostile field cannot become the
202+
log.
203+
204+
A third round found that the completeness anchor could be switched off by
205+
omitting a key — the shape the fix was written against, one layer out — and
206+
that three guards had no test that could tell them from their absence:
207+
208+
- **The anchor is now an invariant of the gate, not a nicety.**
209+
`scripts/regression-gate.sh` knows which baseline it is running, so for the
210+
repo's own it now **fails** when `baseline-meta.json` is missing,
211+
unreadable, or its `corpora` is absent / null / not a non-empty array of
212+
unique strings; an overridden `BESTASR_BASELINE` warns instead. The compare
213+
stage keeps tolerating an absent anchor with a printed NOTE, because a stdin
214+
filter genuinely cannot know whether it was handed a whole sweep — the
215+
invariant belongs where the answer is knowable. This also upgrades the #48
216+
model-artifact pin, which used to be skipped **silently** when that same
217+
file was missing.
218+
- **A satisfied anchor now says so** (`completeness: N corpora pinned by … —
219+
verified`). Its only previous evidence was the *absence* of the warning,
220+
which asks a reader to already know the warning exists — the same argument
221+
this entry makes for rendering the tolerance on passing lines.
222+
- **`error_rate: false`** was the one boolean position still able to flip a
223+
verdict, and nothing tested it: `float(false)` is `0.0`, a boolean laundered
224+
into a *perfect score* rather than a bad one. Every other position became
225+
redundantly covered once `golden` was bounded at 0.5 — which had quietly
226+
made the existing boolean test vacuous.
227+
- `expected_corpora` is shape-checked before use (a bare string would have
228+
been iterated character-wise; an int or a nested list raised a traceback),
229+
the anchor names the file it actually read rather than the default path, its
230+
corpus lists are length-capped, and the sum-bound message renders at 17
231+
significant digits so a rejected value cannot print as *equal* to the
232+
threshold it exceeded.
233+
234+
Every guard is non-vacuous, measured by deleting each independently and
235+
counting failed assertions: effective-threshold bound **9**, completeness
236+
anchor **7**, boolean guard **6**, `isfinite` **5**, `TOLERANCE_MAX` **3**,
237+
`OverflowError` **3**, `ERROR_RATE_MAX` **2**, echo cap **2**, and the gate's
238+
own anchor wiring **1** — a one-token typo in that shell heredoc used to
239+
revert the whole completeness guard with a green suite.
240+
241+
Those numbers come from **`scripts/mutation-check.sh`**, which is committed
242+
alongside them. A count a reader cannot re-run is the same defect as a
243+
threshold a reader cannot see, one level out; the script deletes each guard,
244+
re-runs the suite, restores the tree, and fails if anything scores zero.
162245

163246
Passing lines now render **the tolerance they were judged against**. Auditing
164247
a pass is the whole purpose of this log, and the one number that could expose

0 commit comments

Comments
 (0)