+Prefer deriving the increment from the attempt's terminal state over enumerating failure sites. The stage's existing bounded routes, on both sides of publication, were each written for the case that prompted them, which is exactly how the `run`/`verify` route came to have none, and an enumeration extends that pattern: a failure mode added later escapes instrumentation silently, and the hole reopens without anyone editing the counter. "This attempt ended without a receipt" holds however the attempt died. Three details the design must settle rather than assume, all visible in the stage document: under data-first an attempt owns a *pair* of receipts, analysis and release, with different retirement conditions, so the predicate must say which constitutes publication; a receipt published as pending can still fail later at the replicator or the audit, and those failures are already counted elsewhere, so the counter must distinguish never-published from published-then-failed or it will double-count; and the fresh-attempt transition is already a single shared macro that most repair branches route through, which may make it the natural increment site and the enumeration worry smaller than it first appears. Field experience revised the predicate itself, and the revision is the substance of this entry's recommendation. "Ended without a published, verified receipt" is the intuitive test and it is too narrow: an attempt can verify its analysis receipt and still spend a full build that no counter owns. Three real branches do exactly that — a release producer that fails every iteration re-verifies the analysis member each time; the `no_headline_tags` branch at step 6.5 explicitly increments neither `audit_fix` nor `headline_replication`; and an interrupted-turn owner recovery finding required ledgers absent is not an auditor verdict and belongs to no loop. A receipt-based predicate declines all three and leaves them unbounded. **Test counter responsibility instead: increment whenever the repair transition is entered and no other Stage 3a counter is responsible for the same routing.** Define responsibility as the routing *advancing* a counter toward its cap, or handing off to acceptance. The tempting widening — that a branch which explicitly preserves or freezes a counter also owns the attempt, since the substantive data and method FAIL routes preserve their failing counter while escalating — is wrong, and wrong in the direction that silently reopens the hole. A counter that is not advancing cannot terminate anything, and those FAIL routes re-fire a full build through the same transition, so on repetition the preserved counter never moves and nothing bounds them at all. `core.md`'s escalation non-reset exception is a rule about not zeroing a counter, not a claim that a frozen counter governs a routing; it should not be read as ownership. Defining the counter by what the other counters do rather than by the shape of the failure makes double-counting structurally impossible, catches any future branch that increments nothing without needing to anticipate it, and requires no classification of the failure — which matters, because classification is what the enumeration approach fails at. Two riders emerged from field use and belong in any implementation. The counter is about *failures*, so a transition entered for a requested regeneration of a healthy candidate must not increment whatever counters it does or does not advance — ask first whether an attempt was lost, then whether anything owns it. And the cap route must break its counts down before routing to Stage 2, but must **not** map kind to remedy. Environment interruptions land in this counter and no respecification fixes them — within an hour of going live on eventcal, both counts were interruptions rather than build defects — so a cap route that says flatly "this is a specification problem" is wrong. But the tempting corrective mapping is wrong too, in both directions: an attempt that published a verified receipt and died unowned is usually a step-7.5 substantive data or method FAIL, a real scientific fault that respecification does address, while the never-published class was on tradingdays a scope question needing an operator decision rather than any respecification at all. Only the interruption kind reads reliably. Sizing follows from the same property: a predicate that refuses to classify absorbs interruptions alongside real build failures, so the cap must be looser than a pure build-failure bound would need. Exempting interruptions to win budget back is the wrong trade — the classification-free predicate is exactly what makes the counter unescapable by a branch nobody anticipated. But so is loosening the cap in flight, and that failure is worth recording because it was attempted here and had to be reverted. At 3 of 4, one failure from firing, the cap was raised to 6 on the argument that two of the three counts were environment interruptions and the parameter had been chosen without evidence. Review falsified the premise: the deployment's own debug directory held 45 artifacts for the attempt called an interruption, twelve of them TOOL-FIT-ISSUE reports, and the state entry recorded two producer defects before the interruption that ended it. The true split was two defect-dominated builds and one near-free interruption, the inverse of the claim; the chosen number was inconsistent with even the claimed split; and the change was unnecessary by its own reasoning, since the cap route already carries the interruption breakdown. Hence a structural fence, which any implementation should carry: **a cap may not be changed while its round is within one of it**, taking effect at the next campaign or after a reset. Sizing a cap is a build-versus-respecification cost judgement, not an estimate of an interruption rate from a handful of observations — and a rationale that is a property of the predicate but gets applied to exactly one deployment, the one still running, is an outcome-level argument in disguise. Report the counts alongside each attempt's retirement diagnostic and reason from those. Note also that `loops` records no kind, so the breakdown has to be reconstructed from the registry's retirement reasons. Two further consequences are worth stating because they are counter-intuitive. An attempt that published a verified receipt can still increment, if the routing that killed it is named by no counter — the interrupted-turn owner recovery observed as tradingdays a116 is exactly that shape. And the bound is over a *run* of consecutive unowned failures, not over a campaign: one that alternates an unowned failure with an owned one resets this counter every other attempt and never caps here. Two sites stay outside any such counter and want their own treatment. A requested regeneration of a healthy candidate must be exempt — it is not a failure — but the exempt class is then itself unbounded, since regeneration costs a full build and nothing counts repeats. And Gate 3a-feasibility runs a parallel retry loop ("retire the pending candidate and return to step 1 with K+1") that re-fires full feasibility builds with no counter and no cap, at a site a Stage 3a counter does not reach. One further limit remains: a counter that resets only on success is a rate limiter rather than a terminating cap: `round` grows past `cap` while the build keeps failing, and each cap return to Stage 2 buys exactly one more build with nothing bounding the number of returns. That converts an unbounded re-fire loop into one build per respecification — a large improvement, not a termination proof. Closing it needs a cap-event counter scoped above `theory_version`, which is the same shape #306 asks for, so the two should be designed together. Without any counter here at all, the escalation the stage's other six caps exist to force is bypassable indefinitely by a campaign that never gets far enough to be judged. Deployments already on 2.30.x cannot receive a fix: `update.sh` is same-version-only.
0 commit comments