The P1 deliverable is this measurement, not the code.
Run the fleet 72 hours post-change and record, against the §2.1 and §2.5 baselines:
- rate-limit error count, broken down by class (baseline: 27,662 undifferentiated over 8 days)
- delivery rate (baseline: 56%)
- patch-apply rate (baseline: 25–80%, model-dependent)
review_rejected rate (baseline: 2.5%)
If the apply rate falls to the bottom of the band, that is a finding, not a failure. Record it and re-evaluate the implementer tier before P3 hardcodes the role map. Extra repair rounds can consume the cost saving.
Acceptance: numbers posted in this issue; §2 of the plan updated; explicit go/no-go on the implementer tier.
The P1 deliverable is this measurement, not the code.
Run the fleet 72 hours post-change and record, against the §2.1 and §2.5 baselines:
review_rejectedrate (baseline: 2.5%)If the apply rate falls to the bottom of the band, that is a finding, not a failure. Record it and re-evaluate the implementer tier before P3 hardcodes the role map. Extra repair rounds can consume the cost saving.
Acceptance: numbers posted in this issue; §2 of the plan updated; explicit go/no-go on the implementer tier.