Skip to content

Commit c4aa1e7

Browse files
sturleseclaude
andauthored
fix(metrics): count orphaned-workflow governance incidents in the window rollup (#17)
The window governance counters (blocked_budget, blocked_policy, failed) were incremented inside the per-workflow loop, so a run whose workflow_id is no longer in the config (a pilot deleted/renamed while its historical runs remain) was never counted — while the all-time counters, tallied in a per-run pass with no workflow lookup, DO count it. So for an in-window incident on a since-deleted workflow the window counter read 0 while the all-time counter read 1, breaking the documented "_all is a superset of the window" relationship and understating the window's governance posture. Add a small pass over the window runs for workflows absent from the config, classifying their blocks/failures the same way. Runs of known workflows are already tallied above, so the pass is disjoint — no double counting, and no change for any org whose runs all belong to current workflows. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
1 parent 2fdbaa2 commit c4aa1e7

2 files changed

Lines changed: 33 additions & 0 deletions

File tree

‎src/flightdeck/metrics.py‎

Lines changed: 16 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -321,6 +321,22 @@ def build_report(org: Org, store: Store, ledger: Ledger, days: int = 30, now: da
321321
entry.health = _health(workflow, entry)
322322
report.workflows.append(entry)
323323

324+
# Governance incidents for a workflow since deleted from config still belong in
325+
# the window rollup: the all-time counters already include them, so the window
326+
# counters must too (they have no per-workflow row to be counted under). Runs of
327+
# known workflows were already tallied in the loop above, so this pass is disjoint.
328+
known_workflows = set(org.workflows)
329+
for run in window_runs:
330+
if run.workflow_id in known_workflows:
331+
continue
332+
if run.status == "blocked":
333+
if "budget" in (run.reason or ""):
334+
report.governance.blocked_budget += 1
335+
else:
336+
report.governance.blocked_policy += 1
337+
elif run.status == "failed":
338+
report.governance.failed += 1
339+
324340
completed = [run for run in window_runs if run.status == "completed"]
325341
if completed:
326342
no_training = sum(

‎tests/test_metrics.py‎

Lines changed: 17 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -126,6 +126,23 @@ def test_governance_rollup(self, seeded):
126126
assert gov.no_training_share == 1.0
127127
assert gov.ledger_ok and gov.ledger_entries == 1
128128

129+
def test_window_governance_counts_orphaned_workflow_incidents(self, org, store, ledger):
130+
# A workflow deleted from config leaves its historical runs behind. In-window
131+
# governance incidents for it still belong in the window rollup — the all-time
132+
# counters already include them, so the window counters must too.
133+
when = NOW - timedelta(days=1)
134+
store.add_run(_run("gb", when, workflow_id="deleted-wf", status="blocked",
135+
reason="monthly budget exhausted", model_id="", cost=0))
136+
store.add_run(_run("gp", when, workflow_id="deleted-wf", status="blocked",
137+
reason="no policy-compliant model", model_id="", cost=0))
138+
store.add_run(_run("gf", when, workflow_id="deleted-wf", status="failed",
139+
reason="provider: timeout", model_id="", cost=0))
140+
gov = build_report(org, store, ledger, days=30, now=NOW).governance
141+
# Window counters now match the all-time counters for these in-window incidents.
142+
assert gov.blocked_budget == gov.blocked_budget_all == 1
143+
assert gov.blocked_policy == gov.blocked_policy_all == 1
144+
assert gov.failed == 1
145+
129146
def test_health_flags_the_underperformer(self, seeded):
130147
# acceptance 0.67 vs target 0.80 and ~1.5 weekly actives vs target 6 → worst
131148
# ratio far below 0.75.

0 commit comments

Comments
 (0)