You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
π’ rolling main green at tip (33% β 2 infra/flaky reds cleared) Β· π’ internal 100% (β²30) Β· π’ outerloop 100% Β· π’ PR 80% β β nothing needs you
β Nothing needs you β rolling main recovered: green at the tip since Jul 24 00:05 and quiet ~20h
π§± infra share: ~67% (2 of 3 failed main jobs were environmental) β a Windows hosted-runner drop (Hosting-1, "runner lost communication") + an ubuntu disk-full (Templates, "No space left on device"); the 3rd was a Cli windows timing test that timed out under runner CPU pressure (intermittent flake, not a code regression). Internal AzDO main was 2/2 green β no failed internal build to link. (Release lanes omitted β GH release/13.4 and AzDO release/13.4 had no in-window builds.)
Legend: bar length = sample size / confidence (short = read its % with care) Β· Ξ vs 7d = 36h rate minus the lane's 7-day average Β· Latest build = red/green of the lane's most recent in-window run, linked to it Β· lane names link to the workflow / pipeline.
β Nothing needs active attention right now. Rolling main was intermittently red earlier in the window (2 of 3 runs), but it is green at the tip (Jul 24 00:05) and has been quiet ~20h. All three failing jobs were infra runner drops or a load-induced Cli timing flake β no shared product cause, no deterministic regression. The rerun bot is absorbing PR flakes (PR-only).
What's broken
Who's on it
Since
βͺ
rolling main β two earlier reds, no shared cause: a Windows runner drop + an ubuntu disk-full + a Cli windows timing test that timed out under CPU load; green at the tip now
Looks resolved β recovered to green; every cause is infra or a flaky timeout, not a code regression
last failed ~20h ago (Jul 24 00:03), green since 00:05
π‘
PR validation β a few PR runs flake; 2 branches red at their latest attempt, different owners, no common cause
Rerun bot is absorbing these (PR-only); each branch owner owns their fix
ongoing; bot rescued 8 this window; 2 branches left red
π§± Product vs infra: today's reds are not a product-code regression β 2 of 3 failed main jobs are environmental (Windows hosted-runner drop, ubuntu disk-full) and the 3rd is an intermittent Cli timing test that timed out under runner CPU pressure. Internal AzDO is 2/2 green, so no failed internal build to link.
β±οΈ Build-time note (read lightly, n=3):main's three in-window runs ran long β ~67m incl. the green tip vs ~27m 7-day norm β so the script flagged it, but on only 3 samples (2 of them failed) this is not a confirmed lane-wide slowdown.
exact tests, jobs & counts
GH CI β main, in-window: 1 success / 2 failure = 33%. Tip run 30055153746 (Jul 24 00:05) is green; no main push since (~20h quiet). Roll-up "Final Results" / "Final Test Results" aggregator jobs excluded from the failing-job count.
Tests / Hosting-1 / Hosting-1 (windows-latest) β hosted-runner drop, annotation: "The hosted runner lost communication with the server." Environmental, no test assertion, no failed step.
βͺ run 30055044537 β Jul 24 00:03 (commit "Update dependencies from dcp"):
Tests / Cli / Cli (windows-latest) β test Aspire.Cli.Tests.Projects.ProcessGuestLauncherTests.LaunchAsync_WithGracefulServices_BlockingSignalerDoesNotConsumeGracefulBudget threw System.TimeoutException : The operation has timed out. (~20s) β a timing-sensitive test timing out under runner CPU pressure, i.e. intermittent, not deterministic. No open tracking issue found for this test name.
Tests / Templates-XUnit_Default_NewUpAndBuildSupportProjectTemplatesTests (ubuntu-latest) β runner disk-full, annotation: "System.IO.IOException: No space left on device". Environmental, no failed test step.
βͺ Successful main run in-window (the tip): 30055153746 (Jul 24 00:05, ~2m after the red at 00:03). Auto-rerun is PR-only β this rolling red cleared on the next clean push, not a bot rerun.
PR validation, in-window: 12 success / 15 completed = 80% (excludes 8 cancelled, 12 awaiting-approval, 1 startup_failure). Latest completed run 30091438099 (Jul 24 11:59) is green. Two branches sit red at their latest in-window attempt after the bot exhausted retries β dapine/publisher-output-contract and dapire/security-deps/aspire-lowrisk-batch (distinct owners, no shared cluster; the latter flaked twice); the rerun bot rescued 8 redβgreen (see Β§ π).
ciinsights: returned 0 builds/failures across 36h and 7d for microsoft/aspire this run (see footer), so the per-test cluster join was unavailable β the failures above were reconstructed directly from the failed GH job annotations / logs.
Attention dots (this section): π΄ needs a human now Β· π‘ keep an eye on it Β· βͺ looks resolved β distinct from the lane table's π΄/π’ run-outcome dots.
π Reruns & flaky tax
The auto-rerun bot carried the entire rescue load this window β all 8 PR runs rescued redβgreen were the bot (PR-only); zero human reruns, so manual toil was nil. Low volume overall (quiet window).
rerun actor split (36h, PR validation)
π€ Auto-rerun bot: 22 reruns across 13 PR runs β 8 rescued redβgreen; 3 still red, 2 cancelled.
π§ Manual (human) reruns: 0 reruns across 0 PR runs (0 branches the bot also reran) β on the 0 human-only branches: 0 rescued, 0 still red, 0 cancelled.
π Combined: 13 distinct PR runs needed a rerun in the window (8 went green, 3 still red, 2 cancelled); 0 PR branches needed both bot + human.
Sources: ciinsights MCP + gh api β¦/ci.yml/runs (pass rates + rerun split) + az pipelines build list (def 1602). Window = last 36h. Jul 24 2026. ciinsights degraded this run β the connection reported healthy, yet list_builds / query_test_failures / get_failure_summary / get_health_summary all returned 0 builds/failures for microsoft/aspire across both 36h and 7d (no_data_in_time_window), which blocked the per-test cluster join; the main failures were reconstructed from GH job logs instead. Backlog & trend: see the weekly CI Health report (#18231). Generated 2026-07-24 20:06 UTC.
π’ rolling main green at tip (33% β 2 infra/flaky reds cleared) Β· π’ internal 100% (β²30) Β· π’ outerloop 100% Β· π’ PR 80% β β nothing needs you
Window: Thu Jul 23 08:06 β Fri Jul 24 20:06 UTC (last 36h) Β· Ξ = vs the 7-day average Β· Jul 24 2026.
π Last 36h by lane β activity, pass rate & Ξ vs 7d
π§± infra share: ~67% (2 of 3 failed
mainjobs were environmental) β a Windows hosted-runner drop (Hosting-1, "runner lost communication") + an ubuntu disk-full (Templates, "No space left on device"); the 3rd was a Cli windows timing test that timed out under runner CPU pressure (intermittent flake, not a code regression). Internal AzDO main was 2/2 green β no failed internal build to link. (Release lanes omitted β GHrelease/13.4and AzDOrelease/13.4had no in-window builds.)Legend: bar length = sample size / confidence (short = read its % with care) Β·
Ξ vs 7d= 36h rate minus the lane's 7-day average Β· Latest build = red/green of the lane's most recent in-window run, linked to it Β· lane names link to the workflow / pipeline.π§© What's broken & why
β Nothing needs active attention right now. Rolling
mainwas intermittently red earlier in the window (2 of 3 runs), but it is green at the tip (Jul 24 00:05) and has been quiet ~20h. All three failing jobs were infra runner drops or a load-induced Cli timing flake β no shared product cause, no deterministic regression. The rerun bot is absorbing PR flakes (PR-only).mainβ two earlier reds, no shared cause: a Windows runner drop + an ubuntu disk-full + a Cli windows timing test that timed out under CPU load; green at the tip nowπ§± Product vs infra: today's reds are not a product-code regression β 2 of 3 failed
mainjobs are environmental (Windows hosted-runner drop, ubuntu disk-full) and the 3rd is an intermittent Cli timing test that timed out under runner CPU pressure. Internal AzDO is 2/2 green, so no failed internal build to link.β±οΈ Build-time note (read lightly, n=3):
main's three in-window runs ran long β ~67m incl. the green tip vs ~27m 7-day norm β so the script flagged it, but on only 3 samples (2 of them failed) this is not a confirmed lane-wide slowdown.exact tests, jobs & counts
GH CI β main, in-window: 1 success / 2 failure = 33%. Tip run 30055153746 (Jul 24 00:05) is green; no
mainpush since (~20h quiet). Roll-up "Final Results" / "Final Test Results" aggregator jobs excluded from the failing-job count.Tests / Hosting-1 / Hosting-1 (windows-latest)β hosted-runner drop, annotation: "The hosted runner lost communication with the server." Environmental, no test assertion, no failed step.Tests / Cli / Cli (windows-latest)β testAspire.Cli.Tests.Projects.ProcessGuestLauncherTests.LaunchAsync_WithGracefulServices_BlockingSignalerDoesNotConsumeGracefulBudgetthrewSystem.TimeoutException : The operation has timed out.(~20s) β a timing-sensitive test timing out under runner CPU pressure, i.e. intermittent, not deterministic. No open tracking issue found for this test name.Tests / Templates-XUnit_Default_NewUpAndBuildSupportProjectTemplatesTests (ubuntu-latest)β runner disk-full, annotation: "System.IO.IOException: No space left on device". Environmental, no failed test step.mainrun in-window (the tip): 30055153746 (Jul 24 00:05, ~2m after the red at 00:03). Auto-rerun is PR-only β this rolling red cleared on the next clean push, not a bot rerun.PR validation, in-window: 12 success / 15 completed = 80% (excludes 8 cancelled, 12 awaiting-approval, 1 startup_failure). Latest completed run 30091438099 (Jul 24 11:59) is green. Two branches sit red at their latest in-window attempt after the bot exhausted retries β
dapine/publisher-output-contractanddapire/security-deps/aspire-lowrisk-batch(distinct owners, no shared cluster; the latter flaked twice); the rerun bot rescued 8 redβgreen (see Β§ π).ciinsights: returned 0 builds/failures across 36h and 7d for
microsoft/aspirethis run (see footer), so the per-test cluster join was unavailable β the failures above were reconstructed directly from the failed GH job annotations / logs.Attention dots (this section): π΄ needs a human now Β· π‘ keep an eye on it Β· βͺ looks resolved β distinct from the lane table's π΄/π’ run-outcome dots.
π Reruns & flaky tax
The auto-rerun bot carried the entire rescue load this window β all 8 PR runs rescued redβgreen were the bot (PR-only); zero human reruns, so manual toil was nil. Low volume overall (quiet window).
rerun actor split (36h, PR validation)
Sources: ciinsights MCP +
gh api β¦/ci.yml/runs(pass rates + rerun split) +az pipelines build list(def 1602). Window = last 36h. Jul 24 2026. ciinsights degraded this run β the connection reported healthy, yetlist_builds/query_test_failures/get_failure_summary/get_health_summaryall returned 0 builds/failures formicrosoft/aspireacross both 36h and 7d (no_data_in_time_window), which blocked the per-test cluster join; themainfailures were reconstructed from GH job logs instead. Backlog & trend: see the weekly CI Health report (#18231). Generated 2026-07-24 20:06 UTC.