Skip to content

πŸ“ˆ CI Daily PulseΒ #18232

Description

@radical

🟒 rolling main green at tip (33% β€” 2 infra/flaky reds cleared) Β· 🟒 internal 100% (β–²30) Β· 🟒 outerloop 100% Β· 🟒 PR 80% ⭐ β€” nothing needs you

βœ… Nothing needs you β€” rolling main recovered: green at the tip since Jul 24 00:05 and quiet ~20h

All three earlier main job failures were non-product: a Windows hosted-runner drop (Hosting-1), an ubuntu disk-full (Templates), and one Cli windows timing test that timed out under CPU load β€” no shared cause, no deterministic regression. Internal + outerloop are green at the tip; PR flakes (2 branches) are being absorbed by the rerun bot (PR-only, 8 rescued). See Β§ 🧩 below.

Window: Thu Jul 23 08:06 β†’ Fri Jul 24 20:06 UTC (last 36h) Β· Ξ” = vs the 7-day average Β· Jul 24 2026.

πŸ“Š Last 36h by lane β€” activity, pass rate & Ξ” vs 7d

Lane 36h activity 36h pass rate (Ξ” vs 7d) Latest build
🌳 GH CI β€” main 🟩πŸŸ₯πŸŸ₯πŸŸ₯ πŸ”΄ 33% (1/3, 95% CI 6–79%) Β· β–Ό 19.0 🟒 Jul 24 00:05
πŸ§ͺ GH CI outerloop β€” main 🟩🟩 🟒 100% (1/1, 95% CI 21–100%) Β· β–¬ 0.0 🟒 Jul 24 03:02
πŸ—οΈ Internal (AzDO) β€” main 🟩🟩🟩 🟒 100% (2/2, 95% CI 34–100%) Β· β–² 30.0 🟒 Jul 24 02:18
πŸ”€ GH CI β€” PR validation ⭐ 🟩🟩🟩🟩🟩🟩πŸŸ₯ 🟒 80% (12/15) Β· β–Ό 4.6 🟒 Jul 24 11:59

🧱 infra share: ~67% (2 of 3 failed main jobs were environmental) β€” a Windows hosted-runner drop (Hosting-1, "runner lost communication") + an ubuntu disk-full (Templates, "No space left on device"); the 3rd was a Cli windows timing test that timed out under runner CPU pressure (intermittent flake, not a code regression). Internal AzDO main was 2/2 green β€” no failed internal build to link. (Release lanes omitted β€” GH release/13.4 and AzDO release/13.4 had no in-window builds.)

Legend: bar length = sample size / confidence (short = read its % with care) Β· Ξ” vs 7d = 36h rate minus the lane's 7-day average Β· Latest build = red/green of the lane's most recent in-window run, linked to it Β· lane names link to the workflow / pipeline.

🧩 What's broken & why

βœ… Nothing needs active attention right now. Rolling main was intermittently red earlier in the window (2 of 3 runs), but it is green at the tip (Jul 24 00:05) and has been quiet ~20h. All three failing jobs were infra runner drops or a load-induced Cli timing flake β€” no shared product cause, no deterministic regression. The rerun bot is absorbing PR flakes (PR-only).

What's broken Who's on it Since
βšͺ rolling main β€” two earlier reds, no shared cause: a Windows runner drop + an ubuntu disk-full + a Cli windows timing test that timed out under CPU load; green at the tip now Looks resolved β€” recovered to green; every cause is infra or a flaky timeout, not a code regression last failed ~20h ago (Jul 24 00:03), green since 00:05
🟑 PR validation β€” a few PR runs flake; 2 branches red at their latest attempt, different owners, no common cause Rerun bot is absorbing these (PR-only); each branch owner owns their fix ongoing; bot rescued 8 this window; 2 branches left red

🧱 Product vs infra: today's reds are not a product-code regression β€” 2 of 3 failed main jobs are environmental (Windows hosted-runner drop, ubuntu disk-full) and the 3rd is an intermittent Cli timing test that timed out under runner CPU pressure. Internal AzDO is 2/2 green, so no failed internal build to link.

⏱️ Build-time note (read lightly, n=3): main's three in-window runs ran long β€” ~67m incl. the green tip vs ~27m 7-day norm β€” so the script flagged it, but on only 3 samples (2 of them failed) this is not a confirmed lane-wide slowdown.

exact tests, jobs & counts

GH CI β€” main, in-window: 1 success / 2 failure = 33%. Tip run 30055153746 (Jul 24 00:05) is green; no main push since (~20h quiet). Roll-up "Final Results" / "Final Test Results" aggregator jobs excluded from the failing-job count.

  • βšͺ run 30048462581 β€” Jul 23 22:02 (commit "Add .NET 11 support to project templates" Add .NET 11 support to project templatesΒ #18849):
    • Tests / Hosting-1 / Hosting-1 (windows-latest) β€” hosted-runner drop, annotation: "The hosted runner lost communication with the server." Environmental, no test assertion, no failed step.
  • βšͺ run 30055044537 β€” Jul 24 00:03 (commit "Update dependencies from dcp"):
    • Tests / Cli / Cli (windows-latest) β€” test Aspire.Cli.Tests.Projects.ProcessGuestLauncherTests.LaunchAsync_WithGracefulServices_BlockingSignalerDoesNotConsumeGracefulBudget threw System.TimeoutException : The operation has timed out. (~20s) β€” a timing-sensitive test timing out under runner CPU pressure, i.e. intermittent, not deterministic. No open tracking issue found for this test name.
    • Tests / Templates-XUnit_Default_NewUpAndBuildSupportProjectTemplatesTests (ubuntu-latest) β€” runner disk-full, annotation: "System.IO.IOException: No space left on device". Environmental, no failed test step.
  • βšͺ Successful main run in-window (the tip): 30055153746 (Jul 24 00:05, ~2m after the red at 00:03). Auto-rerun is PR-only β€” this rolling red cleared on the next clean push, not a bot rerun.

PR validation, in-window: 12 success / 15 completed = 80% (excludes 8 cancelled, 12 awaiting-approval, 1 startup_failure). Latest completed run 30091438099 (Jul 24 11:59) is green. Two branches sit red at their latest in-window attempt after the bot exhausted retries β€” dapine/publisher-output-contract and dapire/security-deps/aspire-lowrisk-batch (distinct owners, no shared cluster; the latter flaked twice); the rerun bot rescued 8 redβ†’green (see Β§ πŸ”).

ciinsights: returned 0 builds/failures across 36h and 7d for microsoft/aspire this run (see footer), so the per-test cluster join was unavailable β€” the failures above were reconstructed directly from the failed GH job annotations / logs.

Attention dots (this section): πŸ”΄ needs a human now Β· 🟑 keep an eye on it Β· βšͺ looks resolved β€” distinct from the lane table's πŸ”΄/🟒 run-outcome dots.

πŸ” Reruns & flaky tax

The auto-rerun bot carried the entire rescue load this window — all 8 PR runs rescued red→green were the bot (PR-only); zero human reruns, so manual toil was nil. Low volume overall (quiet window).

rerun actor split (36h, PR validation)
  • πŸ€– Auto-rerun bot: 22 reruns across 13 PR runs β†’ 8 rescued redβ†’green; 3 still red, 2 cancelled.
  • πŸ§‘ Manual (human) reruns: 0 reruns across 0 PR runs (0 branches the bot also reran) β†’ on the 0 human-only branches: 0 rescued, 0 still red, 0 cancelled.
  • πŸ”— Combined: 13 distinct PR runs needed a rerun in the window (8 went green, 3 still red, 2 cancelled); 0 PR branches needed both bot + human.

Sources: ciinsights MCP + gh api …/ci.yml/runs (pass rates + rerun split) + az pipelines build list (def 1602). Window = last 36h. Jul 24 2026. ciinsights degraded this run β€” the connection reported healthy, yet list_builds / query_test_failures / get_failure_summary / get_health_summary all returned 0 builds/failures for microsoft/aspire across both 36h and 7d (no_data_in_time_window), which blocked the per-test cluster join; the main failures were reconstructed from GH job logs instead. Backlog & trend: see the weekly CI Health report (#18231). Generated 2026-07-24 20:06 UTC.

Metadata

Metadata

Assignees

No one assigned

    Labels

    area-engineering-systemsinfrastructure helix infra engineering repo stuffautomatedOpened by bots or toolstriage:bot-seenAspire triage bot has seen this issue

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions