Skip to content

πŸ“ˆ CI Daily PulseΒ #18232

Description

@radical

πŸ”΄ GH CI main red at the tip β€” flaky VS Code E2E suite (0/13) Β· 🟑 outerloop 0% (lone run ~22h) Β· βšͺ internal 38% (green at tip) Β· πŸ”€ PR 53% (red at tip, bot-absorbed) Β· 🟒 release green β€” 1 item for a dev

πŸ”΄ For a dev β€” rolling main CI is red on all 13 runs; the VS Code extension E2E suite is chronically flaky and rolling has no auto-rerun

What: All 13 in-window main runs fail, but the cause is intermittent, not deterministic. The VS Code extension E2E suite flakes in 12 of the 13 runs, with the offending shard rotating run-to-run. The edge-cases (Windows) shard drove the first ~9 runs (the "shows debugger install guidance" test times out after 60s waiting for the Python-debugger CodeLens β€” commandTitles: []) but has been quiet ~10h (passed the last 4 runs); since then each run fails on a different shard (workspace-target-proof / apphost-lifecycle-tools / azure-functions E2E timeouts). No single revert clears it, and rolling main has no auto-rerun, so it stays red until a clean push β€” a dev has to stabilize the flaky suite.
Do: a dev should stabilize the VS Code extension E2E suite β€” the edge-cases install-hint CodeLens break (run, extension/src/test-e2e/edgeCases.e2e.test.ts:162) plus the workspace-target-proof / apphost-lifecycle-tools / azure-functions shard timeouts (run). The very latest red is a one-off NuGet-restore drop, not the suite: tip run.
Secondary LIVE signals: an Azure.Npgsql (Windows) host hang on one run (hang dump β€” test timed out, infra-looking); the tip's red is a transient NuGet dotnet tool restore drop that clears on a clean push. Everything else red = noise / quiet β€” see Β§ 🧩 below.

Window: Tue Aug 25 12:43 β†’ Thu Aug 27 00:43 UTC (last 36h) Β· Ξ” = 36h rate vs the same lane's 7-day average Β· Aug 27, 2026.

πŸ“Š Last 36h by lane β€” activity, pass rate & Ξ” vs 7d

Lane 36h activity 36h pass rate (Ξ” vs 7d) Latest build
🌳 GH CI β€” main πŸŸ₯πŸŸ₯πŸŸ₯πŸŸ₯πŸŸ₯πŸŸ₯πŸŸ₯ πŸ”΄ 0% (0/13) Β· β–Ό 23.2 πŸ”΄ Aug 26 20:09
πŸ§ͺ GH CI outerloop β€” main πŸŸ₯πŸŸ₯ πŸ”΄ 0% (0/1, 95% CI 0–79%) Β· β–Ό 85.7 πŸ”΄ Aug 26 02:29
πŸ—οΈ Internal (AzDO) β€” main 🟩🟩πŸŸ₯πŸŸ₯πŸŸ₯πŸŸ₯ πŸ”΄ 38% (3/8, 95% CI 14–69%) Β· β–Ό 26.1 🟒 Aug 26 20:16
πŸ—οΈ Internal (AzDO) β€” release/13.5 🟩🟩 🟒 100% (1/1, 95% CI 21–100%) Β· β–² 14.3 🟒 Aug 25 19:31
πŸ”€ GH CI β€” PR validation ⭐ 🟩🟩🟩🟩🟩🟩πŸŸ₯πŸŸ₯πŸŸ₯πŸŸ₯πŸŸ₯πŸŸ₯ πŸ”΄ 53% (51/96) Β· β–Ό 12.1 πŸ”΄ Aug 27 00:06

🧱 Mostly flaky-test + infra, not a confirmed live code regression: the rolling-main red is intermittent VS Code E2E-shard flakiness (12/13 runs, offending shard rotates), a now-quiet installer-prep cluster, an Azure.Npgsql host hang, and a transient NuGet-restore network drop at the tip β€” but with no auto-rerun on rolling it stays red until a clean push. Internal AzDO (πŸ—οΈ) had a rough mid-window patch (5 failed) but is partiallySucceeded (green) at the tip; failures are counts-only; latest failed internal build in window: 3057549 (~24h ago).

bar length = sample size / confidence (short = read its % with care) Β· Ξ” vs 7d = 36h rate minus the lane's 7-day average Β· Latest build = red/green of the lane's most recent in-window run, linked to it Β· lane names link to the workflow / pipeline. GH CI release/13.5 push had no runs in the window (row omitted). The internal AzDO link 401s for non-members (expected).

🧩 What's broken & why

One thing needs a human: rolling main is kept red run-after-run by a chronically flaky VS Code extension E2E suite β€” rolling has no auto-rerun, so no bot will clear it. Everything else is flaky noise the rerun bot is absorbing on PRs, a lone outerloop run, or yesterday's fires already gone quiet.

What's broken Who's on it Since
πŸ”΄ GH CI main red on all 13 runs β€” the VS Code extension E2E suite flakes in 12/13, a different shard most runs; the edge-cases install-hint test led early, now other shards time out. No single revert clears it. run Needs a dev to triage β€” no auto-rerun on rolling red all window (36h), still red at tip
🟑 PR validation flaky β€” about half the PR runs go red before the bot reruns them; same E2E flakiness plus assorted test flake. runs Rerun bot is absorbing these (PR-only) ongoing, red at tip
🟑 Outerloop β€” one scheduled run failed on the Dashboard tests (both OSes); lone run, outerloop self-reruns. run Watch β€” one run, too early to call ~22h ago
βšͺ Internal AzDO main β€” rough mid-window patch (5 failed builds), now green at the tip. build Looks recovered β€” green at tip last failed ~24h ago

🧱 Product vs infra: the load-bearing driver is a flaky VS Code E2E test suite (product/test-side, no auto-rerun on rolling) β€” not a confirmed deterministic code regression; the tip red is a transient NuGet dotnet tool restore drop and an early installer-prep cluster that has since gone quiet. Internal AzDO is green at the tip; latest failed internal build 3057549 (~24h ago).

exact tests, jobs & counts

πŸ”΄ GH CI main β€” VS Code extension E2E suite (12 of 13 in-window runs), offending shard rotates:

  • edge-cases (Windows) shard β€” test shows debugger install guidance while the Aspire panel and AppHost source are closed (extension/src/test-e2e/edgeCases.e2e.test.ts:162, DebuggerInstallHintApp): 5 passing, 1 failing β€” Error: Timed out after 60000ms waiting for the Python debugger CodeLens. Last provider result: … "commandTitles":[]. Failed 9 consecutive runs Aug 25 15:11 β†’ Aug 26 14:12 (32864200821 … 32978808010); quiet since (~10h, passed last 4 runs).
  • Later runs rotate shard, one per run: workspace-target-proof (Linux) 32994724407; apphost-lifecycle-tools (Windows) + azure-functions (Linux) 32997866203; azure-functions (Linux) 32999137654 β€” all E2E timeouts.
  • Tip run 33008969799 (Aug 26 20:09): NOT the E2E suite β€” Tests / Hosting.Browsers (windows-latest) Build test project failed on error MSB3073: … dotnet.exe tool restore … exited with code 1 (transient NuGet restore drop, clears on clean push).
  • Now-quiet early cluster (runs Aug 25 15:11 β†’ 20:55 only): Prepare Homebrew installer artifacts / Prepare Homebrew cask + Prepare WinGet installer artifacts / Prepare WinGet manifests.
  • One-off job failures across the window: Dashboard (windows-latest) / Dashboard (ubuntu-latest), Hosting-1 (windows-latest), Azure.Npgsql (windows-latest) (host hang, hang-dump β€” test timed out).

🟑 Outerloop: scheduled run 32922923673 (Aug 26 02:29) β€” outerloop_tests / Dashboard (ubuntu-latest) + Dashboard (windows-latest). Only 1 scheduled run in window (0/1).

βšͺ Internal AzDO main (counts only): 8 completed builds in window β€” 3 partiallySucceeded (green), 5 failed. Failed build IDs 3057169, 3057283, 3057353, 3057434, 3057549 (last, Aug 26 01:11 β€” ~24h ago); green at tip 3058236 (Aug 26 20:16).

Attention dots (action priority, distinct from the lane table's πŸ”΄/🟒 run-outcome dots): πŸ”΄ needs a human now Β· 🟑 keep an eye on it Β· βšͺ looks resolved.

πŸ” Reruns & flaky tax

The auto-rerun bot carries the vast majority of the rerun load (169 reruns rescuing 30 PRs red→green), with a little human toil on top (20 manual reruns) — that bot volume is the flaky tax the broken/flaky VS Code E2E suite is imposing on PRs. (Auto-rerun is PR-only — rolling main gets none, which is why it stays red above.)

rerun actor split (last 36h, PR lane)
  • πŸ€– Auto-rerun bot: 169 reruns across 81 PR runs β†’ 30 rescued redβ†’green; 39 still red, 12 cancelled.
  • πŸ§‘ Manual (human) reruns: 20 reruns across 5 PR runs (3 branches the bot also reran) β†’ on the 1 human-only branches: 0 rescued, 1 still red, 0 cancelled.
  • πŸ”— Combined: 86 distinct PR runs needed a rerun in the window (30 went green, 43 still red, 13 cancelled); 3 PR branches needed both bot + human.

Sources: gh api …/ci.yml/runs (pass rates + rerun split) Β· az pipelines build list def 1602 (internal, counts only). ciinsights unavailable this run β€” cluster detail derived directly from GH failed-job logs (see footer). Window = last 36h. Backlog & trend: see the weekly CI Health report (#18231). Generated Aug 27, 2026 00:43 UTC.

Metadata

Metadata

Assignees

No one assigned

    Labels

    area-engineering-systemsinfrastructure helix infra engineering repo stuffautomatedOpened by bots or toolstriage:bot-seenAspire triage bot has seen this issue

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions