Skip to content

[CI] Restart sidecar↔ambient demo pods so HTTP graph edges appear promptly - #10247

Merged
jshaughn merged 9 commits into
kiali:masterfrom
jshaughn:ci-ambient-http-traffic
Aug 25, 2026
Merged

jshaughn merged 9 commits into
kiali:masterfrom
jshaughn:ci-ambient-http-traffic

Conversation

@jshaughn

@jshaughn jshaughn commented Aug 20, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #10249

Summary

The ambient Cypress job spent ~15–23 extra minutes in [Traffic] Sidecar Ambient traffic waiting for HTTP graph edges that never appeared on the original pods.

On a fresh Sail install, ambient→sidecar HTTP dest-reporter metrics are labeled source_workload=unknown for the life of those pods. Kiali correctly skips that series (IsBadSourceTelemetry case 1), so the graph stays at 2 HTTP / 6 total instead of 4 HTTP / 8. TCP on the same path is labeled curl-client immediately; sidecar→ambient HTTP is also fine. Restarting the demo deployments later, after the mesh is warm, opens new HBONE connections with curl-client labels and the existing graph wait succeeds in ~45s.

This PR:

  • Restarts test-ambient / test-sidecar deployments immediately before the existing sidecar-ambient graph wait (that scenario already runs after other waypoint tests, so the mesh is warm).
  • Names the echo Service ports http (appProtocol: http) and gives the demo curl loops connect/max timeouts so generators do not hang.
  • Lowers the graph-wait maxRetries from 90 (15 min) to 30 (5 min). The pass condition is unchanged.

The Cypress scenario and its expectations are unchanged. waypoint.feature still asserts 8 edges (HTTP+TCP), then 4 (ambient off), then 4 (TCP off). The wait still requires ≥4 HTTP and ≥8 total edges.

Test plan

  • Full Kiali CI on this PR (workflow restored to the normal job graph)
  • Ambient job around the historical ~12–14 min, not ~39–42 min
  • [Traffic] Sidecar Ambient traffic passes on attempt 1 in under a minute (no 15-minute poll, no Cypress retry)

Jay Shaughnessy and others added 2 commits August 20, 2026 17:13
Name the echo Service port HTTP, give curl clients connect timeouts, wait
for sidecar injection, and fail setup if the graph still lacks HTTP edges
so the ambient Cypress job is not blocked on a 15-minute poll.

Co-authored-by: Cursor <cursoragent@cursor.com>
Scope this evaluation branch to initialize, frontend/backend builds, and
the ambient frontend integration tests so duration can be measured in isolation.

Co-authored-by: Cursor <cursoragent@cursor.com>
@jshaughn jshaughn self-assigned this Aug 20, 2026
@jshaughn jshaughn moved this from 📋 Backlog to 🏗 In progress in Kiali Sprint 26-13 | Kiali v2.33 Aug 20, 2026
Jay Shaughnessy and others added 5 commits August 20, 2026 19:43
kubectl jsonpath concatenates container names, so pods with istio-proxy
were treated as missing one and setup aborted before Cypress ran.

Co-authored-by: Cursor <cursoragent@cursor.com>
The graph is immediately 2 HTTP / 6 total with curls already succeeding.
Failing setup here blocked Cypress and hid whether 4/8 ever appears.

Co-authored-by: Cursor <cursoragent@cursor.com>
Dump namespace labels, enrollment, curl HTTP traces, echo-server logs,
Prometheus L4/L7 breakdown, and sidecar proxy stats when the demo graph
is short of expected HTTP edges.

Co-authored-by: Cursor <cursoragent@cursor.com>
Query unknown source_workload and namespace-scoped series, add 1m rate
windows for early install, and document Istio vs graph labeling.

Co-authored-by: Cursor <cursoragent@cursor.com>
Fresh Sail installs can open HBONE with source_workload=unknown for the life of those pods; a rollout on the already-warm mesh used by this Cypress scenario re-establishes connections with curl-client labels.

Co-authored-by: Cursor <cursoragent@cursor.com>
Jay Shaughnessy and others added 2 commits August 24, 2026 15:58
The Cypress scenario already restarts those pods on a warm mesh; the
setup poll always timed out at 2 HTTP / 6 total and added ~2 minutes.
Keep verify-sidecar-ambient-traffic.sh for local Istio repro.

Co-authored-by: Cursor <cursoragent@cursor.com>
Keep the Cypress pod restart, HTTP port names, and curl timeouts. Remove
the unused verify script, sidecar-injector wait, and the temporary
ambient-only workflow so this PR only addresses the graph hang.

Co-authored-by: Cursor <cursoragent@cursor.com>
@jshaughn jshaughn changed the title Make sidecar↔ambient demo traffic produce HTTP telemetry promptly Restart sidecar↔ambient demo pods so HTTP graph edges appear promptly Aug 24, 2026
@jshaughn

Copy link
Copy Markdown
Collaborator Author

How this resolves the hang (test unchanged)

The failing scenario is still [Traffic] Sidecar Ambient traffic in waypoint.feature. We did not change that feature file, its @tags, or its assertions (8 edges with HTTP+TCP, then 4 with ambient off, then 4 with TCP off). The wait still requires ≥4 HTTP and ≥8 total graph edges for test-ambient,test-sidecar. This is not a weakened test.

What actually failed was telemetry, not the UI:

  1. Ambient→sidecar HTTP has no source-reporter L7 (ztunnel is L4). The only HTTP series is dest-reporter on the sidecar.
  2. On a fresh Sail mesh, that dest-reporter series is request_protocol=http, 200, mTLS, SPIFFE set, but source_workload=unknown / source_app=unknown / source_cluster=unknown for the life of those pods. L4 TCP on the same hop is labeled curl-client. Sidecar→ambient HTTP (source reporter) is also curl-client.
  3. Kiali IsBadSourceTelemetry case 1 drops namespace-OK / workload-and-app-unknown series. The graph therefore stays at 2 HTTP / 6 total (HTTP missing on the two ambient→sidecar hops; TCP still draws the nodes). Waiting 15 minutes on the same generation does not rewrite the labels.
  4. The same topology on a warm mesh (or after a rollout restart of those four deployments) emits source_workload=curl-client, and the existing wait sees 4 HTTP / 8 immediately.

The Cypress step the graph page has enough data for sidecar ambient traffic now kubectl rollout restarts the demo deployments in test-sidecar and test-ambient, waits for rollout, then runs the same graph wait as before. That step already runs after the other waypoint scenarios, so the mesh is warm. Measured on this PR: the scenario passed on attempt 1 in ~45–50s; the ambient job dropped from ~39–42 min to ~14 min.

Demo-only extras (not test expectations): echo Service name: http / appProtocol: http, and curl --connect-timeout / --max-time so hung clients cannot stall the generator. Graph-wait maxRetries 90→30 only bounds how long we poll; it does not change what counts as enough data.

@jshaughn
jshaughn marked this pull request as ready for review August 24, 2026 20:47
@jshaughn
jshaughn requested a review from jmazzitelli August 24, 2026 20:47
@jshaughn jshaughn moved this from 🏗 In progress to 👀 In review in Kiali Sprint 26-13 | Kiali v2.33 Aug 24, 2026
@jshaughn jshaughn changed the title Restart sidecar↔ambient demo pods so HTTP graph edges appear promptly [CI] Restart sidecar↔ambient demo pods so HTTP graph edges appear promptly Aug 24, 2026

@jmazzitelli jmazzitelli left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tests are green. ambient tests finished in 11m and 16m which seems like it was much faster.

@jshaughn
jshaughn merged commit d34fe2a into kiali:master Aug 25, 2026
25 checks passed
@github-project-automation github-project-automation Bot moved this from 👀 In review to ✅ Done in Kiali Sprint 26-13 | Kiali v2.33 Aug 25, 2026
@jshaughn
jshaughn deleted the ci-ambient-http-traffic branch August 25, 2026 02:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Development

Successfully merging this pull request may close these issues.

[CI] Sidecar Ambient traffic test very slow

2 participants