[CI] Restart sidecar↔ambient demo pods so HTTP graph edges appear promptly - #10247
Conversation
Name the echo Service port HTTP, give curl clients connect timeouts, wait for sidecar injection, and fail setup if the graph still lacks HTTP edges so the ambient Cypress job is not blocked on a 15-minute poll. Co-authored-by: Cursor <cursoragent@cursor.com>
Scope this evaluation branch to initialize, frontend/backend builds, and the ambient frontend integration tests so duration can be measured in isolation. Co-authored-by: Cursor <cursoragent@cursor.com>
kubectl jsonpath concatenates container names, so pods with istio-proxy were treated as missing one and setup aborted before Cypress ran. Co-authored-by: Cursor <cursoragent@cursor.com>
The graph is immediately 2 HTTP / 6 total with curls already succeeding. Failing setup here blocked Cypress and hid whether 4/8 ever appears. Co-authored-by: Cursor <cursoragent@cursor.com>
Dump namespace labels, enrollment, curl HTTP traces, echo-server logs, Prometheus L4/L7 breakdown, and sidecar proxy stats when the demo graph is short of expected HTTP edges. Co-authored-by: Cursor <cursoragent@cursor.com>
Query unknown source_workload and namespace-scoped series, add 1m rate windows for early install, and document Istio vs graph labeling. Co-authored-by: Cursor <cursoragent@cursor.com>
Fresh Sail installs can open HBONE with source_workload=unknown for the life of those pods; a rollout on the already-warm mesh used by this Cypress scenario re-establishes connections with curl-client labels. Co-authored-by: Cursor <cursoragent@cursor.com>
The Cypress scenario already restarts those pods on a warm mesh; the setup poll always timed out at 2 HTTP / 6 total and added ~2 minutes. Keep verify-sidecar-ambient-traffic.sh for local Istio repro. Co-authored-by: Cursor <cursoragent@cursor.com>
Keep the Cypress pod restart, HTTP port names, and curl timeouts. Remove the unused verify script, sidecar-injector wait, and the temporary ambient-only workflow so this PR only addresses the graph hang. Co-authored-by: Cursor <cursoragent@cursor.com>
How this resolves the hang (test unchanged)The failing scenario is still What actually failed was telemetry, not the UI:
The Cypress step Demo-only extras (not test expectations): echo Service |
jmazzitelli
left a comment
There was a problem hiding this comment.
tests are green. ambient tests finished in 11m and 16m which seems like it was much faster.
Closes #10249
Summary
The ambient Cypress job spent ~15–23 extra minutes in
[Traffic] Sidecar Ambient trafficwaiting for HTTP graph edges that never appeared on the original pods.On a fresh Sail install, ambient→sidecar HTTP dest-reporter metrics are labeled
source_workload=unknownfor the life of those pods. Kiali correctly skips that series (IsBadSourceTelemetrycase 1), so the graph stays at 2 HTTP / 6 total instead of 4 HTTP / 8. TCP on the same path is labeledcurl-clientimmediately; sidecar→ambient HTTP is also fine. Restarting the demo deployments later, after the mesh is warm, opens new HBONE connections withcurl-clientlabels and the existing graph wait succeeds in ~45s.This PR:
test-ambient/test-sidecardeployments immediately before the existing sidecar-ambient graph wait (that scenario already runs after other waypoint tests, so the mesh is warm).http(appProtocol: http) and gives the demo curl loops connect/max timeouts so generators do not hang.maxRetriesfrom 90 (15 min) to 30 (5 min). The pass condition is unchanged.The Cypress scenario and its expectations are unchanged.
waypoint.featurestill asserts 8 edges (HTTP+TCP), then 4 (ambient off), then 4 (TCP off). The wait still requires ≥4 HTTP and ≥8 total edges.Test plan
[Traffic] Sidecar Ambient trafficpasses on attempt 1 in under a minute (no 15-minute poll, no Cypress retry)