Commit d798059
authored
test(e2e): add scheduler policy E2E suite (#2793)
* test(e2e): add scheduler policy E2E suite
Add a Ginkgo E2E suite that asserts deterministic placement for HAMi's
node/GPU binpack and spread policies, comma-separated policy chains, and
per-Pod device scoring weight overrides. Adds a `make e2e-policy-test`
target and `hack/e2e-policy-test.sh` runner.
The suite waits on the scheduler's in-memory cache (via /metrics) so
placement assertions never observe stale utilization: pods are confirmed
released from every scheduler replica's cache after deletion and namespace
cleanup. Metrics are scraped from all running scheduler-extender replicas
rather than a guessed leader, so the suite is correct under HA and across
leader changes. GPU allocation decoding tolerates the trailing-separator
quirk of EncodePodSingleDevice instead of hard-coding the decoded length.
🤖 Assisted-by: Claude Code
Signed-off-by: Tim <tim.wang03@sap.com>
* address latest comment
Signed-off-by: Tim <tim.wang03@sap.com>
* fix(e2e): wait for scheduler GPU registration after device-plugin ready
After labeling a node and waiting for the device-plugin to become Ready,
the scheduler still needs time to process the node update and register
the node's GPU capacity. Without this wait, pods may fail to schedule
with '1 node unregistered' errors.
Add schedulerCanScheduleGPU check in ensureGPUNodeLabeled to verify the
node has GPU resources in its Capacity/Allocatable before proceeding.
Signed-off-by: Tim <tim.wang03@sap.com>
* fix(e2e): abort immediately on terminal pod phases in policy tests
waitForAllocatedPod previously polled for the full 5-minute timeout even
when a pod reached the terminal Failed phase (which can never transition
to Running). This caused the first policy spec to block for 300 seconds
before reporting failure, and the Ordered/Serial suite then skipped all
remaining specs.
Use gomega.StopTrying to abort the Eventually loop as soon as the pod
enters Failed or Succeeded, and surface container exit codes and
termination reasons in the failure message so the root cause is
immediately visible in CI logs.
Signed-off-by: Tim <tim.wang03@sap.com>
* fix(e2e): harden policy suite against device-plugin transitions
The policy E2E suite fails when the previous pod suite's cleanup removes
the gpu=on label and the policy suite re-adds it within ~133ms. The
DaemonSet controller may terminate the device-plugin pod after our
readiness checks pass, causing container startup failures.
Three layered fixes:
1. Add hasTerminatingDevicePlugin check to ensureGPUNodeLabeled and
registeredGPUNodes — blocks test start while a device-plugin pod
is still being torn down.
2. Switch RestartPolicy from Never to OnFailure — kubelet retries
container startup through transient device-plugin unavailability
instead of parking in permanent Failed phase.
3. Add StopTrying + podTerminalReason in waitForAllocatedPod — if a
pod still reaches terminal Failed, abort in ~2s with exit codes
and termination reasons instead of burning the full 5min timeout.
Signed-off-by: Tim <tim.wang03@sap.com>
* fix(e2e): wait for device-plugin healthy devices before creating GPU pods
The Pod E2E suite creates GPU pods immediately after adding the gpu=on
label, without waiting for the device-plugin to register healthy devices
with kubelet. This causes intermittent 'no healthy devices present'
admission errors when the device-plugin hasn't finished initializing.
Add WaitForDevicePluginReady to test/utils that polls until:
- At least one device-plugin pod is Running and Ready on the node
- No device-plugin pods are terminating (DaemonSet revision transition)
- The node reports GPU resources in Allocatable (healthy device registration)
Call it from the Pod E2E BeforeAll after labeling the node.
Signed-off-by: Tim <tim.wang03@sap.com>
---------
Signed-off-by: Tim <tim.wang03@sap.com>1 parent 9754da5 commit d798059
6 files changed
Lines changed: 1022 additions & 13 deletions
File tree
- hack
- test
- e2e
- pod
- policy
- utils
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
122 | 122 | | |
123 | 123 | | |
124 | 124 | | |
| 125 | + | |
| 126 | + | |
| 127 | + | |
| 128 | + | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
| 1 | + | |
| 2 | + | |
| 3 | + | |
| 4 | + | |
| 5 | + | |
| 6 | + | |
| 7 | + | |
| 8 | + | |
| 9 | + | |
| 10 | + | |
| 11 | + | |
| 12 | + | |
| 13 | + | |
| 14 | + | |
| 15 | + | |
| 16 | + | |
| 17 | + | |
| 18 | + | |
| 19 | + | |
| 20 | + | |
| 21 | + | |
| 22 | + | |
| 23 | + | |
| 24 | + | |
| 25 | + | |
| 26 | + | |
| 27 | + | |
| 28 | + | |
| 29 | + | |
| 30 | + | |
| 31 | + | |
| 32 | + | |
| 33 | + | |
| 34 | + | |
| 35 | + | |
| 36 | + | |
| 37 | + | |
| 38 | + | |
| 39 | + | |
| 40 | + | |
| 41 | + | |
| 42 | + | |
| 43 | + | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
57 | 57 | | |
58 | 58 | | |
59 | 59 | | |
| 60 | + | |
| 61 | + | |
| 62 | + | |
| 63 | + | |
60 | 64 | | |
61 | 65 | | |
62 | 66 | | |
| |||
0 commit comments