Skip to content

PERF: reduce Xavier GPU snapshot eviction overhead - #5624

Merged
qinxuye merged 4 commits into
xorbitsai:mainfrom
qinxuye:perf/xavier-packed-snapshot-demotion
Oct 4, 2026
Merged

qinxuye merged 4 commits into
xorbitsai:mainfrom
qinxuye:perf/xavier-packed-snapshot-demotion

Conversation

@qinxuye

@qinxuye qinxuye commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

GPU-first Xavier snapshots retain evicted GPU blocks on CPU. As the cold-prefix history grows, admitting each new block scans the CPU history to find a GPU victim and performs a synchronous device-to-host copy for every layer. This delays new prefills even when downstream consumers hit GPU snapshots.

Current review range after #5623 merged

Rebased only this PR onto main 35165f552ea0eb36da23715abe592cb11b15ac72 (the #5623 merge). Current head: 234bcdb30dffa077c81f498e5ec3f1b90220f9d5. All preceding dependencies are merged; the normal Files changed tab now contains only this increment.

Current review diff.

The rebase retains the merged async-load, lease-cleanup, slab-selection, shutdown and failure-reporting fixes. Both sides of the appended-test conflict are preserved. The reused-layer eviction assertion now covers both CPU/GPU snapshot placements from the merged parameterized test. Shutdown additionally clears the new GPU-only LRU index, checked for both successful and failed fences.

Current validation:

  • Local Xavier regression: 211 passed, 10 skipped, 1 deselected (the known baseline test_block_tracker mismatch).
  • CUDA-enabled test_gpu_transfer.py, test_tiered_snapshot.py and test_request_transfer.py on gpu_remote: 103 passed, including packed mixed-dtype demotion and failed-copy atomicity.
  • All modified-file pre-commit checks and git diff --check passed.
  • Existing explicit torch.long indices are retained. The partial-hit gather optimization is tracked in the existing perf: reduce Xavier cold-prefill snapshot overhead #5627 follow-up, which includes rechecking hits after preceding admissions can evict them; it is not pulled into this PR. The suggested zero-length replacement would still divide by zero, while Tensor.element_size() already avoids allocations, so that suggestion was not applied.

Historical performance measurements below remain attributed to their original pre-rebase heads; they were not rerun as new throughput claims. Only #5624 was rebased/pushed; later PRs are unchanged to limit CI usage.

Follow-up test coverage

The latest review fixes only tests. Both packed-demotion tests now run CPU and CUDA variants: the CPU case substitutes device metadata only during dispatch, exercising the actual grouping, cat, split, reshape, shape/dtype checks, copy independence and failure atomicity without CUDA. The physical CUDA variants remain intact. A shared GPU-LRU invariant checks its order against the GPU subsequence of global LRU, placement keys and tier counts; it now covers CPU-full/leased eviction, unpublished staging cleanup after OOM/runtime failures, failed packed copies and ordinary hits/leases.

Validation: 211 passed, 10 skipped, 1 known-baseline deselection locally; test_tiered_snapshot.py on gpu_remote: 20 passed, including both CPU and real CUDA variants. Modified-file pre-commit and diff checks passed. No runtime code or later PRs changed.

Previous implementation and validation record

This change:

  • Keeps a GPU-only LRU index aligned with the existing global order, including staging reuse, leases, demotion and failed copies.
  • Packs compatible layers of each evicted block into one host copy per device/dtype group. Host allocations belong to one block, so eviction does not retain unrelated snapshots.
  • Skips redundant LRU refreshes when consecutive cached layers reuse the same block sequence. New copies and changed sequences still refresh the order.

CPU overflow, logical dtype metadata, lease protection, and copy-failure atomicity remain intact. Temporary packing storage is bounded by one block per group and remains outside the retained-snapshot budget.

Historical measurement baseline: #5623 revision ad9c5b8a5; measured candidate commits: 1800311b7 and bf1cc5a25. The current review range is the rebased increment linked above.

Validation:

  • Local Xavier tests: 151 passed, 9 skipped, 1 known baseline test deselected (test_block_tracker, which expects retained empty tracker entries).
  • Real CUDA transfer, tiered-snapshot and request-transfer tests: 74 passed.
  • All modified-file pre-commit checks passed.
  • Final benchmark and CUDA-test source hashes match bf1cc5a25.
  • All 6072 final benchmark requests succeeded, with no server tracebacks. All 72 initial responses match the baseline exactly. Concurrent text identity is not asserted because batching differs.

Performance on two RTX 3090 Ti GPUs with NVLink, Qwen2.5-0.5B BF16, eager vLLM 0.21.0, torch 2.11.0, xoscar 0.11.1 and NIXL 1.1.0:

Interleaved workload metric #5623 baseline This change Native vLLM NIXL
Incoming long-prompt TTFT p50 4677.3 ms 2494.5 ms 369.0 ms
Incoming long-prompt TTFT p95 12209.8 ms 4346.0 ms 1036.1 ms
Aggregate output tokens/s 552.5 965.9 1241.4
Ongoing-stream client chunk gap p99 15.3 ms 15.6 ms 13.1 ms

Incoming TTFT median decreases 46.7%, p95 decreases 64.4%, and workload throughput increases 74.8%. The throughput gap to native NIXL shrinks from 55.5% to 22.2%; incoming TTFT remains substantially higher. These measurements describe this interleaved workload, not steady-state throughput for every workload. Chunk gaps describe client-visible chunks, not individual GPU token timing.

Warm-cache results retain a small tradeoff:

Concurrency Tokens/s baseline → candidate Change TTFT p50 / p95 baseline → candidate TPOT p95 baseline → candidate
16 1601.0 → 1580.1 -1.3% 152.0 / 249.8 → 169.2 / 252.8 ms 8.01 → 8.10 ms
32 2341.2 → 2298.7 -1.8% 326.9 / 548.8 → 307.2 / 567.1 ms 10.10 → 12.04 ms

Method: ABBA (baseline, candidate, candidate, baseline), followed by native NIXL and an 8 MiB GPU-budget candidate. All six deployments start fresh with the same 12 initial prompts. GPU ABBA deployments run 600 warm requests each at C16/C32; native and mixed validation deployments run 120 at each concurrency using the same unique prompts. Every deployment then runs the identical three interleaved rounds: eight ongoing streams, followed by 32 unique approximately 4400-token prompts staggered by 100 ms. Incoming output is capped at 64 tokens; ongoing output is capped at 1024, with actual lengths affected by EOS. Prefix caching is enabled, model length is 8192, and profiling is disabled. GPU-first runs use a 256 MiB snapshot budget.

Both baseline and candidate retain 1365 GPU and 26452 CPU snapshots after 26452 demotions. The 8 MiB candidate successfully exercises 98 GPU and 549 CPU batches and retains 42 GPU and 27775 CPU blocks, preserving CPU overflow. The test services were shut down after validation.

@XprobeBot XprobeBot added enhancement New feature or request gpu labels Oct 3, 2026
@XprobeBot XprobeBot added this to the v3.x milestone Oct 3, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces an experimental GPU-first Xavier cache feature for the vLLM transfer backend, leveraging xoscar NIXL and CUDA IPC to optimize KV cache transfers. It includes documentation, frontend options, backend validation, and lifecycle management for the GPU cache budget, alongside comprehensive unit tests. The review feedback highlights opportunities to optimize staging by filtering out already-cached blocks to prevent redundant GPU copies, ensuring robust indexing by explicitly specifying dtype=torch.long, and safely calculating block sizes using dtype.itemsize to avoid potential indexing errors.

Comment thread xinference/model/llm/vllm/xavier/gpu_transfer.py Outdated
Comment thread xinference/model/llm/vllm/xavier/gpu_transfer.py
Comment thread xinference/model/llm/vllm/xavier/gpu_transfer.py
@qinxuye
qinxuye force-pushed the perf/xavier-packed-snapshot-demotion branch from bf1cc5a to 2a42c96 Compare October 3, 2026 14:51
@qinxuye
qinxuye force-pushed the perf/xavier-packed-snapshot-demotion branch from 2a42c96 to 1a72cad Compare October 4, 2026 07:12
@qinxuye
qinxuye requested a review from rogercloud October 4, 2026 07:14

@rogercloud rogercloud left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor

  • xinference/model/llm/vllm/xavier/test/test_tiered_snapshot.py:206 — the packed demotion path (_copy_to_cpu cat/split/reshape and failed-copy atomicity) is only covered by CUDA-gated tests, which no automated CI job runs. Make the packing branch reachable on CPU in tests so CPU CI covers it.
  • xinference/model/llm/vllm/xavier/test/test_tiered_snapshot.py:280 — the _gpu_lru vs. global-order invariant is asserted in only one test; the drop-when-CPU-full-and-leased branch (tiered_snapshot.py:116) and GPUTransfer.stage failure cleanup never check it. Add a shared invariant helper and call it in those tests.

Blocking: no — recommended event: APPROVE

Comment thread xinference/model/llm/vllm/xavier/test/test_tiered_snapshot.py Outdated
Comment thread xinference/model/llm/vllm/xavier/test/test_tiered_snapshot.py Outdated
@qinxuye
qinxuye merged commit b9dc5ce into xorbitsai:main Oct 4, 2026
13 of 15 checks passed
@qinxuye
qinxuye deleted the perf/xavier-packed-snapshot-demotion branch October 4, 2026 09:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request gpu

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants