Repository navigation
PERF: reduce Xavier GPU snapshot eviction overhead - #5624
Merged
qinxuye merged 4 commits intoOct 4, 2026
Merged
Conversation
Contributor
There was a problem hiding this comment.
Code Review
This pull request introduces an experimental GPU-first Xavier cache feature for the vLLM transfer backend, leveraging xoscar NIXL and CUDA IPC to optimize KV cache transfers. It includes documentation, frontend options, backend validation, and lifecycle management for the GPU cache budget, alongside comprehensive unit tests. The review feedback highlights opportunities to optimize staging by filtering out already-cached blocks to prevent redundant GPU copies, ensuring robust indexing by explicitly specifying dtype=torch.long, and safely calculating block sizes using dtype.itemsize to avoid potential indexing errors.
qinxuye
force-pushed
the
perf/xavier-packed-snapshot-demotion
branch
from
October 3, 2026 14:51
bf1cc5a to
2a42c96
Compare
qinxuye
force-pushed
the
perf/xavier-packed-snapshot-demotion
branch
from
October 4, 2026 07:12
2a42c96 to
1a72cad
Compare
rogercloud
approved these changes
Oct 4, 2026
rogercloud
left a comment
Contributor
There was a problem hiding this comment.
Minor
xinference/model/llm/vllm/xavier/test/test_tiered_snapshot.py:206— the packed demotion path (_copy_to_cpucat/split/reshape and failed-copy atomicity) is only covered by CUDA-gated tests, which no automated CI job runs. Make the packing branch reachable on CPU in tests so CPU CI covers it.xinference/model/llm/vllm/xavier/test/test_tiered_snapshot.py:280— the_gpu_lruvs. global-order invariant is asserted in only one test; the drop-when-CPU-full-and-leased branch (tiered_snapshot.py:116) andGPUTransfer.stagefailure cleanup never check it. Add a shared invariant helper and call it in those tests.
Blocking: no — recommended event: APPROVE
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
GPU-first Xavier snapshots retain evicted GPU blocks on CPU. As the cold-prefix history grows, admitting each new block scans the CPU history to find a GPU victim and performs a synchronous device-to-host copy for every layer. This delays new prefills even when downstream consumers hit GPU snapshots.
Current review range after #5623 merged
Rebased only this PR onto main
35165f552ea0eb36da23715abe592cb11b15ac72(the #5623 merge). Current head:234bcdb30dffa077c81f498e5ec3f1b90220f9d5. All preceding dependencies are merged; the normal Files changed tab now contains only this increment.Current review diff.
The rebase retains the merged async-load, lease-cleanup, slab-selection, shutdown and failure-reporting fixes. Both sides of the appended-test conflict are preserved. The reused-layer eviction assertion now covers both CPU/GPU snapshot placements from the merged parameterized test. Shutdown additionally clears the new GPU-only LRU index, checked for both successful and failed fences.
Current validation:
test_block_trackermismatch).test_gpu_transfer.py,test_tiered_snapshot.pyandtest_request_transfer.pyongpu_remote: 103 passed, including packed mixed-dtype demotion and failed-copy atomicity.git diff --checkpassed.torch.longindices are retained. The partial-hit gather optimization is tracked in the existing perf: reduce Xavier cold-prefill snapshot overhead #5627 follow-up, which includes rechecking hits after preceding admissions can evict them; it is not pulled into this PR. The suggested zero-length replacement would still divide by zero, whileTensor.element_size()already avoids allocations, so that suggestion was not applied.Historical performance measurements below remain attributed to their original pre-rebase heads; they were not rerun as new throughput claims. Only #5624 was rebased/pushed; later PRs are unchanged to limit CI usage.
Follow-up test coverage
The latest review fixes only tests. Both packed-demotion tests now run CPU and CUDA variants: the CPU case substitutes device metadata only during dispatch, exercising the actual grouping, cat, split, reshape, shape/dtype checks, copy independence and failure atomicity without CUDA. The physical CUDA variants remain intact. A shared GPU-LRU invariant checks its order against the GPU subsequence of global LRU, placement keys and tier counts; it now covers CPU-full/leased eviction, unpublished staging cleanup after OOM/runtime failures, failed packed copies and ordinary hits/leases.
Validation: 211 passed, 10 skipped, 1 known-baseline deselection locally;
test_tiered_snapshot.pyongpu_remote: 20 passed, including both CPU and real CUDA variants. Modified-file pre-commit and diff checks passed. No runtime code or later PRs changed.Previous implementation and validation record
This change:
CPU overflow, logical dtype metadata, lease protection, and copy-failure atomicity remain intact. Temporary packing storage is bounded by one block per group and remains outside the retained-snapshot budget.
Historical measurement baseline: #5623 revision
ad9c5b8a5; measured candidate commits:1800311b7andbf1cc5a25. The current review range is the rebased increment linked above.Validation:
test_block_tracker, which expects retained empty tracker entries).bf1cc5a25.Performance on two RTX 3090 Ti GPUs with NVLink, Qwen2.5-0.5B BF16, eager vLLM 0.21.0, torch 2.11.0, xoscar 0.11.1 and NIXL 1.1.0:
Incoming TTFT median decreases 46.7%, p95 decreases 64.4%, and workload throughput increases 74.8%. The throughput gap to native NIXL shrinks from 55.5% to 22.2%; incoming TTFT remains substantially higher. These measurements describe this interleaved workload, not steady-state throughput for every workload. Chunk gaps describe client-visible chunks, not individual GPU token timing.
Warm-cache results retain a small tradeoff:
Method: ABBA (baseline, candidate, candidate, baseline), followed by native NIXL and an 8 MiB GPU-budget candidate. All six deployments start fresh with the same 12 initial prompts. GPU ABBA deployments run 600 warm requests each at C16/C32; native and mixed validation deployments run 120 at each concurrency using the same unique prompts. Every deployment then runs the identical three interleaved rounds: eight ongoing streams, followed by 32 unique approximately 4400-token prompts staggered by 100 ms. Incoming output is capped at 64 tokens; ongoing output is capped at 1024, with actual lengths affected by EOS. Prefix caching is enabled, model length is 8192, and profiling is disabled. GPU-first runs use a 256 MiB snapshot budget.
Both baseline and candidate retain 1365 GPU and 26452 CPU snapshots after 26452 demotions. The 8 MiB candidate successfully exercises 98 GPU and 549 CPU batches and retains 42 GPU and 27775 CPU blocks, preserving CPU overflow. The test services were shut down after validation.