Repository navigation
perf: reduce Xavier cold-prefill snapshot overhead - #5627
Merged
Merged
Conversation
This was referenced Oct 3, 2026
qinxuye
force-pushed
the
perf/xavier-cold-request-staging
branch
from
October 4, 2026 10:43
8274258 to
26388f2
Compare
rogercloud
approved these changes
Oct 4, 2026
rogercloud
left a comment
Contributor
There was a problem hiding this comment.
Minor
xinference/model/llm/vllm/xavier/gpu_transfer.py:241— Two entries in onestage()call that share a new prefix: the second entry falls back to the per-layer path for all of its keys (see inline).xinference/model/llm/vllm/xavier/test/test_tiered_snapshot.py:364/test_gpu_transfer.py:1317— The following are untested: theMAX_REQUEST_BLOCKScap (the bound test hits only the byte cap), the oversized-blockcount == 1case, a failure in a later chunk after an earlier chunk succeeded (earlier-chunk keys must be dropped by the outer cleanup), and same-dtype layers with different per-block shapes (the non-uniformbranch ofstage_blocksis never exercised). Add cases for these.xinference/model/llm/vllm/xavier/test/test_gpu_transfer.py:1363—stage()swallows the stub'sAssertionError(see inline).xinference/model/llm/vllm/xavier/gpu_transfer.py:260— Optional: a packed chunk is copied three times on the device (index_selectper layer, thentorch.cat, then the per-block copy), and the gathered dict and the cat result are both alive during the per-block copies. Gathering straight into one packed buffer, or droppingvaluesafterstage_blocksconsumes it, would save one copy and halve the temporary memory.
Simplification
tiered_snapshot.pyL197: shrink: theuniformflag and theunbindbranch duplicate what split+reshape already does. Always use[v.reshape(s) for v, s in zip(block.split(sizes), shapes)], the same as_copy_to_cpu.
net: -8 lines possible
Blocking: no — recommended event: APPROVE
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Cold requests synchronously snapshot their KV data before prefill completes. Copying each layer of each block separately delays subsequent arrivals. Pack compatible layers into complete blocks and copy each block once per dtype, including requests with a cached chat-template prefix.
Current review range after #5625 merged
Rebased only this PR onto main
1be7706141846dda3ef0ef4e5155ddf60151614e(the #5625 merge). Current head:c568eee6802f6919817e704726112effb244e5f3. All preceding dependencies are merged; this is now independently reviewable through the normal Files changed tab.Current review diff.
The rebase preserves the merged async/lease cleanup, packed-demotion CPU coverage, GPU-LRU invariants, size-accounting failure assertions and cancellation logging. Test conflicts were resolved by retaining both sets of coverage: packed and layer staging failures now both check LRU and physical byte-count cleanup. The newly introduced packed-gather index explicitly uses
torch.long, matching the earlier index fixes.Current validation:
test_block_trackermismatch).gpu_remote: 128 passed.git diff --checkpassed.Historical performance measurements below remain attributed to their original heads. They were not rerun as new throughput claims for this rebase.
Changes
Latest review follow-up
Validation: local Xavier 231 passed, 13 skipped, 1 known-baseline deselection; CUDA-enabled transfer/tiered/request tests on
gpu_remote128 passed; pre-commit and diff checks passed. No new end-to-end throughput claim is made for these review changes.The optional gather-directly-into-packed-storage optimization is deferred. It would change the staging buffer layout and lifetime; deleting the gathered dictionary only after
stage_blocksreturns would not reduce peak memory during its per-block copies. Existing chunk limits remain in place, and no temporary-memory reduction is claimed in this update.Historical validation and benchmark revision
Measured head
827425856; baseline #5625c98d91ad9. Benchmark runtime hashes match that original commit; current-head correctness validation is listed above.Measured performance
Two RTX 3090 Ti GPUs with NVLink; Qwen2.5-0.5B BF16, eager, vLLM 0.21, xoscar 0.11.1/NIXL 1.1. Maximum model length 8192; GPU memory utilization 0.6; prefix caching enabled. Xavier snapshot budget 256 MiB. Profiling disabled for comparisons.
Fresh deployments in baseline/candidate/candidate/baseline order, with 600 warm requests at each of C16 and C32 per deployment. Each deployment also runs three rounds with eight background decode streams (up to 1024 output tokens), then 32 distinct long prompts injected at 100 ms intervals (64 output tokens). Token throughput uses actual output lengths. The native anchor is one fresh deployment with three identical arrival rounds.
The measured throughput gap to native shrinks from 21.9% to 11.5%; cold latency remains substantially higher. Warm throughput changes are small: C16 1602.0 -> 1583.5 token/s (-1.2%), C32 2260.5 -> 2289.3 (+1.3%). This is a cold-arrival improvement with a small increase in background token gaps, not a general warm-throughput claim. These results cover this model and topology.
CPU-overflow control
An additional matched baseline/candidate run uses an 8 MiB GPU snapshot budget. Incoming TTFT p50 improves 2469.9 -> 1667.7 ms, p95 3648.1 -> 2276.6 ms, and arrival-workload throughput 1045.2 -> 1118.3 token/s (+7.0%). Background token gap p95 changes 24.31 -> 25.85 ms.
A short 120-request C32 trial showed a 7% warm-throughput drop. A separate longer ABBA check, 1,200 warm requests per version, did not reproduce it: 2245.6 -> 2234.9 token/s (-0.5%); mean trial TTFT p95 559.0 -> 551.5 ms. No mixed-tier warm-throughput gain is claimed.
Across the primary suite and both controls: 9,864 successful requests, four intentional cancellations, and 144 identical serial initial outputs. Cancellation recovery includes 600 successful requests; transfer-actor FD counts remain 124/121 through the recorded recovery samples. Both GPUs return to 554 MiB / 0% utilization after cleanup.