Skip to content

perf: reduce Xavier cold-prefill snapshot overhead - #5627

Merged
qinxuye merged 3 commits into
xorbitsai:mainfrom
qinxuye:perf/xavier-cold-request-staging
Oct 4, 2026
Merged

qinxuye merged 3 commits into
xorbitsai:mainfrom
qinxuye:perf/xavier-cold-request-staging

Conversation

@qinxuye

@qinxuye qinxuye commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Cold requests synchronously snapshot their KV data before prefill completes. Copying each layer of each block separately delays subsequent arrivals. Pack compatible layers into complete blocks and copy each block once per dtype, including requests with a cached chat-template prefix.

Current review range after #5625 merged

Rebased only this PR onto main 1be7706141846dda3ef0ef4e5155ddf60151614e (the #5625 merge). Current head: c568eee6802f6919817e704726112effb244e5f3. All preceding dependencies are merged; this is now independently reviewable through the normal Files changed tab.

Current review diff.

The rebase preserves the merged async/lease cleanup, packed-demotion CPU coverage, GPU-LRU invariants, size-accounting failure assertions and cancellation logging. Test conflicts were resolved by retaining both sets of coverage: packed and layer staging failures now both check LRU and physical byte-count cleanup. The newly introduced packed-gather index explicitly uses torch.long, matching the earlier index fixes.

Current validation:

Historical performance measurements below remain attributed to their original heads. They were not rerun as new throughput claims for this rebase.

Changes

  • Gather only missing blocks, in chunks bounded by 64 blocks and 16 MiB of useful KV data (or one block if larger).
  • Compatible layers share a single block's storage; separate blocks own separate allocations. Evicting a block cannot retain the entire staging batch.
  • Preserve GPU-first placement, CPU overflow, read leases, publication fencing, and failed-stage cleanup. Recheck cached hits after preceding admissions because those admissions may evict them.
  • Use the existing layer path for partially staged blocks or layers with different block mappings.

Latest review follow-up

  • A stage call now tracks complete blocks it packed itself. Subsequent entries sharing a new prefix reuse those snapshots without waiting for publication or falling back to per-layer staging. The set is local to the call; unpublished blocks from another call retain the existing fallback. Hits are rechecked after preceding admissions can evict them.
  • Tests cover same-call prefix reuse, separate-call unpublished protection, and an admission-evicted call-local hit; gather assertions verify only missing blocks are gathered in the shared-prefix case.
  • Bounded staging now has separate byte-cap, block-count-cap and oversized-single-block cases. A failure injected into a later chunk verifies outer cleanup drops unpublished keys from both earlier and failing chunks while preserving published content, byte counts and GPU-LRU consistency.
  • Removed the redundant uniform/unbind branch. All layer shapes use split+reshape; CPU/CUDA tests now cover same-dtype layers with different shapes, shared per-block storage and independent block allocations.
  • The fallback test records packed calls and asserts none occurred after stage returns, so its assertion cannot be swallowed by production error handling.

Validation: local Xavier 231 passed, 13 skipped, 1 known-baseline deselection; CUDA-enabled transfer/tiered/request tests on gpu_remote 128 passed; pre-commit and diff checks passed. No new end-to-end throughput claim is made for these review changes.

The optional gather-directly-into-packed-storage optimization is deferred. It would change the staging buffer layout and lifetime; deleting the gathered dictionary only after stage_blocks returns would not reduce peak memory during its per-block copies. Existing chunk limits remain in place, and no temporary-memory reduction is claimed in this update.

Historical validation and benchmark revision

Measured head 827425856; baseline #5625 c98d91ad9. Benchmark runtime hashes match that original commit; current-head correctness validation is listed above.

  • Pre-commit passed.
  • Local Xavier/PD/model/launch regression: 238 passed, 11 skipped, one known baseline tracker test deselected.
  • GPU-host Xavier regression: 183 passed, two skipped, the same baseline test deselected.
  • New tests cover all BF16 bit patterns, mixed dtypes/shapes, independent block ownership, CPU overflow, leases, shared-prefix hits, admission-time eviction, bounded gathering, and cleanup after copy failures.
  • Primary end-to-end suite: 6,672 successful requests plus four intentional streaming cancellations followed by 600 successful recovery requests. All 72 serial initial outputs match. Concurrent text is not deterministic even between baseline repeats (24/120 differ; candidate repeats 19/120); KV correctness is checked directly by the tensor/bit-pattern regressions.

Measured performance

Two RTX 3090 Ti GPUs with NVLink; Qwen2.5-0.5B BF16, eager, vLLM 0.21, xoscar 0.11.1/NIXL 1.1. Maximum model length 8192; GPU memory utilization 0.6; prefix caching enabled. Xavier snapshot budget 256 MiB. Profiling disabled for comparisons.

Fresh deployments in baseline/candidate/candidate/baseline order, with 600 warm requests at each of C16 and C32 per deployment. Each deployment also runs three rounds with eight background decode streams (up to 1024 output tokens), then 32 distinct long prompts injected at 100 ms intervals (64 output tokens). Token throughput uses actual output lengths. The native anchor is one fresh deployment with three identical arrival rounds.

Arrival workload metric Baseline Candidate Native NIXL anchor
Incoming TTFT p50 2463.2 ms 1782.3 ms (-27.6%) 342.5 ms
Incoming TTFT p95 4115.5 ms 2738.0 ms (-33.5%) 835.4 ms
Output throughput 971.1 token/s 1099.9 token/s (+13.3%) 1242.7 token/s
Background token gap p95 12.79 ms 13.48 ms 10.49 ms
Background token gap p99 15.60 ms 16.82 ms 13.35 ms

The measured throughput gap to native shrinks from 21.9% to 11.5%; cold latency remains substantially higher. Warm throughput changes are small: C16 1602.0 -> 1583.5 token/s (-1.2%), C32 2260.5 -> 2289.3 (+1.3%). This is a cold-arrival improvement with a small increase in background token gaps, not a general warm-throughput claim. These results cover this model and topology.

CPU-overflow control

An additional matched baseline/candidate run uses an 8 MiB GPU snapshot budget. Incoming TTFT p50 improves 2469.9 -> 1667.7 ms, p95 3648.1 -> 2276.6 ms, and arrival-workload throughput 1045.2 -> 1118.3 token/s (+7.0%). Background token gap p95 changes 24.31 -> 25.85 ms.

A short 120-request C32 trial showed a 7% warm-throughput drop. A separate longer ABBA check, 1,200 warm requests per version, did not reproduce it: 2245.6 -> 2234.9 token/s (-0.5%); mean trial TTFT p95 559.0 -> 551.5 ms. No mixed-tier warm-throughput gain is claimed.

Across the primary suite and both controls: 9,864 successful requests, four intentional cancellations, and 144 identical serial initial outputs. Cancellation recovery includes 600 successful requests; transfer-actor FD counts remain 124/121 through the recorded recovery samples. Both GPUs return to 554 MiB / 0% utilization after cleanup.

@qinxuye
qinxuye force-pushed the perf/xavier-cold-request-staging branch from 8274258 to 26388f2 Compare October 4, 2026 10:43
@qinxuye
qinxuye requested a review from rogercloud October 4, 2026 10:45

@rogercloud rogercloud left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor

  • xinference/model/llm/vllm/xavier/gpu_transfer.py:241 — Two entries in one stage() call that share a new prefix: the second entry falls back to the per-layer path for all of its keys (see inline).
  • xinference/model/llm/vllm/xavier/test/test_tiered_snapshot.py:364 / test_gpu_transfer.py:1317 — The following are untested: the MAX_REQUEST_BLOCKS cap (the bound test hits only the byte cap), the oversized-block count == 1 case, a failure in a later chunk after an earlier chunk succeeded (earlier-chunk keys must be dropped by the outer cleanup), and same-dtype layers with different per-block shapes (the non-uniform branch of stage_blocks is never exercised). Add cases for these.
  • xinference/model/llm/vllm/xavier/test/test_gpu_transfer.py:1363 — stage() swallows the stub's AssertionError (see inline).
  • xinference/model/llm/vllm/xavier/gpu_transfer.py:260 — Optional: a packed chunk is copied three times on the device (index_select per layer, then torch.cat, then the per-block copy), and the gathered dict and the cat result are both alive during the per-block copies. Gathering straight into one packed buffer, or dropping values after stage_blocks consumes it, would save one copy and halve the temporary memory.

Simplification

  • tiered_snapshot.py L197: shrink: the uniform flag and the unbind branch duplicate what split+reshape already does. Always use [v.reshape(s) for v, s in zip(block.split(sizes), shapes)], the same as _copy_to_cpu.
    net: -8 lines possible

Blocking: no — recommended event: APPROVE

Comment thread xinference/model/llm/vllm/xavier/gpu_transfer.py Outdated
Comment thread xinference/model/llm/vllm/xavier/test/test_gpu_transfer.py
Comment thread xinference/model/llm/vllm/xavier/tiered_snapshot.py Outdated
@qinxuye
qinxuye merged commit e0db7ad into xorbitsai:main Oct 4, 2026
13 of 15 checks passed
@qinxuye
qinxuye deleted the perf/xavier-cold-request-staging branch October 4, 2026 12:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request gpu

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants