Repository navigation
perf: default Xavier PD to GPU handoff with tiered KV history - #5628
Merged
qinxuye merged 9 commits intoOct 4, 2026
Merged
Conversation
This was referenced Oct 4, 2026
qinxuye
force-pushed
the
perf/xavier-direct-tiered-history
branch
from
October 4, 2026 12:35
13925a3 to
fdaffdc
Compare
rogercloud
reviewed
Oct 4, 2026
rogercloud
left a comment
Contributor
There was a problem hiding this comment.
Major
xinference/model/llm/vllm/xavier/v1_connector.py:406— Decode re-issues the direct load after preemption, against a ticket that has already been released, and the failure crashes D's EngineCore.
Trigger: in vLLM 0.21,_preempt_requestresetsnum_computed_tokens=0and prepends the request to the waiting queue, which callsget_num_new_matched_tokensagain withxavier_directstill present.send_directthen raises, andget_finishedcallstask.result()and re-raises inside the worker. Impact: every in-flight request on D fails.xinference/model/llm/vllm/xavier/direct_handoff.py:88— The 120 s producer deadline keeps running while D is still queued, so a D backlog longer than 120 s ends on the same fatal path instead of a local recompute.
Trigger: D is saturated (max_num_seqs/KV) for more than 120 s after P finishes. Impact: D EngineCore crash.xinference/model/llm/vllm/xavier/direct_history.py:114— Themin(..., 64 MiB)pending cap skips history for whole requests. On an 8B GQA model (2 MiB/block) that is any prompt over about 512 tokens, regardless ofxavier_gpu_cache_bytes. Retain the leading prefix instead of skipping, or scale the cap with the budget. Count this skip separately from busy-writer skips and document it.
Minor
xinference/model/llm/vllm/xavier/v1_connector.py:580— If D aborts beforeget_num_new_matched_tokens, or the router fails or cancels after P succeeds, nobody releases the ticket (free_prefill_model_cacheis a no-op for direct handoff,xinference/core/pd_model.py:254), so P blocks stay pinned for 120 s. Releasekv_transfer_params["xavier_direct"]in both paths.xinference/model/llm/vllm/xavier/direct_handoff.py:170— In a multi-requestload_direct/load_history, the first failure skips releasing the remaining tickets and leases. The release infinallycan also mask the original error. This contradicts the README's "attempts every lease release" promise; mirrorload_gpu_requests_v1'sgather(..., return_exceptions=True).xinference/model/llm/vllm/xavier/direct_history.py:96—assert self.history.reserve(...)is the only reserve call, sopython -Odrops it. Call it first, then check the result.xinference/model/llm/vllm/xavier/v1_connector.py:237/:321— The producer does apoll_direct_gpu_v1RPC and atorch.cuda.synchronize()on every step, even with no direct requests pending. Base returned early here. Gate both on outstanding tickets or store requests.xinference/model/llm/vllm/xavier/direct_history.py:193— Retentions abandoned at the 10 ms soft deadline, and candidates dropped once the limit is reached (:133), update no metric, so missing history is invisible.- Tests: no coverage for decode re-entry after load/release (preemption), a D-side load after expiry, decode abort before load releasing the ticket, multi-request partial-failure cleanup, the producer
request_finishedhistory branch (hashes plusheld_blocks), or theget_finished→poll_direct_gpu_v1wiring. - Docs:
pd_separation.rstdoes not mention the 120 s handoff deadline, the 64 MiB history skip, or that history is restored only on P. Lines ~130-141 still describe the pre-direct CPU/GPU-first cache. xinference/core/pd_model.py:254/:357— For supervisor-launched P/D, direct handoff is now always on, so the snapshotfree_model_cachegather and theset_unpin_handlerbranch are unreachable.test_direct_handoff.py:284sets an env var that nothing reads.
Simplification
- v1_connector.py L549-565: shrink: two
register_direct_gpu_v1calls differ only in the history args. Build the args once and make a single call. - direct_history.py L155-164: stdlib: hand-built
exc_infotuple. Uselogger.error(..., exc_info=exc). - direct_history.py L37: delete:
cpu_capacityparam of_init_historyis test-only. Tests can setstore.cpu_capacityfirst.
net: -10 lines possible
Blocking: yes — recommended event: COMMENT
xinference/model/llm/vllm/xavier/v1_connector.py:406— major — D EngineCore crash on preemption re-load [new]xinference/model/llm/vllm/xavier/direct_handoff.py:88— major — D EngineCore crash when queued >120 s [new]
rogercloud
approved these changes
Oct 4, 2026
rogercloud
left a comment
Contributor
There was a problem hiding this comment.
Major
xinference/model/llm/vllm/xavier/direct_handoff.py:128— Once D claims a ticket, nothing on P can release it any more. See the inline comment.
Minor
xinference/model/llm/vllm/xavier/direct_history.py:123— History LRU order evicts the head of each prefix chain first. See the inline comment.xinference/model/llm/vllm/xavier/direct_history.py:133— Admission is not prefix-closed, so blocks can be stored that can never be reserved. See the inline comment.xinference/model/llm/vllm/xavier/test/test_pd_gpu.py:178— The default-history e2e case never shows that history is used. See the inline comment.- Tests: these new branches have no tests:
- D's full-local-hit release (
v1_connector.py:403-412,tokens <= num_computed_tokens); - P's re-reserve, which releases the previous lease (
v1_connector.py:357-362); - the router's stream-success pop of
_direct_transfers(xinference/core/pd_model.py:364).test_pd_model.pycovers only the non-stream path.
- D's full-local-hit release (
xinference/model/llm/vllm/xavier/direct_history.py:308-313—history_hit_blocksand the per-tier hit counters count the whole lease.TieredKVSnapshotStore.reservealso counts hits again on every scheduling retry, so the counters overstate restored blocks. Count hits once, after a successful load, over the blocks actually written.
Simplification
v1_connector.pyL133: delete: the_gpu_budget is Noneand role checks are already guaranteed byuses_direct_handoff(transport.py:100). Keep onlylen(kv_cache_config.kv_cache_groups) != 1.v1_connector.pyL356, L574: delete:getattr(self, "_history_enabled", False).__init__always sets the attribute, so useself._history_enabled.direct_handoff.pyL196: shrink: the innerrelease()re-importsTransferActorand rebuilds the actor ref. Callself.actor.release_remote_direct_gpu_v1(rank, ticket)instead.direct_history.pyL194: shrink: the deadline/closing drop is duplicated at L195-200 and L202-207. Fold both into onewhile True:loop that checks closing/deadline, then the locks, then sleeps.
net: -13 lines possible
Blocking: no — recommended event: APPROVE
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Xavier P/D currently stages transferable KV into a snapshot cache before decode can consume it. Make P/D use request-scoped GPU handoff by default, and retain independent GPU-first history after the handoff, spilling older unleased history to CPU when GPU capacity is exhausted.
xavier_gpu_cache_bytessetting. No new flag or environment variable.0disables history while retaining direct transfer.xoscar[nixl]>=0.11.1and vLLM >=0.21; missing NIXL fails launch. P-to-D transport uses xoscar NIXL GPU copies. CPU history is a capacity tier, restored locally to GPU in per-layer batches.n=1.#5627 and all preceding PRs are merged. Rebased onto main
e0db7adf2798cd4ec7f20f55102b4e5b15fdfd27; this is the next independent PR for review. Focused comparison: https://github.com/xorbitsai/inference/pull/5628/files .Follow-up review fixes (
e590a45ac166fc7a4216910ae02d54efdf3a323a)Validation: 428 local tests passed, 22 hardware/live tests skipped and the distributed tracker integration deselected. The two-GPU run passed 88 tests, including all three real P/D configurations and positive history restores; after the final counter-only adjustment, all 48 history tests passed again on the GPU host. Pre-commit, msgfmt/catalog compilation, English plus nine translated builds, and rendered idle-lease translations passed. Throughput and cache-reuse benchmark tables remain historical; they were not rerun for these changes.
Review fixes (2026-10-04,
121c3f9fa236b463df9bf71d47b8d20cf0ea0498)Validation on this revision: 417 local tests passed, 22 hardware/live cases skipped, distributed tracker integration deselected; 79 tests passed on the two RTX3090Ti/NVLink GPUs, including default Xavier history, history disabled, native NIXL end-to-end deployments and CUDA history/raw-bit tests. The history reserve/restore regression also passes under
python -O. Pre-commit, all nine msgfmt checks/catalog compilation, English plus nine translated page builds, and rendered new-paragraph checks passed.The historical performance tables below are unchanged and were not rerun after these lifecycle fixes. The atomic ticket claim adds a scheduler RPC; these correctness checks do not establish its throughput cost or a new performance comparison.
Rebase validation (2026-10-04, head
fdaffdc3a500ad00c07b1dfc6047ff9871bac144):8 == 4from counting both P and D completions. Integration includes streaming, repeated prompts, four concurrent requests and shutdown; local engine prefix caching is disabled. Fixed-prefix/raw-bit CUDA tests validate history separately.Historical performance validation below predates this rebase; throughput and cache-reuse benchmarks were not rerun for the new head. The native comparison and cache-reuse measurements explicitly identify original head
13925a3b31707e595fc64db634e0e0997e4e849c.A separate 1 MiB GPU-budget run disables local prefix caching to exercise history: 132 successful requests, 7,324 CPU-history block hits, 20 GPU-history block hits, 3,307 demotions, and zero CPU network-transfer batches. All direct handoffs finished without expiry or history-save failure. Total across the controlled comparisons and this capacity-tier check: 2,740 successful requests, zero request errors.
The normal-prefix performance trials hit vLLM's local cache and do not establish a throughput benefit from CPU history itself. These measurements are specific to this model, workload and hardware. The prior output-divergence diagnosis also showed that BF16 greedy outputs must be compared at the same reused-prefix length; KV bit checks and fixed-length comparisons are distinct from throughput measurements.
Native vLLM NIXL comparison (fresh paired run on
13925a3b31707e595fc64db634e0e0997e4e849c): same hardware, model, engine configuration and request sequence as above. Both backends run through the Xinference P/D route; the native column selects vLLM's native NIXL connector, while Xavier uses xoscar NIXL with its default 256 MiB GPU history and CPU overflow. Each backend uses a fresh server/model and handles 652 requests (12 serial initialization, 300 C16, 300 C32, and 40 overlap requests including 32 cold arrivals).All 1,304 requests succeeded. Xavier registered and finished all 652 handoffs, with zero expiry, history-save failures or CPU network batches. History remained enabled: 1,835 saved blocks and 470 GPU-to-CPU demotions. History hits were zero with the normal local-prefix-cache workload, so this compares the direct path including history-write cost, not history-hit benefit. This is one paired trial, not a statistically established P95 advantage; native NIXL still leads throughput and median cold TTFT on this workload. It is a connector comparison through the same Xinference API, not a standalone
vllm servebenchmark.Cache-reuse benchmark (2026-10-04)
These tests keep the P hop: the scheduler still calls P, history is restored on P, and D consumes the handoff. No P-bypass routing is involved. They use the same commit and runtime versions as the native comparison above, one P and one D, Qwen2.5-Instruct 0.5B BF16/eager, and both backends have vLLM local prefix caching enabled. Requests are serial (C1), about 4,000 input tokens and 16 output tokens, temperature 0,
ignore_eos=true. All numbers below are client-observed TTFT in milliseconds.To induce eviction with a short benchmark,
num_gpu_blocks_override=1024limits each engine to 16,384 cached tokens (192 MiB KV). Xavier retains its default 256 MiB GPU history plus a CPU history capacity of 1,024 blocks (192 MiB on P in this test). Native NIXL is also tested with 2,389 engine blocks: 192 MiB + the same 1,365 whole blocks that fit Xavier's 256 MiB history budget. This matches GPU KV storage capacity, not total process GPU usage or total host memory; Xavier still has the additional CPU tier. This is an intentionally constrained-cache microbenchmark, not the default engine capacity or a production-size-model result.Fixed hot set: initialize four different documents once, then repeat six cycles of: immediately repeat all four documents; issue 12 new documents with different prefixes; revisit the same four hot documents. Each backend/configuration starts a fresh server/model and handles 125 requests including one startup request. Each immediate/revisit distribution contains 24 observations. A document is a distinct identifying prefix followed by
Water evaporates in sunlight and condenses into clouds.repeated 360 times andExplain this process.; pressure-document identifiers change every cycle. Pressure exceeds even the expanded native engine's 38,224-token capacity.Changing working set / negative control: a separate pair of 73-request runs uses three cycles with four new target documents per cycle, immediate repeat, 12 new pressure documents, then target revisit. It uses 1,024 engine blocks for both backends. This is not the stable-hot-set workload above.
Here history still loads 12 times, but only 1,486 blocks are hit in total (about 124 blocks per load), versus 249 per load in the fixed-hot-set test; 355 hits are CPU and 1,131 GPU. History reaches capacity, with admission rejections and bounded retention, so cache presence does not imply nearly complete prefix reuse. Revisit median latency does not improve overall, although the mean and P95 do. The first cycle's Xavier revisit TTFTs are 88–106 ms; later cycles are about 128–144 ms. This limits the claim to workloads with sufficient reusable history, not arbitrary churn.
All six cache-reuse runs total 646 requests, zero request errors (500 fixed-hot-set/control requests and 146 changing-working-set requests). The runs are sequential on the same machine; the six cycles within a run are not six independent fresh-server trials. P95 has only 24 samples per hot-set configuration, and Xavier had one 249.9 ms revisit outlier, so the lower P95 does not mean every request is faster. DEBUG logging was enabled equally for these diagnostic runs; these results should not be combined directly with the earlier C16/C32 throughput measurements.
Conclusion: tiered history demonstrably reduces TTFT for recurring long prefixes after engine-cache eviction, including against a native NIXL control with matched GPU KV capacity, at the cost of additional CPU history storage. It does not establish a benefit for immediate engine-cache hits, arbitrary working-set churn, saturated throughput, or other model sizes. Keep the simple P route; prioritize history admission/retention and CPU-restore efficiency, then validate on larger models/concurrent reuse before making broader claims.