Skip to content

perf: default Xavier PD to GPU handoff with tiered KV history - #5628

Merged
qinxuye merged 9 commits into
xorbitsai:mainfrom
qinxuye:perf/xavier-direct-tiered-history
Oct 4, 2026
Merged

qinxuye merged 9 commits into
xorbitsai:mainfrom
qinxuye:perf/xavier-direct-tiered-history

Conversation

@qinxuye

@qinxuye qinxuye commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Xavier P/D currently stages transferable KV into a snapshot cache before decode can consume it. Make P/D use request-scoped GPU handoff by default, and retain independent GPU-first history after the handoff, spilling older unleased history to CPU when GPU capacity is exhausted.

  • P/D defaults to a 256 MiB GPU history budget per replica through the existing xavier_gpu_cache_bytes setting. No new flag or environment variable. 0 disables history while retaining direct transfer.
  • Require xoscar[nixl]>=0.11.1 and vLLM >=0.21; missing NIXL fails launch. P-to-D transport uses xoscar NIXL GPU copies. CPU history is a capacity tier, restored locally to GPU in per-layer batches.
  • Retain producer ownership through decoder reads and optional history copies. Bound history staging to one writer and a soft deadline, protect read leases, and require observed reuse before replacing content when both tiers are full. CPU history remains bounded by engine KV block capacity.
  • Add fixed-prefix (16/2576/4432/4440 tokens) mapping and raw BF16 bit regressions, CPU spill/restore and lease tests, route/config tests, and CPU/GPU CI coverage. Update the guide and all nine translations. Direct handoff currently supports n=1.

#5627 and all preceding PRs are merged. Rebased onto main e0db7adf2798cd4ec7f20f55102b4e5b15fdfd27; this is the next independent PR for review. Focused comparison: https://github.com/xorbitsai/inference/pull/5628/files .

Follow-up review fixes (e590a45ac166fc7a4216910ae02d54efdf3a323a)

  • Claimed handoffs now use a 600-second idle lease refreshed by scheduling retries and slab transfers; active reads cannot expire. Orphaned claims are reclaimed. Expiry before or between slabs reports all request destinations via vLLM's invalid-block callback together with receive completion, using Xavier's default recompute policy. Real transfer/layout errors still propagate.
  • Keep prefix heads newest during reservation, retention and probation. Admission stops at the first rejected/capacity-limited block and never retains a suffix past that gap. Only new missing candidates count as admission drops.
  • Count unique history keys actually written once after successful loading, including per-tier counters. Scheduling retries and unused lease keys no longer inflate hits. The default-history e2e case now requires a positive P-side restore on both repeated streaming and non-streaming requests with engine prefix caching disabled.
  • Added orphan/active lease, inter-slab expiry, invalid-block callback, prefix eviction/admission, actual-hit accounting, full-local-hit release, P re-reservation and successful-stream cleanup regressions. Applied all four suggested simplifications. Updated the user guide, README and nine translations.

Validation: 428 local tests passed, 22 hardware/live tests skipped and the distributed tracker integration deselected. The two-GPU run passed 88 tests, including all three real P/D configurations and positive history restores; after the final counter-only adjustment, all 48 history tests passed again on the GPU host. Pre-commit, msgfmt/catalog compilation, English plus nine translated builds, and rendered idle-lease translations passed. Throughput and cache-reuse benchmark tables remain historical; they were not rerun for these changes.

Review fixes (2026-10-04, 121c3f9fa236b463df9bf71d47b8d20cf0ea0498)

  • Consume the single-use handoff at D allocation; preempted decode requests recompute locally. D atomically claims a live ticket before reporting a remote hit. Unclaimed tickets expire after 120 seconds and become a scheduler miss; claimed tickets now use the bounded idle lease described in the follow-up fixes above.
  • Release tickets on D abort before allocation and on router failure/cancellation after P returns. Router abandonment cannot release a claimed/in-flight read, and successful requests avoid an extra cleanup RPC. Remove the unreachable snapshot/unpin router path and redundant direct-handoff constructor option.
  • Clean every ticket/history lease after a batch failure without masking the original exception. Reserve history outside assertions so optimized Python keeps lease protection.
  • Retain a bounded leading candidate prefix for long requests instead of rejecting the whole request above 64 MiB; remove obsolete held-block plumbing. Count capacity, deadline and closing drops separately from a busy writer. Gate producer polling on outstanding sends and CUDA fences on scheduled KV writes.
  • Apply the review simplifications and update the user guide, developer README and all nine translations. The guide explains P-side history, partial retention and expiry behavior; implementation limits/counters stay in the README.

Validation on this revision: 417 local tests passed, 22 hardware/live cases skipped, distributed tracker integration deselected; 79 tests passed on the two RTX3090Ti/NVLink GPUs, including default Xavier history, history disabled, native NIXL end-to-end deployments and CUDA history/raw-bit tests. The history reserve/restore regression also passes under python -O. Pre-commit, all nine msgfmt checks/catalog compilation, English plus nine translated page builds, and rendered new-paragraph checks passed.

The historical performance tables below are unchanged and were not rerun after these lifecycle fixes. The atomic ticket claim adds a scheduler RPC; these correctness checks do not establish its throughput cost or a new performance comparison.

Rebase validation (2026-10-04, head fdaffdc3a500ad00c07b1dfc6047ff9871bac144):

  • Preserved merged mapping serialization, schema checks, async completion logging and lease cleanup. Direct tickets and history leases now use their corresponding cleanup APIs when submission fails before task creation; both cases have regression tests.
  • 402 focused local tests passed; 22 hardware/live tests skipped and the distributed tracker integration deselected. Includes the request-limit regressions.
  • Real two-GPU validation (RTX3090Ti/NVLink, vLLM 0.21.0, xoscar 0.11.1, NIXL 1.1.0): 63 direct/history tests and native NIXL integration passed. All three integration configurations (Xavier default tiered history, Xavier history disabled, and native NIXL) passed again after fixing the completion-log assertion to exclude P-side history restores and require all four D-side handoffs. The initial default-history failure was 8 == 4 from counting both P and D completions. Integration includes streaming, repeated prompts, four concurrent requests and shutdown; local engine prefix caching is disabled. Fixed-prefix/raw-bit CUDA tests validate history separately.
  • Pre-commit passed. All nine catalogs passed msgfmt checks and compilation; the changed page built in English and all nine languages. The three new default-handoff/history paragraphs were verified in every rendered translation. User guidance stays in the deployment page; implementation and validation details are in Xavier's README.

Historical performance validation below predates this rebase; throughput and cache-reuse benchmarks were not rerun for the new head. The native comparison and cache-reuse measurements explicitly identify original head 13925a3b31707e595fc64db634e0e0997e4e849c.

  • Two fresh-server trials each for baseline perf: reduce Xavier cold-prefill snapshot overhead #5627 and this implementation: Qwen2.5-Instruct 0.5B, BF16, eager, one P and one D on two RTX3090Ti GPUs with NVLink, vLLM 0.21.0 / xoscar 0.11.1 / NIXL 1.1.0. Each trial includes 300 C16 requests, 300 C32 requests, serial initialization and cold arrivals during background decode. Local prefix caching is enabled in these performance trials.
Metric #5627 snapshot baseline Direct + tiered history
C16 mean output tokens/s (two trials) 1568 1736 (+10.7%)
C32 mean output tokens/s (two trials) 2316 2514 (+8.6%)
Cold-arrival TTFT P50, trial A / B 1836 / 1524 ms 295 / 289 ms
Cold-arrival TTFT P95, trial A / B 2630 / 2561 ms 565 / 332 ms

A separate 1 MiB GPU-budget run disables local prefix caching to exercise history: 132 successful requests, 7,324 CPU-history block hits, 20 GPU-history block hits, 3,307 demotions, and zero CPU network-transfer batches. All direct handoffs finished without expiry or history-save failure. Total across the controlled comparisons and this capacity-tier check: 2,740 successful requests, zero request errors.

The normal-prefix performance trials hit vLLM's local cache and do not establish a throughput benefit from CPU history itself. These measurements are specific to this model, workload and hardware. The prior output-divergence diagnosis also showed that BF16 greedy outputs must be compared at the same reused-prefix length; KV bit checks and fixed-length comparisons are distinct from throughput measurements.

Native vLLM NIXL comparison (fresh paired run on 13925a3b31707e595fc64db634e0e0997e4e849c): same hardware, model, engine configuration and request sequence as above. Both backends run through the Xinference P/D route; the native column selects vLLM's native NIXL connector, while Xavier uses xoscar NIXL with its default 256 MiB GPU history and CPU overflow. Each backend uses a fresh server/model and handles 652 requests (12 serial initialization, 300 C16, 300 C32, and 40 overlap requests including 32 cold arrivals).

Metric Native vLLM NIXL Xavier direct + tiered history
C16 output tokens/s 1782.7 1750.6 (-1.8%)
C32 output tokens/s 2695.7 2593.2 (-3.8%)
Cold-arrival TTFT P50 247.2 ms 281.0 ms
Cold-arrival TTFT P95 358.8 ms 313.6 ms

All 1,304 requests succeeded. Xavier registered and finished all 652 handoffs, with zero expiry, history-save failures or CPU network batches. History remained enabled: 1,835 saved blocks and 470 GPU-to-CPU demotions. History hits were zero with the normal local-prefix-cache workload, so this compares the direct path including history-write cost, not history-hit benefit. This is one paired trial, not a statistically established P95 advantage; native NIXL still leads throughput and median cold TTFT on this workload. It is a connector comparison through the same Xinference API, not a standalone vllm serve benchmark.

Cache-reuse benchmark (2026-10-04)

These tests keep the P hop: the scheduler still calls P, history is restored on P, and D consumes the handoff. No P-bypass routing is involved. They use the same commit and runtime versions as the native comparison above, one P and one D, Qwen2.5-Instruct 0.5B BF16/eager, and both backends have vLLM local prefix caching enabled. Requests are serial (C1), about 4,000 input tokens and 16 output tokens, temperature 0, ignore_eos=true. All numbers below are client-observed TTFT in milliseconds.

To induce eviction with a short benchmark, num_gpu_blocks_override=1024 limits each engine to 16,384 cached tokens (192 MiB KV). Xavier retains its default 256 MiB GPU history plus a CPU history capacity of 1,024 blocks (192 MiB on P in this test). Native NIXL is also tested with 2,389 engine blocks: 192 MiB + the same 1,365 whole blocks that fit Xavier's 256 MiB history budget. This matches GPU KV storage capacity, not total process GPU usage or total host memory; Xavier still has the additional CPU tier. This is an intentionally constrained-cache microbenchmark, not the default engine capacity or a production-size-model result.

Fixed hot set: initialize four different documents once, then repeat six cycles of: immediately repeat all four documents; issue 12 new documents with different prefixes; revisit the same four hot documents. Each backend/configuration starts a fresh server/model and handles 125 requests including one startup request. Each immediate/revisit distribution contains 24 observations. A document is a distinct identifying prefix followed by Water evaporates in sunlight and condenses into clouds. repeated 360 times and Explain this process.; pressure-document identifiers change every cycle. Pressure exceeds even the expanded native engine's 38,224-token capacity.

Configuration Immediate repeat P50 / P95 After eviction P50 / P95 After eviction mean
Native NIXL, 1,024 engine blocks 60.1 / 64.9 134.9 / 212.4 150.7
Native NIXL, 2,389 engine blocks (matched GPU KV capacity) 61.2 / 67.6 143.9 / 159.7 147.2
Xavier direct, history disabled, 1,024 engine blocks 66.7 / 75.4 145.8 / 153.2 146.5
Xavier direct + default tiered history, 1,024 engine blocks 68.4 / 77.8 96.5 / 129.9 103.7
  • Against native NIXL with the same engine capacity, Xavier history reduces revisit P50/P95 by 28.5%/38.8%. Against the expanded native engine with matched GPU KV capacity, reductions are 33.0%/18.7%.
  • The Xavier history-off control isolates the history effect: enabling history reduces revisit P50 from 145.8 to 96.5 ms (33.8%) and mean TTFT from 146.5 to 103.7 ms (29.2%). The direct-transfer route itself stays enabled in both cases.
  • Hot-set history counters show 24 history loads and 5,976 block hits (249 blocks per revisit on average), including 3,414 GPU and 2,562 CPU hits. All 125 direct handoffs finished, with zero expiry, history-save failures or CPU network-transfer batches. The disabled-history control reports zero history loads/hits. Both GPU reuse and CPU recovery participated; this run does not isolate their individual speedups.
  • P-call timing supports the mechanism: median P-call time for revisits is about 33.0 ms with history versus 81.3/81.5 ms for native NIXL with the original/expanded engine capacity. These are P actor-call durations including its work, not isolated GPU-kernel timings. Saving prefill work is sufficient to produce the observed benefit while still going through P.
  • Immediate repeats already hit the engine's own prefix cache. Native NIXL is faster there: 60.1 ms P50 versus Xavier's 68.4 ms. Xavier without history is 66.7 ms, so the whole difference must not be attributed to history. There is no general claim that Xavier wins whenever any cache hits.

Changing working set / negative control: a separate pair of 73-request runs uses three cycles with four new target documents per cycle, immediate repeat, 12 new pressure documents, then target revisit. It uses 1,024 engine blocks for both backends. This is not the stable-hot-set workload above.

Metric Native NIXL Xavier with history
Immediate-repeat P50 / P95 54.9 / 105.4 61.8 / 110.2
Revisit P50 / P95 128.9 / 206.2 131.7 / 143.8
Revisit mean 140.8 121.8

Here history still loads 12 times, but only 1,486 blocks are hit in total (about 124 blocks per load), versus 249 per load in the fixed-hot-set test; 355 hits are CPU and 1,131 GPU. History reaches capacity, with admission rejections and bounded retention, so cache presence does not imply nearly complete prefix reuse. Revisit median latency does not improve overall, although the mean and P95 do. The first cycle's Xavier revisit TTFTs are 88–106 ms; later cycles are about 128–144 ms. This limits the claim to workloads with sufficient reusable history, not arbitrary churn.

All six cache-reuse runs total 646 requests, zero request errors (500 fixed-hot-set/control requests and 146 changing-working-set requests). The runs are sequential on the same machine; the six cycles within a run are not six independent fresh-server trials. P95 has only 24 samples per hot-set configuration, and Xavier had one 249.9 ms revisit outlier, so the lower P95 does not mean every request is faster. DEBUG logging was enabled equally for these diagnostic runs; these results should not be combined directly with the earlier C16/C32 throughput measurements.

Conclusion: tiered history demonstrably reduces TTFT for recurring long prefixes after engine-cache eviction, including against a native NIXL control with matched GPU KV capacity, at the cost of additional CPU history storage. It does not establish a benefit for immediate engine-cache hits, arbitrary working-set churn, saturated throughput, or other model sizes. Keep the simple P route; prioritize history admission/retention and CPU-restore efficiency, then validate on larger models/concurrent reuse before making broader claims.

@qinxuye
qinxuye force-pushed the perf/xavier-direct-tiered-history branch from 13925a3 to fdaffdc Compare October 4, 2026 12:35
@qinxuye
qinxuye requested a review from rogercloud October 4, 2026 12:37

@rogercloud rogercloud left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Major

  • xinference/model/llm/vllm/xavier/v1_connector.py:406 — Decode re-issues the direct load after preemption, against a ticket that has already been released, and the failure crashes D's EngineCore.
    Trigger: in vLLM 0.21, _preempt_request resets num_computed_tokens=0 and prepends the request to the waiting queue, which calls get_num_new_matched_tokens again with xavier_direct still present. send_direct then raises, and get_finished calls task.result() and re-raises inside the worker. Impact: every in-flight request on D fails.
  • xinference/model/llm/vllm/xavier/direct_handoff.py:88 — The 120 s producer deadline keeps running while D is still queued, so a D backlog longer than 120 s ends on the same fatal path instead of a local recompute.
    Trigger: D is saturated (max_num_seqs/KV) for more than 120 s after P finishes. Impact: D EngineCore crash.
  • xinference/model/llm/vllm/xavier/direct_history.py:114 — The min(..., 64 MiB) pending cap skips history for whole requests. On an 8B GQA model (2 MiB/block) that is any prompt over about 512 tokens, regardless of xavier_gpu_cache_bytes. Retain the leading prefix instead of skipping, or scale the cap with the budget. Count this skip separately from busy-writer skips and document it.

Minor

  • xinference/model/llm/vllm/xavier/v1_connector.py:580 — If D aborts before get_num_new_matched_tokens, or the router fails or cancels after P succeeds, nobody releases the ticket (free_prefill_model_cache is a no-op for direct handoff, xinference/core/pd_model.py:254), so P blocks stay pinned for 120 s. Release kv_transfer_params["xavier_direct"] in both paths.
  • xinference/model/llm/vllm/xavier/direct_handoff.py:170 — In a multi-request load_direct/load_history, the first failure skips releasing the remaining tickets and leases. The release in finally can also mask the original error. This contradicts the README's "attempts every lease release" promise; mirror load_gpu_requests_v1's gather(..., return_exceptions=True).
  • xinference/model/llm/vllm/xavier/direct_history.py:96 — assert self.history.reserve(...) is the only reserve call, so python -O drops it. Call it first, then check the result.
  • xinference/model/llm/vllm/xavier/v1_connector.py:237 / :321 — The producer does a poll_direct_gpu_v1 RPC and a torch.cuda.synchronize() on every step, even with no direct requests pending. Base returned early here. Gate both on outstanding tickets or store requests.
  • xinference/model/llm/vllm/xavier/direct_history.py:193 — Retentions abandoned at the 10 ms soft deadline, and candidates dropped once the limit is reached (:133), update no metric, so missing history is invisible.
  • Tests: no coverage for decode re-entry after load/release (preemption), a D-side load after expiry, decode abort before load releasing the ticket, multi-request partial-failure cleanup, the producer request_finished history branch (hashes plus held_blocks), or the get_finished → poll_direct_gpu_v1 wiring.
  • Docs: pd_separation.rst does not mention the 120 s handoff deadline, the 64 MiB history skip, or that history is restored only on P. Lines ~130-141 still describe the pre-direct CPU/GPU-first cache.
  • xinference/core/pd_model.py:254 / :357 — For supervisor-launched P/D, direct handoff is now always on, so the snapshot free_model_cache gather and the set_unpin_handler branch are unreachable. test_direct_handoff.py:284 sets an env var that nothing reads.

Simplification

  • v1_connector.py L549-565: shrink: two register_direct_gpu_v1 calls differ only in the history args. Build the args once and make a single call.
  • direct_history.py L155-164: stdlib: hand-built exc_info tuple. Use logger.error(..., exc_info=exc).
  • direct_history.py L37: delete: cpu_capacity param of _init_history is test-only. Tests can set store.cpu_capacity first.
    net: -10 lines possible

Blocking: yes — recommended event: COMMENT

  • xinference/model/llm/vllm/xavier/v1_connector.py:406 — major — D EngineCore crash on preemption re-load [new]
  • xinference/model/llm/vllm/xavier/direct_handoff.py:88 — major — D EngineCore crash when queued >120 s [new]

Comment thread xinference/model/llm/vllm/xavier/v1_connector.py
Comment thread xinference/model/llm/vllm/xavier/direct_handoff.py
Comment thread xinference/model/llm/vllm/xavier/direct_history.py Outdated
Comment thread xinference/model/llm/vllm/xavier/v1_connector.py
Comment thread xinference/model/llm/vllm/xavier/direct_handoff.py
Comment thread xinference/model/llm/vllm/xavier/direct_history.py Outdated
Comment thread xinference/model/llm/vllm/xavier/v1_connector.py
Comment thread xinference/model/llm/vllm/xavier/direct_history.py Outdated
Comment thread xinference/model/llm/vllm/xavier/test/test_direct_handoff.py Outdated
@qinxuye
qinxuye requested a review from rogercloud October 4, 2026 13:12

@rogercloud rogercloud left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Major

  • xinference/model/llm/vllm/xavier/direct_handoff.py:128 — Once D claims a ticket, nothing on P can release it any more. See the inline comment.

Minor

  • xinference/model/llm/vllm/xavier/direct_history.py:123 — History LRU order evicts the head of each prefix chain first. See the inline comment.
  • xinference/model/llm/vllm/xavier/direct_history.py:133 — Admission is not prefix-closed, so blocks can be stored that can never be reserved. See the inline comment.
  • xinference/model/llm/vllm/xavier/test/test_pd_gpu.py:178 — The default-history e2e case never shows that history is used. See the inline comment.
  • Tests: these new branches have no tests:
    • D's full-local-hit release (v1_connector.py:403-412, tokens <= num_computed_tokens);
    • P's re-reserve, which releases the previous lease (v1_connector.py:357-362);
    • the router's stream-success pop of _direct_transfers (xinference/core/pd_model.py:364). test_pd_model.py covers only the non-stream path.
  • xinference/model/llm/vllm/xavier/direct_history.py:308-313 — history_hit_blocks and the per-tier hit counters count the whole lease. TieredKVSnapshotStore.reserve also counts hits again on every scheduling retry, so the counters overstate restored blocks. Count hits once, after a successful load, over the blocks actually written.

Simplification

  • v1_connector.py L133: delete: the _gpu_budget is None and role checks are already guaranteed by uses_direct_handoff (transport.py:100). Keep only len(kv_cache_config.kv_cache_groups) != 1.
  • v1_connector.py L356, L574: delete: getattr(self, "_history_enabled", False). __init__ always sets the attribute, so use self._history_enabled.
  • direct_handoff.py L196: shrink: the inner release() re-imports TransferActor and rebuilds the actor ref. Call self.actor.release_remote_direct_gpu_v1(rank, ticket) instead.
  • direct_history.py L194: shrink: the deadline/closing drop is duplicated at L195-200 and L202-207. Fold both into one while True: loop that checks closing/deadline, then the locks, then sleeps.

net: -13 lines possible

Blocking: no — recommended event: APPROVE

Comment thread xinference/model/llm/vllm/xavier/direct_handoff.py Outdated
Comment thread xinference/model/llm/vllm/xavier/direct_history.py
Comment thread xinference/model/llm/vllm/xavier/direct_history.py Outdated
Comment thread xinference/model/llm/vllm/xavier/test/test_pd_gpu.py
@qinxuye
qinxuye merged commit f90f673 into xorbitsai:main Oct 4, 2026
15 checks passed
@qinxuye
qinxuye deleted the perf/xavier-direct-tiered-history branch October 4, 2026 14:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request gpu

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants