Repository navigation
PERF: reduce Xavier V1 cache overhead and allow CUDA Graph - #5643
Merged
Merged
Conversation
…ectors Merge current main to validate the integrated Xavier connector paths. Skip CPU export polling when there are no owned payloads, and cover exportless completion hooks for both blocking and nonblocking polls.
rogercloud
approved these changes
Oct 7, 2026
rogercloud
left a comment
Contributor
There was a problem hiding this comment.
Minor
xinference/model/llm/vllm/xavier/v1_connector.py:1226— The refresh path publishes and registers outside_snapshot_export_lock_v1, so itsregister_snapshot_blockscan land after the actor'sexport()has evicted and unregistered the same key, leaving a stale tracker/directory positive (consumers failreserveand recompute). Route the refresh through the actor under the same lock, or unregister requested-but-unavailable keys for this rank.xinference/model/llm/vllm/xavier/v1_connector.py:1964— A failed or unknown-ticket optional CPU snapshot export raises out ofget_finished/wait_for_saveand kills EngineCore, although these exports are documented as optional. For arena tickets, log, release slots and reset_exported_keyswithout raising.xinference/model/llm/vllm/xavier/v1_connector.py:1108— When the directory can never attach (cross-hostboot_idmismatch, or the tracker returnsNone), discovery is retried every second forever, adding a tracker RPC to the scheduler query path. Remember permanent non-attachable results and retry only on transient errors or invalidation.xinference/model/llm/xavier/backends/torch/local_read.py:56— If theread_request_blocks_local_v1reply or the release RPC is lost,_leaseis never cleared and all later reads silently fall back to the RPC path. Add a lease TTL/owner reset, or at least log the fallback.xinference/model/llm/vllm/xavier/v1_connector.py:450— The arena budget floormax(2 * _MAX_EXPORT_BYTES, free // 8)reserves at least 64 MiB after KV sizing even when little memory is free, and CUDA graph capture now runs after it by default. Skip the arena whenfree // 8 < 2 * _MAX_EXPORT_BYTES.
Simplification
v1_connector.pyL594-603: delete: the arena branch always returns, socount/per_batch/slotsand theself._gpu_export_arena is not None and slots > ...clause are dead. Keep only the byte-budget check.v1_connector.pyL1205: delete:save_kv_layeris now a no-op, so_stage_kv_layer_for_request,_stage_layer_blocks,_request_staged_layersand this grouping branch are reachable only from tests. Remove them with their tests.vllm/xavier/gc_lifecycle.pyL8: delete: verbatim copy ofsglang/gc_lifecycle.py(already imported cross-package byxavier/backends/torch/pd.py). Import or hoist the shared class.local_read.pyL19: shrink:_boot_id()duplicatesxavier/local_directory.py. Import it.snapshot.pyL193: shrink: re-testskey not in self.blocks. Useif fresh:.v1_connector.pyL1699: shrink:_cpu_export_sizerebuilds per-layer views 2-3 times perwait_for_save. Cache per-block bytes inregister_kv_caches.
net: -110 lines possible
Blocking: no — recommended event: APPROVE
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Xavier V1 CPU snapshots added synchronous control RPCs, repeated hash/metadata work and per-layer allocations to ordinary replica serving, even when native prefix caching had already handled the request. This change preserves native APC and CUDA Graph while moving snapshot publication off the EngineCore critical path and bounding its storage and transfer ownership.
additional_configchanged the compilation hash on every launch; the user's compilation configuration is now preserved and both replicas reuse the same compiled graph cache. Attention-only models retain the requested eager setting; recurrent models retain their eager default.Validation on
9b0ce1d29713abd56faef9db06672e1657979a0f:pre-commit run --all-filespassed. All 77 files in the tested PR/main-integration scope match the committed source; controlled benchmarks use the same frozen source.Current-head controlled measurements (
9b0ce1d29713abd56faef9db06672e1657979a0f):Qwen2.5-0.5B-Instruct FP16, two RTX 3090 Ti cards on one host, two ordinary hybrid replicas, TP=PP=1, vLLM 0.21.0 / Torch 2.11. Native APC, asynchronous scheduling and CUDA Graph are enabled in both modes, with the same 2,048-token chunked-prefill budget. Cold long prompts contain 4,454 input tokens and measured requests generate 32 tokens. Each round completes 6,612 requests, including 2,048 cold short, 256 cold long and 4,096 warm long C16 requests. The two pairs run Xavier → native, then native → Xavier, cooling both cards to <=65 C before each mode and recording temperatures/clocks. All four rounds passed per-second GPU ownership/memory audits and completed 26,448 requests without errors.
The runtime includes the merged Gloo GIL fix xorbitsai/xoscar#213, native binary SHA-256
248280991d074d78953d11d217e416c4839b920870ca1406cb5ee0d306c73e2b. These results depend on that runtime fix as well as this PR.Excluding warmups, the largest observed throughput loss across both pairs is 3.29%, in cold long C16 of the first pair; that stage is +0.42% in the reverse pair. Cold short C16 loses 1.21% / 2.27%, warm long C16 loses 1.11% / 1.51%, and cold long C1 loses 1.56% / 1.46%. Treat small gains as variation rather than a demonstrated general speedup.
A separate clean control on preceding commit
c99a6eca0measured cold long C1 TTFT p50/p95 at 92.06 / 102.12 ms. The current head measures 88.05 / 89.46 ms and 87.80 / 90.02 ms: its single-request first-token cost is now about 1.5–1.6 ms above native. This control uses the same workload/runtime and no profiling wrappers; it is a separate preceding-head run.There is no consistent cold-long C16 TTFT improvement relative to native: the first pair improves p50 by 15.23 ms, while the reverse pair adds 48.42 ms and 63.61 ms at p95. Native-only C16 p50 itself varies by 21.97 ms between repeats. Warm C16 p50 still adds 3.60–3.73 ms. Better C1 latency and small throughput loss do not establish negligible first-token cost at concurrency 16.
Each Xavier round retains 16 genuine cross-replica hits at the same depth: one hit reuses 4,437 tokens and the other fifteen reuse 4,421 tokens. All paired input/output token counts match. All measured texts outside cold-long C16, including cross-replica hits and warm requests, match native exactly. Cold-long C16 has 13/256 and 8/256 differing texts, while native-only repeats differ on 6/256; no Xavier external hits occur in that stage. Byte-level CUDA and immutable-snapshot tests provide separate transfer correctness coverage; the benchmark does not establish deterministic greedy text for every concurrent batch.
Post-workload process-tree memory, decimal GB:
c99a6eca0, XavierOrdinary CPU history reduces retained memory by about 7.2 GB in this workload, but total process memory remains much higher than native. RSS sums shared pages across processes; PSS apportions them. Model weights and native KV capacity are unchanged. CPU history capacity still follows the GPU KV block count; an independent CPU byte budget is not part of this change.
Xavier remains opt-in. The measured C1 improvement and 1–3.3% throughput cost do not justify global default enablement given the remaining C16 TTFT and CPU-history memory costs. Larger models and cross-host performance have not been established by this single-host, small-model test.
Snapshot robustness and validation:
The performance measurements above were gathered before the robustness and CI updates; the full serving benchmark was not repeated for them.
CI follow-up on
143460177:c0e2e1ce6run passed all Ubuntu Python 3.10–3.14 jobs, GPU T4, Metal and lint. macOS 3.14 failed in the llama.cpp native Metal image decoder; Windows 3.14 had three cluster-startup errors and a model-spec subprocess timeout, before executing the Xavier slice. No changed Xavier runtime path was reached by those failures.cb6c1ad2694515f42b23abd61ddbf51758226389:asyncio.wait_forcan return an already-completed RPC instead of propagating cancellation, leaving release waiting indefinitely. Renewal now checks that it still owns the room after release removes that owner. Added success/error race regressions and verbose test names in the extra CI slice.