Skip to content

PERF: reduce Xavier V1 cache overhead and allow CUDA Graph - #5643

Merged
qinxuye merged 7 commits into
xorbitsai:mainfrom
qinxuye:perf/xavier-cache-hot-path
Oct 8, 2026
Merged

qinxuye merged 7 commits into
xorbitsai:mainfrom
qinxuye:perf/xavier-cache-hot-path

Conversation

@qinxuye

@qinxuye qinxuye commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Xavier V1 CPU snapshots added synchronous control RPCs, repeated hash/metadata work and per-layer allocations to ordinary replica serving, even when native prefix caching had already handled the request. This change preserves native APC and CUDA Graph while moving snapshot publication off the EngineCore critical path and bounding its storage and transfer ownership.

  • Export complete computed blocks during chunked prefill through a fused gather and a reusable 32 MiB GPU arena. Gather on a side CUDA stream and fence the next model write or KV load, so immutable packing can overlap logits/sampling while source slots remain protected. Ordinary CPU history export admits only available arena/byte credits and skips optional work under congestion; skipped blocks remain eligible for later export. Required direct P/D handoff and GPU-history paths retain their existing behavior.
  • Copy IPC gathers directly into ordinary CPU allocations in a background thread, keeping allocation and pageable D2H off the actor loop shared by streaming requests. Preserve both arena readiness and legacy CUDA IPC stream dependencies. Repeated cancellation waits for the copy thread, and failures drain its stream before releasing GPU ownership. Packed slabs retain independent bytes with lazy layer views and whole-slab accounting/eviction, including BF16 bit preservation. Batched pinned H2D and leased same-host receive buffers keep the existing read path and remote/busy fallback.
  • Use a bounded, coherent same-host negative directory, bounded hash caches and compact tracker indices to skip unnecessary lookup work. Publish/evict under a shared lock and reject stale owners. Refresh mixed-prefix publication once per interval rather than querying on every cold prefill after an expired deadline; recovery still retries authoritative discovery.
  • Freeze initialized library objects only in the ordinary V1 CPU-snapshot actor process. GC stays enabled for new request cycles, and cleanup preserves an external freeze owner. EngineCore GC behavior is unchanged.
  • Keep actor addresses, IDs and ranks solely in connector transport configuration. Including them in vLLM additional_config changed the compilation hash on every launch; the user's compilation configuration is now preserved and both replicas reuse the same compiled graph cache. Attention-only models retain the requested eager setting; recurrent models retain their eager default.

Validation on 9b0ce1d29713abd56faef9db06672e1657979a0f:

  • Exact second CPU test group from the merged CI workflow: 781 passed, 48 skipped.
  • Linux CUDA/vLLM, full V1 and shared Xavier suites: 723 passed, 20 skipped. Coverage includes computed-only chunk publication, preemption/resume, optional-export congestion, delayed side-stream gathers and source overwrites, FP16/BF16 bytes, legacy CUDA IPC readiness, actor recovery, repeated cancellation, failure cleanup, leases and cross-engine completion.
  • Full pre-commit run --all-files passed. All 77 files in the tested PR/main-integration scope match the committed source; controlled benchmarks use the same frozen source.
  • GitHub CI for this head: lint and both documentation builds have passed; platform, Metal and GPU jobs are still running. Local validation is separate from CI completion.

Current-head controlled measurements (9b0ce1d29713abd56faef9db06672e1657979a0f):

Qwen2.5-0.5B-Instruct FP16, two RTX 3090 Ti cards on one host, two ordinary hybrid replicas, TP=PP=1, vLLM 0.21.0 / Torch 2.11. Native APC, asynchronous scheduling and CUDA Graph are enabled in both modes, with the same 2,048-token chunked-prefill budget. Cold long prompts contain 4,454 input tokens and measured requests generate 32 tokens. Each round completes 6,612 requests, including 2,048 cold short, 256 cold long and 4,096 warm long C16 requests. The two pairs run Xavier → native, then native → Xavier, cooling both cards to <=65 C before each mode and recording temperatures/clocks. All four rounds passed per-second GPU ownership/memory audits and completed 26,448 requests without errors.

The runtime includes the merged Gloo GIL fix xorbitsai/xoscar#213, native binary SHA-256 248280991d074d78953d11d217e416c4839b920870ca1406cb5ee0d306c73e2b. These results depend on that runtime fix as well as this PR.

Pair / order Mode Cold short C16 req/s Cold long C16 req/s Warm long C16 req/s
1: Xavier then native Native APC 112.06 22.21 106.57
1: Xavier then native Xavier CPU snapshots 110.71 21.48 105.39
2: native then Xavier Native APC 112.89 22.18 106.69
2: native then Xavier Xavier CPU snapshots 110.33 22.27 105.09

Excluding warmups, the largest observed throughput loss across both pairs is 3.29%, in cold long C16 of the first pair; that stage is +0.42% in the reverse pair. Cold short C16 loses 1.21% / 2.27%, warm long C16 loses 1.11% / 1.51%, and cold long C1 loses 1.56% / 1.46%. Treat small gains as variation rather than a demonstrated general speedup.

A separate clean control on preceding commit c99a6eca0 measured cold long C1 TTFT p50/p95 at 92.06 / 102.12 ms. The current head measures 88.05 / 89.46 ms and 87.80 / 90.02 ms: its single-request first-token cost is now about 1.5–1.6 ms above native. This control uses the same workload/runtime and no profiling wrappers; it is a separate preceding-head run.

TTFT (ms), native → current Xavier Pair 1 Pair 2, reversed
Cold long C1 p50 86.46 → 88.05 86.33 → 87.80
Cold long C1 p95 88.73 → 89.46 88.86 → 90.02
Cold long C16 p50 245.77 → 230.54 223.80 → 272.23
Cold long C16 p95 452.81 → 414.46 425.87 → 489.48
Cold long C16 p99 629.90 → 632.48 629.00 → 640.49
Warm long C16 p50 43.97 → 47.70 44.87 → 48.48

There is no consistent cold-long C16 TTFT improvement relative to native: the first pair improves p50 by 15.23 ms, while the reverse pair adds 48.42 ms and 63.61 ms at p95. Native-only C16 p50 itself varies by 21.97 ms between repeats. Warm C16 p50 still adds 3.60–3.73 ms. Better C1 latency and small throughput loss do not establish negligible first-token cost at concurrency 16.

Each Xavier round retains 16 genuine cross-replica hits at the same depth: one hit reuses 4,437 tokens and the other fifteen reuse 4,421 tokens. All paired input/output token counts match. All measured texts outside cold-long C16, including cross-replica hits and warm requests, match native exactly. Cold-long C16 has 13/256 and 8/256 differing texts, while native-only repeats differ on 6/256; no Xavier external hits occur in that stage. Byte-level CUDA and immutable-snapshot tests provide separate transfer correctness coverage; the benchmark does not establish deterministic greedy text for every concurrent batch.

Post-workload process-tree memory, decimal GB:

Mode RSS PSS
Preceding c99a6eca0, Xavier 44.06 41.65
Current Xavier, both rounds 36.88 34.46
Native APC, both rounds 9.31–9.32 7.30–7.31

Ordinary CPU history reduces retained memory by about 7.2 GB in this workload, but total process memory remains much higher than native. RSS sums shared pages across processes; PSS apportions them. Model weights and native KV capacity are unchanged. CPU history capacity still follows the GPU KV block count; an independent CPU byte budget is not part of this change.

Xavier remains opt-in. The measured C1 improvement and 1–3.3% throughput cost do not justify global default enablement given the remaining C16 TTFT and CPU-history memory costs. Larger models and cross-host performance have not been established by this single-host, small-model test.

Snapshot robustness and validation:

  • Serialize periodic refresh and background-export publication inside the TransferActor lock, unregistering requested keys that this rank no longer owns.
  • Treat completed optional export failures and unknown tickets as cache misses, releasing completed ownership and invalidating the mirror. Preserve pending-copy fencing.
  • Remember permanently unavailable same-host directory metadata, retry transient discovery errors, and warn once per busy shared-read lease when using the tensor RPC fallback.
  • Respect free-memory headroom before arena allocation; cache per-block export bytes at KV registration and remove unused staging and duplicate GC code.
  • Complete Xavier regression slice: 476 passed, 66 skipped on macOS; 531 passed, 11 skipped on Linux with vLLM 0.21.0 and real CUDA IPC, actor-recovery and same-host directory/read tests. Linux tested-source hashes match all 12 changed files. Full pre-commit and commit hooks passed.

The performance measurements above were gathered before the robustness and CI updates; the full serving benchmark was not repeated for them.

CI follow-up on 143460177:

  • The c0e2e1ce6 run passed all Ubuntu Python 3.10–3.14 jobs, GPU T4, Metal and lint. macOS 3.14 failed in the llama.cpp native Metal image decoder; Windows 3.14 had three cluster-startup errors and a model-spec subprocess timeout, before executing the Xavier slice. No changed Xavier runtime path was reached by those failures.
  • Windows 3.10 completed the broad suite (3787 passed), then hung in the SGLang lease-renewal test. A deterministic Python 3.10 reproduction shows the same cancellation race on base cb6c1ad2694515f42b23abd61ddbf51758226389: asyncio.wait_for can return an already-completed RPC instead of propagating cancellation, leaving release waiting indefinitely. Renewal now checks that it still owns the room after release removes that owner. Added success/error race regressions and verbose test names in the extra CI slice.
  • Both new race regressions fail with the base renewal implementation. The fixed SGLang and model-spec slice passes 147 tests on macOS and 147 on Linux; the native Python 3.10 race reproduction now terminates normally. Full pre-commit and commit hooks passed. Fresh Windows and macOS 3.14 CI results are pending.

@XprobeBot XprobeBot added enhancement New feature or request gpu labels Oct 6, 2026
@XprobeBot XprobeBot added this to the v3.x milestone Oct 6, 2026
…ectors

Merge current main to validate the integrated Xavier connector paths. Skip CPU export polling when there are no owned payloads, and cover exportless completion hooks for both blocking and nonblocking polls.
@qinxuye
qinxuye requested a review from rogercloud October 7, 2026 14:40

@rogercloud rogercloud left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor

  • xinference/model/llm/vllm/xavier/v1_connector.py:1226 — The refresh path publishes and registers outside _snapshot_export_lock_v1, so its register_snapshot_blocks can land after the actor's export() has evicted and unregistered the same key, leaving a stale tracker/directory positive (consumers fail reserve and recompute). Route the refresh through the actor under the same lock, or unregister requested-but-unavailable keys for this rank.
  • xinference/model/llm/vllm/xavier/v1_connector.py:1964 — A failed or unknown-ticket optional CPU snapshot export raises out of get_finished/wait_for_save and kills EngineCore, although these exports are documented as optional. For arena tickets, log, release slots and reset _exported_keys without raising.
  • xinference/model/llm/vllm/xavier/v1_connector.py:1108 — When the directory can never attach (cross-host boot_id mismatch, or the tracker returns None), discovery is retried every second forever, adding a tracker RPC to the scheduler query path. Remember permanent non-attachable results and retry only on transient errors or invalidation.
  • xinference/model/llm/xavier/backends/torch/local_read.py:56 — If the read_request_blocks_local_v1 reply or the release RPC is lost, _lease is never cleared and all later reads silently fall back to the RPC path. Add a lease TTL/owner reset, or at least log the fallback.
  • xinference/model/llm/vllm/xavier/v1_connector.py:450 — The arena budget floor max(2 * _MAX_EXPORT_BYTES, free // 8) reserves at least 64 MiB after KV sizing even when little memory is free, and CUDA graph capture now runs after it by default. Skip the arena when free // 8 < 2 * _MAX_EXPORT_BYTES.

Simplification

  • v1_connector.py L594-603: delete: the arena branch always returns, so count/per_batch/slots and the self._gpu_export_arena is not None and slots > ... clause are dead. Keep only the byte-budget check.
  • v1_connector.py L1205: delete: save_kv_layer is now a no-op, so _stage_kv_layer_for_request, _stage_layer_blocks, _request_staged_layers and this grouping branch are reachable only from tests. Remove them with their tests.
  • vllm/xavier/gc_lifecycle.py L8: delete: verbatim copy of sglang/gc_lifecycle.py (already imported cross-package by xavier/backends/torch/pd.py). Import or hoist the shared class.
  • local_read.py L19: shrink: _boot_id() duplicates xavier/local_directory.py. Import it.
  • snapshot.py L193: shrink: re-tests key not in self.blocks. Use if fresh:.
  • v1_connector.py L1699: shrink: _cpu_export_size rebuilds per-layer views 2-3 times per wait_for_save. Cache per-block bytes in register_kv_caches.
    net: -110 lines possible

Blocking: no — recommended event: APPROVE

Comment thread xinference/model/llm/vllm/xavier/v1_connector.py Outdated
Comment thread xinference/model/llm/vllm/xavier/v1_connector.py Outdated
Comment thread xinference/model/llm/vllm/xavier/v1_connector.py
Comment thread xinference/model/llm/xavier/backends/torch/local_read.py Outdated
Comment thread xinference/model/llm/vllm/xavier/v1_connector.py Outdated
Comment thread xinference/model/llm/vllm/xavier/v1_connector.py Outdated
Comment thread xinference/model/llm/vllm/xavier/v1_connector.py Outdated
Comment thread xinference/model/llm/vllm/xavier/gc_lifecycle.py Outdated
@qinxuye
qinxuye merged commit ec9d674 into xorbitsai:main Oct 8, 2026
14 of 15 checks passed
@qinxuye
qinxuye deleted the perf/xavier-cache-hot-path branch October 8, 2026 00:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request gpu

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants