Skip to content

feat: add MLX Xavier KV cache sharing and PD separation - #5638

Merged
qinxuye merged 4 commits into
xorbitsai:mainfrom
qinxuye:feat/mlx-xavier-cache
Oct 6, 2026
Merged

qinxuye merged 4 commits into
xorbitsai:mainfrom
qinxuye:feat/mlx-xavier-cache

Conversation

@qinxuye

@qinxuye qinxuye commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Adds Xavier prefix KV sharing and explicit prefill/decode replicas to MLX. Prefill evaluates all but the final prompt token and exports FP16 KV pages; decode imports the prefix, evaluates the final token and generates the response. Both streaming and non-streaming requests use the existing APIs.

Based on merged #5631 and current main (10f91bdcd). There are no unmerged PR dependencies.

  • Adds bounded CPU bytes storage under the shared Xavier package, load-time model/tokenizer/attention compatibility checks, cached weight fingerprints, eviction, reserved handoff capacity and cancellation/failure cleanup.
  • Supports enable_xavier=True and explicit prefill/decode placement through the API, clients, CLI and UI. Requires Apple silicon, mlx-lm >= 0.31.2, one worker per replica and unquantized (none/fp16/bf16 labels) full-attention Qwen2/Qwen3/Llama text models with FP16 weights and KV cache. Vision models fail before loading.
  • Handles MLX replica recovery and isolates cache publication/cleanup failures per request. Ordinary single replicas warn and disable Xavier. Publication sends only uncached pages and skips prefixes exceeding the configured page budget or protocol limit. Incremental writes refresh aged prefixes before eviction.
  • Adds Metal regression coverage, a reproducible API benchmark and deployment documentation in all nine maintained locales, including cache sizing and reservation deadlines.

Final API measurements at 7a329378b: Apple M5 Pro / 64 GiB, Python 3.12.3, MLX 0.32.0, mlx-lm 0.31.3 and published xoscar 0.11.1. Identical FP16 Qwen2.5-0.5B-Instruct weights, two independent model processes per mode, native local prompt cache enabled, eight serial warm requests and 32 output tokens per request.

Input tokens Mode Cold document TTFT (ms) Warm TTFT median (ms) Warm output tokens/s
1138 Ordinary MLX 199.4 78.9 144.6
1138 Xavier shared 114.3 77.9 144.5
1138 Xavier P/D 206.5 87.2 134.9
4438 Ordinary MLX 261.9 242.8 79.8
4438 Xavier shared 282.1 205.4 88.2
4438 Xavier P/D 616.6 190.8 89.6

All 60 streamed/non-streamed responses at this revision match within each context. Warm Xavier requests reuse 1137/4437 tokens, with actual cache reads and zero remaining handoffs. These are single-Mac measurements on a shared Metal GPU; cold-document requests retain the 24-token chat prefix from warm-up. Multi-host performance, native NIXL and cross-engine/NVIDIA-to-Mac transfers are outside this PR.

Validation: 543 focused regression tests passed, 10 skipped for platform/dependency conditions. Coverage includes worker recovery, request cancellation, failed publication/release, page-budget/protocol limits, aged-prefix LRU, handoff reservation eviction, load-time contracts, prepared-cache forwarding, chat prefill and real Metal full/partial KV reuse. Changed Python files pass pre-commit; benchmark passes Black/isort/Ruff; the changed UI file passes ESLint/Prettier. All 18 catalogs validate and compile, and both changed pages render with verified MLX translations in English plus all nine locales. Full platform CI runs on the PR.

@rogercloud rogercloud left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Major

  • xinference/core/supervisor.py:3299: MLX P/D replica auto-recovery always fails (see inline; worker.py:6754-6763). [new]
  • xinference/core/supervisor.py:3111: built-in unquantized MLX specs are rejected (see inline). [new]

Minor

  • xinference/model/llm/mlx/core.py:352: an exception in publish() drops the whole step's results (see inline).
  • xinference/model/llm/mlx/xavier.py:273: flush() gathers the shared write set, so one cancelled request cancels others (see inline).
  • xinference/model/llm/mlx/xavier.py:247: unguarded release_handoff fails or masks the real result (see inline); same at L333.
  • xinference/model/llm/xavier/backends/bytes/storage.py:88: LRU order evicts chain heads first and oversized puts self-evict (see inline).
  • xinference/model/llm/mlx/xavier.py:94: weights are fully re-hashed on every load (see inline).
  • xinference/model/llm/mlx/xavier.py:129: contract configure is deferred to the first RPC (see inline).
  • xinference/core/supervisor.py:3098-3117: MLX enable_xavier with replica<=1 still creates the cache actor and publishes every prompt; mirror the SGLang/vLLM warn-and-disable.
  • xinference/model/llm/mlx/core.py:461-470: publish re-encodes and ships the whole prefix, including pages just fetched remotely. Publish only pages beyond the cached prefix.
  • xinference/model/llm/mlx/tests/test_xavier.py:291: the Metal test drives the kv_transfer_params branch production never uses; no test asserts prepared_cache/prompt_token_ids reach generate_stream, and the MLXChatModel prefill metadata copy (core.py:1739-1744) is untested. Test the prepared_cache path and chat prefill.

Simplification

  • xinference/model/llm/mlx/core.py L453: dead fallback xavier.fetch and kv_transfer_params param on generate_stream/generate (L428, L621, L643); production always passes prepared_cache. Delete and rework test_xavier.py:291.
  • xinference/model/llm/mlx/xavier.py L76: redundant not tokenizer_files or (any() implies non-empty). Drop it and hoist the filename tuples to module constants.
    net: -7 lines possible

Blocking: yes — recommended event: COMMENT

Blocking issues

  • xinference/core/supervisor.py:3299 (worker.py:6754-6763), major, MLX prefill/decode crash (e.g. Metal OOM) makes recover_model raise KeyError on cache_config["rank"], so the replica is never relaunched and P/D requests fail until manual relaunch. [new]

Comment thread xinference/core/supervisor.py
Comment thread xinference/core/supervisor.py Outdated
Comment thread xinference/model/llm/mlx/core.py Outdated
Comment thread xinference/model/llm/mlx/xavier.py Outdated
Comment thread xinference/model/llm/mlx/xavier.py Outdated
Comment thread xinference/model/llm/xavier/backends/bytes/storage.py Outdated
Comment thread xinference/model/llm/mlx/xavier.py Outdated
Comment thread xinference/model/llm/mlx/xavier.py Outdated
@qinxuye
qinxuye requested a review from rogercloud October 6, 2026 15:25

@rogercloud rogercloud left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Major

  • xinference/model/llm/mlx/xavier.py:273 — publish encodes prompts larger than the storage capacity on the batch event loop, then storage drops every page; this repeats on every long request. Fix: see inline.

Minor

  • xinference/core/supervisor.py:3121 — PD launch of an MLX vision model is accepted, but every request then fails. Fix: see inline.
  • doc/source/user_guide/pd_separation.rst:100 — no guidance on sizing xavier_cache_bytes, and the five-minute expiry is described inaccurately. Fix: see inline.
  • doc/source/user_guide/pd_separation.rst:49 — the doc says to "select vLLM or SGLang" in the Web UI, but replica-placement-config.tsx:43 now allows placement for MLX too. Fix: add MLX.
  • xinference/core/tests/test_pd_launch.py:179 — the n_worker: 2 case passes for the wrong reason, and the MLX n_worker != 1 check at xinference/core/supervisor.py:3112 can never be reached. Fix: see inline.
  • xinference/model/llm/mlx/xavier.py:330 — the partial-hit P prefill path (0 < cached < len(prefix)) is untested; current tests only cover the cold case and the fully warm case. xinference/model/llm/xavier/backends/bytes/storage.py:119 — eviction of unpinned pages in prepare_handoff is also untested. Fix: add a test for each.

Simplification

  • xinference/model/llm/mlx/xavier.py L292: delete: the writes=None branch in flush. Its only caller (core.py:617) and the tests always pass a list, so make writes required.

net: -1 lines possible

Blocking: no — recommended event: APPROVE

Comment thread xinference/model/llm/mlx/xavier.py Outdated
Comment thread xinference/core/supervisor.py
Comment thread doc/source/user_guide/pd_separation.rst Outdated
Comment thread xinference/core/tests/test_pd_launch.py Outdated
@qinxuye
qinxuye force-pushed the feat/mlx-xavier-cache branch from 0d67672 to 7a32937 Compare October 6, 2026 16:29
@qinxuye
qinxuye merged commit 2c3ab8a into xorbitsai:main Oct 6, 2026
15 checks passed
@qinxuye
qinxuye deleted the feat/mlx-xavier-cache branch October 6, 2026 17:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants