Repository navigation
feat: add MLX Xavier KV cache sharing and PD separation - #5638
Merged
Merged
Conversation
This was referenced Oct 6, 2026
qinxuye
force-pushed
the
feat/mlx-xavier-cache
branch
from
October 6, 2026 11:05
0c38afc to
64f07ab
Compare
rogercloud
reviewed
Oct 6, 2026
rogercloud
left a comment
Contributor
There was a problem hiding this comment.
Major
- xinference/core/supervisor.py:3299: MLX P/D replica auto-recovery always fails (see inline; worker.py:6754-6763). [new]
- xinference/core/supervisor.py:3111: built-in unquantized MLX specs are rejected (see inline). [new]
Minor
- xinference/model/llm/mlx/core.py:352: an exception in
publish()drops the whole step's results (see inline). - xinference/model/llm/mlx/xavier.py:273:
flush()gathers the shared write set, so one cancelled request cancels others (see inline). - xinference/model/llm/mlx/xavier.py:247: unguarded
release_handofffails or masks the real result (see inline); same at L333. - xinference/model/llm/xavier/backends/bytes/storage.py:88: LRU order evicts chain heads first and oversized puts self-evict (see inline).
- xinference/model/llm/mlx/xavier.py:94: weights are fully re-hashed on every load (see inline).
- xinference/model/llm/mlx/xavier.py:129: contract configure is deferred to the first RPC (see inline).
- xinference/core/supervisor.py:3098-3117: MLX
enable_xavierwithreplica<=1still creates the cache actor and publishes every prompt; mirror the SGLang/vLLM warn-and-disable. - xinference/model/llm/mlx/core.py:461-470: publish re-encodes and ships the whole prefix, including pages just fetched remotely. Publish only pages beyond the cached prefix.
- xinference/model/llm/mlx/tests/test_xavier.py:291: the Metal test drives the
kv_transfer_paramsbranch production never uses; no test assertsprepared_cache/prompt_token_idsreachgenerate_stream, and theMLXChatModelprefill metadata copy (core.py:1739-1744) is untested. Test theprepared_cachepath and chat prefill.
Simplification
- xinference/model/llm/mlx/core.py L453: dead fallback
xavier.fetchandkv_transfer_paramsparam on generate_stream/generate (L428, L621, L643); production always passesprepared_cache. Delete and rework test_xavier.py:291. - xinference/model/llm/mlx/xavier.py L76: redundant
not tokenizer_files or(any()implies non-empty). Drop it and hoist the filename tuples to module constants.
net: -7 lines possible
Blocking: yes — recommended event: COMMENT
Blocking issues
- xinference/core/supervisor.py:3299 (worker.py:6754-6763), major, MLX prefill/decode crash (e.g. Metal OOM) makes
recover_modelraise KeyError oncache_config["rank"], so the replica is never relaunched and P/D requests fail until manual relaunch. [new]
rogercloud
approved these changes
Oct 6, 2026
rogercloud
left a comment
Contributor
There was a problem hiding this comment.
Major
xinference/model/llm/mlx/xavier.py:273—publishencodes prompts larger than the storage capacity on the batch event loop, then storage drops every page; this repeats on every long request. Fix: see inline.
Minor
xinference/core/supervisor.py:3121— PD launch of an MLX vision model is accepted, but every request then fails. Fix: see inline.doc/source/user_guide/pd_separation.rst:100— no guidance on sizingxavier_cache_bytes, and the five-minute expiry is described inaccurately. Fix: see inline.doc/source/user_guide/pd_separation.rst:49— the doc says to "select vLLM or SGLang" in the Web UI, butreplica-placement-config.tsx:43now allows placement for MLX too. Fix: add MLX.xinference/core/tests/test_pd_launch.py:179— then_worker: 2case passes for the wrong reason, and the MLXn_worker != 1check atxinference/core/supervisor.py:3112can never be reached. Fix: see inline.xinference/model/llm/mlx/xavier.py:330— the partial-hit P prefill path (0 < cached < len(prefix)) is untested; current tests only cover the cold case and the fully warm case.xinference/model/llm/xavier/backends/bytes/storage.py:119— eviction of unpinned pages inprepare_handoffis also untested. Fix: add a test for each.
Simplification
xinference/model/llm/mlx/xavier.pyL292: delete: thewrites=Nonebranch inflush. Its only caller (core.py:617) and the tests always pass a list, so makewritesrequired.
net: -1 lines possible
Blocking: no — recommended event: APPROVE
qinxuye
force-pushed
the
feat/mlx-xavier-cache
branch
from
October 6, 2026 16:29
0d67672 to
7a32937
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds Xavier prefix KV sharing and explicit prefill/decode replicas to MLX. Prefill evaluates all but the final prompt token and exports FP16 KV pages; decode imports the prefix, evaluates the final token and generates the response. Both streaming and non-streaming requests use the existing APIs.
Based on merged #5631 and current
main(10f91bdcd). There are no unmerged PR dependencies.enable_xavier=Trueand explicit prefill/decode placement through the API, clients, CLI and UI. Requires Apple silicon, mlx-lm >= 0.31.2, one worker per replica and unquantized (none/fp16/bf16labels) full-attention Qwen2/Qwen3/Llama text models with FP16 weights and KV cache. Vision models fail before loading.Final API measurements at
7a329378b: Apple M5 Pro / 64 GiB, Python 3.12.3, MLX 0.32.0, mlx-lm 0.31.3 and published xoscar 0.11.1. Identical FP16 Qwen2.5-0.5B-Instruct weights, two independent model processes per mode, native local prompt cache enabled, eight serial warm requests and 32 output tokens per request.All 60 streamed/non-streamed responses at this revision match within each context. Warm Xavier requests reuse 1137/4437 tokens, with actual cache reads and zero remaining handoffs. These are single-Mac measurements on a shared Metal GPU; cold-document requests retain the 24-token chat prefix from warm-up. Multi-host performance, native NIXL and cross-engine/NVIDIA-to-Mac transfers are outside this PR.
Validation: 543 focused regression tests passed, 10 skipped for platform/dependency conditions. Coverage includes worker recovery, request cancellation, failed publication/release, page-budget/protocol limits, aged-prefix LRU, handoff reservation eviction, load-time contracts, prepared-cache forwarding, chat prefill and real Metal full/partial KV reuse. Changed Python files pass pre-commit; benchmark passes Black/isort/Ruff; the changed UI file passes ESLint/Prettier. All 18 catalogs validate and compile, and both changed pages render with verified MLX translations in English plus all nine locales. Full platform CI runs on the PR.