Skip to content

FEAT: report cache hit tokens across vLLM, SGLang and MLX - #5647

Merged
qinxuye merged 3 commits into
xorbitsai:mainfrom
qinxuye:feat/xavier-cached-token-usage
Oct 7, 2026
Merged

qinxuye merged 3 commits into
xorbitsai:mainfrom
qinxuye:feat/xavier-cached-token-usage

Conversation

@qinxuye

@qinxuye qinxuye commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

vLLM and SGLang currently discard the engine's per-request cache hit count when building completion responses, so clients cannot see native prefix-cache or Xavier cache reuse. This PR exposes that count as usage.prompt_tokens_details.cached_tokens in completions and chat completions, including streaming responses. The existing Responses API adapter forwards it as usage.input_tokens_details.cached_tokens.

MLX already reports cached tokens, but its continuous batching path does not emit the final usage-only chunk requested by stream_options.include_usage. This PR adds that chunk and verifies actual Xavier reuse through the MLX model adapter.

In heterogeneous P/D, the router forwards the decode engine's usage unchanged. The count includes KV imported from the prefill engine, including KV computed for this request; it does not separately measure historical prefix-cache hits on the prefill engine.

  • Preserve the count in the final finish chunk and the optional choices=[] usage chunk across all three engines.
  • Pass through existing engine counts without additional cache queries, RPCs, or cache scheduling changes. SGLang's aggregate count is used directly, avoiding double counting its device/host/storage breakdown; vLLM counts the cached prompt once when generating multiple outputs.
  • Omit prompt_tokens_details when older engines return no count, while preserving reported zeroes. The Responses API retains its existing zero fallback.
  • Document the fields and streaming option in the client guide, with translations for all nine maintained locales.

Validation:

  • 33 cached-token regression cases passed on macOS and Linux with vLLM 0.21.0: cache hit/miss, missing/None counters, completions, chat conversion, Responses conversion, streaming finish/usage chunks, and multiple vLLM outputs.
  • macOS regression slice: 465 passed, 9 skipped, 10 deselected. This covers the engine adapters, shared chat utilities, MLX/Xavier, and Responses API. The deselections are three full-model launch tests and seven failures independently reproduced on the base revision.
  • Real Metal MLX/Xavier tests passed in both streaming and non-streaming modes: cold requests report 0, fresh remote-prefix reuse and P/D handoff report 69 cached tokens out of 70 prompt tokens, with identical generated text.
  • Linux runtime slice with vLLM 0.21.0: 341 passed, 8 skipped, plus three failures reproduced on the base revision (listed below).
  • Heterogeneous P/D regression slice: 188 passed, including eight new router cases covering chat/completions, streaming/non-streaming and GPU/host modes. The router preserves D usage without adding or replacing it with P usage. Four existing real Metal host-import cases also passed.
  • Review follow-up: 211 passed on macOS and 33 passed on Linux with vLLM 0.21.0. Missing/None counters are tested only for vLLM/SGLang, including the Responses zero default. Six tools-streaming hit/miss cases cover the Qwen parser, usage-only chunk and Responses mapping across all three engines. The P/D count semantics are documented and rendered in English and all nine locales.
  • Full pre-commit run --all-files and commit hooks passed.
  • All nine modified catalogs passed msgfmt --check --check-format; python doc/build_i18n.py --all passed. The changed page built in English and all nine locales, and both translated paragraphs were verified in the rendered HTML.

Existing failures reproduced on base cb6c1ad2694515f42b23abd61ddbf51758226389:

  • Two vLLM tool-stream tests have stale expected chunks (test_async_to_tool_completion_chunks_without_thinking and test_async_to_tool_completion_chunks_with_parser).
  • Five vLLM configuration tests fail on macOS without vLLM installed because they compare VLLM_VERSION=None; they pass in the installed vLLM 0.21.0 Linux environment.
  • The old Linux distributed-executor integration test imports vllm.executor.executor_base, removed by vLLM 0.21.0.

No full-model serving benchmark, full heterogeneous serving run, or real SGLang/Xavier GPU run was performed for this response-mapping change; the available SGLang environment is 0.5.20, below Xavier's 0.5.21 requirement.

@XprobeBot XprobeBot added this to the v3.x milestone Oct 7, 2026
@qinxuye
qinxuye requested a review from rogercloud October 7, 2026 15:25

@rogercloud rogercloud left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor

  • doc/source/user_guide/client_api.rst:99 — P/D semantics of cached_tokens are undocumented (decode reports imported KV as cached, so it is close to prompt_tokens even for cold requests).
  • xinference/model/llm/tests/test_cached_token_usage.py:212 — MLX MISSING/None cases test a state production MLX cannot produce.
  • xinference/model/llm/tests/test_cached_token_usage.py:218 — no coverage for the tools path or the Responses zero default.

Blocking: no — recommended event: APPROVE

Comment thread doc/source/user_guide/client_api.rst
Comment thread xinference/model/llm/tests/test_cached_token_usage.py
Comment thread xinference/model/llm/tests/test_cached_token_usage.py
@qinxuye
qinxuye merged commit 9a4a900 into xorbitsai:main Oct 7, 2026
15 checks passed
@qinxuye
qinxuye deleted the feat/xavier-cached-token-usage branch October 7, 2026 18:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants