Repository navigation
FEAT: report cache hit tokens across vLLM, SGLang and MLX - #5647
Merged
Merged
Conversation
rogercloud
approved these changes
Oct 7, 2026
rogercloud
left a comment
Contributor
There was a problem hiding this comment.
Minor
doc/source/user_guide/client_api.rst:99— P/D semantics ofcached_tokensare undocumented (decode reports imported KV as cached, so it is close toprompt_tokenseven for cold requests).xinference/model/llm/tests/test_cached_token_usage.py:212— MLXMISSING/Nonecases test a state production MLX cannot produce.xinference/model/llm/tests/test_cached_token_usage.py:218— no coverage for the tools path or the Responses zero default.
Blocking: no — recommended event: APPROVE
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
vLLM and SGLang currently discard the engine's per-request cache hit count when building completion responses, so clients cannot see native prefix-cache or Xavier cache reuse. This PR exposes that count as
usage.prompt_tokens_details.cached_tokensin completions and chat completions, including streaming responses. The existing Responses API adapter forwards it asusage.input_tokens_details.cached_tokens.MLX already reports cached tokens, but its continuous batching path does not emit the final usage-only chunk requested by
stream_options.include_usage. This PR adds that chunk and verifies actual Xavier reuse through the MLX model adapter.In heterogeneous P/D, the router forwards the decode engine's usage unchanged. The count includes KV imported from the prefill engine, including KV computed for this request; it does not separately measure historical prefix-cache hits on the prefill engine.
choices=[]usage chunk across all three engines.prompt_tokens_detailswhen older engines return no count, while preserving reported zeroes. The Responses API retains its existing zero fallback.Validation:
pre-commit run --all-filesand commit hooks passed.msgfmt --check --check-format;python doc/build_i18n.py --allpassed. The changed page built in English and all nine locales, and both translated paragraphs were verified in the rendered HTML.Existing failures reproduced on base
cb6c1ad2694515f42b23abd61ddbf51758226389:test_async_to_tool_completion_chunks_without_thinkingandtest_async_to_tool_completion_chunks_with_parser).VLLM_VERSION=None; they pass in the installed vLLM 0.21.0 Linux environment.vllm.executor.executor_base, removed by vLLM 0.21.0.No full-model serving benchmark, full heterogeneous serving run, or real SGLang/Xavier GPU run was performed for this response-mapping change; the available SGLang environment is 0.5.20, below Xavier's 0.5.21 requirement.