feat(npu): add DeepSeek V4 async CAM W4A8 layered FFN path - #387
Merged
jiangkuaixue123 merged 8 commits intoSep 24, 2026
Merged
jiangkuaixue123 merged 8 commits into
jiangkuaixue123 merged 8 commits into
Conversation
Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
jiangkuaixue123
commented
Sep 23, 2026
jiangkuaixue123
commented
Sep 23, 2026
jiangkuaixue123
commented
Sep 23, 2026
jiangkuaixue123
commented
Sep 23, 2026
jiangkuaixue123
commented
Sep 23, 2026
Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
Rename the legacy transfer state after the layered branch so mypy keeps the two types separate. Type the AST test namespace and event list, add the required SPDX headers, and spell PR vllm-project#384 as prose so markdownlint does not parse it as a heading. Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
Allow the opt-in layered W4A8 executor to start with the checkpoint limit of 10.0 while the fused GMM kernel lacks clamp support. Log the actual limit and ignored status, document that accuracy is not equivalent, and cover the startup behavior in the CPU test. Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
The model loader exposes a logical [E,K,N] INT32 view over five-dimensional NZ storage, while the layered op required a five-dimensional view and expanded the packed INT8 storage as though it were INT32. Accept the logical view when its NZ storage matches the expected shape and expand the A8W4 physical last dimension by two INT4 values per INT8. Update the NPU fixture to use the loader packing. This changes metadata validation and interpretation only, with negligible runtime overhead; contribute the contract fix upstream after broader shape coverage. Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
Repeated identical W4A8 layered calls can differ slightly on NPU, so an exact comparison between two invocations incorrectly fails even when both match the FP32 reference. Compare the tail-perturbed valid rows against that reference with the same tolerance, and record the eight passing Ascend 910C cases plus the NZ op API scope in the testing guide. Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
The typos hook rewrote dimension names such as WEIGHT_ND_DIM_LIMIT in the layered W4A8 operator, causing PR pre-commit CI to fail after the storage-shape fix. Ignore the specific ND identifier forms used by the Ascend operator code so the hook leaves these semantic names intact. Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
The async CAM metadata carries a device-side layer index in a slice of a larger tensor. CANN host tiling validates its physical storage shape and rejected that slice during Prefill only startup. Clone the one-element slice for contiguous model layer IDs; noncontiguous IDs already use index_select, which materializes a one-element tensor. Extend the device regression to both layer ID layouts and assert the standalone storage in the CPU contract test. Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Add an opt-in DeepSeek V4 W4A8 path for Async CAM FFN that keeps received layer IDs and routed token counts on device and calls layered W13/W2 GMM. The existing path remains the default.
Scope and implementation
dynamicQuant=1, eager ModelRunnerV1, static routed experts, and per-channel or per-group loaded weights.AFD_ASYNC_CAM_LAYERED_GMM=1enables the FFN path. Other roles and unsupported configurations fail early.[E,K,N]INT32 view over five-dimensional NZ INT8-packed storage. The change is limited to shape validation and packed storage interpretation; the compute kernel is unchanged.swiglu_limit=10.0, but the current layered fused GMM does not implement that clamp. The opt-in path temporarily ignores this limit, logs that decision, and is not numerically equivalent by contract. Full-model accuracy was checked below; implementing the clamp remains follow-up work.group_listinterface from fix(npu): accept device group lists in layered GMM #384.Validation
Docs
Updated
docs/npu/CAM_ASYNC_CONNECTOR_USER_GUIDE.mdanddocs/npu/TESTING.mdfor the opt-in path, scope, and validation.Essential PR Checklist