Skip to content

feat(npu): add DeepSeek V4 async CAM W4A8 layered FFN path - #387

Merged
jiangkuaixue123 merged 8 commits into
vllm-project:mainfrom
jiangkuaixue123:codex/async-cam-w4a8-layered-gmm
Sep 24, 2026
Merged

jiangkuaixue123 merged 8 commits into
vllm-project:mainfrom
jiangkuaixue123:codex/async-cam-w4a8-layered-gmm

Conversation

@jiangkuaixue123

@jiangkuaixue123 jiangkuaixue123 commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Add an opt-in DeepSeek V4 W4A8 path for Async CAM FFN that keeps received layer IDs and routed token counts on device and calls layered W13/W2 GMM. The existing path remains the default.

Scope and implementation

  • Ascend 910C, DeepSeek V4 W4A8, Attention-side gate, dynamicQuant=1, eager ModelRunnerV1, static routed experts, and per-channel or per-group loaded weights.
  • AFD_ASYNC_CAM_LAYERED_GMM=1 enables the FFN path. Other roles and unsupported configurations fail early.
  • Build per-layer weight and scale TensorLists after model load; map global MoE layer IDs to compact device indices. Materialize a one-element layer-index slice so CANN sees standalone storage.
  • Run layered W13 and W2 on capacity-sized activations, then send compact CAM metadata to combine. No layer ID or token-count host read is introduced in the execution path.
  • Adapt the vendored GMM op API to the checkpoint loader's logical [E,K,N] INT32 view over five-dimensional NZ INT8-packed storage. The change is limited to shape validation and packed storage interpretation; the compute kernel is unchanged.
  • The model specifies swiglu_limit=10.0, but the current layered fused GMM does not implement that clamp. The opt-in path temporarily ignores this limit, logs that decision, and is not numerically equivalent by contract. Full-model accuracy was checked below; implementing the clamp remains follow-up work.
  • Builds on the device group_list interface from fix(npu): accept device group lists in layered GMM #384.

Validation

  • Pre-commit and DCO checks pass; targeted CPU/unit tests pass.
  • CANN extension rebuilt on Ascend 910C. Synthetic layered W4A8 tests covered contiguous and noncontiguous layer IDs, NZ-packed weights, per-channel/per-group scales, and capacity tails. One full-capacity comparison fluctuated on its first run and passed on targeted rerun; this is retained as a numerical validation caveat.
  • Prefill-only Attention DP2/TP4 + FFN EP8 with the layered path returned a correct chat response. GSM8K 5-shot, 64 concurrent, 300 samples: strict/flexible exact match 97.00% (291/300).
  • PD deployment: Prefill Attention DP2/TP2 + FFN EP4 with layered W4A8; ordinary Decode DP8/TP1/EP8. End-to-end chat via the PD proxy returned HTTP 200 and the correct answer.
  • PD GSM8K 5-shot, 64 concurrent, all 1319 test samples: strict/flexible exact match 95.15% (1255/1319); no request errors. The earlier PD configuration scored 95.22% strict and 95.30% flexible on the same 1319 samples. This is a model-level accuracy check, not proof of exact SwiGLU-limit equivalence.
  • Decode must use the system CAM operator directory. Inheriting this PR's custom CAM directory caused the native Decode W4A8 fused GMM to fail shape validation during startup; restoring the system directory resolved it. This is a deployment environment setting, not a code change in this PR.
  • The Async CAM FFN receive kernel can time out after roughly ten minutes without requests. The validation deployment uses a four-minute keepalive; idle-time handling remains open.
  • No performance improvement is claimed yet.

Docs

Updated docs/npu/CAM_ASYNC_CONNECTOR_USER_GUIDE.md and docs/npu/TESTING.md for the opt-in path, scope, and validation.


Essential PR Checklist
  • Purpose and supported scope are clear.
  • vLLM v0.26.0 compatibility is considered.
  • Imports remain CPU-safe; NPU operator loading is delayed until a qualified FFN path starts.
  • CPU, NPU, deployment, and accuracy evidence is included.
  • Numerical and idle-time limitations are documented.
  • Documentation impact is stated.

Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
Comment thread afd_plugin/compat/npu/feature_validation.py Outdated
Comment thread afd_plugin/model_executor/models/npu/deepseek_v2_attention_gate.py Outdated
Comment thread docs/npu/CAM_ASYNC_CONNECTOR_USER_GUIDE.md Outdated
Comment thread docs/npu/TESTING.md Outdated
Comment thread afd_plugin/model_executor/models/deepseek_v2.py Outdated
Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
@jiangkuaixue123 jiangkuaixue123 changed the title feat(npu): add optional async CAM W4A8 layered FFN path feat(npu): add DeepSeek V4 async CAM W4A8 layered FFN path Sep 23, 2026
Rename the legacy transfer state after the layered branch so mypy keeps the two types separate. Type the AST test namespace and event list, add the required SPDX headers, and spell PR vllm-project#384 as prose so markdownlint does not parse it as a heading.

Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
Allow the opt-in layered W4A8 executor to start with the checkpoint limit of 10.0 while the fused GMM kernel lacks clamp support. Log the actual limit and ignored status, document that accuracy is not equivalent, and cover the startup behavior in the CPU test.

Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
The model loader exposes a logical [E,K,N] INT32 view over five-dimensional NZ storage, while the layered op required a five-dimensional view and expanded the packed INT8 storage as though it were INT32. Accept the logical view when its NZ storage matches the expected shape and expand the A8W4 physical last dimension by two INT4 values per INT8. Update the NPU fixture to use the loader packing. This changes metadata validation and interpretation only, with negligible runtime overhead; contribute the contract fix upstream after broader shape coverage.

Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
Repeated identical W4A8 layered calls can differ slightly on NPU, so an exact comparison between two invocations incorrectly fails even when both match the FP32 reference. Compare the tail-perturbed valid rows against that reference with the same tolerance, and record the eight passing Ascend 910C cases plus the NZ op API scope in the testing guide.

Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
The typos hook rewrote dimension names such as WEIGHT_ND_DIM_LIMIT in the layered W4A8 operator, causing PR pre-commit CI to fail after the storage-shape fix.

Ignore the specific ND identifier forms used by the Ascend operator code so the hook leaves these semantic names intact.

Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
The async CAM metadata carries a device-side layer index in a slice of a larger tensor. CANN host tiling validates its physical storage shape and rejected that slice during Prefill only startup.

Clone the one-element slice for contiguous model layer IDs; noncontiguous IDs already use index_select, which materializes a one-element tensor. Extend the device regression to both layer ID layouts and assert the standalone storage in the CPU contract test.

Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
@jiangkuaixue123
jiangkuaixue123 merged commit 205bd77 into vllm-project:main Sep 24, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant