Skip to content

perf(npu): reuse dispatch prefixes in async combine recv - #403

Merged
jiangkuaixue123 merged 2 commits into
vllm-project:mainfrom
jiangkuaixue123:codex/optimize-combine-recv-prefix
Sep 29, 2026
Merged

jiangkuaixue123 merged 2 commits into
vllm-project:mainfrom
jiangkuaixue123:codex/optimize-combine-recv-prefix

Conversation

@jiangkuaixue123

Copy link
Copy Markdown
Collaborator

Purpose

Async combine recv currently scans routing entries from token zero on every AIV to reconstruct its starting expert occurrence indices. Reuse dispatch send's existing per-AIV cumulative counts so each core starts at its own token partition for batches larger than the AIV count.

Small batches keep the original compiled path (keys 100/101); large batches use separate template specializations (102/103). FP16/BF16 small-key executable bytes were verified identical to baseline.

Issue

Standalone performance improvement; no issue is closed.

Scope

  • In scope: combine recv host tiling and kernel specialization, plus reproducible experiment results and per-rank samples.
  • Other communication operators and the test harness are unchanged.

Implementation Notes

This changes plugin-owned Ascend code, without upstream monkey patches or interface changes. The extra prefix load reuses an existing UB buffer before its final contents are populated. UB layout, top-k accumulation order and the communication completion protocol stay unchanged.

Correctness requires dispatch send and combine recv to retain identical AIV counts and token partition rules. Keep that contract when changing either partition.

Test Plan

Build the native operator package and run the existing tests/npu/async_cam_precision.py and async_cam_performance.py with four ranks. Profile with multi-rank msprof, excluding precision preflight and five warmups; take the mean of five per-iteration maxima across Attention ranks. Full reproduction commands and sample data are committed in docs/npu/COMBINE_RECV_PREFIX_OPTIMIZATION.md and its linked JSON.

Test Result

Validated on jcz_afd2: Ascend910_9382, CANN 9.0.1, torch 2.10.0+cpu, torch-npu 2.10.0.post2, 2 Attention + 2 MoE ranks, TP=2.

Scenario Baseline → optimized device task Change
Lengths 1025/1024, hidden 256, 8 experts, top-k 3, BF16/quantized 140.036 → 39.560 µs −71.75%
Lengths 1025/1024, hidden 256, 256 experts, top-k 2, BF16/unquantized 105.668 → 42.064 µs −60.19%
Lengths 49/48, hidden 256, top-k 3, BF16/quantized 17.592 → 16.676 µs −5.21%
  • Seven precision configurations passed with zero mismatches: both dtypes, quantization on/off, sparse routes, multiple chunks, uneven lengths, 49/48 mixed-key boundary and 256 experts. The performance harness also passed its precision preflight.
  • Runtime profiles confirm keys 100/102 at the mixed boundary and key 103 for large FP16 inputs. Installed host/kernel binaries were checked.
  • Small-batch hidden-7168 A/B: 14.368→14.768 µs and 14.392→13.788 µs, with opposite directions and byte-identical small-key code; no stable regression observed.
  • These are device task results, not end-to-end model speedups. Large-shape candidate results each represent one profile group. Event timing was noisier and is not used for the improvement claims.
  • git diff --check, JSON validation and recomputation of all committed cross-rank sample aggregates passed. Full pre-commit tooling is not installed locally.

Hardware results were collected against baseline 205bd770113aee7c9c098c55068ca2ce7c766d9c. The PR is rebased onto current main; the three tested source-file hashes are unchanged, and intervening upstream commits do not touch these operators.

Docs Impact

  • docs/npu/COMBINE_RECV_PREFIX_OPTIMIZATION.md: algorithm, contract, precision matrix, timing method, limitations and reproduction commands.
  • docs/npu/experiments/combine_recv_prefix_20260928.json: per-rank samples, precision results and small-key executable hashes.

Essential PR Checklist
  • Purpose is clear and linked to public context when possible.
  • Scope is bounded.
  • Compatibility with vLLM v0.26.0 is considered.
  • No changes are made to the vLLM source checkout.
  • Plugin-owned classes or explicit dotted class paths are preferred over monkey patches.
  • Any compat shim or monkey patch is isolated, idempotent, version-guarded, documented, and tested (none added).
  • Imports remain CPU-safe; CUDA-heavy work is delayed or GPU-gated (no Python import changes).
  • Validation evidence is included, including skipped GPU tests when applicable (NPU operator validation).
  • Documentation impact is stated.

Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
@jiangkuaixue123
jiangkuaixue123 force-pushed the codex/optimize-combine-recv-prefix branch from 9521918 to 9d48b09 Compare September 29, 2026 09:12
@jiangkuaixue123
jiangkuaixue123 merged commit 4c2541a into vllm-project:main Sep 29, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants