perf(npu): reuse dispatch prefixes in async combine recv - #403
Merged
jiangkuaixue123 merged 2 commits intoSep 29, 2026
Merged
jiangkuaixue123 merged 2 commits into
jiangkuaixue123 merged 2 commits into
Conversation
Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
GuangyuZhu04
approved these changes
Sep 29, 2026
Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
jiangkuaixue123
force-pushed
the
codex/optimize-combine-recv-prefix
branch
from
September 29, 2026 09:12
9521918 to
9d48b09
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Async combine recv currently scans routing entries from token zero on every AIV to reconstruct its starting expert occurrence indices. Reuse dispatch send's existing per-AIV cumulative counts so each core starts at its own token partition for batches larger than the AIV count.
Small batches keep the original compiled path (keys 100/101); large batches use separate template specializations (102/103). FP16/BF16 small-key executable bytes were verified identical to baseline.
Issue
Standalone performance improvement; no issue is closed.
Scope
Implementation Notes
This changes plugin-owned Ascend code, without upstream monkey patches or interface changes. The extra prefix load reuses an existing UB buffer before its final contents are populated. UB layout, top-k accumulation order and the communication completion protocol stay unchanged.
Correctness requires dispatch send and combine recv to retain identical AIV counts and token partition rules. Keep that contract when changing either partition.
Test Plan
Build the native operator package and run the existing
tests/npu/async_cam_precision.pyandasync_cam_performance.pywith four ranks. Profile with multi-rankmsprof, excluding precision preflight and five warmups; take the mean of five per-iteration maxima across Attention ranks. Full reproduction commands and sample data are committed indocs/npu/COMBINE_RECV_PREFIX_OPTIMIZATION.mdand its linked JSON.Test Result
Validated on
jcz_afd2: Ascend910_9382, CANN 9.0.1, torch 2.10.0+cpu, torch-npu 2.10.0.post2, 2 Attention + 2 MoE ranks, TP=2.git diff --check, JSON validation and recomputation of all committed cross-rank sample aggregates passed. Full pre-commit tooling is not installed locally.Hardware results were collected against baseline
205bd770113aee7c9c098c55068ca2ce7c766d9c. The PR is rebased onto current main; the three tested source-file hashes are unchanged, and intervening upstream commits do not touch these operators.Docs Impact
docs/npu/COMBINE_RECV_PREFIX_OPTIMIZATION.md: algorithm, contract, precision matrix, timing method, limitations and reproduction commands.docs/npu/experiments/combine_recv_prefix_20260928.json: per-rank samples, precision results and small-key executable hashes.Essential PR Checklist