Conversation
matmul_grad was routed through GeneralBinaryGradInferMeta, which only shares x/y meta with dx/dy and skips the contraction (K) dimension check that the forward MatmulInferMeta enforces. As a result the backward op silently accepted K-mismatched inputs that the forward rejects (returning results instead of raising). Add a dedicated MatmulGradInferMeta that replicates the forward K check and keeps the original share_meta behavior, and route the dense matmul_grad (dygraph / PIR / legacy static) and legacy_matmul_grad to it. Sparse matmul_grad has no transpose args and is left unchanged.
The K-dim check keyed on config.is_runtime, but the backward InferMeta is invoked with the default MetaConfig (is_runtime=true) even at graph-build time, so it wrongly rejected legal dynamic-shape graphs with a -1 dim (e.g. InputSpec([-1, -1])). A real runtime tensor never carries a -1 dim, so enforcing only when both contraction dims are concrete still catches every genuine K mismatch while allowing symbolic shapes.
Contributor
Paddle-Bot Review Board (review完成)
Powered by Nyanpasu claude with Opus 4.8 默认推理级别, please check the suggestions carefully. |
risemeup1111
left a comment
Contributor
There was a problem hiding this comment.
PR 描述与实际改动完全不符:描述通篇讲的是修复
topk 在空规约轴(reduction axis 长度为 0)返回 NaN 的问题,涉及 CPU/GPU/XPU 三个 TopkKernel;但本 PR 改动的 6 个文件全部是 matmul_grad 的 K 维一致性校验(新增 MatmulGradInferMeta、更新 dygraph_backward.yaml / static_backward.yaml / legacy/static_backward.yaml 及 test/legacy_test/test_matmul_0_size_op.py),未包含任何 topk 相关代码。作为 cherry-pick 到 release/3.4 的 PR,描述与实际 port 内容不一致会影响 release 追溯与评审核对,建议将描述更新为与实际改动(matmul_grad K 校验,对应 devPR)一致。
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PR Category
Operator Mechanism
PR Types
Bug fixes
Description
本 PR 修复
topk在规约轴长度为 0 且k >= 1时静默返回全NaN的问题,覆盖 CPU / GPU / XPU 三个TopkKernel(topk与topk_v1共用同一实现)。核心问题和结果可以概括为:paddle.empty([1024, 0])沿axis=-1取k=10)时,输出在该轴上的期望长度为k >= 1,输出并非空张量。三个 kernel 在真正的k越界校验(PADDLE_ENFORCE_GE(x.numel(), k))之前,先用x.numel() == 0短路,直接把输出填成NaN、索引填成0并返回,导致非法请求被静默接受、返回无意义结果,而不是报错。selected index k out of range。Paddle 前向的这类"轴无法提供 k 个元素"应当报错,而不是产出 NaN。paddle.empty([0, 5])沿axis=-1取k=3,输出形状[0, 3]、numel == 0)走的是更早的out->numel() == 0空输出分支,本 PR 不触碰该路径。1. 空规约轴静默返回 NaN
问题
三个 kernel 的开头顺序都是:先处理"输出为空"(
out->numel() == 0)与 0-D 输入,再处理x.numel() == 0,最后才做k的范围校验。问题出在x.numel() == 0这一分支:能走到
x.numel() == 0这一步,说明输出已通过前面的out->numel() == 0判断(即输出非空)。输入为空而输出非空,只可能是 topk 轴长度为 0——此时该轴根本无法提供k >= 1个元素,属于非法请求。旧实现却把它当成"空输入"用Full(NAN)兜底:CPU kernel 的
FullTopK里其实有正确的PADDLE_ENFORCE_LE(k, input_width),但对空输入这条分支到不了;GPU/XPU 也在该分支之后才有k的范围校验。三处是同一根因。改动
TopkKernel中x.numel() == 0的 NaN 兜底分支替换为显式报错:topk_v1复用TopkKernel,无需单独改动。修改后结果
k >= 1的请求在进入计算前明确抛出InvalidArgument,与 torch 的selected index k out of range语义一致;NaN的"看似成功"结果;修改后的总体行为
InvalidArgument,对齐 torch[0,5]沿axis=-1取k=3,输出[0,3])合法输入的结果不变;修改只影响原本被静默接受、返回 NaN 的非法边界请求。
devPR:#79799
是否引起精度变化
否