Skip to content

[Cherry-Pick][Operator Mechanism] Fix topk returning NaN for an empty reduction axis - #79817

Closed
feixi139 wants to merge 3 commits into
PaddlePaddle:release/3.4from
feixi139:cp79800_release34
Closed

feixi139 wants to merge 3 commits into
PaddlePaddle:release/3.4from
feixi139:cp79800_release34

Conversation

@feixi139

Copy link
Copy Markdown
Contributor

PR Category

Operator Mechanism

PR Types

Bug fixes

Description

本 PR 修复 topk 在规约轴长度为 0 且 k >= 1 时静默返回全 NaN 的问题,覆盖 CPU / GPU / XPU 三个 TopkKernel(topk 与 topk_v1 共用同一实现)。核心问题和结果可以概括为:

  1. 空规约轴被短路成 NaN 输出:当输入在 topk 轴上的长度为 0(如 paddle.empty([1024, 0]) 沿 axis=-1 取 k=10)时,输出在该轴上的期望长度为 k >= 1,输出并非空张量。三个 kernel 在真正的 k 越界校验(PADDLE_ENFORCE_GE(x.numel(), k))之前,先用 x.numel() == 0 短路,直接把输出填成 NaN、索引填成 0 并返回,导致非法请求被静默接受、返回无意义结果,而不是报错。
  2. 与 torch 行为不一致:torch 对同样的输入抛出 selected index k out of range。Paddle 前向的这类"轴无法提供 k 个元素"应当报错,而不是产出 NaN。
  3. 合法空张量不受影响:真正合法的空输出(如 paddle.empty([0, 5]) 沿 axis=-1 取 k=3,输出形状 [0, 3]、numel == 0)走的是更早的 out->numel() == 0 空输出分支,本 PR 不触碰该路径。

1. 空规约轴静默返回 NaN

问题

三个 kernel 的开头顺序都是:先处理"输出为空"(out->numel() == 0)与 0-D 输入,再处理 x.numel() == 0,最后才做 k 的范围校验。问题出在 x.numel() == 0 这一分支:

import paddle
# 规约轴长度为 0,k>=1:输出期望形状 [1024, 10],非空
x = paddle.empty([1024, 0], dtype='float32')
v, i = paddle.topk(x, 10, axis=-1)
# 期望:抛 InvalidArgument(对齐 torch "selected index k out of range")
# 实际(修复前):v 为形状 [1024, 10] 的全 NaN,无任何报错

能走到 x.numel() == 0 这一步,说明输出已通过前面的 out->numel() == 0 判断(即输出非空)。输入为空而输出非空,只可能是 topk 轴长度为 0——此时该轴根本无法提供 k >= 1 个元素,属于非法请求。旧实现却把它当成"空输入"用 Full(NAN) 兜底:

if (x.numel() == 0) {
  Full<T, Context>(dev_ctx, out->dims(), NAN, out);
  Full<int64_t, Context>(dev_ctx, indices->dims(), 0, indices);
  return;
}

CPU kernel 的 FullTopK 里其实有正确的 PADDLE_ENFORCE_LE(k, input_width),但对空输入这条分支到不了;GPU/XPU 也在该分支之后才有 k 的范围校验。三处是同一根因。

改动
  • 把 CPU / GPU / XPU 三个 TopkKernel 中 x.numel() == 0 的 NaN 兜底分支替换为显式报错:
if (x.numel() == 0) {
  // Reaching here with a non-empty output implies the topk axis has size 0,
  // which cannot supply the requested k (>=1) elements. Align with the
  // "selected index k out of range" semantics instead of returning NaN.
  PADDLE_THROW(errors::InvalidArgument(
      "topk cannot select k = %d elements from an axis of size 0 "
      "(selected index k out of range).",
      static_cast<int>(k)));
}
  • 该分支位于"输出为空"早返回之后,因此只拦截"输入空 + 输出非空"这一必然非法的组合;合法的空输出请求在更早的分支已经返回,不受影响。
  • topk_v1 复用 TopkKernel,无需单独改动。
修改后结果
  • 规约轴长度为 0 且 k >= 1 的请求在进入计算前明确抛出 InvalidArgument,与 torch 的 selected index k out of range 语义一致;
  • 不再返回全 NaN 的"看似成功"结果;
  • CPU / GPU / XPU 三个后端行为统一。

修改后的总体行为

类别 之前 之后
规约轴长度为 0 且 k>=1(输出非空) 静默返回全 NaN,不报错 抛 InvalidArgument,对齐 torch
合法空输出(如 [0,5] 沿 axis=-1 取 k=3,输出 [0,3]) 返回空张量 返回空张量(不变)
0-D 输入 原样拷贝、索引置 0 原样拷贝、索引置 0(不变)
非空且 k 合法 正常 topk 正常 topk(不变)

合法输入的结果不变;修改只影响原本被静默接受、返回 NaN 的非法边界请求。

devPR:#79799

是否引起精度变化

否

matmul_grad was routed through GeneralBinaryGradInferMeta, which only
shares x/y meta with dx/dy and skips the contraction (K) dimension
check that the forward MatmulInferMeta enforces. As a result the
backward op silently accepted K-mismatched inputs that the forward
rejects (returning results instead of raising).

Add a dedicated MatmulGradInferMeta that replicates the forward K
check and keeps the original share_meta behavior, and route the dense
matmul_grad (dygraph / PIR / legacy static) and legacy_matmul_grad to
it. Sparse matmul_grad has no transpose args and is left unchanged.
The K-dim check keyed on config.is_runtime, but the backward InferMeta is
invoked with the default MetaConfig (is_runtime=true) even at graph-build
time, so it wrongly rejected legal dynamic-shape graphs with a -1 dim (e.g.
InputSpec([-1, -1])). A real runtime tensor never carries a -1 dim, so
enforcing only when both contraction dims are concrete still catches every
genuine K mismatch while allowing symbolic shapes.
@feixi139 feixi139 changed the title [Operator Mechanism] Fix topk returning NaN for an empty reduction axis [Cherry-Pick][Operator Mechanism] Fix topk returning NaN for an empty reduction axis Sep 28, 2026
@risemeup1111

risemeup1111 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Paddle-Bot Review Board (review完成)

序号 位置 优先级 规则来源 状态
1 PR 描述与实际改动不符(描述为 topk,实为 matmul_grad K 校验) P2 默认规则 🚧

Powered by Nyanpasu claude with Opus 4.8 默认推理级别, please check the suggestions carefully.

@risemeup1111 risemeup1111 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 PR 描述与实际改动完全不符:描述通篇讲的是修复 topk 在空规约轴(reduction axis 长度为 0)返回 NaN 的问题,涉及 CPU/GPU/XPU 三个 TopkKernel;但本 PR 改动的 6 个文件全部是 matmul_grad 的 K 维一致性校验(新增 MatmulGradInferMeta、更新 dygraph_backward.yaml / static_backward.yaml / legacy/static_backward.yaml 及 test/legacy_test/test_matmul_0_size_op.py),未包含任何 topk 相关代码。作为 cherry-pick 到 release/3.4 的 PR,描述与实际 port 内容不一致会影响 release 追溯与评审核对,建议将描述更新为与实际改动(matmul_grad K 校验,对应 devPR)一致。

@feixi139 feixi139 closed this Sep 28, 2026
@feixi139
feixi139 deleted the cp79800_release34 branch September 30, 2026 06:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants