Skip to content

[Operator Mechanism] Fix topk returning NaN for an empty reduction axis - #79799

Open
feixi139 wants to merge 4 commits into
PaddlePaddle:developfrom
feixi139:fix_topk_empty_axis
Open

feixi139 wants to merge 4 commits into
PaddlePaddle:developfrom
feixi139:fix_topk_empty_axis

Conversation

@feixi139

@feixi139 feixi139 commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

PR Category

Operator Mechanism

PR Types

Bug fixes

Description

本 PR 修复 topk 在规约轴长度为 0 且 k >= 1 时静默返回全 NaN 的问题,覆盖 CPU / GPU / XPU 三个 TopkKernel(topk 与 topk_v1 共用同一实现)。核心问题和结果可以概括为:

  1. 空规约轴被短路成 NaN 输出:当输入在 topk 轴上的长度为 0(如 paddle.empty([1024, 0]) 沿 axis=-1 取 k=10)时,输出在该轴上的期望长度为 k >= 1,输出并非空张量。三个 kernel 在真正的 k 越界校验(PADDLE_ENFORCE_GE(x.numel(), k))之前,先用 x.numel() == 0 短路,直接把输出填成 NaN、索引填成 0 并返回,导致非法请求被静默接受、返回无意义结果,而不是报错。
  2. 与 torch 行为不一致:torch 对同样的输入抛出 selected index k out of range。Paddle 前向的这类"轴无法提供 k 个元素"应当报错,而不是产出 NaN。
  3. 合法空张量不受影响:真正合法的空输出(如 paddle.empty([0, 5]) 沿 axis=-1 取 k=3,输出形状 [0, 3]、numel == 0)走的是更早的 out->numel() == 0 空输出分支,本 PR 不触碰该路径。

1. 空规约轴静默返回 NaN

问题

三个 kernel 的开头顺序都是:先处理"输出为空"(out->numel() == 0)与 0-D 输入,再处理 x.numel() == 0,最后才做 k 的范围校验。问题出在 x.numel() == 0 这一分支:

import paddle
# 规约轴长度为 0,k>=1:输出期望形状 [1024, 10],非空
x = paddle.empty([1024, 0], dtype='float32')
v, i = paddle.topk(x, 10, axis=-1)
# 期望:抛 InvalidArgument(对齐 torch "selected index k out of range")
# 实际(修复前):v 为形状 [1024, 10] 的全 NaN,无任何报错

能走到 x.numel() == 0 这一步,说明输出已通过前面的 out->numel() == 0 判断(即输出非空)。输入为空而输出非空,只可能是 topk 轴长度为 0——此时该轴根本无法提供 k >= 1 个元素,属于非法请求。旧实现却把它当成"空输入"用 Full(NAN) 兜底:

if (x.numel() == 0) {
  Full<T, Context>(dev_ctx, out->dims(), NAN, out);
  Full<int64_t, Context>(dev_ctx, indices->dims(), 0, indices);
  return;
}

CPU kernel 的 FullTopK 里其实有正确的 PADDLE_ENFORCE_LE(k, input_width),但对空输入这条分支到不了;GPU/XPU 也在该分支之后才有 k 的范围校验。三处是同一根因。

改动
  • 把 CPU / GPU / XPU 三个 TopkKernel 中 x.numel() == 0 的 NaN 兜底分支替换为显式报错:
if (x.numel() == 0) {
  // Reaching here with a non-empty output implies the topk axis has size 0,
  // which cannot supply the requested k (>=1) elements. Align with the
  // "selected index k out of range" semantics instead of returning NaN.
  PADDLE_THROW(errors::InvalidArgument(
      "topk cannot select k = %d elements from an axis of size 0 "
      "(selected index k out of range).",
      static_cast<int>(k)));
}
  • 该分支位于"输出为空"早返回之后,因此只拦截"输入空 + 输出非空"这一必然非法的组合;合法的空输出请求在更早的分支已经返回,不受影响。
  • topk_v1 复用 TopkKernel,无需单独改动。
修改后结果
  • 规约轴长度为 0 且 k >= 1 的请求在进入计算前明确抛出 InvalidArgument,与 torch 的 selected index k out of range 语义一致;
  • 不再返回全 NaN 的"看似成功"结果;
  • CPU / GPU / XPU 三个后端行为统一。

修改后的总体行为

类别 之前 之后
规约轴长度为 0 且 k>=1(输出非空) 静默返回全 NaN,不报错 抛 InvalidArgument,对齐 torch
合法空输出(如 [0,5] 沿 axis=-1 取 k=3,输出 [0,3]) 返回空张量 返回空张量(不变)
0-D 输入 原样拷贝、索引置 0 原样拷贝、索引置 0(不变)
非空且 k 合法 正常 topk 正常 topk(不变)

合法输入的结果不变;修改只影响原本被静默接受、返回 NaN 的非法边界请求。

是否引起精度变化

否

When the topk axis has size 0 while k >= 1, the CPU/GPU/XPU kernels
short-circuited the empty input to a NaN-filled output and returned
before the k-range check, silently producing NaN instead of raising.
Align with torch's "selected index k out of range" semantics by
throwing InvalidArgument in the empty-input branch of all three kernels.
@risemeup1111

risemeup1111 commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Paddle-Bot Review Board (review完成)

序号 位置 优先级 规则来源 状态
1 空规约轴异常路径缺少测试覆盖 P2 仓库规则:测试质量与验证 ✅

Powered by Nyanpasu claude with Opus 4.8 默认推理级别, please check the suggestions carefully.

Comment on lines +172 to +178
// Reaching here with a non-empty output implies the topk axis has size 0,
// which cannot supply the requested k (>=1) elements. Align with the
// "selected index k out of range" semantics instead of returning NaN.
PADDLE_THROW(errors::InvalidArgument(
"topk cannot select k = %d elements from an axis of size 0 "
"(selected index k out of range).",
static_cast<int>(k)));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2

本 PR 将「空规约轴(x.numel()==0 且输出非空)」从静默返回全 NaN 改为抛出 InvalidArgument,属于用户可见的行为变化,但未新增覆盖该异常路径的测试。

现有 test/legacy_test/test_top_k_v2_op.py 中 [0, 20]、axis=1 的用例,其规约轴长度为 20(非空),空的是 batch 维,走的是更早的 out->numel()==0 合法空输出分支,并不能触发本次改动的分支。

建议在 test/legacy_test/test_top_k_v2_op.py(及/或 test_top_k_op.py)补充异常路径断言,例如 x = paddle.empty([1024, 0]); paddle.topk(x, 10, axis=-1) 应抛 InvalidArgument,以固化 CPU/GPU/XPU 三端一致的新行为并防止回归。

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

新提交 018d129 已在 test/legacy_test/test_top_k_v2_op.py::TestTopKAPI.test_errors 补充异常路径断言:paddle.empty([1024, 0]) 取 k=10, axis=-1 断言抛 ValueError,并新增 [0, 5] 合法空输出返回 [0, 3] 的回归断言。与相邻 k=0/k=-1 用例一致,覆盖了本次行为变化。该问题已解决。

@codecov-commenter

codecov-commenter commented Sep 22, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 0% with 1 line in your changes missing coverage. Please review.
⚠️ Please upload report for BASE (develop@c817668). Learn more about missing BASE report.

Files with missing lines Patch % Lines
paddle/phi/kernels/cpu/top_k_kernel.cc 0.00% 1 Missing ⚠️

❌ Your patch status has failed because the patch coverage (0.00%) is below the target coverage (90.00%). You can increase the patch coverage or adjust the target coverage.

Additional details and impacted files
@@            Coverage Diff             @@
##             develop   #79799   +/-   ##
==========================================
  Coverage           ?    0.00%           
==========================================
  Files              ?        1           
  Lines              ?        1           
  Branches           ?        0           
==========================================
  Hits               ?        0           
  Misses             ?        1           
  Partials           ?        0           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

The empty-reduction-axis assertion ran on the default place, so on the GPU
coverage build the CPU kernel throw branch was never executed and reported
as uncovered. Force CPUPlace (always available) plus GPU/XPU when compiled so
each backend's throw branch is exercised, and keep the legit empty-output
control on every place.
On CUDA builds topk registers TopkKernelCuda, not TopkKernel, so its own
empty-input branch still filled the output with NaN when the reduction
axis had size 0 and k >= 2 (k == 1 already delegates to TopkKernel).
Throw InvalidArgument there as well so CPU/GPU/XPU and ROCm all report
the same "selected index k out of range" error.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants