Conversation
When the topk axis has size 0 while k >= 1, the CPU/GPU/XPU kernels short-circuited the empty input to a NaN-filled output and returned before the k-range check, silently producing NaN instead of raising. Align with torch's "selected index k out of range" semantics by throwing InvalidArgument in the empty-input branch of all three kernels.
Paddle-Bot Review Board (review完成)
Powered by Nyanpasu claude with Opus 4.8 默认推理级别, please check the suggestions carefully. |
| // Reaching here with a non-empty output implies the topk axis has size 0, | ||
| // which cannot supply the requested k (>=1) elements. Align with the | ||
| // "selected index k out of range" semantics instead of returning NaN. | ||
| PADDLE_THROW(errors::InvalidArgument( | ||
| "topk cannot select k = %d elements from an axis of size 0 " | ||
| "(selected index k out of range).", | ||
| static_cast<int>(k))); |
There was a problem hiding this comment.
本 PR 将「空规约轴(x.numel()==0 且输出非空)」从静默返回全 NaN 改为抛出 InvalidArgument,属于用户可见的行为变化,但未新增覆盖该异常路径的测试。
现有 test/legacy_test/test_top_k_v2_op.py 中 [0, 20]、axis=1 的用例,其规约轴长度为 20(非空),空的是 batch 维,走的是更早的 out->numel()==0 合法空输出分支,并不能触发本次改动的分支。
建议在 test/legacy_test/test_top_k_v2_op.py(及/或 test_top_k_op.py)补充异常路径断言,例如 x = paddle.empty([1024, 0]); paddle.topk(x, 10, axis=-1) 应抛 InvalidArgument,以固化 CPU/GPU/XPU 三端一致的新行为并防止回归。
There was a problem hiding this comment.
新提交 018d129 已在 test/legacy_test/test_top_k_v2_op.py::TestTopKAPI.test_errors 补充异常路径断言:paddle.empty([1024, 0]) 取 k=10, axis=-1 断言抛 ValueError,并新增 [0, 5] 合法空输出返回 [0, 3] 的回归断言。与相邻 k=0/k=-1 用例一致,覆盖了本次行为变化。该问题已解决。
Codecov Report❌ Patch coverage is
❌ Your patch status has failed because the patch coverage (0.00%) is below the target coverage (90.00%). You can increase the patch coverage or adjust the target coverage. Additional details and impacted files@@ Coverage Diff @@
## develop #79799 +/- ##
==========================================
Coverage ? 0.00%
==========================================
Files ? 1
Lines ? 1
Branches ? 0
==========================================
Hits ? 0
Misses ? 1
Partials ? 0 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
The empty-reduction-axis assertion ran on the default place, so on the GPU coverage build the CPU kernel throw branch was never executed and reported as uncovered. Force CPUPlace (always available) plus GPU/XPU when compiled so each backend's throw branch is exercised, and keep the legit empty-output control on every place.
On CUDA builds topk registers TopkKernelCuda, not TopkKernel, so its own empty-input branch still filled the output with NaN when the reduction axis had size 0 and k >= 2 (k == 1 already delegates to TopkKernel). Throw InvalidArgument there as well so CPU/GPU/XPU and ROCm all report the same "selected index k out of range" error.
PR Category
Operator Mechanism
PR Types
Bug fixes
Description
本 PR 修复
topk在规约轴长度为 0 且k >= 1时静默返回全NaN的问题,覆盖 CPU / GPU / XPU 三个TopkKernel(topk与topk_v1共用同一实现)。核心问题和结果可以概括为:paddle.empty([1024, 0])沿axis=-1取k=10)时,输出在该轴上的期望长度为k >= 1,输出并非空张量。三个 kernel 在真正的k越界校验(PADDLE_ENFORCE_GE(x.numel(), k))之前,先用x.numel() == 0短路,直接把输出填成NaN、索引填成0并返回,导致非法请求被静默接受、返回无意义结果,而不是报错。selected index k out of range。Paddle 前向的这类"轴无法提供 k 个元素"应当报错,而不是产出 NaN。paddle.empty([0, 5])沿axis=-1取k=3,输出形状[0, 3]、numel == 0)走的是更早的out->numel() == 0空输出分支,本 PR 不触碰该路径。1. 空规约轴静默返回 NaN
问题
三个 kernel 的开头顺序都是:先处理"输出为空"(
out->numel() == 0)与 0-D 输入,再处理x.numel() == 0,最后才做k的范围校验。问题出在x.numel() == 0这一分支:能走到
x.numel() == 0这一步,说明输出已通过前面的out->numel() == 0判断(即输出非空)。输入为空而输出非空,只可能是 topk 轴长度为 0——此时该轴根本无法提供k >= 1个元素,属于非法请求。旧实现却把它当成"空输入"用Full(NAN)兜底:CPU kernel 的
FullTopK里其实有正确的PADDLE_ENFORCE_LE(k, input_width),但对空输入这条分支到不了;GPU/XPU 也在该分支之后才有k的范围校验。三处是同一根因。改动
TopkKernel中x.numel() == 0的 NaN 兜底分支替换为显式报错:topk_v1复用TopkKernel,无需单独改动。修改后结果
k >= 1的请求在进入计算前明确抛出InvalidArgument,与 torch 的selected index k out of range语义一致;NaN的"看似成功"结果;修改后的总体行为
InvalidArgument,对齐 torch[0,5]沿axis=-1取k=3,输出[0,3])合法输入的结果不变;修改只影响原本被静默接受、返回 NaN 的非法边界请求。
是否引起精度变化
否