Describe the bug
For floating-point operands, Triton's % lowers through the floating remainder operation. On the tested build, the exact-zero result loses the sign of the dividend: -0.0 % 3.0 is stored as +0.0 for both f32 and f64. Other finite remainder values in the same probe matched the reference.
Minimal f32 reproducer:
import torch
import triton
import triton.language as tl
@triton.jit
def kernel(x_ptr, y_ptr, out_ptr):
i = tl.arange(0, 2)
x = tl.load(x_ptr + i)
y = tl.load(y_ptr + i)
tl.store(out_ptr + i, x % y)
x = torch.tensor([-0.0, 0.0], device="cuda", dtype=torch.float32)
y = torch.tensor([3.0, 3.0], device="cuda", dtype=torch.float32)
out = torch.empty_like(x)
kernel[(1,)](x, y, out)
torch.cuda.synchronize()
print(out, out.view(torch.int32))
Expected first result: -0.0 (f32 bits 0x80000000). Observed: +0.0 (bits 0x00000000). The same mismatch was observed for f64. TRITON_INTERPRET=1 passes this isolated % test.
Environment details
Triton 4fd7cc5 on A800
Describe the bug
For floating-point operands, Triton's
%lowers through the floating remainder operation. On the tested build, the exact-zero result loses the sign of the dividend:-0.0 % 3.0is stored as+0.0for both f32 and f64. Other finite remainder values in the same probe matched the reference.Minimal f32 reproducer:
Expected first result:
-0.0(f32 bits0x80000000). Observed:+0.0(bits0x00000000). The same mismatch was observed for f64.TRITON_INTERPRET=1passes this isolated%test.Environment details
Triton 4fd7cc5 on A800