Skip to content

fix(gguf): validate every dim/stride crossing into int kernel params (fixes #127) - #128

Open
x14ngch3n wants to merge 2 commits into
vllm-project:mainfrom
x14ngch3n:fix/int32-range-kernel-params
Open

x14ngch3n wants to merge 2 commits into
vllm-project:mainfrom
x14ngch3n:fix/int32-range-kernel-params

Conversation

@x14ngch3n

Copy link
Copy Markdown

Fixes #127.

What

One-file fix: check_kernel_int_range (TORCH_CHECK-class guard via
header-only c10::Error) at each public op entry — 24 checks covering
every dim / derived padded / stride / slot-count that crosses into an
int-typed kernel parameter (ggml_dequantize, ggml_mul_mat_vec_a8,
ggml_mul_mat_a8, ggml_moe_a8, ggml_moe_a8_vec → the
mmvq/mmq/moe/moe_vec dispatchers and quantize_row_q8_1_cuda).

Kernel signatures are unchanged on purpose: today every write path is
self-consistent (guard and write address share the same narrowed
value), which is why the observable impact is OOB read / wrong results
rather than OOB write — a partial widening that desynchronizes guard
and addressing would turn this into an OOB-write primitive. If you
prefer the int64_t widening class end-to-end (the class of the
existing CVE-2026-53923 fix), the same differential applies — happy to
rework.

Differential verification

Plugin @ d4c1f0d, unmodified csrc/, pre/post this patch, RTX 3090,
torch 2.14.0+cu130:

case pre-fix post-fix
mmvq control col=64 err 0.158 (within the repo's own test atol) identical
attacker col=2^32+64 output bitwise-equal to a 64-column control (silent truncation) RuntimeError: gguf: kernel parameter out of int32 range: mul_mat_vec col = 4294967360
moe control stride=32256 y = 31.984 (correct) identical
attacker W.stride(0)=3·2^30 CUDA illegal memory access (reads W − 1 GiB, moe.cuh:36) RuntimeError: ... moe W.stride(0) (expert stride) = 3221227008

Test

tests/e2e_repro_issue127.py — drives the real dispatchers through
vllm_gguf_plugin.ops with synthetic Q4_0 weights; control paths must
stay bit-identical, the two attacker inputs must be rejected at the op
boundary. ~18 GiB GPU:

GGUF_BUILD=/path/to/plugin python tests/e2e_repro_issue127.py

Build note for the patch itself: it throws via header-only
c10::Error (c10/util/Exception.h); TORCH_CHECK /
torch/csrc/Exceptions.h pulls pybind11, which does not compile in
this TU under nvcc.

@x14ngch3n
x14ngch3n force-pushed the fix/int32-range-kernel-params branch from 6d50553 to cae4dc9 Compare September 10, 2026 01:43
x14ngch3n and others added 2 commits September 10, 2026 10:32
All GGUF kernel dispatchers receive tensor dims/strides at 'const int'
parameters (mmvq/mmq/moe/moe_vec clones + quantize_row_q8_1_cuda). A
model-declared dim >= 2^31 (legal at int64, unvalidated anywhere on the
path) silently narrows there; the kernels use the value directly as
tile extents and index strides -> GPU OOB read (moe exp_stride =
W.stride(0) reading W-1GiB, CUDA illegal memory access) and silent
wrong results (mmvq extent = truncated col).

Fix: check_kernel_int_range at each public op entry, 24 checks covering
every crossing. Kernel signatures unchanged (guard and addressing stay
self-consistent). Verified on RTX 3090: control paths bit-identical;
attacker inputs rejected at the op boundary with a named error.

Co-Authored-By: Claude <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

GPU OOB read + silent wrong results: unvalidated int64→int narrowing of tensor dims in all GGUF kernel dispatchers

1 participant