Conversation
x14ngch3n
force-pushed
the
fix/int32-range-kernel-params
branch
from
September 10, 2026 01:43
6d50553 to
cae4dc9
Compare
All GGUF kernel dispatchers receive tensor dims/strides at 'const int' parameters (mmvq/mmq/moe/moe_vec clones + quantize_row_q8_1_cuda). A model-declared dim >= 2^31 (legal at int64, unvalidated anywhere on the path) silently narrows there; the kernels use the value directly as tile extents and index strides -> GPU OOB read (moe exp_stride = W.stride(0) reading W-1GiB, CUDA illegal memory access) and silent wrong results (mmvq extent = truncated col). Fix: check_kernel_int_range at each public op entry, 24 checks covering every crossing. Kernel signatures unchanged (guard and addressing stay self-consistent). Verified on RTX 3090: control paths bit-identical; attacker inputs rejected at the op boundary with a named error. Co-Authored-By: Claude <noreply@anthropic.com>
…truncation + moe stride OOB)
x14ngch3n
force-pushed
the
fix/int32-range-kernel-params
branch
from
September 10, 2026 02:32
cae4dc9 to
f4b646e
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #127.
What
One-file fix:
check_kernel_int_range(TORCH_CHECK-class guard viaheader-only
c10::Error) at each public op entry — 24 checks coveringevery dim / derived
padded/ stride / slot-count that crosses into anint-typed kernel parameter (ggml_dequantize,ggml_mul_mat_vec_a8,ggml_mul_mat_a8,ggml_moe_a8,ggml_moe_a8_vec→ themmvq/mmq/moe/moe_vec dispatchers and
quantize_row_q8_1_cuda).Kernel signatures are unchanged on purpose: today every write path is
self-consistent (guard and write address share the same narrowed
value), which is why the observable impact is OOB read / wrong results
rather than OOB write — a partial widening that desynchronizes guard
and addressing would turn this into an OOB-write primitive. If you
prefer the
int64_twidening class end-to-end (the class of theexisting CVE-2026-53923 fix), the same differential applies — happy to
rework.
Differential verification
Plugin @
d4c1f0d, unmodifiedcsrc/, pre/post this patch, RTX 3090,torch 2.14.0+cu130:
col=64col=2^32+64RuntimeError: gguf: kernel parameter out of int32 range: mul_mat_vec col = 4294967360stride=32256W.stride(0)=3·2^30illegal memory access(readsW − 1 GiB, moe.cuh:36)RuntimeError: ... moe W.stride(0) (expert stride) = 3221227008Test
tests/e2e_repro_issue127.py— drives the real dispatchers throughvllm_gguf_plugin.opswith synthetic Q4_0 weights; control paths muststay bit-identical, the two attacker inputs must be rejected at the op
boundary. ~18 GiB GPU:
Build note for the patch itself: it throws via header-only
c10::Error(c10/util/Exception.h);TORCH_CHECK/torch/csrc/Exceptions.hpulls pybind11, which does not compile inthis TU under nvcc.