Skip to content

Fix/gdn conv seq fp32 state#500

Draft
liu-shaojun wants to merge 1 commit into
upgrade/vllm-xpu-v0.21.0from
fix/gdn-conv-seq-fp32-state
Draft

Fix/gdn conv seq fp32 state#500
liu-shaojun wants to merge 1 commit into
upgrade/vllm-xpu-v0.21.0from
fix/gdn-conv-seq-fp32-state

Conversation

@liu-shaojun

Copy link
Copy Markdown
Contributor

No description provided.

v0.21.0 sets mamba_ssm_dtype="float32" (from HF config), so ssm_state
is float32 instead of fp16. The kernel was reading/writing it as fp16,
producing silent data corruption (all-zero state writes, NaN output).

Fix: change ssm_state_ptr from fp16* to float* in gdn_conv_fused_seq.h
and the binding in esimd_kernel_lgrf.sycl. lsc_load/store helpers now
operate on float32 directly (computation was already fp32 internally).

The interleaved variant (gdn_conv_fused.h) is NOT changed — it serves
FP8 models where ssm_state remains fp16.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@liu-shaojun
liu-shaojun changed the base branch from main to upgrade/vllm-xpu-v0.21.0 June 29, 2026 08:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant