Hi team,
Thank you for the impressive work on h3.c!
I'd like to kindly inquire if there are plans to officially support Sparse Attention (or standard interface hooks for custom attention mechanisms) in h3.c?
For reference, I recently implemented SolAttn on Apple Silicon (MPS / Metal) and evaluated it on MiniMax-H3 Turbo. On an M3 Ultra, the SolAttn Metal kernel achieves a 5.18x – 7.36x speedup on long-sequence self-attention over native MPS SDPA (yielding a 1.39x – 1.68x end-to-end generation speedup across 480p and 720p workflows):
👉 yshenaw/ComfyUI-SolAttn-MPS
If sparse attention fits into your roadmap, I would be very happy to share our benchmark insights or help with adapting sparse attention support for the project.
Thanks again for your time and contribution!
Best regards,
Hi team,
Thank you for the impressive work on h3.c!
I'd like to kindly inquire if there are plans to officially support Sparse Attention (or standard interface hooks for custom attention mechanisms) in h3.c?
For reference, I recently implemented SolAttn on Apple Silicon (MPS / Metal) and evaluated it on MiniMax-H3 Turbo. On an M3 Ultra, the SolAttn Metal kernel achieves a 5.18x – 7.36x speedup on long-sequence self-attention over native MPS SDPA (yielding a 1.39x – 1.68x end-to-end generation speedup across 480p and 720p workflows):
👉 yshenaw/ComfyUI-SolAttn-MPS
If sparse attention fits into your roadmap, I would be very happy to share our benchmark insights or help with adapting sparse attention support for the project.
Thanks again for your time and contribution!
Best regards,