Production-grade FlashAttention FP8 e4m3 forward kernel for NVIDIA Blackwell consumer GPUs (sm_120a, e.g. RTX PRO 6000). 647–652 TFLOPS at hd=128, sl=8192. Multi-kernel dispatcher, C library with Go and Python bindings
cuda transformer attention gpu-kernels tensor-cores blackwell fp8 flash-attention fp8e4m3 sm120 rtx-pro-6000 sm-120a
-
Updated
Jun 12, 2026 - Cuda