CUDA sparse triangular solve (SpTRSV) with four progressively optimized kernels - level scheduling and caching reach a ~132x geomean speedup over CPU.
-
Updated
Jul 18, 2026 - Cuda
CUDA sparse triangular solve (SpTRSV) with four progressively optimized kernels - level scheduling and caching reach a ~132x geomean speedup over CPU.
Add a description, image, and links to the sptrsv topic page so that developers can more easily learn about it.
To associate your repository with the sptrsv topic, visit your repo's landing page and select "manage topics."