@@ -20,7 +20,8 @@ sudo apt install libvulkan-dev vulkan-tools
2020# sudo apt install glslc-tools # not available on all distros
2121
2222# Optional: CUDA backend (NVIDIA GPUs)
23- # Install a CUDA Toolkit supported by your compiler and driver. CI uses CUDA 12.6.
23+ # Install a CUDA Toolkit supported by your compiler and driver. CI uses CUDA 12.6.3.
24+ # The prebuilt CUDA packages target Turing (CC 7.5) and newer GPUs.
2425# https://developer.nvidia.com/cuda-downloads
2526
2627# Optional: ccache for faster rebuilds
@@ -53,6 +54,10 @@ brew install cmake ccache
5354
5455# Vulkan SDK (optional, for Vulkan backend)
5556# https://vulkan.lunarg.com/sdk/home
57+ #
58+ # CUDA Toolkit 12.6.x (optional, for CUDA backend)
59+ # https://developer.nvidia.com/cuda-downloads
60+ # CUDA 12.6 supports Visual Studio 2022 / MSVC 193x.
5661```
5762
5863## Quick start
@@ -100,7 +105,7 @@ cmake -B build -DCMAKE_BUILD_TYPE=Release \
100105 -DGAME_GGML_BUILD_CLI=ON \
101106 -DGAME_GGML_CUDA=ON \
102107 -DGGML_NATIVE=OFF \
103- -DCMAKE_CUDA_ARCHITECTURES=" 75;80;86;89"
108+ -DCMAKE_CUDA_ARCHITECTURES=" 75-real ;80-real ;86-real ;89-real "
104109
105110# macOS Apple Silicon + Metal
106111cmake -B build -DCMAKE_BUILD_TYPE=Release \
@@ -117,6 +122,13 @@ cmake -B build -DCMAKE_BUILD_TYPE=Release \
117122cmake -B build -DCMAKE_BUILD_TYPE=Release `
118123 -DGAME_GGML_BUILD_CLI=ON `
119124 -DGAME_GGML_VULKAN=ON
125+
126+ # Windows + CUDA (from Visual Studio 2022 Developer PowerShell)
127+ cmake -B build -DCMAKE_BUILD_TYPE=Release `
128+ -DGAME_GGML_BUILD_CLI=ON `
129+ -DGAME_GGML_CUDA=ON `
130+ -DGGML_NATIVE=OFF `
131+ -DCMAKE_CUDA_ARCHITECTURES=" 75-real;80-real;86-real;89-real"
120132```
121133
122134## Converting a PyTorch checkpoint to GGUF
@@ -163,14 +175,33 @@ build/bin/game_ggml_cli serve game_medium.gguf
163175# Then write binary request frames to stdin (see src/cli/main.cpp for protocol)
164176```
165177
166- ## CI CUDA scope
167-
168- The hosted CI builds and packages the Linux x64 CUDA backend with CUDA Toolkit
169- 12.6. It verifies Toolkit discovery, CUDA compilation, linking, and that the CLI
170- starts with ` --version ` . GitHub-hosted runners do not provide an NVIDIA GPU, so
171- actual CUDA inference must still be smoke-tested on an NVIDIA system. The
172- packaged CUDA backend uses the CUDA runtime and cuBLAS libraries supplied by the
173- installed NVIDIA CUDA runtime/toolkit.
178+ ## CUDA compatibility and CI scope
179+
180+ The hosted CI builds Linux x64 and Windows x64 CUDA packages with CUDA Toolkit
181+ 12.6.3 and Visual Studio 2022 on Windows. It verifies Toolkit discovery, CUDA
182+ compilation, linking, and that the CLI starts with ` --version ` . GitHub-hosted
183+ runners do not provide an NVIDIA GPU, so actual CUDA inference must still be
184+ smoke-tested on an NVIDIA system.
185+
186+ The release architecture list is ` 75-real;80-real;86-real;89-real ` , covering:
187+
188+ - CC 7.5: Turing (for example, GeForce RTX 20 series)
189+ - CC 8.0/8.6: Ampere (A100 and GeForce RTX 30 series)
190+ - CC 8.9: Ada (GeForce RTX 40 series)
191+
192+ ` -real ` emits native SASS for each target and avoids requiring the display driver
193+ to JIT PTX generated by CUDA 12.6. Pascal and Volta are not included in the
194+ prebuilt package. Source builds that need these older GPUs can use CUDA 12.x and
195+ add ` 61-real ` and/or ` 70-real ` . CUDA 13.0 removed NVCC offline compilation for
196+ architectures older than CC 7.5; use CUDA 12.9 or earlier when maintaining such
197+ builds.
198+
199+ CUDA 12.x minor-version compatibility requires at least NVIDIA driver
200+ 525.60.13 on Linux or 528.33 on Windows, subject to the limitations documented
201+ in NVIDIA's CUDA Compatibility Guide. Using the current production driver is
202+ recommended. The packages currently expect the CUDA 12 runtime and cuBLAS
203+ libraries to be installed on the target system; they do not bundle NVIDIA's
204+ runtime libraries.
174205
175206## Troubleshooting
176207
0 commit comments