Working recipes and measured benchmarks for serving large LLMs on pre-Ampere NVIDIA GPUs — Tesla V100 (sm_70) and RTX 2080 Ti (sm_75). vLLM forks, llama.cpp tensor parallelism, AWQ/MoE gotchas.
-
Updated
Jul 22, 2026
Working recipes and measured benchmarks for serving large LLMs on pre-Ampere NVIDIA GPUs — Tesla V100 (sm_70) and RTX 2080 Ti (sm_75). vLLM forks, llama.cpp tensor parallelism, AWQ/MoE gotchas.
ACE-Step 1.5 XL Optimized Fork: 4-Bit (INT4) Windows support + RTX 2080 Ti (Turing) stability fixes. Turing architecture compatible. Runs XL SFT on 11GB VRAM without OOM.
Add a description, image, and links to the rtx-2080ti topic page so that developers can more easily learn about it.
To associate your repository with the rtx-2080ti topic, visit your repo's landing page and select "manage topics."