GLM-5.2-NVFP4-REAP-469B serving on SM120 (4× RTX PRO 6000 Blackwell) — one-command vLLM launch recipe, 250K context, DeepSeek Sparse Attention + MTP speculative decode
-
Updated
Jun 19, 2026 - Shell
GLM-5.2-NVFP4-REAP-469B serving on SM120 (4× RTX PRO 6000 Blackwell) — one-command vLLM launch recipe, 250K context, DeepSeek Sparse Attention + MTP speculative decode
Optimized SGLang runtime for Qwen3.8-27B FP8 with DFlash2 and Qwen3.8 Flash-Next NVFP4 with FR-Spec on one NVIDIA RTX PRO 6000 Blackwell 96 GB GPU (SM120): 524K context, HiCache and NIXL.
Validated GLM-5.3 Flash recipe for 2x NVIDIA RTX PRO 6000 Blackwell 96GB: 262K context, EXL3/TR3, adaptive MTP, tools, and vision.
Reproducible SGLang recipe + public prebuilt image (ghcr.io) for DeepSeek-V4-Flash-0731 on 4x RTX PRO 6000 Blackwell (SM120): TP4/DP4/EP4, 1M ctx, benchmarks, and the DSPARK draft-depth corruption boundary
SGLang 0.5.20 for Qwen3.8-Flash-Next on one RTX PRO 6000: 262K context, exact sampled MTP, validated hot-token/QSA patches, reproducible 8K-token benchmarks and real-task checks.
Systematic 24-hour benchmark study of Qwen3.6-27B inference on dual NVIDIA RTX PRO 6000 Blackwell SM120 (TP=2). 8 experiments comparing repne/vllm fork vs upstream vLLM across FP8/BF16/NVFP4/Q8_0 quants and MTP/DFlash speculative decoding. Peak: 2,083 tok/s at c=32. Quality: KLD vs BF16 = 0.0018 (noise floor).
Fish Audio OpenAudio S2-Pro on vLLM-Omni. low-latency ~100ms TTFA, OpenAI-compatible, runs on NVIDIA Blackwell (RTX 5090 / RTX PRO 6000). Self-hosted streaming TTS & voice cloning.
Image-to-3D-Video-Asset-Generator is an all-in-one generative 3D pipeline that transitions smoothly from textual concepts or reference images into fully realized 3D mesh assets (.glb), dynamic camera movements in 5-second MP4 videos, and clean bundle exports (.zip).
MiniMax-M3 (428B MoE) running on 3× RTX PRO 6000 Blackwell at TP=3 with 240K context, FP8 KV cache, and working multimodal vision input. Includes dist_utils.py patch for non-divisible attention heads.
Production-grade FlashAttention FP8 e4m3 forward kernel for NVIDIA Blackwell consumer GPUs (sm_120a, e.g. RTX PRO 6000). 647–652 TFLOPS at hd=128, sl=8192. Multi-kernel dispatcher, C library with Go and Python bindings
An opinionated single-GPU vLLM recipe for Qwen3.8-Flash-Next on an RTX PRO 6000 Blackwell: vendored Qwen-Sharp chat template, MTP-3 speculative decoding, and a bundled local Prometheus + Grafana stack. Every non-default flag benchmarked.
Taming Qwen3.8-Flash-Next on a single RTX 6000 Pro 96GB — KV pool 374K → 1,111,168, 1M ctx, native NVFP4 MTP | 单卡服役实录:96G 显存 + 62G 内存,SSD-Stream 双通 PLE,768K 长文实测过闸,全数字带日志出处
Rust + CUDA LLM inference engine for Blackwell (Tuned specifically on RTX PRO 6000, RTX 5090, B200): OpenAI-compatible (+converse and ant) serving, per-model X hardware exactness gates. NVFP4/mixed (fp8 hybrid, 4o6, etc - correctness, performance, hardware specific adapted) main quant support.
Hub for ongoing Qwen inference benchmarks on NVIDIA Blackwell. Indexes all studies, hosts the rolling SOTA leaderboard, points to the toolchain.
Prolepsis is a speculative decoding implementation for Qwen3 draft-target models with Hugging Face and vLLM backends. On an RTX PRO 6000 at batch size 1, it measured 1.72x throughput with vLLM FP8 and 1.32x with Hugging Face BF16, with complete latency and response artifacts.
The shared toolchain for ComfyUI workflow packages on any NVIDIA GPU host: one installer, pinned packs, declared models from Hugging Face or GitHub, the boot, the driver, and Comfy MCP so an agent can drive the machine. MIT, maintained by SoulCraft.
GLM-5.3-Flash UD-IQ3_XXS on one RTX PRO 6000: 56.84 tok/s across 2K-token probes, 512K context, native vision, RAM expert caching and DFlash2. Pinned source and reproducible benchmarks.
Prebuilt spconv v2.3.8 wheels for CUDA 12.8 / 13.0 with native Blackwell (RTX 50-series, sm_120) kernels — for default PyPI torch (cu130) or torch +cu128
To associate your repository with the rtx-pro-6000 topic, visit your repo's landing page and select "manage topics."