Popular repositories Loading
-
vllm-qwen3.8-flash-next-rtx-pro-6000-sharp-monitoring
vllm-qwen3.8-flash-next-rtx-pro-6000-sharp-monitoring PublicAn opinionated single-GPU vLLM recipe for Qwen3.8-Flash-Next on an RTX PRO 6000 Blackwell: vendored Qwen-Sharp chat template, MTP-3 speculative decoding, and a bundled local Prometheus + Grafana st…
Jinja 2
Repositories
Showing 1 of 1 repositories
- vllm-qwen3.8-flash-next-rtx-pro-6000-sharp-monitoring Public
An opinionated single-GPU vLLM recipe for Qwen3.8-Flash-Next on an RTX PRO 6000 Blackwell: vendored Qwen-Sharp chat template, MTP-3 speculative decoding, and a bundled local Prometheus + Grafana stack. Every non-default flag benchmarked.
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…