Pinned Loading
Repositories
- vllm Public Forked from vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
- uccl Public Forked from uccl-project/uccl
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
- dgx-spark-wheels Public
pip-installable PEP 503 index of prebuilt GB10 (DGX Spark / sm_121) Python wheels for flash-attn, sageattention, sageattn3, nunchaku, onnxruntime-gpu -- built in CI, glibc-2.35-floored, sigstore-attested.
- DeepGEMM Public Forked from vllm-project/DeepGEMM
DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
- flextensor Public Forked from ai-dynamo/flextensor
FlexTensor is a tensor offloading and management library for PyTorch that enables running large models on limited GPU memory by intelligently offloading tensors between GPU and CPU memory.
- LMCache Public Forked from LMCache/LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
- Mooncake Public Forked from kvcache-ai/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
- causal-conv1d Public Forked from Dao-AILab/causal-conv1d
Causal depthwise conv1d in CUDA, with a PyTorch interface
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…