Skip to content

All

    Repositories list

    • vllm

      Public
      A high-throughput and memory-efficient inference and serving engine for LLMs
      Python
      Apache License 2.0
      22k000Updated Sep 12, 2026Sep 12, 2026
    • uccl

      Public
      UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
      C++
      Apache License 2.0
      172000Updated Sep 12, 2026Sep 12, 2026
    • pip-installable PEP 503 index of prebuilt GB10 (DGX Spark / sm_121) Python wheels for flash-attn, sageattention, sageattn3, nunchaku, onnxruntime-gpu -- built i…
      Shell
      Apache License 2.0
      0000Updated Sep 9, 2026Sep 9, 2026
    • FlashInfer: Kernel Library for LLM Serving
      Python
      Apache License 2.0
      1.4k000Updated Sep 9, 2026Sep 9, 2026
    • DeepGEMM

      Public
      DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
      Cuda
      MIT License
      1.3k000Updated Sep 5, 2026Sep 5, 2026
    • FlexTensor is a tensor offloading and management library for PyTorch that enables running large models on limited GPU memory by intelligently offloading tensors…
      Python
      Apache License 2.0
      14000Updated Sep 1, 2026Sep 1, 2026
    • LMCache

      Public
      LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
      Python
      Apache License 2.0
      1.9k000Updated Aug 29, 2026Aug 29, 2026
    • nixl

      Public
      NVIDIA Inference Xfer Library (NIXL)
      C++
      Other
      440000Updated Aug 27, 2026Aug 27, 2026
    • Mooncake

      Public
      Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
      C++
      Apache License 2.0
      1.2k000Updated Aug 21, 2026Aug 21, 2026
    • Causal depthwise conv1d in CUDA, with a PyTorch interface
      Python
      BSD 3-Clause "New" or "Revised" License
      210000Updated Aug 17, 2026Aug 17, 2026
    • mamba

      Public
      Mamba SSM architecture
      Python
      Apache License 2.0
      1.8k000Updated Aug 17, 2026Aug 17, 2026
    • Branded PDF report toolkit — ReportLab + matplotlib chassis with markdown, charts, Mermaid, and LaTeX math support.
      Python
      Apache License 2.0
      0100Updated Aug 17, 2026Aug 17, 2026
    • pytorch

      Public
      Tensors and Dynamic neural networks in Python with strong GPU acceleration
      Python
      Other
      29k000Updated Aug 16, 2026Aug 16, 2026
    • vision

      Public
      Datasets, Transforms and Models specific to Computer Vision
      Python
      BSD 3-Clause "New" or "Revised" License
      7.3k000Updated Aug 16, 2026Aug 16, 2026
    • audio

      Public
      Data manipulation and transformation for audio signal processing, powered by PyTorch
      Python
      BSD 2-Clause "Simplified" License
      796000Updated Aug 16, 2026Aug 16, 2026
    • triton

      Public
      Development repository for the Triton language and compiler
      MLIR
      MIT License
      3.2k000Updated Aug 16, 2026Aug 16, 2026
    • kilocode

      Public
      Kilo is the all-in-one agentic engineering platform. Build, ship, and iterate faster with the most popular open source coding agent.
      TypeScript
      MIT License
      3.2k000Updated Aug 9, 2026Aug 9, 2026
    • ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
      C++
      MIT License
      4.2k000Updated Aug 5, 2026Aug 5, 2026
    • Recursive subfolder listing for ComfyUI's LoadImage/LoadImageMask/LoadAudio/LoadVideo pickers, plus a fix for nested-file basename collisions in GET /view.
      Python
      MIT License
      0000Updated Aug 5, 2026Aug 5, 2026
    • [ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across la…
      Cuda
      Apache License 2.0
      506000Updated Jul 28, 2026Jul 28, 2026
    • nunchaku

      Public
      [ICLR2025 Spotlight] SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models
      Python
      Apache License 2.0
      277000Updated Jul 28, 2026Jul 28, 2026
    • Fast and memory-efficient exact attention
      Python
      BSD 3-Clause "New" or "Revised" License
      3.1k000Updated Jul 28, 2026Jul 28, 2026
    • A code editor view written in Swift powered by tree-sitter.
      Swift
      MIT License
      160000Updated Jul 6, 2026Jul 6, 2026
    • Swift
      15000Updated Jul 6, 2026Jul 6, 2026
    • A text editor specialized for displaying and editing code documents. Written in pure Swift.
      Swift
      MIT License
      54000Updated Jul 4, 2026Jul 4, 2026
    • A Collection of Tree-Sitter Parsers for Syntax Highlighting
      Swift
      55000Updated Jul 4, 2026Jul 4, 2026
    • Axum HTTP server serving BGE-M3 dense and sparse embeddings via ONNX Runtime
      Rust
      Apache License 2.0
      15012Updated Jun 29, 2026Jun 29, 2026
    • Transparent HTTP reverse proxy for bge-m3-embedding-server: routes between GPU and CPU embedding pools based on health and queue depth
      Rust
      Apache License 2.0
      01010Updated Jun 16, 2026Jun 16, 2026
    • Synology Active Backup for Business Agent - Kernel 6.15+ Patches
      C
      8000Updated Jun 11, 2026Jun 11, 2026
    • ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
      C++
      MIT License
      4.2k000Updated Feb 28, 2026Feb 28, 2026
    ProTip! When viewing an organization's repositories, you can use the props. filter to filter by custom property.