Repositories list
70 repositories
Magpie
PublicA lightweight, general-purpose framework for evaluating GPU kernel and benchmark.Primus
PublicA flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUsInfera
PublicMore token goodput from frontier models. A distributed, SLA-aware serving mesh — disaggregated prefill/decode, KV-aware routing, and cache offload, tuned to you…TraceLens
PublicAutomating analysis from trace filesGEAK
PublicHyperloom
Public- A high-performance acceleration library dedicated to large-scale model training on AMD GPUs
- Diffusion model inference benchmarking, profiling and optimizing
Primus-SaFE
PublicPrimus-SaFE(Stability and Fault Endurance)PrimusClaw
PublicLLM agent orchestration on Kubernetes: autonomous coding-agent sessions in per-session sandboxes.AgentKernelArena
PublicAgentKernelArena provides an end-to-end siloed-benchmarking environment where different LLM-powered agents—such as Cursor Agent, Claude Code, Codex, SWE-agent, …ALTO
PublicALTO: Advanced Low-precision Training and Optimizationvllm-2026
PublicAMD-Hybrid-Models
PublicOfficial repo for AMD hybrid models training and inference workflowApex
PublicAgents, and RL environment, for optimizing GPU kernels on AMD ROCm using LLM agents. Benchmarks LLM serving workloads end-to-end, profiles bottleneck kernels, o…Instella-MoE
Publichip_kernel_llm_lab
PublicAMDLongContextServing
PublicKimi-Linear long-context FP8 benchmark tooling for vLLM on AMD MI355X (AITER MLA head-padding + patched decode/prefill kernels).pr_pundit
PublicAn agent for OSS contributionsmlperf-common
PublicFarSkip-Collective
PublicTraining and inference implementation of FarSkip-Collective models enabling communication-computation overlapmaxtext-slurm
PublicHummingbirdXT
PublicThis repository presents an efficient acceleration pipeline for Diffusion Transformer (DiT) based video generation models, optimized for AMD client-grade GPUs, …torchtitan-amd
PublicA PyTorch native platform for training generative AI modelsgpt-fast
PublicInstella-Math
PublicAMD-LLM
PublicTraining code and resources for AMD-135M language models on AMD GPUs.m3d_rocm
PublicThis project is an optimized version of Matrix3D. It has better compatibility with ROCm ecosystem.prime_amd
PublicPARD
PublicPARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation (ICLR 26)
ProTip! When viewing an organization's repositories, you can use the
props. filter to filter by custom property.