Skip to content
#

multi-head-latent-attention

Here are 13 public repositories matching this topic...

MythosForge

🔧 Recurrent-Depth Transformer Research Lab — LTI-stable looped inference, switchable MLA/GQA attention, MoE routing & adaptive halting (ACT). Independent research, not affiliated with Anthropic.

  • Updated Jul 27, 2026
  • Python

This repository shows how to build a DeepSeek language model from scratch using PyTorch. It includes clean, well-structured implementations of advanced attention techniques such as key–value caching for fast decoding, multi-query attention, grouped-query attention, and multi-head latent attention.

  • Updated Jan 10, 2026
  • Jupyter Notebook

Add this topic to your repo

To associate your repository with the multi-head-latent-attention topic, visit your repo's landing page and select "manage topics."

Learn more