StableAdamW with Muon orthogonalisation for matrix parameters.
-
Updated
Jun 9, 2026 - Python
StableAdamW with Muon orthogonalisation for matrix parameters.
OpenAI parameter-golf challenge: train the smallest LM that fits in 16MB. My SP8192 frontier submission (openai/parameter-golf#1887): 11L x 512d w/ targeted middle recurrence, MuonEq-R optimizer, int6 GPTQ+SDClip, Brotli-11 compression, legal score-first TTT. 8xH100 DDP.
Comprehensive resources and scripts for training and fine-tuning Large Language Models (LLMs) from scratch using Hugging Face Transformers, litGPT, and custom GPT implementations with PyTorch and PyTorch Lightning.
Add a description, image, and links to the language-model-training topic page so that developers can more easily learn about it.
To associate your repository with the language-model-training topic, visit your repo's landing page and select "manage topics."