A 197M-param decoder-only LM trained from scratch on FineWeb-Edu — PyTorch with RoPE, GQA, SwiGLU, 32k BPE, built for 6GB VRAM
-
Updated
Sep 10, 2026 - Python
A 197M-param decoder-only LM trained from scratch on FineWeb-Edu — PyTorch with RoPE, GQA, SwiGLU, 32k BPE, built for 6GB VRAM
GPT-style LM pretrained from scratch on FineWeb-Edu under 100M params: RoPE, RMSNorm, QK-norm, SwiGLU, Muon
从 0 训练约 1.15B 参数(激活约 0.16B)的 Kimi-K3 风格 MoE 模型:KDA + MLA 混合注意力,4×V100 分布式预训练(训练中)
Implementation of ELECTRA pretraining architecture on a logic-focused corpus. Achieved 96.42% discriminator accuracy with 80% reduction in discriminator loss. Includes generator-discriminator training pipeline, domain adaptive pretraining, and comparison with RoBERTa baseline
Architecture-agnostic data profiling via field physics — improves LLM pretraining by 1.0–1.9% on Transformers
Unofficial reproduction of Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training (NeurIPS 2025, arXiv:2504.13161) — search-found mixtures beat uniform baselines +0.014–0.031 STEM at d28; novel finding: selection-mechanism winner's curse; Ascend NPU backend.
To associate your repository with the llm-pretraining topic, visit your repo's landing page and select "manage topics."