Skip to content

Pull requests: ISEEKYAN/Megatron-LM

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

Bound vocab-parallel cross entropy memory by token chunks
#132 opened Jul 25, 2026 by ISEEKYAN Owner Loading…
[lite] feat: plug-and-play LoRA primitive (apply_lora_to_chunks)
#131 opened Jul 25, 2026 by ISEEKYAN Owner Loading…
Fix M-FSDP DCP checkpoint parameter scope
#125 opened Jul 22, 2026 by conver334 Loading…
Add Muon comparison study: verl + Megatron vs. Megatron-Lite
#99 opened Jul 13, 2026 by ISEEKYAN Owner Loading…
Document CUDA Graph design for THD-only MLite training
#98 opened Jul 10, 2026 by ISEEKYAN Owner Loading…
Add Hopper blockwise FP8 precision profiles
#96 opened Jul 10, 2026 by ISEEKYAN Owner Loading…
Document Muon post-training configuration
#95 opened Jul 10, 2026 by ISEEKYAN Owner Loading…
Analyze incremental RL weight synchronization
#93 opened Jul 9, 2026 by ISEEKYAN Owner Loading…
Document MLite FP8 training architecture
#92 opened Jul 9, 2026 by ISEEKYAN Owner Loading…
Document Muon optimizer integration design
#88 opened Jul 9, 2026 by ISEEKYAN Owner Loading…
ProTip! Find all pull requests that aren't related to any open issues with -linked:issue.