A curated list of ICML 2026 (main conference) papers on training and serving large language models in a decentralized / collaborative dfdf way.
Full author lists and abstracts are in papers.md.
In scope: decentralized/distributed methods for training across the full pipeline (pretraining, mid-training, SFT, RL) and for serving/inference. Out of scope: downstream task-specific fine-tuning / adaptation, single-node-only efficiency tricks, and datacenter tensor/pipeline parallelism that assumes high-bandwidth interconnects.
This list was compiled with the help of Claude: candidate papers were surfaced from the ICML 2026 accepted-paper metadata, summarized, and curated to the scope above. Summaries and categorization may contain errors β corrections and additions welcome (see Contributing).
Legend: π€ oral Β· π poster session & board number. Title links go to OpenReview; [abstract] links to the abstract + full author list in papers.md. Sections ordered roughly by closeness to the topic's core.
- Low-communication & decentralized training
- Asynchronous, fault-tolerant & distributed post-training
- Decentralized & distributed optimization (methods & theory)
- Communication-efficient & low-precision training
- Distributed & disaggregated serving / inference
- Position & vision
- Contributing
- License
- Decentralized Instruction Tuning: Conflict-Aware Splitting and Weight Merging β Splits instruction-tuning data by gradient conflict, fine-tunes each partition communication-free, then merges once by weighted averaging.
Minsik Choi et al. Β· π Poster Session 1, #2107 Β· abstract - Factored Gossip DiLoCo: Reducing Blocking Communication within DiLoCo β Factorizes DiLoCo's outer sync into a non-blocking gossip step plus a blocking mixing step; degrades gracefully under stragglers and comm failures.
Chamin Hewa Koneputugodage et al. Β· π Poster Session 1, #3811 Β· abstract - LoRDO: Distributed Low-Rank Optimization with Infrequent Communication β Unifies low-rank optimization with infrequent communication for distributed foundation-model training, cutting communication ~10Γ.
Andrej JovanoviΔ et al. Β· π Poster Session 3, #3705 Β· abstract - MuLoCo: Muon is a Practical Inner Optimizer for DiLoCo β Uses Muon as DiLoCo's inner optimizer, yielding more directionally-correct pseudogradients and beating DiLoCo across 150Mβ3.1B models.
Benjamin ThΓ©rien et al. Β· π Poster Session 5, #3600 Β· abstract
Asynchronous pipeline parallelism
- Mitigating Staleness in Asynchronous Pipeline Parallelism via Basis Rotation β Rectifies stale gradients in asynchronous pipeline parallelism via basis rotation, restoring scalable async training (1B LLM).
Hyunji Jung et al. Β· π Poster Session 1, #3713 Β· abstract - One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining β Shows one-step-delay async pipeline (PipeDream-2BW) trains well with Muon + an error-feedback fix, up to 10B parameters.
Philip Zmushko et al. Β· π Poster Session 1, #3610 Β· abstract
Fault tolerance at scale
- Identifying and Mitigating Errors in Gradient Aggregation of Distributed Data Parallel Training β Detects and mitigates silent data corruption in distributed data-parallel gradient aggregation via dynamic asynchronous parameter sync.
Zhenheng Tang et al. Β· π Poster Session 4, #605 Β· abstract - SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs β Fault tolerance for 100k+-GPU LLM pretraining: stacked-parallelism redundancy and adaptive reordering mask node failures at near-constant overhead.
Jin Lee et al. Β· π Poster Session 5, #4101 Β· abstract
Distributed RL post-training
- D-ARL: A Distribution-Matched Asynchronous Reinforcement Learning Framework for Language Reasoning β Distribution-matched asynchronous RL for LLM post-training: selects stale rollout samples well-aligned with the current policy.
η½ ε― ε² et al. Β· π Poster Session 1, #2111 Β· abstract - DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training β Fully distributed, multi-controller RL framework for LLM post-training; decouples data from control, scaling near-linearly to 512 GPUs.
zhixin wang et al. Β· π Poster Session 1, #3809 Β· abstract
Decentralized SGD: convergence & topology
- Accelerated Dual Method for Distributed Optimization: An Inexact-Gradient View of Local Updates β Accelerated dual method with multi-step local SGD; attains optimal communication complexity when local computation suffices.
Junchi Yang et al. Β· π Poster Session 4, #3800 Β· abstract - High-Probability Convergence Guarantees of Decentralized SGD β First high-probability convergence guarantees (and linear speed-up) for Decentralized SGD under light-tailed noise.
Aleksandar Armacki et al. Β· π Poster Session 1, #3815 Β· abstract - Improved Convergence Analysis of Topology Dependence in Decentralized SGD β Tighter analysis showing all mixing-matrix eigenvalues β not just the spectral gap β govern Decentralized SGD convergence.
Yuki Takezawa et al. Β· π Poster Session 7, #3713 Β· abstract
Communication compression & error-feedback
- On the Convergence of Decentralized Stochastic Minimax Optimization Algorithm with Compressed Communication β Communication-efficient decentralized stochastic minimax (SGDA) with an error-feedback compression mechanism and convergence guarantees.
Yihan Zhang et al. Β· π poster TBD Β· abstract - On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach β An SDE framework for distributed compressed / sign SGD under (L0,L1)-smoothness and heavy-tailed noise.
Enea Monzio Compagnoni et al. Β· π Poster Session 4, #3600 Β· abstract - Temporal Difference Learning with Compressed Updates: Error-Feedback meets Reinforcement Learning β Brings error-feedback gradient compression to TD/Q-learning; multi-agent linear speed-up while communicating ~O(1) bits per iteration.
Aritra Mitra et al. Β· π Poster Session 3, #118 Β· abstract
Distributed bilevel / minimax / multi-level
- Convergence Analysis of Decentralized Hessian-/Jacobian-Free Algorithm for Nonconvex Stochastic Bilevel Optimization β First-order decentralized algorithm for nonconvex stochastic bilevel optimization under the PL condition (no Hessians/Jacobians).
Yihan Zhang et al. Β· π poster TBD Β· abstract - Distributed Stochastic $K$-Level Optimization Over Networks β Decentralized variance-reduced algorithm for stochastic K-level (K>2) optimization over networks.
Xinwen Zhang et al. Β· π poster TBD Β· abstract - FAB: A First-Order AB-based Gradient Algorithm for Distributed Bilevel Optimization over Time-Varying Directed Graphs β First-order PushβPull (AB) algorithm for distributed bilevel optimization over time-varying directed graphs.
Yaoshuai Ma et al. Β· π Poster Session 1, #3810 Β· abstract
Decentralized unlearning
- Trajectory-Aware Certified Decentralized Unlearning via SGD Stability β Certified unlearning of a client's influence from a decentralized-trained model via SGD-stability sensitivity analysis.
Hengliang Wu et al. Β· π Poster Session 7, #3100 Β· abstract
Low-precision / fully-quantized training
- AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs β 4-bit activation + 8-bit gradient quantization with precision-preserving 8-bit all-reduce for memory-efficient distributed LLM training.
WenXiang Lin et al. Β· π Poster Session 2, #2402 Β· abstract - Clover: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation β Fully-NVFP4 LLM pretraining with an unbiased micro-scaled quantizer (>2Γ lower error than stochastic rounding).
Andrei Panferov et al. Β· π Poster Session 3, #2901 Β· abstract - ECO: Quantized Training without Full-Precision Master Weights β Quantized training without high-precision master weights: injects quantization error into optimizer momentum (error-feedback).
Mahdi Nikdan et al. Β· π Poster Session 6, #3410 Β· abstract - TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier Control β End-to-end NVFP4 LLM training with oscillation suppression and outlier control, halving the accuracy gap to BF16.
Yuxiang Chen et al. Β· π Poster Session 3, #2809 Β· abstract - WinQ: Accelerating Quantization-Aware Training of Large Language Models around Saddle Points β Accelerates quantization-aware LLM training by escaping saddle points via weight interpolation and noise injection.
Dongyue Li et al. Β· π Poster Session 5, #1703 Β· abstract
Communication-efficient optimizers
- AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping β Adaptive per-tensor gradient clipping that stabilizes LLM pretraining and lowers communication vs global clipping.
Guoxia Wang et al. Β· π Poster Session 3, #2311 Β· abstract - BAS: Bridging Adam and SignSGD for Memory-Efficient LLM Training β Block-Adaptive Signum bridging Adam and SignSGD: low optimizer-state memory plus a communication-efficient variant.
Yijie Zhou et al. Β· π Poster Session 7, #3910 Β· abstract - Softsignum: Smooth Your Signum For Better Heterogeneity Handling β A smooth signβSGD transition that handles parameter heterogeneity, with convergence guarantees incl. LLM pretraining.
Dmitrii Feoktistov et al. Β· π Poster Session 4, #3708 Β· abstract
Communication-efficient MoE parallelism
- Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism β MoE parallelism with O(1), fully-balanced, deterministic communication regardless of the number of activated experts.
Chenwei Cui et al. Β· π Poster Session 4, #2010 Β· abstract
Disaggregated & distributed serving
- EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference β Cross-layer load balancing for expert-parallel distributed MoE inference, with no expert-device remapping.
Yize Wu et al. Β· π Poster Session 7, #1912 Β· abstract - Efficient Multi-round LLM Inference over Disaggregated Serving β Disaggregated serving tuned for multi-round / agentic workloads via adaptive prefill placement and scheduling.
Wenhao He et al. Β· π Poster Session 3, #3706 Β· abstract - HexGen-3: A Fully Disaggregated LLM Serving Framework with Fine-Grained Heterogeneous Resource Autoscaling β Fully disaggregated LLM serving with fine-grained heterogeneous-resource autoscaling.
Youhe Jiang et al. Β· π Poster Session 2, #3707 Β· abstract - Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM Serving β Prefill-capable decode nodes reuse cached KV for multi-turn serving, cutting KV-transfer congestion.
Zongze Li et al. Β· π Poster Session 5, #4003 Β· abstract - Semantic Cache Distillation: Efficient State Transfer via Reuse and Selective Patching β Replaces raw KV-cache transfer in disaggregated serving with compact semantic codes (up to 2.6Γ less transfer).
Qianli Ma et al. Β· π Poster Session 2, #1906 Β· abstract
Edgeβcloud collaborative inference
- Bridging On-Device and Cloud LLMs for Collaborative Reasoning: A Unified Methodology for Local Routing and Post-Training β An on-device LLM learns via RL when to offload a query to a cloud LLM for collaborative reasoning.
Wenzhi Fang et al. Β· π Poster Session 5, #2613 Β· abstract - HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference β Edgeβcloud LLM inference with dependency-aware subtask routing for parallel, budget-adaptive execution.
Jiangwen Dong et al. Β· π Poster Session 1, #3816 Β· abstract
Low-bandwidth multi-device inference
- ASTRA: Communication-Efficient Acceleration for Multi-Device Transformer Inference β Communication-efficient multi-device transformer inference at bandwidths as low as 10 Mbps via quantized token transfer.
Xiao Liu et al. Β· π poster TBD Β· abstract - CoCoQuant: Breaking the Bandwidth Wall via Co-Optimized Communication and Computation Quantization β Co-optimizes communication and computation quantization to break the bandwidth wall in distributed LLM inference.
Haojie Duanmu et al. Β· π Poster Session 8, #2701 Β· abstract
- Position: Enabling Fair Revenue Sharing for Data Providers in GenAI Systems β Calls for decentralized, contribution-aware revenue sharing with data providers in GenAI systems.
Gengrui (Edward) Zhang Β· π Poster Session 6, #3308 Β· abstract - Position: Federated Learning is a Lens towards a Democratized Future for the Scaling Law Era β Argues federated learning still points toward a decentralized, democratized ML future in the scaling-law era.
Harry Jiang et al. Β· π Poster Session 1, #3308 Β· abstract - Position: Zeroth-Order Optimization in Deep Learning Is Underexplored, Not Underpowered β Argues zeroth-order optimization is underexplored; its forward-only updates suit communication-efficient, pipeline-friendly training.
Sijia Liu et al. Β· π Poster Session 6, #3705 Β· abstract
Contributions welcome!
- Is your ICML 2026 paper missing? If it fits the scope above but isn't listed, please open a PR to add it β this list was curated semi-automatically (with Claude) from the accepted-paper metadata and may well have missed papers. No need to be shy: if it belongs here, we'd like it here.
- Corrections to descriptions, categories, poster sessions, or board numbers are equally welcome.
- Format: a one-line entry in
README.md(- [Title](OpenReview link) β description. <br><sub>First Author et al. Β· π poster location Β· [abstract](papers.md#anchor)</sub>) plus a matching section inpapers.mdwith the full author list and abstract.
