Feat/task21 hybrid async multimodal perf - #183
Closed
leelingrui wants to merge 1 commit into
Closed
leelingrui wants to merge 1 commit into
leelingrui wants to merge 1 commit into
Conversation
# ⭐ Feature ## Push actor->rollout weights via DCS in hybrid mode - Replace hybrid mode's CUDA-IPC UpdateWeightFromTensor push with the same DCS (NCCL/GLOO device-direct broadcast) checkpoint_engine_client path pure fully-async already uses, unifying the two weight-sync mechanisms - Fix 5 real-hardware-only bugs found validating this on actual Megatron-Bridge multimodal weights on GPU: stale weight_updater init-time call under DCS, feeding CPU-pinned tensors straight into an NCCL/GLOO broadcast that requires device tensors, and Megatron's dynamically-attached TP attributes (tensor_model_parallel/ partition_dim/partition_stride/parallel_mode) being silently dropped across two separate tensor-reallocation points (TensorBackuper's CPU-snapshot torch.empty_like() and the H2D tensor.to(device) copy) - Measured ~2% slower step_time than the CUDA-IPC push on 2xH20 — kept for architectural unification with fully-async, not as a perf win ## Make --num-iters-per-train-update actually control train_hybrid - train_hybrid's forward-phase sub-batching was already driven by --num-steps-per-rollout (via build_rollout_minibatch_plan), not by --num-iters-per-train-update as the docs described; the flag was a silent no-op in hybrid mode and the merged training step always ran as a single optimizer update regardless of its value - Re-chunk the fully merged batch into num_iters_per_train_update equal pieces before building the data iterator, so the existing get_data_iterator/train() multi-step machinery (ROLLOUT_MINI_LOCAL_SAMPLE_COUNTS_KEY) runs that many separate optimizer steps instead of one, independent of however many sub-batches the forward phase used - Default num_iters_per_train_update=1 reproduces prior behavior exactly; measured as a ~5.5% step_time regression on 2xH20 when set above 1, so not recommended as a default ## Add --hybrid-async-weight-sync (opt-in, default off) - Push the DCS weight sync from a background thread (double-buffered CPU snapshot, ThreadPoolExecutor), joined at the top of the next iteration's push instead of blocking on it immediately - Validated correct on real 2xH20 hardware (no deadlocks/races, correct weight updates every rollout) but measured no faster than the synchronous push — the DCS push's H2D copy + TP all-gather + NCCL broadcast still run on the actor's own GPU and contend with the next step's compute either way. Kept as an opt-in switch rather than deleted, since larger-scale/slower-interconnect setups may see a different tradeoff - Automatically falls back to sync when --disable-weights-backuper or --keep-old-actor is set ## Add --dedup-multimodal-preprocess and --pin-multimodal-h2d-copy (opt-in, default off) - Rollout-side: when every sample in a prompt group shares the same multimodal_inputs object, run the image processor once per group instead of once per sample, caching the result as _pre_encoded_image - Training-side: pin multimodal_train_inputs CPU tensors in get_batch() so the H2D copy in move_tensors_to_device can go non_blocking instead of a synchronous pageable-memory copy every microbatch - Both validated correct on single-GPU colocate (20/20 rollout_ids, zero tracebacks, non-degenerate reward/loss) but without a quantified before/after timing comparison yet, hence off by default --- # 📝 Documentation ## Record real-GPU validation results on 2xH20 - exps/hybrid_async_perf_h20/README.md: full history of every attempt on this perf investigation — the reverted ThreadPoolExecutor CUDA-IPC overlap, the DCS weight-sync swap and its 5 bugs, the num-iters-per-train-update split (both sync- and async-push variants), with root-cause analysis for why each measured as a regression on this 2xH20 setup rather than the expected improvement
leelingrui
requested review from
Aurelius84,
NINGBENZHE and
Yangruipis
as code owners
July 30, 2026 04:56
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
概述
本 PR 针对 GitHub issue #21「Hybrid-Async 多模态性能」,在 2×H20(Qwen3-VL-4B,
--hybrid模式)上完成了权重同步/训练流水线的性能验证,修复了一个文档描述与实际代码行为不一致的 bug,并把此前验证过的一项性能尝试(DCS push 异步化)保留为默认关闭的开关,而不是直接删除。背景
Hybrid 模式下,actor 与 rollout 使用独立 GPU,actor 内部通过
TensorBackuper+_switch_model在actor/ref/teacher等 tag 间切换权重完成 forward,训练完成后需要把最新权重推送给 rollout。docs/en/guide/hybrid-training.md的 "Next Steps" 记录了两个待办:UpdateWeightFromTensor推送--num-iters-per-train-update真正把训练步拆成多次,和 TransferQueue 数据消费形成流水线本 PR 是对这两项的完整验证与收尾。
做了什么
1. 多模态预处理去重 + pin memory 异步 H2D 拷贝(新增开关,默认关闭)
--dedup-multimodal-preprocess,默认False):开启后,rollout 侧同一个 prompt group 内的n_samples_per_prompt份样本共享同一张图片时,只跑一次_run_image_processor(),结果缓存在_pre_encoded_image/_pre_encoded_image_elapsed上给其余样本复用,省掉n_samples_per_prompt - 1份重复的 HF processor 调用(relax/engine/rollout/sglang_rollout.py)。--pin-multimodal-h2d-copy,默认False):开启后,get_batch()组装multimodal_train_inputs时对 CPU tensor 做pin_memory(),配合move_tensors_to_device()里恒定的non_blocking=True拷贝,避免forward_step()每个 microbatch 都同步阻塞等 pixel tensor 从 pageable memory 搬到 GPU(relax/backends/megatron/data.py)。关闭时non_blocking=True对未 pin 的 tensor 是安全的 no-op,等价于原来的同步拷贝。2. DCS 权重同步
把 hybrid 模式的权重推送从 CUDA-IPC 换成和纯 fully-async 共用的 DCS(
checkpoint_engine_client.update_weights_for_rollout)广播机制。3. 在hybrid模式下实现
--num-iters-per-train-update功能train_hybrid的 Phase 3现在会把完整 batch 重新切分成num_iters_per_train_update份,复用get_data_iterator/train()已有的多步训练机制,让每一份都作为独立的optimizer.step()训练,和 Phase 1 用了几个 forward 子批次完全解耦。num_iters_per_train_update=1(默认值)行为与改动前完全一致。4. DCS push 异步化
DCS push 改成异步处理 + 双缓冲 CPU snapshot, 使用--hybrid-async-weight-sync`,默认关闭
改动影响
--dedup-multimodal-preprocess、--pin-multimodal-h2d-copy,均默认False。不传这两个 flag 时,relax/engine/rollout/sglang_rollout.pyrelax/backends/megatron/data.py和这次改动之前完全一致;--hybrid-async-weight-sync默认False,不传这个 flag 的所有现有脚本/流程,权重推送仍然是同步调用,和这次改动之前一致。relax/backends/megatron/actor.py、relax/backends/megatron/data.py、relax/engine/rollout/sglang_rollout.py、relax/utils/arguments.py。--num-iters-per-train-update的行为变化:仅对--hybrid模式生效;默认值 1 时行为不变,设为 >1 时会真正拆分训练步。tests/backends/megatron/test_data_vpp.py、tests/utils/test_arguments_opd_teacher_colocate.py、tests/engine/rollout/test_sglang_rollout_diagnostics.py等(20/20 通过)和 2×H20 30 步测试上验证,详见exps/hybrid_async_perf_h20/README.md。测试
scripts/training/multimodal/run-qwen3-vl-4B-2xgpu-openr1mm-hybrid-async.sh
在一台 H20 * 2 主机上运行








baseline1
baseline2
optimized1
optimized2