Skip to content

Feat/task21 hybrid async multimodal perf - #183

Closed
leelingrui wants to merge 1 commit into
redai-studio:mainfrom
leelingrui:feat/task21-hybrid-async-multimodal-perf
Closed

leelingrui wants to merge 1 commit into
redai-studio:mainfrom
leelingrui:feat/task21-hybrid-async-multimodal-perf

Conversation

@leelingrui

Copy link
Copy Markdown

概述

本 PR 针对 GitHub issue #21「Hybrid-Async 多模态性能」,在 2×H20(Qwen3-VL-4B,--hybrid 模式)上完成了权重同步/训练流水线的性能验证,修复了一个文档描述与实际代码行为不一致的 bug,并把此前验证过的一项性能尝试(DCS push 异步化)保留为默认关闭的开关,而不是直接删除。

背景

Hybrid 模式下,actor 与 rollout 使用独立 GPU,actor 内部通过 TensorBackuper + _switch_model 在 actor/ref/teacher 等 tag 间切换权重完成 forward,训练完成后需要把最新权重推送给 rollout。docs/en/guide/hybrid-training.md 的 "Next Steps" 记录了两个待办:

  1. 用 DCS(NCCL/GLOO 广播)替换旧的 CUDA-IPC UpdateWeightFromTensor 推送
  2. 让 --num-iters-per-train-update 真正把训练步拆成多次,和 TransferQueue 数据消费形成流水线

本 PR 是对这两项的完整验证与收尾。

做了什么

1. 多模态预处理去重 + pin memory 异步 H2D 拷贝(新增开关,默认关闭)

  • 去重预处理(--dedup-multimodal-preprocess,默认 False):开启后,rollout 侧同一个 prompt group 内的 n_samples_per_prompt 份样本共享同一张图片时,只跑一次 _run_image_processor(),结果缓存在 _pre_encoded_image/_pre_encoded_image_elapsed 上给其余样本复用,省掉 n_samples_per_prompt - 1 份重复的 HF processor 调用(relax/engine/rollout/sglang_rollout.py)。
  • pin memory + non_blocking H2D(--pin-multimodal-h2d-copy,默认 False):开启后,get_batch() 组装 multimodal_train_inputs 时对 CPU tensor 做 pin_memory(),配合 move_tensors_to_device() 里恒定的 non_blocking=True 拷贝,避免 forward_step() 每个 microbatch 都同步阻塞等 pixel tensor 从 pageable memory 搬到 GPU(relax/backends/megatron/data.py)。关闭时 non_blocking=True 对未 pin 的 tensor 是安全的 no-op,等价于原来的同步拷贝。

2. DCS 权重同步

把 hybrid 模式的权重推送从 CUDA-IPC 换成和纯 fully-async 共用的 DCS(checkpoint_engine_client.update_weights_for_rollout)广播机制。

3. 在hybrid模式下实现 --num-iters-per-train-update 功能

train_hybrid 的 Phase 3现在会把完整 batch 重新切分成 num_iters_per_train_update 份,复用 get_data_iterator/train() 已有的多步训练机制,让每一份都作为独立的 optimizer.step() 训练,和 Phase 1 用了几个 forward 子批次完全解耦。num_iters_per_train_update=1(默认值)行为与改动前完全一致。

4. DCS push 异步化

DCS push 改成异步处理 + 双缓冲 CPU snapshot, 使用--hybrid-async-weight-sync`,默认关闭

改动影响

  • 新增参数 --dedup-multimodal-preprocess、--pin-multimodal-h2d-copy,均默认 False。不传这两个 flag 时,relax/engine/rollout/sglang_rollout.py relax/backends/megatron/data.py和这次改动之前完全一致;
  • 新增参数 --hybrid-async-weight-sync 默认 False,不传这个 flag 的所有现有脚本/流程,权重推送仍然是同步调用,和这次改动之前一致。
  • 改动范围:relax/backends/megatron/actor.py、relax/backends/megatron/data.py、relax/engine/rollout/sglang_rollout.py、relax/utils/arguments.py。
  • --num-iters-per-train-update 的行为变化:仅对 --hybrid 模式生效;默认值 1 时行为不变,设为 >1 时会真正拆分训练步。
  • 四项改动均已在 tests/backends/megatron/test_data_vpp.py、tests/utils/test_arguments_opd_teacher_colocate.py、tests/engine/rollout/test_sglang_rollout_diagnostics.py 等(20/20 通过)和 2×H20 30 步测试上验证,详见 exps/hybrid_async_perf_h20/README.md。

测试

scripts/training/multimodal/run-qwen3-vl-4B-2xgpu-openr1mm-hybrid-async.sh

在一台 H20 * 2 主机上运行
baseline1
train (2)
perf (2)
baseline2
train (3)
perf (1)
optimized1
train (4)
perf (3)
optimized2
train (5)
perf (4)

# ⭐ Feature

## Push actor->rollout weights via DCS in hybrid mode

- Replace hybrid mode's CUDA-IPC UpdateWeightFromTensor push with the
  same DCS (NCCL/GLOO device-direct broadcast) checkpoint_engine_client
  path pure fully-async already uses, unifying the two weight-sync
  mechanisms
- Fix 5 real-hardware-only bugs found validating this on actual
  Megatron-Bridge multimodal weights on GPU: stale weight_updater
  init-time call under DCS, feeding CPU-pinned tensors straight into an
  NCCL/GLOO broadcast that requires device tensors, and Megatron's
  dynamically-attached TP attributes (tensor_model_parallel/
  partition_dim/partition_stride/parallel_mode) being silently dropped
  across two separate tensor-reallocation points (TensorBackuper's
  CPU-snapshot torch.empty_like() and the H2D tensor.to(device) copy)
- Measured ~2% slower step_time than the CUDA-IPC push on 2xH20 — kept
  for architectural unification with fully-async, not as a perf win

## Make --num-iters-per-train-update actually control train_hybrid

- train_hybrid's forward-phase sub-batching was already driven by
  --num-steps-per-rollout (via build_rollout_minibatch_plan), not by
  --num-iters-per-train-update as the docs described; the flag was a
  silent no-op in hybrid mode and the merged training step always ran
  as a single optimizer update regardless of its value
- Re-chunk the fully merged batch into num_iters_per_train_update
  equal pieces before building the data iterator, so the existing
  get_data_iterator/train() multi-step machinery
  (ROLLOUT_MINI_LOCAL_SAMPLE_COUNTS_KEY) runs that many separate
  optimizer steps instead of one, independent of however many
  sub-batches the forward phase used
- Default num_iters_per_train_update=1 reproduces prior behavior exactly;
  measured as a ~5.5% step_time regression on 2xH20 when set above 1, so
  not recommended as a default

## Add --hybrid-async-weight-sync (opt-in, default off)

- Push the DCS weight sync from a background thread (double-buffered
  CPU snapshot, ThreadPoolExecutor), joined at the top of the next
  iteration's push instead of blocking on it immediately
- Validated correct on real 2xH20 hardware (no deadlocks/races, correct
  weight updates every rollout) but measured no faster than the
  synchronous push — the DCS push's H2D copy + TP all-gather + NCCL
  broadcast still run on the actor's own GPU and contend with the next
  step's compute either way. Kept as an opt-in switch rather than
  deleted, since larger-scale/slower-interconnect setups may see a
  different tradeoff
- Automatically falls back to sync when --disable-weights-backuper or
  --keep-old-actor is set

## Add --dedup-multimodal-preprocess and --pin-multimodal-h2d-copy (opt-in, default off)

- Rollout-side: when every sample in a prompt group shares the same
  multimodal_inputs object, run the image processor once per group
  instead of once per sample, caching the result as _pre_encoded_image
- Training-side: pin multimodal_train_inputs CPU tensors in get_batch()
  so the H2D copy in move_tensors_to_device can go non_blocking instead
  of a synchronous pageable-memory copy every microbatch
- Both validated correct on single-GPU colocate (20/20 rollout_ids, zero
  tracebacks, non-degenerate reward/loss) but without a quantified
  before/after timing comparison yet, hence off by default

---

# 📝 Documentation

## Record real-GPU validation results on 2xH20

- exps/hybrid_async_perf_h20/README.md: full history of every attempt on
  this perf investigation — the reverted ThreadPoolExecutor CUDA-IPC
  overlap, the DCS weight-sync swap and its 5 bugs, the
  num-iters-per-train-update split (both sync- and async-push variants),
  with root-cause analysis for why each measured as a regression on this
  2xH20 setup rather than the expected improvement
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant