Conversation
|
这个正确性验证是如何做的?有跑 >150 个 step 比对 loss、grad norm、reward、mismatch 吗? |
|
另外, |
|
@xiaoliang0601 感谢review,以下是两个问题的回答~
端到端逐 step 精确对拍我试过,但是这条路径 run-to-run 本身就不确定(FLA 的 gated delta rule 反向用 atomics,加上 NCCL 规约顺序不固定),同一份配置、同一个 seed 跑两遍,loss 就能差 1.5e-4 ~ 5.3e-4,grad_norm 相对差能到 6% ~ 58%;(主要是grad_norm抖动大,单纯比loss和有效token的话,本 PR chunkwise 对比 Stage 0 老镜像老代码headwise 220步中: loss 逐步绝对差 max 2.08e-4,mean 6.29e-5,中位 5.81e-5,逐步有效token数完全相同) 我主要比较的是:step 0 的 loss 逐 bit 相等。我比了四组:同配置跑两遍、chunkwise 改动前后、chunkwise vs headwise、chunkwise vs all_gather;他们step 0 的 loss 全部相同,grad_norm 相对差 0.01%。 二是单测,现在的单测会涵盖cpu端的route计算(确保我们迁移后的新的route计算和老版本一致),以及gpu端的一次实际的cp chunkwise forward正确性结果检查。
现在这条路径上是不会被静默修改的;Relax 和 MCore 里所有对 刚刚提交了新的commit,新加了一些守卫是:身份之外再比一下 autograd 的版本计数器 这样实现不知道可以不,或者你有什么建议吗? |
绝对值差多少?如果绝对值太小的话,grad_norm 的相对值意义并不大。 我希望你能贴一个 reward、loss 和 grad norm 的 A/A 和 A/B 对比的曲线(可以用绝对值,也可以取 log),这样直观地判断训练是否有问题。 |
|
如果 A/B 能落到 A/A 的误差范围内的话,我觉得是没问题的,不需要 bitwise 对齐。 |
|
好,我没什么问题了。 |
Pre-commit was not run before pushing the previous commit, so CI's ruff-format and docformatter hooks failed on the new test file. Formatting only -- no test logic changed. Co-authored-by: Cursor <cursoragent@cursor.com>
# ⚡ Performance - Adapt the rebased Task 32 work to current main using reviewer PR redai-studio#348 (b2cd009). - Preserve the pinned dependencies, chunkwise default and native selective GDN recompute while keeping all_gather as an explicit fallback. - Keep route and boundary caches tied to final packed metadata and validate tensor versions, inference tensors and runtime CP geometry. # 🐛 Bug Fix - Move pure GDN mode validation into a standard-library-only module so model providers do not import the optimizer and training argument stack. - Retain provider-level packing validation without breaking the existing VPP dependency-isolation tests. # ✅ Tests - Port the current CPU and CUDA/NCCL regression suites from redai-studio#348. - Pass 159 CPU regressions against exact pinned MCore sources, including all 17 VPP cases that fail on the reference branch. - Pass all pre-commit checks and verify the complete cumulative patch applies to the pinned Bridge/MCore assembly using Dockerfile's patch command. References: redai-studio#213, redai-studio#273, redai-studio#348.
# ✅ Tests - Preserve reviewer PR redai-studio#348 production code, including the GDN validator in arguments.py and the original model-provider import. - Complete fake Megatron optimizer, parser, and tokenizer dependencies for isolated VPP tests without replacing the real GDN validator. - Restore arguments/model-provider module caches and package attributes after each test, including when those modules were initially absent. - Restore the reference stage2 test import strategy. - Pass 159 CPU regressions, 17 standalone VPP tests, and 46 combined VPP/FP16 tests with zero skips, plus real production imports and all pre-commit hooks. References: redai-studio#273, redai-studio#348.
# ✅ Tests - Restore the model-provider test loader to reviewer PR redai-studio#348 unchanged. - Limit the reference diff to 13 lines completing fake Megatron import dependencies in _install_fake_megatron. - Preserve all production code and existing test assertions from redai-studio#348. - Pass 159 CPU regressions, 17 isolated VPP tests, 46 combined VPP/FP16 tests, and all pre-commit hooks with zero skipped tests. References: redai-studio#273, redai-studio#348.
# ✅ Tests - Replace the special raising helper with the existing lambda placeholder style for fake parser, validator, and tokenizer imports. - Keep the reference diff limited to ten fake-dependency additions. - Pass 17 isolated VPP tests, 46 VPP/FP16 tests, and all pre-commit hooks.
de5a26f to
6c190fe
Compare
Nyanpasu 审查看板审查状态: 💬 已完成 · 有补充意见 审查版本: e09e8f5 已恢复真实GDN CP1/CP2 chunkwise与all_gather输出、输入梯度和参数梯度回归,F3覆盖缺失已解决;静态核查无新增发现。F1已解决,正文F2仍待处理。本机无CUDA,新GPU测试收集后跳过;当前head尚无CI结果,GPU数值和H20耗时未验证。
Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.
|
rai-studio-bot
left a comment
There was a problem hiding this comment.
需要先修复行级意见中的 NPU GDN CP 初始化兼容性回归。
CPU 隔离检查通过 117 项,模型入口测试通过 17 项;随机路由的输出、peer splits 与梯度检查通过。当前 head 无 CI 结果;本机缺少 CUDA/NPU 硬件,未重跑 GPU/NCCL、NPU 或多节点训练验证,完整 MCore 配置/CLI 初始化也未运行。
优先级:P3 非行级:PR 正文范围与当前 diff 不符。请更新 §2/改动清单中的“relax/ 零改动”、旧补丁路径,以及仅按对象身份缓存的描述;当前实际包含 Relax 配置与分派修改,并检查版本计数。也请标明性能及 GPU 验证对应的提交/依赖版本,区分历史结果与本次 MCore 适配后的验证。
| or torch.distributed.get_rank(group=torch.distributed.group.WORLD) == 0 | ||
| ): | ||
| logger.info( | ||
| f"[GDN CP] role={role} linear_cp_mode={model_config.linear_cp_mode} " |
There was a problem hiding this comment.
请保留 NPU 固定依赖栈的兼容性。docker/Dockerfile.npu 固定 MCore core_v0.16.1 / MindSpeed-Bridge v0.3.1,没有应用这里新增的 GPU MCore 补丁;旧 TransformerConfig 和 Qwen3.5 provider 均没有 linear_cp_mode。仓库现有 scripts/training/text/run-qwen35-9B-16xnpu-cp.sh 使用 CP=4,其 provider 的 experimental_attention_variant 是 gated_delta_net,因此 rank 0 会进入这里,并在构造日志字符串时抛出 AttributeError,尚未开始训练就失败。对旧配置字段执行这段初始化代码也能复现异常。
请对该日志做字段兼容处理,并按后端能力隔离依赖新版 GDN 接口的分派,增加“不含 linear_cp_mode 的 NPU GDN 配置 + CP>1”的初始化回归覆盖。
There was a problem hiding this comment.
已复核 41d5296:这里改用 getattr 后,旧配置不再因日志读取缺失字段而失败。对实际初始化分支检查了旧配置/三种模式 × rank 0/非 rank 0,共 8 组均通过。现有 NPU Qwen3.5 provider 使用独立的 MindSpeed GDN 类,不进入新版 MCore 的 forward 分派;原 P1 标记为已解决。此验证不替代真实 NPU 训练。
# 🐛 Bug Fix - Read linear_cp_mode with getattr when logging GDN CP initialization so legacy NPU model configurations without this field do not raise. - Preserve existing forward dispatch and report None for an absent field. # ✅ Tests - Check the actual log expression with a missing field and all three existing modes on CPU; existing mode output is unchanged. - Verify the repository Qwen3.5 NPU CP4 provider uses its own Bridge GDN class. - Run all pre-commit hooks. References: redai-studio#273.
|
@xiaoliang0601 已按建议 rebase 并参考 #348 更新了 #273。当前生产代码和 MCore patch 与 #348 完全一致,唯一额外修改是 tests/backends/megatron/test_model_provider_vpp.py:主要是补齐假 Megatron 依赖。 这边跑了几组端到端SFT实验验证,配置是Qwen3.5-9B、OpenMathReasoning-mini、8×H200,覆盖静态 CP4、动态 CP1/2/4,以及 chunkwise、headwise、all_gather 和 main 基线。 排除前 20 步预热后,chunkwise 的 actor 训练吞吐如下;
正确性方面,静态 chunkwise A/A 出现过一次梯度尖峰,原始动态 headwise 出现过两次;其余全部对齐;不知道这是否正常,我将对比图片都放在这里:
|
|
CI问题我看是因为没加载当前 PR 对应的 patched MCore 源码镜像;我这边本地测试是没问题的 |
@rai-studio-bot 定位下 GPU Unit Tests / Unit Tests 时间增加的原因 |
rai-studio-bot
left a comment
There was a problem hiding this comment.
回复 @SigureMo:已用两次运行上传的 junit.xml 对齐到具体用例。增加的约 21 分钟几乎全部来自本 PR 新增的 11 项 GDN GPU 测试,原有测试没有整体变慢。
| 口径 | main b7ad7b49 |
PR 5144565 |
|---|---|---|
| pytest 总耗时 | 385.83 s(6:26) | 1651.81 s(27:32) |
| 两边共有用例的 JUnit 耗时合计 | 346.52 s | 347.10 s |
| PR 新增 137 项用例 | — | 1265.75 s |
其中 test_gdn_chunkwise_cp_gpu.py 的 11 项 |
— | 1264.31 s(21:04) |
数据来源:main 基线运行、本次 PR 运行。pytest 总时间包含收集等开销,所以不等于 testcase 时间之和。本次镜像准备仅 27 秒、GPU 分配在同一秒完成,均不在上述 pytest 时间内;本次测试最终 2861 passed / 30 skipped,没有超时失败。
新增文件里最慢的五项如下,合计占该文件耗时约 85%:
| 用例 | 秒 |
|---|---|
| TP2/CP2 vs CP1 | 282.91 |
| GDN 模块对齐 fp32 | 280.34 |
| GDN 模块对齐 bf16 | 206.81 |
| FLA kernel 对齐 fp32 | 172.52 |
| FLA kernel 对齐 bf16 | 130.39 |
测试实现对 11 个 case 分别调用 mp.spawn,总共启动 24 个 worker(10×2 + 1×4),重复导入依赖、初始化 CUDA/NCCL/模型。最慢几项还执行真实 FLA 前反向,并对比 CP1、headwise、chunkwise,涉及不同 dtype、head 数与 TP 配置;这比 CPU route/cache 单测重得多。新增其余 126 项合计只有约 1.44 秒。
可以确定的是耗时集中在哪些测试;尚不能确定这 21 分钟内编译、autotune、初始化、collective 各占多少。 当前日志没有这些阶段的计时。固定的 FLA 0.4.2 在 Triton 支持时默认允许 autotune 磁盘缓存,不能说每次 spawn 都重新编译全部内核;但当前 CI 容器只挂载 uv 缓存,没有显式跨运行保存 Triton/Inductor 缓存。
建议先对这五项增加 worker 内的阶段计时(导入/分布式初始化/模型构建/首次及再次 forward-backward,GPU 阶段同步后计时),做同配置冷/热对照。随后优先尝试让相同 world size 的用例复用 worker,保留完整正确性断言;若确认首用内核开销占主导,再评估按依赖版本和 GPU 架构隔离的编译缓存。需要控制每次 PR 的测试预算时,可以保留代表性的必跑 smoke,将完整 dtype × TP/CP × recompute 矩阵放到独立完整验证任务,不建议简单删除覆盖。
本轮只做日志、JUnit 与源码分析,未修改测试或在本地重跑 GPU profiling。
新增测试增加这么长时间我觉得也是不合理的,需要看看能不能降低测试成本呢? |
嗯嗯好的,我这边看看怎么改一下 |
# ✅ Tests - Keep one real NCCL THD/SBHD layout round-trip test and remove the ten expensive full GPU validation cases and their dedicated helpers. - Merge CPU partition, route, and cache coverage into test_gdn_cp_layout. - Group configuration and dispatch coverage in test_gdn_cp_mode, removing historical stage names while preserving all 126 core CPU cases. - Pass 159 related CPU regressions, verify the GPU file collects one test, and run all pre-commit hooks. No new GPU execution is claimed.
rai-studio-bot
left a comment
There was a problem hiding this comment.
测试重组后的 CPU 隔离检查通过 117 项;生产代码未变,原 P1 保持已解决。本轮新增一项非阻塞测试覆盖意见,见行级讨论;正文更新建议仍待处理。
当前 GPU 测试仅保留布局往返,不能将测试删除后的提速视为 GDN 前反向验证。新 head 尚无 CI 结果,本机缺少 CUDA,未重跑 GPU/NCCL 或多节点验证。
| def test_layout_round_trip_over_real_cp_group(): | ||
| """RFC 5.1 at the collective level: the layout swap is a pure | ||
| permutation.""" | ||
| _spawn(_worker_layout_round_trip, "n/a", 29544) |
There was a problem hiding this comment.
请保留至少一项真实 GDN 前反向集成回归,或将原数值矩阵移到默认 CI 之外的独立验证入口。本次删掉 test_gdn_cp_matches_cp1、recompute 和 TP2/CP2 等检查后,唯一 GPU 测试只调用布局转换,输入也没有 requires_grad;CPU mode 测试则把 GatedDeltaNet.forward 替换成 spy。现存 test_gdn_cp_reassembly 使用 matmul/cumsum 的 toy scan,覆盖的是 all-gather,不能替代这里的 chunkwise GDN。
因此,即使 GDN 的 route 接入、CP context 或重计算反向接线出错,这些保留下来的测试仍可能全部通过,而这些正是本 PR 修改的生产路径。可以将单一 dtype、CP=2 对 CP=1 的输出/输入梯度/参数梯度检查压缩为一次 spawn,并把较重矩阵保留为可单独运行的测试;无需恢复每次 CI 都执行完整 21 分钟矩阵。
There was a problem hiding this comment.
e09e8f5f 已解决这里的覆盖缺口:新的 test_gdn_cp_gpu.py 在一次双 GPU spawn 中,用相同权重比较 CP=1 与 CP=2 chunkwise / all_gather 的输出、输入梯度和逐参数梯度;参数梯度在 CP 组内求和后与基准比较,all_gather 同时执行 checkpoint 前反向,原布局往返检查也保留。
本轮核对了实际 GDN 调用、动态 CP 组和梯度汇总路径,未发现新增问题。本机无 CUDA,测试收集成功后按条件跳过;“覆盖已恢复”不代表已验证 GPU 数值通过或 H20 耗时低于 4 分钟,这两项仍需当前 head 的 GPU CI 确认。
|
@SigureMo hi,我这边把大部分的gpu测试都删掉了,只保留了布局转换的,这个只用时10s左右 |
|
CI 里 Python3.11 的报错问题还需要看一下,可以在 CPU 环境测一下 |
|
@rai-studio-bot 分析一下 |
rai-studio-bot
left a comment
There was a problem hiding this comment.
回复 @SigureMo:不是本 PR 独有,也不只发生在 Python 3.11。 这次补查了其他分支的实际失败日志,找到两个相同断言的反例:
| PR / 运行 | Python | 失败位置 |
|---|---|---|
| #323,9 月 19 日 | 3.11 | [prefetch],第 980 行,警告捕获为空 |
| #257,9 月 21 日 | 3.10 | 同上 |
| 本 PR,当前 head 第 3 次运行 | 3.10 | 同上;3.11、3.12 通过 |
而同一个 e9304afc 的第 2 次运行,三个 Python 版本全部通过。因此“本 PR 连续出现”确实值得检查,但不能据此归为本 PR 特有或版本特有问题。
根因仍是测试的日志捕获时序。 测试第 963–974 行先 ds.shuffle(0),后 monkeypatch logger.warning;shuffle 已经启动预取线程。上述失败日志都实际打印了缺失图片的警告,且补齐样本的断言已通过,只是后台线程可能在安装捕获前发出警告。相关测试、SFT 实现与预取器相对 main 均未改变。
本机 Python 3.10 的隔离受控复现中,让预取线程先完成,原测试 5/5 失败;将捕获安装移到 shuffle 之前,同样调度下 5/5 通过。这是对竞态机制的验证,不是自然失败率测量,也未运行完整 CI 环境。
统计边界:检查了最近 100 条 CI 运行记录中最新结果为失败的 23 条日志,其中 5 条命中这一断言(本 PR 3 条、其他 PR 2 条),没有穷尽历史重跑;CI 还使用 -x,其他失败可能提前终止。因此目前能确认“并非本 PR 独有”,不能量化本 PR 是否更易触发。建议直接把 ds.shuffle(0) 移到 monkeypatch 之后、try 内,消除捕获窗口;不建议依靠重跑或加 sleep。
@ShigureNyako 参考这个单独提一个 PR 优化下吧 |
rai-studio-bot
left a comment
There was a problem hiding this comment.
已核验 83a096ad 合入的 #358:logger.warning 捕获先于 ds.shuffle(0) 安装,且启动预取位于 try/finally 内,修复了此前确认的日志捕获竞态。使用当前测试代码,在 Python 3.10 隔离环境强制预取线程先完成,5/5 通过;这不是完整 CI 验证,当前 head 尚无 CI 结果。
相对 main 的 PR 差异与上轮一致,无新增发现;真实 GDN 前反向覆盖意见及正文更新建议仍待处理。
|
#358 合入后看起来稳定性高了不少,我去推进下合入 |
|
@jambow0320 补个 gpu 测试吧,可以小点儿,主要测真实 GDN 在 CP=2 下,输出和梯度是否与 CP=1 一致。CP=1、CP=2 chunkwise、CP=2 all gather 数值一致。按照 @SigureMo 说的那样,在 H20 CI 上尽量 < 4 分钟,不要弄太大的计算。 |
嗯嗯,现在 CI 限定 timeout 10min,如果超时(或者濒临超时),这个 PR 里可以将 10min 改为 15min |
@xiaoliang0601 可以的,我这边先调一下规模本地测试一下; |
精准测试其实是有风险的,一旦某块逻辑经常不跑,保护作用会失效的,根据之前的经验,如非必要,不引入这种机制 |
|
了解~那我看看这块怎么设置测试更好 |
# ✅ Tests - Compare real BF16 GDN CP=1 with CP=2 chunkwise and all_gather outputs, input gradients, and all parameter gradients using 128-wide heads and unequal packed sequences. - Run the existing exact THD/SBHD layout checks in the same two-GPU spawn. - Select one FLA autotune candidate only inside test workers, retaining real JIT kernels and NCCL while avoiding repeated candidate benchmarks. - Update references to the consolidated GPU test filename. ## Validation - Two H200s, empty compilation caches: 1 passed, 0 skipped; 60.224 seconds for the combined regression (61.740 seconds including pytest overhead). - Python 3.11 CPU-only collection skips the GPU case as intended. - pre-commit run --all-files --show-diff-on-failure passed.
rai-studio-bot
left a comment
There was a problem hiding this comment.
真实 GDN 前反向回归入口已恢复,原覆盖意见已在对应讨论更新。本轮仅测试变化,未发现新增问题;正文范围与验证版本的更新建议仍待处理。
当前 head 尚无 CI 结果。本机测试收集正常,但因无 CUDA 跳过,未验证 GPU/NCCL 数值或 H20 执行耗时。
|
@xiaoliang0601 已补充 GDN 的 CP=1、CP=2 chunkwise/all_gather 输出与梯度对照并合并布局测试;通过固定 FLA 配置跳过 autotune 后,在h20上CI耗时49秒通过 |
|
我没问题了,cc @SigureMo 走合入流程吧~ |




Task 32 第三阶段:迁移 MCore #5664 的 THD route 预构建,让 Chunkwise 产生正向收益
对应 RFC:redai-infra/Relax#213
第一阶段 PR(FLA 0.4.2 + MCore #3282 backport):redai-infra/Relax#251
第二阶段 PR(Relax 侧 GDN CP 静态路由接入):redai-infra/Relax#254
结论:接入后 chunkwise 从"比 headwise 慢 17%"变成静态 CP 下快 45–52%、dynamic CP 下快 23%,在测过的四种 recompute × sequence length × 静态/动态 CP 组合下都是最快的模式。同 harness A/B 显示 chunkwise 训练吞吐 23,371 → 43,397 tok/s(+85.7%),headwise 在噪声范围内不变(+1.3%)。GDN 布局转换引入的 device-host 同步从 65,280 次 / 2.69 s 降到 0,GPU 忙碌率 56.2% → 85.1%。
9.22 UPD:
参考 #348,rebase了最新的main分支,PR范围改为聚焦于复用GDN CP 路由信息,以加速CP速度
1. 要解决什么问题
headwise 和 chunkwise 的模型算力是相同的——headwise 按 head 切(每卡算全序列 × 1/cp 的头),chunkwise 按时间切(每卡算 1/cp 序列 × 全部头),总 FLOPs 一样,两者都不重复计算;只有 all_gather 每卡跑全序列 × 全部头,重复 CP 倍。trace 直接证实了这点:两种模式的 GEMM kernel 名字、调用次数、耗时全都对得上(例如
nvjet_tss_128x256_64x4_2x1_v_badd_coopA_NTN是 630.8 ms × 963 对 631.4 ms × 963)。然而 Stage 2(#254)里朴素实现的 chunkwise,性能既不如 all_gather 也不如 headwise。
既然算力相同,chunkwise 慢就只能是开销。把 GPU kernel 时间拆成 NCCL 与非 NCCL 两部分看(同为
--profile-with-stack抓取,rank 0):通过 trace 分析,热点集中在
_zigzag_contiguous_thd_swap。每做一次 zigzag↔contiguous 转换,它都要现场推导 all-to-all 的路由:2 × cp_size次get_thd_context_parallel_rank_indices,每次内部有cu[0].item()、cu[-1].item()和两次torch.any(...)判断 → CP=4 时 32 次同步;cp_size次nonzero()和cp_size次布尔索引 —— 输出形状依赖数据,必然同步;arange/bucketize/argsort/scatter的小 kernel,规模是O(T_global)。而这套路由只依赖
cu_seqlens和(cp_size, cp_rank):一个 micro-batch 内所有 GDN 层、两个转换方向、以及 full recompute 的重放,算出来的结果完全一样。Stage 1 backport 的代码里本来就留着这个 TODO:NVIDIA/Megatron-LM#5664 是一个比较大的 PR 改动,其中包含这个路由缓存的优化。本 PR 只最小化迁移这部分的相关代码。
2. 实施方案
从上游 MCore PR 迁移代码,改动只落在 pinned MCore patch 的两个文件,
relax/目录零改动。2.1 核心思路(来自 #5664)
两种 layout 都是全局 token 区间的并集——contiguous 每 rank 一段连续区间,zigzag 每 rank
2 × 序列数段。所以整条 all-to-all 路由可以用区间求交得到,不需要逐 token 的索引张量:把cu_seqlens一次性.tolist()到 CPU,之后全部用 Python int 做双指针求交,产出send_rows/recv_rows/input_split_sizes/output_split_sizes。相比于之前的朴素实现(逐 token 计算索引),复杂度从
O(cp_size × T_global)(GPU)降到O(cp_size × 序列条数)(CPU)。与此同时把这四个对象存进 route 缓存,之后 forward 时只需取出直接做 all-to-all,不必每层重算一遍 CPU/GPU 上的索引:
2.2 原封不动迁移过来的函数(6 个)
_cp_layout_nvtx_range_compact_thd_cu_seqlens_to_listcu_seqlens一次性.tolist()+ 去掉重复边界(padding 空槽)_append_rangerows.extend(range(start, start + length))_row_list_is_identity_pack_thd_cp_route_send_bufferindex_select_scatter_thd_cp_route_recv_bufferindex_copy_2.3 只改了命名和报错文案,核心逻辑不修改(4 个)
_validate_thd_route_partitioning_build_thd_layout_segmentscp_partition_mode→layout;加了 docstring_intersect_thd_layout_segments_thd_cp_partition_route_attr_name为什么改名:#5664 在它自己的重构里把
layout全局重命名成了cp_partition_mode。跟着改会污染 Stage 1 已经暴露出去的接口(get_thd_context_parallel_rank_indices(..., layout=...)以及已合入的单测),所以保留 Stage 1 的命名。2.4 迁移过来但做了一定改动(3 个)
build_thd_cp_partition_route求交主体(区间构建 → 双指针求交 → 展开行号 → 两个完整性 assert)与上游逐字节相同,nvtx range 也保留。唯一实质差异是返回类型:
torch.Tensor,每次转换再decode_thd_cp_partition_route()解开;ThdCpPartitionRoute。原因:上游那层编码是为了让 route 能当 CUDA graph 的捕获输入,而 decode 里有 3 处
.cpu(),有一定性能影响,而且 Megatron 本身对 THD 布局转换是拒绝 full-iteration CUDA graph 的。兼容性:
ThdCpPartitionRoute的前 6 个字段就是上游decode_thd_cp_partition_route()的返回元组,连"恒等用None表示"这个约定都一样(上游_encode把恒等的 payload 存成空列表,decode 出来就是None)。这条偏离已写进函数 docstring,并注明:若将来 GDN 需要在 graph capture 下运行,这里必须重新考虑。get_thd_cp_partition_route/prebuild_thd_cp_partition_routes接口与上游一致,函数体不同。上游假定 route 由数据侧(
get_batch)主动 prebuild,所以查找是一次裸getattr,build-on-miss 只是兼容后路并会warnings.warn(FutureWarning)。我们这边 build-on-miss 是预期路径(原因见 2.5),warn 只会变成噪声;取而代之的是每次复用前先过一道有效性校验。2.5 Relax 所必须要的特有逻辑(2 个)
当前不采用上游"在数据准备侧提前算好 route"的做法,改成了懒构建 + 缓存。
上游 #5664 的用法是:数据侧在
get_batch里调prebuild_thd_cp_partition_routes(packed_seq_params),把两个方向的 route 一次性写到PackedSeqParams上;模型里所有消费点只管getattr取用。当前 PR 考虑到两个问题暂时没有使用这套机制:
GDN.forward里(#3282 的位置),route 的消费点在模型内部而不是数据侧;preprocess_packed_seqs会在 embedding 之后重打包出新的PackedSeqParams,GDN 拿到的未必是数据侧建的那个,提前 prebuild 有可能挂在一个随后被替换掉的对象上。所以改成:第一个 GDN 层用到时才构建,缓存到本次 forward 实际收到的那个对象上,同 micro-batch 的后续层与 recompute 重放直接命中。
上游那套单写多读天然不会陈旧,我们这套懒构建 + 缓存复用则必须自己证明缓存没过期(下一个 micro-batch 换边界、dynamic CP 换
cp_size/cp_rank)。校验要拿当前请求和 route 的来源逐项比对,route 因此得自带来源信息——这就引出两个 Relax 特有的符号:ThdCpPartitionRoute(NamedTuple)decode_thd_cp_partition_route()的返回值,参与计算;后 5 个(cu_seqlens/cp_size/cp_rank/source_layout/target_layout)是 Relax 加的来源凭据,不参与任何计算_thd_cp_partition_route_is_reusablecu_seqlens用对象身份(is)而不是数值比较prebuild_thd_cp_partition_routes()也按上游接口导出了,但本 PR 未调用;等 Bridge/VLM 的最终PackedSeqParams对象归属理清后,可以直接切回上游那套数据侧 prebuild 的用法。2.6 老代码的改动面
context_parallel_layout.py里 Stage 1 已有的函数:| 函数 | 状态 |
|
_zigzag_contiguous_thd_swap| 96 → 53 行。头(cp_size==1短路、movedim、contiguous)、尾(movedim回去、contiguous)和all_to_all的调用形式全部没动,只把中间 60 行的路由推导换成"取 route + 三步执行" |2.8 一个不属于 #5664 的附带优化
route 优化之后,profile 里第一名同步点变成
_resolve_cu_seqlens(943 ms / 2,880 次)。它每层都要做cu_seqlens[-1].item()、(seq_lengths % cp_size != 0).any()和torch.equal(cu_q, cu_kv)。新增
_resolve_thd_cu_seqlens(),用与 route 相同的缓存键(四个源张量的对象身份 +seq_len_global+cp_size)把这套校验收敛到每 micro-batch 一次。影响面:对 headwise 也生效(A/B 实测 +1.3%,噪声量级);对 all_gather 不生效(Relax fallback 整个替换了 MCore 的 forward)
2.8 改动清单
docker/patch/megatron/20260805-85bced0ae.patchcontext_parallel_layout.py307 → 691 行;gated_delta_net.py的forward+ 新增_resolve_thd_cu_seqlenstests/backends/megatron/test_gdn_chunkwise_cp_route.py3. Trace 对比
同一 recipe(8×H200 / Qwen3.5-9B / TP2 / 静态 CP4 / SP / full recompute /
max-tokens-per-gpu 8192),seed 1234,用 Relax 自带的--use-pytorch-profiler --profile-target train_overall --profile-with-stack抓 step 5,取 rank 0。优化前(Stage 2)
优化后(本 PR)
3.1 同步热点
context_parallel_layout.get_thd_context_parallel_rank_indicesgated_delta_net._resolve_cu_seqlensTensor.nonzero_resolve_cu_seqlens剩下的 60 次 = 10 个 micro-batch × 2(q/kv)× 3 次检查,正好是"每 micro-batch 一次"。3.2 整步指标
gpu_memset计算 kernel 少掉的 1.49 s 是 route 构建原先在 GPU 上跑的那批索引小算子——GEMM kernel 的名字、调用次数、耗时前后完全不变,模型算力一点没动。优化后 chunkwise 的计算 kernel(4.88 s)从"比 headwise 多 0.97 s"变成"比它少 0.52 s";这 0.52 s 同样不是算力差异,而是 headwise 自己的开销:它按 packed 序列逐条做 a2a 再
torch.cat拼回,光CatArrayBatchedCopy就是 262 ms / 15,984 次,chunkwise 优化后的前 8 大 kernel 则全是 GEMM。NCCL 略升是因为 CPU 不再拖后腿,collective 发得更密,单次 kernel 里等待 peer 的时间变长。
4. 端到端性能
4.1 口径
8×H200,Qwen3.5-9B,OpenMathReasoning-mini SFT,TP2 + SP,seed 1234,35 step,取 step 10–34 共 25 step。为让每一步都是训练步,关掉了 eval。每组 25 步的 token 数完全相同(7,270,905),吞吐可直接比。
个别 step 会出现 2–4 倍的偶发 stall(TFLOPs 同步塌陷,说明是同样的工作被拉长的环境噪声)。下表统一剔除
> 1.6 ×中位数的步并给出剔除数量——最终表中所有配置的剔除数均为 0。S3 整组复跑过一次以排除噪声,两轮吻合在 2.3% 以内(chunkwise 48,087 / 46,998,headwise 31,692 / 32,306,all_gather 23,043 / 23,186),表中用的是干净的复跑。4.2 S1 — Stage 2 的配置(full recompute,
max-tokens-per-gpu 8192)4.3 S2 — 关掉 recompute(
max-tokens-per-gpu 8192)all_gather 在这个场景跑不了:它在每张卡上重放完整序列的 scan,GDN 激活按全局上下文长度增长,Relax 的
_assert_gdn_full_recompute()因此强制要求 full recompute:chunkwise 和 headwise 全程保持 1/cp 分片,所以两者在这个配置下都能跑;差别在速度——chunkwise 快 50%,而且是全部实验里绝对吞吐最高的一组(54,254 tok/s,MFU 0.381)。
4.4 S3 — 拉长 pack(full recompute,
max-tokens-per-gpu 24576,即每 micro-batch 全局 98,304 token)序列拉长后 all_gather 从"略快于 headwise"掉到"慢 28%"——它重复 CP 倍的 scan,代价随上下文长度线性增长。
4.5 S4 — dynamic CP(
--dynamic-context-parallel,max CP=4,full recompute,8192)CP 逐 micro-batch 在 {1,2,4} 中选择,本轮实际用到 CP1 × 8、CP2 × 41、CP4 × 7 个 micro-batch。
优势收窄符合预期:CP=1 的 micro-batch 根本不做布局转换,CP=2 时 headwise 的 all-to-all 也便宜得多。这一组同时验证了 route 缓存在
cp_size/cp_rank逐 micro-batch 变化下的正确失效。4.6 A/B:收益确实来自本次改动
同一 harness、同一 spec,把 MCore 的两个文件换回 Stage 1 版本(无 route 预构建)再跑一遍:
两点交叉验证:
5. 正确性
test_gdn_chunkwise_cp_route.py,10 个测试函数、参数化展开 33 项,分三组:cp_size ∈ {1,2,4,8}× 两个方向 × 三种 packed 边界(单序列、不等长多序列、含重复边界即空 padding 槽),在本地模拟全体 rank 的 all-to-all,要求 route 产出的每一行都与 Stage 1 未改动的get_thd_context_parallel_rank_indices分区完全相同,同时断言相邻 rank 的 split sizes 互相对得上。2×cp整除、cu_seqlens不从 0 开始、非单调、非法方向,新旧实现必须同样报ValueError。cu_seqlens对象重建、dynamic CP 换cp_size/cp_rank重建、prebuild填充两个方向、非 THD / CP=1 时prebuild为 no-op。test_gdn_chunkwise_cp_layout.py(51)+test_gdn_cp_mode_stage2.py(19)全过。三个文件合计 103 项全过。test_gdn_chunkwise_cp_gpu.py10 项全过(20 分 36 秒),含真实 CP 组上 zigzag→contiguous→zigzag 的 token 级往返、CP=2 vs CP=1 的 fp32/bf16 前反向对齐、state_dict/sharded_state_dict一致性。mean train/loss全部落在 0.35111–0.35119,跨模式、跨 recompute、跨 sequence length、跨静态/动态 CP 一致。6. 已知边界
dev,对应 main 的是 #6233)。本 PR 的定位是"提前迁移其核心优化并验证收益",正确性由"与 Stage 1 已合入实现逐 token 对拍"保证。prebuild_thd_cp_partition_routes()已导出,待对象流理清后可直接切换。build_thd_cp_partition_route的返回类型要换回上游的编码张量形式。ThdCpPartitionRoute、_thd_cp_partition_route_is_reusable和 GDN forward 里的粘合代码可整体删除,改用上游的 block-level 调度器 + 数据侧 prebuild。