Skip to content

[Task32] feat: fla mcore chunkwise cp stage 3 - #273

Merged
SigureMo merged 14 commits into
redai-studio:mainfrom
jambow0320:task32-fla-mcore-chunkwise-cp-stage3
Sep 23, 2026
Merged

SigureMo merged 14 commits into
redai-studio:mainfrom
jambow0320:task32-fla-mcore-chunkwise-cp-stage3

Conversation

@jambow0320

@jambow0320 jambow0320 commented Aug 13, 2026 •

Copy link
Copy Markdown
Contributor

Task 32 第三阶段:迁移 MCore #5664 的 THD route 预构建,让 Chunkwise 产生正向收益

对应 RFC:redai-infra/Relax#213
第一阶段 PR(FLA 0.4.2 + MCore #3282 backport):redai-infra/Relax#251
第二阶段 PR(Relax 侧 GDN CP 静态路由接入):redai-infra/Relax#254

结论:接入后 chunkwise 从"比 headwise 慢 17%"变成静态 CP 下快 45–52%、dynamic CP 下快 23%,在测过的四种 recompute × sequence length × 静态/动态 CP 组合下都是最快的模式。同 harness A/B 显示 chunkwise 训练吞吐 23,371 → 43,397 tok/s(+85.7%),headwise 在噪声范围内不变(+1.3%)。GDN 布局转换引入的 device-host 同步从 65,280 次 / 2.69 s 降到 0,GPU 忙碌率 56.2% → 85.1%。

9.22 UPD:
参考 #348,rebase了最新的main分支,PR范围改为聚焦于复用GDN CP 路由信息,以加速CP速度


1. 要解决什么问题

headwise 和 chunkwise 的模型算力是相同的——headwise 按 head 切(每卡算全序列 × 1/cp 的头),chunkwise 按时间切(每卡算 1/cp 序列 × 全部头),总 FLOPs 一样,两者都不重复计算;只有 all_gather 每卡跑全序列 × 全部头,重复 CP 倍。trace 直接证实了这点:两种模式的 GEMM kernel 名字、调用次数、耗时全都对得上(例如 nvjet_tss_128x256_64x4_2x1_v_badd_coopA_NTN 是 630.8 ms × 963 对 631.4 ms × 963)。

然而 Stage 2(#254)里朴素实现的 chunkwise,性能既不如 all_gather 也不如 headwise。

既然算力相同,chunkwise 慢就只能是开销。把 GPU kernel 时间拆成 NCCL 与非 NCCL 两部分看(同为 --profile-with-stack 抓取,rank 0):

模式 span GPU 忙碌率 kernel 总计 NCCL 非 NCCL 计算
headwise 21.45 s 75.0% 15.66 s 10.27 s 5.40 s
chunkwise 22.73 s 56.2% 12.13 s 5.76 s 6.37 s
all_gather 12.66 s 88.3% 10.56 s 2.65 s 7.92 s

通过 trace 分析,热点集中在 _zigzag_contiguous_thd_swap。每做一次 zigzag↔contiguous 转换,它都要现场推导 all-to-all 的路由:

  • 2 × cp_size 次 get_thd_context_parallel_rank_indices,每次内部有 cu[0].item()、cu[-1].item() 和两次 torch.any(...) 判断 → CP=4 时 32 次同步;
  • cp_size 次 nonzero() 和 cp_size 次布尔索引 —— 输出形状依赖数据,必然同步;
  • 一批 arange / bucketize / argsort / scatter 的小 kernel,规模是 O(T_global)。

而这套路由只依赖 cu_seqlens 和 (cp_size, cp_rank):一个 micro-batch 内所有 GDN 层、两个转换方向、以及 full recompute 的重放,算出来的结果完全一样。Stage 1 backport 的代码里本来就留着这个 TODO:

# TODO: Let a future CP layout scheduler precompute this routing once per
# microbatch from immutable cu_seqlens and pass it through both THD swaps.
# Do not cache it across microbatches because packed sequence boundaries change.

NVIDIA/Megatron-LM#5664 是一个比较大的 PR 改动,其中包含这个路由缓存的优化。本 PR 只最小化迁移这部分的相关代码。


2. 实施方案

从上游 MCore PR 迁移代码,改动只落在 pinned MCore patch 的两个文件,relax/ 目录零改动。

2.1 核心思路(来自 #5664)

两种 layout 都是全局 token 区间的并集——contiguous 每 rank 一段连续区间,zigzag 每 rank 2 × 序列数 段。所以整条 all-to-all 路由可以用区间求交得到,不需要逐 token 的索引张量:把 cu_seqlens 一次性 .tolist() 到 CPU,之后全部用 Python int 做双指针求交,产出 send_rows / recv_rows / input_split_sizes / output_split_sizes。

相比于之前的朴素实现(逐 token 计算索引),复杂度从 O(cp_size × T_global)(GPU)降到 O(cp_size × 序列条数)(CPU)。

与此同时把这四个对象存进 route 缓存,之后 forward 时只需取出直接做 all-to-all,不必每层重算一遍 CPU/GPU 上的索引:

send_buf = _pack_thd_cp_route_send_buffer(x, route.local_source_length, route.send_rows)
recv_buf = all_to_all(cp_group, send_buf, route.output_split_sizes, route.input_split_sizes)
out      = _scatter_thd_cp_route_recv_buffer(recv_buf, route.recv_rows, out_shape)

2.2 原封不动迁移过来的函数(6 个)

符号 作用
_cp_layout_nvtx_range nvtx range 的 contextmanager
_compact_thd_cu_seqlens_to_list cu_seqlens 一次性 .tolist() + 去掉重复边界(padding 空槽)
_append_range rows.extend(range(start, start + length))
_row_list_is_identity 判断置换是否恒等
_pack_thd_cp_route_send_buffer 恒等则零拷贝,否则 index_select
_scatter_thd_cp_route_recv_buffer 恒等则零拷贝,否则 index_copy_

2.3 只改了命名和报错文案,核心逻辑不修改(4 个)

符号 具体差异
_validate_thd_route_partitioning 报错文案加了 "/contiguous" 一个词
_build_thd_layout_segments 参数 cp_partition_mode → layout;加了 docstring
_intersect_thd_layout_segments 加了 1 行 docstring
_thd_cp_partition_route_attr_name 同样的参数重命名

为什么改名:#5664 在它自己的重构里把 layout 全局重命名成了 cp_partition_mode。跟着改会污染 Stage 1 已经暴露出去的接口(get_thd_context_parallel_rank_indices(..., layout=...) 以及已合入的单测),所以保留 Stage 1 的命名。

2.4 迁移过来但做了一定改动(3 个)

build_thd_cp_partition_route

求交主体(区间构建 → 双指针求交 → 展开行号 → 两个完整性 assert)与上游逐字节相同,nvtx range 也保留。唯一实质差异是返回类型:

  • 上游把结果序列化成一个扁平 torch.Tensor,每次转换再 decode_thd_cp_partition_route() 解开;
  • 我们直接返回解码后的 ThdCpPartitionRoute。

原因:上游那层编码是为了让 route 能当 CUDA graph 的捕获输入,而 decode 里有 3 处 .cpu(),有一定性能影响,而且 Megatron 本身对 THD 布局转换是拒绝 full-iteration CUDA graph 的。

兼容性:ThdCpPartitionRoute 的前 6 个字段就是上游 decode_thd_cp_partition_route() 的返回元组,连"恒等用 None 表示"这个约定都一样(上游 _encode 把恒等的 payload 存成空列表,decode 出来就是 None)。这条偏离已写进函数 docstring,并注明:若将来 GDN 需要在 graph capture 下运行,这里必须重新考虑。

get_thd_cp_partition_route / prebuild_thd_cp_partition_routes

接口与上游一致,函数体不同。上游假定 route 由数据侧(get_batch)主动 prebuild,所以查找是一次裸 getattr,build-on-miss 只是兼容后路并会 warnings.warn(FutureWarning)。我们这边 build-on-miss 是预期路径(原因见 2.5),warn 只会变成噪声;取而代之的是每次复用前先过一道有效性校验。

2.5 Relax 所必须要的特有逻辑(2 个)

当前不采用上游"在数据准备侧提前算好 route"的做法,改成了懒构建 + 缓存。

上游 #5664 的用法是:数据侧在 get_batch 里调 prebuild_thd_cp_partition_routes(packed_seq_params),把两个方向的 route 一次性写到 PackedSeqParams 上;模型里所有消费点只管 getattr 取用。

当前 PR 考虑到两个问题暂时没有使用这套机制:

  1. 不大规模重构成 #5664 的 block-level layout 调度,布局转换仍留在 GDN.forward 里(#3282 的位置),route 的消费点在模型内部而不是数据侧;
  2. 数据侧 prebuild 在 Relax 这边不可靠——Bridge/VLM 路径的 preprocess_packed_seqs 会在 embedding 之后重打包出新的 PackedSeqParams,GDN 拿到的未必是数据侧建的那个,提前 prebuild 有可能挂在一个随后被替换掉的对象上。

所以改成:第一个 GDN 层用到时才构建,缓存到本次 forward 实际收到的那个对象上,同 micro-batch 的后续层与 recompute 重放直接命中。

上游那套单写多读天然不会陈旧,我们这套懒构建 + 缓存复用则必须自己证明缓存没过期(下一个 micro-batch 换边界、dynamic CP 换 cp_size/cp_rank)。校验要拿当前请求和 route 的来源逐项比对,route 因此得自带来源信息——这就引出两个 Relax 特有的符号:

符号 说明
ThdCpPartitionRoute(NamedTuple) 前 6 个字段 = 上游 decode_thd_cp_partition_route() 的返回值,参与计算;后 5 个(cu_seqlens / cp_size / cp_rank / source_layout / target_layout)是 Relax 加的来源凭据,不参与任何计算
_thd_cp_partition_route_is_reusable 复用前逐项核对上述凭据。其中 cu_seqlens 用对象身份(is)而不是数值比较

prebuild_thd_cp_partition_routes() 也按上游接口导出了,但本 PR 未调用;等 Bridge/VLM 的最终 PackedSeqParams 对象归属理清后,可以直接切回上游那套数据侧 prebuild 的用法。

2.6 老代码的改动面

context_parallel_layout.py 里 Stage 1 已有的函数:

| 函数 | 状态 |
| _zigzag_contiguous_thd_swap | 96 → 53 行。头(cp_size==1 短路、movedim、contiguous)、尾(movedim 回去、contiguous)和 all_to_all 的调用形式全部没动,只把中间 60 行的路由推导换成"取 route + 三步执行" |

2.8 一个不属于 #5664 的附带优化

route 优化之后,profile 里第一名同步点变成 _resolve_cu_seqlens(943 ms / 2,880 次)。它每层都要做 cu_seqlens[-1].item()、(seq_lengths % cp_size != 0).any() 和 torch.equal(cu_q, cu_kv)。

新增 _resolve_thd_cu_seqlens(),用与 route 相同的缓存键(四个源张量的对象身份 + seq_len_global + cp_size)把这套校验收敛到每 micro-batch 一次。

影响面:对 headwise 也生效(A/B 实测 +1.3%,噪声量级);对 all_gather 不生效(Relax fallback 整个替换了 MCore 的 forward)

2.8 改动清单

文件 改动
docker/patch/megatron/20260805-85bced0ae.patch context_parallel_layout.py 307 → 691 行;gated_delta_net.py 的 forward + 新增 _resolve_thd_cu_seqlens
tests/backends/megatron/test_gdn_chunkwise_cp_route.py 新增,10 个测试函数(参数化展开 33 项)

3. Trace 对比

同一 recipe(8×H200 / Qwen3.5-9B / TP2 / 静态 CP4 / SP / full recompute / max-tokens-per-gpu 8192),seed 1234,用 Relax 自带的 --use-pytorch-profiler --profile-target train_overall --profile-with-stack 抓 step 5,取 rank 0。

优化前(Stage 2)

image

优化后(本 PR)

image

3.1 同步热点

调用点 优化前 优化后
context_parallel_layout.get_thd_context_parallel_rank_indices 2,417.6 ms / 61,440 次 0
gated_delta_net._resolve_cu_seqlens 943.5 ms / 2,880 次 1.5 ms / 60 次
裸 Tensor.nonzero 267.4 ms / 3,840 次 0
全部 sync 类算子 73,771 次 871 次

_resolve_cu_seqlens 剩下的 60 次 = 10 个 micro-batch × 2(q/kv)× 3 次检查,正好是"每 micro-batch 一次"。

3.2 整步指标

指标 优化前 优化后 变化
profiled step 墙钟 22.73 s 13.36 s −41.2%
GPU 忙碌率 56.2% 85.1% +28.9 pt
kernel 总耗时 12.13 s 10.94 s −9.8%
其中计算 kernel(非 NCCL) 6.37 s 4.88 s −23.4%
gpu_memset 47.0 ms 5.2 ms −89%
NCCL 5.76 s 6.06 s +5%

计算 kernel 少掉的 1.49 s 是 route 构建原先在 GPU 上跑的那批索引小算子——GEMM kernel 的名字、调用次数、耗时前后完全不变,模型算力一点没动。优化后 chunkwise 的计算 kernel(4.88 s)从"比 headwise 多 0.97 s"变成"比它少 0.52 s";这 0.52 s 同样不是算力差异,而是 headwise 自己的开销:它按 packed 序列逐条做 a2a 再 torch.cat 拼回,光 CatArrayBatchedCopy 就是 262 ms / 15,984 次,chunkwise 优化后的前 8 大 kernel 则全是 GEMM。

NCCL 略升是因为 CPU 不再拖后腿,collective 发得更密,单次 kernel 里等待 peer 的时间变长。


4. 端到端性能

4.1 口径

8×H200,Qwen3.5-9B,OpenMathReasoning-mini SFT,TP2 + SP,seed 1234,35 step,取 step 10–34 共 25 step。为让每一步都是训练步,关掉了 eval。每组 25 步的 token 数完全相同(7,270,905),吞吐可直接比。

个别 step 会出现 2–4 倍的偶发 stall(TFLOPs 同步塌陷,说明是同样的工作被拉长的环境噪声)。下表统一剔除 > 1.6 × 中位数的步并给出剔除数量——最终表中所有配置的剔除数均为 0。S3 整组复跑过一次以排除噪声,两轮吻合在 2.3% 以内(chunkwise 48,087 / 46,998,headwise 31,692 / 32,306,all_gather 23,043 / 23,186),表中用的是干净的复跑。

4.2 S1 — Stage 2 的配置(full recompute,max-tokens-per-gpu 8192)

模式 train 时间 train 吞吐 MFU 相对 headwise mean loss
chunkwise 6.70 s 43,397 tok/s 0.304 1.524 0.35117
all_gather 9.89 s 29,416 tok/s 0.205 1.033 0.35116
headwise 10.21 s 28,483 tok/s 0.199 1.000 0.35115

4.3 S2 — 关掉 recompute(max-tokens-per-gpu 8192)

模式 train 时间 train 吞吐 MFU 相对 headwise mean loss
chunkwise 5.36 s 54,254 tok/s 0.381 1.499 0.35114
headwise 8.04 s 36,185 tok/s 0.254 1.000 0.35119
all_gather 启动即失败 — — — —

all_gather 在这个场景跑不了:它在每张卡上重放完整序列的 scan,GDN 激活按全局上下文长度增长,Relax 的 _assert_gdn_full_recompute() 因此强制要求 full recompute:

GatedDeltaNet context-parallel (cp>1) requires whole-layer activation recompute:
pass `--recompute-granularity full --recompute-method uniform --recompute-num-layers 1`.
Got recompute_granularity=None.

chunkwise 和 headwise 全程保持 1/cp 分片,所以两者在这个配置下都能跑;差别在速度——chunkwise 快 50%,而且是全部实验里绝对吞吐最高的一组(54,254 tok/s,MFU 0.381)。

4.4 S3 — 拉长 pack(full recompute,max-tokens-per-gpu 24576,即每 micro-batch 全局 98,304 token)

模式 train 时间 train 吞吐 MFU 相对 headwise mean loss
chunkwise 6.19 s 46,998 tok/s 0.329 1.455 0.35112
headwise 9.00 s 32,306 tok/s 0.225 1.000 0.35119
all_gather 12.54 s 23,186 tok/s 0.162 0.718 0.35113

序列拉长后 all_gather 从"略快于 headwise"掉到"慢 28%"——它重复 CP 倍的 scan,代价随上下文长度线性增长。

4.5 S4 — dynamic CP(--dynamic-context-parallel,max CP=4,full recompute,8192)

CP 逐 micro-batch 在 {1,2,4} 中选择,本轮实际用到 CP1 × 8、CP2 × 41、CP4 × 7 个 micro-batch。

模式 train 时间 train 吞吐 MFU 相对 headwise mean loss
chunkwise 6.23 s 46,671 tok/s 0.326 1.228 0.35113
headwise 7.65 s 38,004 tok/s 0.267 1.000 0.35114

优势收窄符合预期:CP=1 的 micro-batch 根本不做布局转换,CP=2 时 headwise 的 all-to-all 也便宜得多。这一组同时验证了 route 缓存在 cp_size/cp_rank 逐 micro-batch 变化下的正确失效。

4.6 A/B:收益确实来自本次改动

同一 harness、同一 spec,把 MCore 的两个文件换回 Stage 1 版本(无 route 预构建)再跑一遍:

配置 模式 Stage 1 本 PR 变化
S1(8192) chunkwise 23,371 tok/s 43,397 tok/s +85.7%
S1(8192) headwise 28,124 tok/s 28,483 tok/s +1.3%
S3(24576) chunkwise 34,947 tok/s 46,998 tok/s +34.5%
S3(24576) headwise 32,090 tok/s 32,306 tok/s +0.7%

两点交叉验证:

  1. Stage 1 在 S1 下 chunkwise/headwise = 0.831,与 Stage 2 §12.3 报的 19,802/23,081 = 0.858 一致,说明这套更短的 harness 复现了原结论;
  2. headwise 前后在 ±1.3% 内,收益是 chunkwise 专属的。

5. 正确性

  • CPU 单测:新增 test_gdn_chunkwise_cp_route.py,10 个测试函数、参数化展开 33 项,分三组:
    • 逐 token 等价(1 个函数,24 项):对 cp_size ∈ {1,2,4,8} × 两个方向 × 三种 packed 边界(单序列、不等长多序列、含重复边界即空 padding 槽),在本地模拟全体 rank 的 all-to-all,要求 route 产出的每一行都与 Stage 1 未改动的 get_thd_context_parallel_rank_indices 分区完全相同,同时断言相邻 rank 的 split sizes 互相对得上。
    • fail-fast 对拍(3 个):长度不能被 2×cp 整除、cu_seqlens 不从 0 开始、非单调、非法方向,新旧实现必须同样报 ValueError。
    • 缓存有效性(6 个):这是 Relax 特有、上游没有对应实现的部分,也是最需要覆盖的——同对象命中、两个方向分开缓存、换 cu_seqlens 对象重建、dynamic CP 换 cp_size/cp_rank 重建、prebuild 填充两个方向、非 THD / CP=1 时 prebuild 为 no-op。
  • Stage 1 + Stage 2 回归:test_gdn_chunkwise_cp_layout.py(51)+ test_gdn_cp_mode_stage2.py(19)全过。三个文件合计 103 项全过。
  • GPU / NCCL:test_gdn_chunkwise_cp_gpu.py 10 项全过(20 分 36 秒),含真实 CP 组上 zigzag→contiguous→zigzag 的 token 级往返、CP=2 vs CP=1 的 fp32/bf16 前反向对齐、state_dict / sharded_state_dict 一致性。
  • 训练 loss:上述 14 组 run 的 mean train/loss 全部落在 0.35111–0.35119,跨模式、跨 recompute、跨 sequence length、跨静态/动态 CP 一致。

6. 已知边界

  1. #5664 尚未合入上游(open,目标分支 dev,对应 main 的是 #6233)。本 PR 的定位是"提前迁移其核心优化并验证收益",正确性由"与 Stage 1 已合入实现逐 token 对拍"保证。
  2. route 走懒构建而非上游推荐的数据侧 prebuild,prebuild_thd_cp_partition_routes() 已导出,待对象流理清后可直接切换。
  3. CUDA graph:本 PR 不支持在 full-iteration graph capture 下运行 THD 布局转换(Megatron 本身也拒绝这个组合)。若将来需要,build_thd_cp_partition_route 的返回类型要换回上游的编码张量形式。
  4. #5664 合入后的清理路径:ThdCpPartitionRoute、_thd_cp_partition_route_is_reusable 和 GDN forward 里的粘合代码可整体删除,改用上游的 block-level 调度器 + 数据侧 prebuild。

@xiaoliang0601

Copy link
Copy Markdown
Contributor

这个正确性验证是如何做的?有跑 >150 个 step 比对 loss、grad norm、reward、mismatch 吗?

@xiaoliang0601

Copy link
Copy Markdown
Contributor

另外,cu_seqlens、_resolve_thd_cu_seqlens() 作为缓存对象,只检查是不是同一个 Tensor 对象,是不是不够?还应该检查这个 Tensor 对象 有没有被修改过?或者你能判断这个值没有被修改的可能吗?

@jambow0320

Copy link
Copy Markdown
Contributor Author

@xiaoliang0601 感谢review,以下是两个问题的回答~

这个正确性验证是如何做的?有跑 >150 个 step 比对 loss、grad norm、reward、mismatch 吗?

端到端逐 step 精确对拍我试过,但是这条路径 run-to-run 本身就不确定(FLA 的 gated delta rule 反向用 atomics,加上 NCCL 规约顺序不固定),同一份配置、同一个 seed 跑两遍,loss 就能差 1.5e-4 ~ 5.3e-4,grad_norm 相对差能到 6% ~ 58%;(主要是grad_norm抖动大,单纯比loss和有效token的话,本 PR chunkwise 对比 Stage 0 老镜像老代码headwise 220步中: loss 逐步绝对差 max 2.08e-4,mean 6.29e-5,中位 5.81e-5,逐步有效token数完全相同)

我主要比较的是:step 0 的 loss 逐 bit 相等。我比了四组:同配置跑两遍、chunkwise 改动前后、chunkwise vs headwise、chunkwise vs all_gather;他们step 0 的 loss 全部相同,grad_norm 相对差 0.01%。

二是单测,现在的单测会涵盖cpu端的route计算(确保我们迁移后的新的route计算和老版本一致),以及gpu端的一次实际的cp chunkwise forward正确性结果检查。

另外,cu_seqlens、_resolve_thd_cu_seqlens() 作为缓存对象,只检查是不是同一个 Tensor 对象,是不是不够?还应该检查这个 Tensor 对象 有没有被修改过?或者你能判断这个值没有被修改的可能吗?

现在这条路径上是不会被静默修改的;Relax 和 MCore 里所有对 cu_seqlens 的原地写全是 torch.zeros 出来立刻 cu[1:] = cumsum(...) 填一次,发生在张量进 PackedSeqParams 之前。但确实可能出现,比如 MCore 有这种写法,mamba_metadata.py:228 的 self._cu_seqlens_buffer[0] = 0 就是常驻 buffer 原地写(为了 CUDA graph 复用缓冲区),这类模式哪天进到 GDN 路径,就会静默拿上一批的 route 去重排这一批 token而且不报错。

刚刚提交了新的commit,新加了一些守卫是:身份之外再比一下 autograd 的版本计数器 tensor._version,避免这个被静默改写不重新构建route。_resolve_thd_cu_seqlens() 同样处理,键改成四个源张量的 (id, _version) 指纹。

这样实现不知道可以不,或者你有什么建议吗?

@xiaoliang0601

Copy link
Copy Markdown
Contributor

grad_norm 相对差能到 6% ~ 58%

绝对值差多少?如果绝对值太小的话,grad_norm 的相对值意义并不大。

我希望你能贴一个 reward、loss 和 grad norm 的 A/A 和 A/B 对比的曲线(可以用绝对值,也可以取 log),这样直观地判断训练是否有问题。

@xiaoliang0601

xiaoliang0601 commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

如果 A/B 能落到 A/A 的误差范围内的话,我觉得是没问题的,不需要 bitwise 对齐。

@jambow0320

Copy link
Copy Markdown
Contributor Author

grad_norm 相对差能到 6% ~ 58%

绝对值差多少?如果绝对值太小的话,grad_norm 的相对值意义并不大。

我希望你能贴一个 reward、loss 和 grad norm 的 A/A 和 A/B 对比的曲线(可以用绝对值,也可以取 log),这样直观地判断训练是否有问题。

ok,我拿之前存的日志信息画了个图,我认为基本上都是在允许误差范围内的;A是当前pr的chunkwise,B是stage0就是当前Relax main分支headwise跑出来的
image

@xiaoliang0601

Copy link
Copy Markdown
Contributor

好,我没什么问题了。

jambow0320 and others added 8 commits September 21, 2026 21:13
Pre-commit was not run before pushing the previous commit, so CI's
ruff-format and docformatter hooks failed on the new test file.
Formatting only -- no test logic changed.

Co-authored-by: Cursor <cursoragent@cursor.com>
# ⚡ Performance

- Adapt the rebased Task 32 work to current main using reviewer PR redai-studio#348
  (b2cd009).
- Preserve the pinned dependencies, chunkwise default and native selective
  GDN recompute while keeping all_gather as an explicit fallback.
- Keep route and boundary caches tied to final packed metadata and validate
  tensor versions, inference tensors and runtime CP geometry.

# 🐛 Bug Fix

- Move pure GDN mode validation into a standard-library-only module so model
  providers do not import the optimizer and training argument stack.
- Retain provider-level packing validation without breaking the existing
  VPP dependency-isolation tests.

# ✅ Tests

- Port the current CPU and CUDA/NCCL regression suites from redai-studio#348.
- Pass 159 CPU regressions against exact pinned MCore sources, including all
  17 VPP cases that fail on the reference branch.
- Pass all pre-commit checks and verify the complete cumulative patch applies
  to the pinned Bridge/MCore assembly using Dockerfile's patch command.

References: redai-studio#213, redai-studio#273,
redai-studio#348.
# ✅ Tests

- Preserve reviewer PR redai-studio#348 production code, including the GDN validator in
  arguments.py and the original model-provider import.
- Complete fake Megatron optimizer, parser, and tokenizer dependencies for
  isolated VPP tests without replacing the real GDN validator.
- Restore arguments/model-provider module caches and package attributes after
  each test, including when those modules were initially absent.
- Restore the reference stage2 test import strategy.
- Pass 159 CPU regressions, 17 standalone VPP tests, and 46 combined VPP/FP16
  tests with zero skips, plus real production imports and all pre-commit hooks.

References: redai-studio#273, redai-studio#348.
# ✅ Tests

- Restore the model-provider test loader to reviewer PR redai-studio#348 unchanged.
- Limit the reference diff to 13 lines completing fake Megatron import
  dependencies in _install_fake_megatron.
- Preserve all production code and existing test assertions from redai-studio#348.
- Pass 159 CPU regressions, 17 isolated VPP tests, 46 combined VPP/FP16
  tests, and all pre-commit hooks with zero skipped tests.

References: redai-studio#273, redai-studio#348.
# ✅ Tests

- Replace the special raising helper with the existing lambda placeholder
  style for fake parser, validator, and tokenizer imports.
- Keep the reference diff limited to ten fake-dependency additions.
- Pass 17 isolated VPP tests, 46 VPP/FP16 tests, and all pre-commit hooks.
@jambow0320
jambow0320 force-pushed the task32-fla-mcore-chunkwise-cp-stage3 branch from de5a26f to 6c190fe Compare September 22, 2026 07:39
@rai-studio-bot

rai-studio-bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Nyanpasu 审查看板

审查状态: 💬 已完成 · 有补充意见

审查版本: e09e8f5f1f481680d1d50177b4c7f3e945add59c

e09e8f5 已恢复真实GDN CP1/CP2 chunkwise与all_gather输出、输入梯度和参数梯度回归,F3覆盖缺失已解决;静态核查无新增发现。F1已解决,正文F2仍待处理。本机无CUDA,新GPU测试收集后跳过;当前head尚无CI结果,GPU数值和H20耗时未验证。

编号 问题 优先级 状态 规则来源
F1 兼容 NPU 旧配置,避免 GDN CP 初始化日志触发 AttributeError P1 ✅ 已解决 —
F2 更新 PR 正文的变更范围、缓存说明与验证版本 P3 🚧 未解决 —
F3 保留真实 GDN 前反向集成回归入口 P2 ✅ 已解决 —
Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

需要先修复行级意见中的 NPU GDN CP 初始化兼容性回归。

CPU 隔离检查通过 117 项,模型入口测试通过 17 项;随机路由的输出、peer splits 与梯度检查通过。当前 head 无 CI 结果;本机缺少 CUDA/NPU 硬件,未重跑 GPU/NCCL、NPU 或多节点训练验证,完整 MCore 配置/CLI 初始化也未运行。

  • P3 优先级:P3 非行级:PR 正文范围与当前 diff 不符。请更新 §2/改动清单中的“relax/ 零改动”、旧补丁路径,以及仅按对象身份缓存的描述;当前实际包含 Relax 配置与分派修改,并检查版本计数。也请标明性能及 GPU 验证对应的提交/依赖版本,区分历史结果与本次 MCore 适配后的验证。
Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.

Comment thread relax/backends/megatron/model.py Outdated
or torch.distributed.get_rank(group=torch.distributed.group.WORLD) == 0
):
logger.info(
f"[GDN CP] role={role} linear_cp_mode={model_config.linear_cp_mode} "

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 优先级:P1

请保留 NPU 固定依赖栈的兼容性。docker/Dockerfile.npu 固定 MCore core_v0.16.1 / MindSpeed-Bridge v0.3.1,没有应用这里新增的 GPU MCore 补丁;旧 TransformerConfig 和 Qwen3.5 provider 均没有 linear_cp_mode。仓库现有 scripts/training/text/run-qwen35-9B-16xnpu-cp.sh 使用 CP=4,其 provider 的 experimental_attention_variant 是 gated_delta_net,因此 rank 0 会进入这里,并在构造日志字符串时抛出 AttributeError,尚未开始训练就失败。对旧配置字段执行这段初始化代码也能复现异常。

请对该日志做字段兼容处理,并按后端能力隔离依赖新版 GDN 接口的分派,增加“不含 linear_cp_mode 的 NPU GDN 配置 + CP>1”的初始化回归覆盖。

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fixed

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已复核 41d5296:这里改用 getattr 后,旧配置不再因日志读取缺失字段而失败。对实际初始化分支检查了旧配置/三种模式 × rank 0/非 rank 0,共 8 组均通过。现有 NPU Qwen3.5 provider 使用独立的 MindSpeed GDN 类,不进入新版 MCore 的 forward 分派;原 P1 标记为已解决。此验证不替代真实 NPU 训练。

# 🐛 Bug Fix

- Read linear_cp_mode with getattr when logging GDN CP initialization so
  legacy NPU model configurations without this field do not raise.
- Preserve existing forward dispatch and report None for an absent field.

# ✅ Tests

- Check the actual log expression with a missing field and all three existing
  modes on CPU; existing mode output is unchanged.
- Verify the repository Qwen3.5 NPU CP4 provider uses its own Bridge GDN class.
- Run all pre-commit hooks.

References: redai-studio#273.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已复核 41d5296,原 NPU 初始化阻塞问题已解决,确认记录见原讨论。本次单行修改未发现新问题;上一轮 PR 正文更新建议仍待处理,为非阻塞意见。

本轮初始化分支 8 组检查通过。当前 head 无 CI 结果;缺少相应硬件,未重跑 GPU/NPU 或多节点训练验证。

Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.

@jambow0320

Copy link
Copy Markdown
Contributor Author

@xiaoliang0601 已按建议 rebase 并参考 #348 更新了 #273。当前生产代码和 MCore patch 与 #348 完全一致,唯一额外修改是 tests/backends/megatron/test_model_provider_vpp.py:主要是补齐假 Megatron 依赖。

这边跑了几组端到端SFT实验验证,配置是Qwen3.5-9B、OpenMathReasoning-mini、8×H200,覆盖静态 CP4、动态 CP1/2/4,以及 chunkwise、headwise、all_gather 和 main 基线。

排除前 20 步预热后,chunkwise 的 actor 训练吞吐如下;

配置 main PR 提升
静态 CP4 24,114 token/s 45,548 token/s 88.89%
动态 CP 32,327 token/s 47,307 token/s 46.34%

正确性方面,静态 chunkwise A/A 出现过一次梯度尖峰,原始动态 headwise 出现过两次;其余全部对齐;不知道这是否正常,我将对比图片都放在这里:

image image image

@jambow0320

Copy link
Copy Markdown
Contributor Author

CI问题我看是因为没加载当前 PR 对应的 patched MCore 源码镜像;我这边本地测试是没问题的

@SigureMo

Copy link
Copy Markdown
Member

@jambow0320 麻烦注意下,GPU Unit Tests 之前平均测试时间(不考虑 GPU、镜像准备)只有 6min 左右,这个 PR 我之前就注意到会明显高于这个时间,我觉得应该是不合理的,麻烦看看呢?

@rai-studio-bot 定位下 GPU Unit Tests / Unit Tests 时间增加的原因

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

回复 @SigureMo:已用两次运行上传的 junit.xml 对齐到具体用例。增加的约 21 分钟几乎全部来自本 PR 新增的 11 项 GDN GPU 测试,原有测试没有整体变慢。

口径 main b7ad7b49 PR 5144565
pytest 总耗时 385.83 s(6:26) 1651.81 s(27:32)
两边共有用例的 JUnit 耗时合计 346.52 s 347.10 s
PR 新增 137 项用例 — 1265.75 s
其中 test_gdn_chunkwise_cp_gpu.py 的 11 项 — 1264.31 s(21:04)

数据来源:main 基线运行、本次 PR 运行。pytest 总时间包含收集等开销,所以不等于 testcase 时间之和。本次镜像准备仅 27 秒、GPU 分配在同一秒完成,均不在上述 pytest 时间内;本次测试最终 2861 passed / 30 skipped,没有超时失败。

新增文件里最慢的五项如下,合计占该文件耗时约 85%:

用例 秒
TP2/CP2 vs CP1 282.91
GDN 模块对齐 fp32 280.34
GDN 模块对齐 bf16 206.81
FLA kernel 对齐 fp32 172.52
FLA kernel 对齐 bf16 130.39

测试实现对 11 个 case 分别调用 mp.spawn,总共启动 24 个 worker(10×2 + 1×4),重复导入依赖、初始化 CUDA/NCCL/模型。最慢几项还执行真实 FLA 前反向,并对比 CP1、headwise、chunkwise,涉及不同 dtype、head 数与 TP 配置;这比 CPU route/cache 单测重得多。新增其余 126 项合计只有约 1.44 秒。

可以确定的是耗时集中在哪些测试;尚不能确定这 21 分钟内编译、autotune、初始化、collective 各占多少。 当前日志没有这些阶段的计时。固定的 FLA 0.4.2 在 Triton 支持时默认允许 autotune 磁盘缓存,不能说每次 spawn 都重新编译全部内核;但当前 CI 容器只挂载 uv 缓存,没有显式跨运行保存 Triton/Inductor 缓存。

建议先对这五项增加 worker 内的阶段计时(导入/分布式初始化/模型构建/首次及再次 forward-backward,GPU 阶段同步后计时),做同配置冷/热对照。随后优先尝试让相同 world size 的用例复用 worker,保留完整正确性断言;若确认首用内核开销占主导,再评估按依赖版本和 GPU 架构隔离的编译缓存。需要控制每次 PR 的测试预算时,可以保留代表性的必跑 smoke,将完整 dtype × TP/CP × recompute 矩阵放到独立完整验证任务,不建议简单删除覆盖。

本轮只做日志、JUnit 与源码分析,未修改测试或在本地重跑 GPU profiling。

Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.

@SigureMo

Copy link
Copy Markdown
Member

增加的约 21 分钟几乎全部来自本 PR 新增的 11 项 GDN GPU 测试,原有测试没有整体变慢。

新增测试增加这么长时间我觉得也是不合理的,需要看看能不能降低测试成本呢?

@jambow0320

Copy link
Copy Markdown
Contributor Author

增加的约 21 分钟几乎全部来自本 PR 新增的 11 项 GDN GPU 测试,原有测试没有整体变慢。

新增测试增加这么长时间我觉得也是不合理的,需要看看能不能降低测试成本呢?

嗯嗯好的,我这边看看怎么改一下

# ✅ Tests

- Keep one real NCCL THD/SBHD layout round-trip test and remove the ten
  expensive full GPU validation cases and their dedicated helpers.
- Merge CPU partition, route, and cache coverage into test_gdn_cp_layout.
- Group configuration and dispatch coverage in test_gdn_cp_mode, removing
  historical stage names while preserving all 126 core CPU cases.
- Pass 159 related CPU regressions, verify the GPU file collects one test,
  and run all pre-commit hooks. No new GPU execution is claimed.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

测试重组后的 CPU 隔离检查通过 117 项;生产代码未变,原 P1 保持已解决。本轮新增一项非阻塞测试覆盖意见,见行级讨论;正文更新建议仍待处理。

当前 GPU 测试仅保留布局往返,不能将测试删除后的提速视为 GDN 前反向验证。新 head 尚无 CI 结果,本机缺少 CUDA,未重跑 GPU/NCCL 或多节点验证。

Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.

def test_layout_round_trip_over_real_cp_group():
"""RFC 5.1 at the collective level: the layout swap is a pure
permutation."""
_spawn(_worker_layout_round_trip, "n/a", 29544)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 优先级:P2

请保留至少一项真实 GDN 前反向集成回归,或将原数值矩阵移到默认 CI 之外的独立验证入口。本次删掉 test_gdn_cp_matches_cp1、recompute 和 TP2/CP2 等检查后,唯一 GPU 测试只调用布局转换,输入也没有 requires_grad;CPU mode 测试则把 GatedDeltaNet.forward 替换成 spy。现存 test_gdn_cp_reassembly 使用 matmul/cumsum 的 toy scan,覆盖的是 all-gather,不能替代这里的 chunkwise GDN。

因此,即使 GDN 的 route 接入、CP context 或重计算反向接线出错,这些保留下来的测试仍可能全部通过,而这些正是本 PR 修改的生产路径。可以将单一 dtype、CP=2 对 CP=1 的输出/输入梯度/参数梯度检查压缩为一次 spawn,并把较重矩阵保留为可单独运行的测试;无需恢复每次 CI 都执行完整 21 分钟矩阵。

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

e09e8f5f 已解决这里的覆盖缺口:新的 test_gdn_cp_gpu.py 在一次双 GPU spawn 中,用相同权重比较 CP=1 与 CP=2 chunkwise / all_gather 的输出、输入梯度和逐参数梯度;参数梯度在 CP 组内求和后与基准比较,all_gather 同时执行 checkpoint 前反向,原布局往返检查也保留。

本轮核对了实际 GDN 调用、动态 CP 组和梯度汇总路径,未发现新增问题。本机无 CUDA,测试收集成功后按条件跳过;“覆盖已恢复”不代表已验证 GPU 数值通过或 H20 耗时低于 4 分钟,这两项仍需当前 head 的 GPU CI 确认。

@jambow0320

Copy link
Copy Markdown
Contributor Author

@SigureMo hi,我这边把大部分的gpu测试都删掉了,只保留了布局转换的,这个只用时10s左右

@SigureMo

SigureMo commented Sep 22, 2026 •

Copy link
Copy Markdown
Member

CI 里 Python3.11 的报错问题还需要看一下,可以在 CPU 环境测一下

@SigureMo

Copy link
Copy Markdown
Member

@rai-studio-bot 分析一下 test_streaming_dataset_missing_media_skip_refills_batch 失败问题是否只有本 PR 随机出现,因为据我观察,其他 PR 基本很稳定不会出现这一问题

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

回复 @SigureMo:不是本 PR 独有,也不只发生在 Python 3.11。 这次补查了其他分支的实际失败日志,找到两个相同断言的反例:

PR / 运行 Python 失败位置
#323,9 月 19 日 3.11 [prefetch],第 980 行,警告捕获为空
#257,9 月 21 日 3.10 同上
本 PR,当前 head 第 3 次运行 3.10 同上;3.11、3.12 通过

而同一个 e9304afc 的第 2 次运行,三个 Python 版本全部通过。因此“本 PR 连续出现”确实值得检查,但不能据此归为本 PR 特有或版本特有问题。

根因仍是测试的日志捕获时序。 测试第 963–974 行先 ds.shuffle(0),后 monkeypatch logger.warning;shuffle 已经启动预取线程。上述失败日志都实际打印了缺失图片的警告,且补齐样本的断言已通过,只是后台线程可能在安装捕获前发出警告。相关测试、SFT 实现与预取器相对 main 均未改变。

本机 Python 3.10 的隔离受控复现中,让预取线程先完成,原测试 5/5 失败;将捕获安装移到 shuffle 之前,同样调度下 5/5 通过。这是对竞态机制的验证,不是自然失败率测量,也未运行完整 CI 环境。

统计边界:检查了最近 100 条 CI 运行记录中最新结果为失败的 23 条日志,其中 5 条命中这一断言(本 PR 3 条、其他 PR 2 条),没有穷尽历史重跑;CI 还使用 -x,其他失败可能提前终止。因此目前能确认“并非本 PR 独有”,不能量化本 PR 是否更易触发。建议直接把 ds.shuffle(0) 移到 monkeypatch 之后、try 内,消除捕获窗口;不建议依靠重跑或加 sleep。

Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.

@SigureMo

Copy link
Copy Markdown
Member

回复 @SigureMo:不是本 PR 独有,也不只发生在 Python 3.11。 这次补查了其他分支的实际失败日志,找到两个相同断言的反例:

PR / 运行 Python 失败位置
#323,9 月 19 日 3.11 [prefetch],第 980 行,警告捕获为空
#257,9 月 21 日 3.10 同上
本 PR,当前 head 第 3 次运行 3.10 同上;3.11、3.12 通过
而同一个 e9304afc 的第 2 次运行,三个 Python 版本全部通过。因此“本 PR 连续出现”确实值得检查,但不能据此归为本 PR 特有或版本特有问题。

根因仍是测试的日志捕获时序。 测试第 963–974 行先 ds.shuffle(0),后 monkeypatch logger.warning;shuffle 已经启动预取线程。上述失败日志都实际打印了缺失图片的警告,且补齐样本的断言已通过,只是后台线程可能在安装捕获前发出警告。相关测试、SFT 实现与预取器相对 main 均未改变。

本机 Python 3.10 的隔离受控复现中,让预取线程先完成,原测试 5/5 失败;将捕获安装移到 shuffle 之前,同样调度下 5/5 通过。这是对竞态机制的验证,不是自然失败率测量,也未运行完整 CI 环境。

统计边界:检查了最近 100 条 CI 运行记录中最新结果为失败的 23 条日志,其中 5 条命中这一断言(本 PR 3 条、其他 PR 2 条),没有穷尽历史重跑;CI 还使用 -x,其他失败可能提前终止。因此目前能确认“并非本 PR 独有”,不能量化本 PR 是否更易触发。建议直接把 ds.shuffle(0) 移到 monkeypatch 之后、try 内,消除捕获窗口;不建议依靠重跑或加 sleep。

Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.

@ShigureNyako 参考这个单独提一个 PR 优化下吧

@ShigureNyako

Copy link
Copy Markdown
Contributor

@SigureMo 已按这条请求单独提交 PR #358。仅修改测试:将 ds.shuffle(0) 移到 logger.warning 捕获安装之后、try 内,保留 finally 中的 ds.stop(),不改生产逻辑。

本地 Python 3.10 验证:test_streaming.py 57 passed;SFT dataset 与共享 streaming dataset 相关测试 142 passed / 1 skipped(缺少 pyarrow)。pre-commit 全通过;PR CI 尚未全部完成,目前 Pre-commit Checks 已通过,其余仍 pending。

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已核验 83a096ad 合入的 #358:logger.warning 捕获先于 ds.shuffle(0) 安装,且启动预取位于 try/finally 内,修复了此前确认的日志捕获竞态。使用当前测试代码,在 Python 3.10 隔离环境强制预取线程先完成,5/5 通过;这不是完整 CI 验证,当前 head 尚无 CI 结果。

相对 main 的 PR 差异与上轮一致,无新增发现;真实 GDN 前反向覆盖意见及正文更新建议仍待处理。

Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.

@SigureMo

Copy link
Copy Markdown
Member

#358 合入后看起来稳定性高了不少,我去推进下合入

@xiaoliang0601

Copy link
Copy Markdown
Contributor

@jambow0320 补个 gpu 测试吧,可以小点儿,主要测真实 GDN 在 CP=2 下,输出和梯度是否与 CP=1 一致。CP=1、CP=2 chunkwise、CP=2 all gather 数值一致。按照 @SigureMo 说的那样,在 H20 CI 上尽量 < 4 分钟,不要弄太大的计算。

@SigureMo

Copy link
Copy Markdown
Member

在 H20 CI 上尽量 < 4 分钟

嗯嗯,现在 CI 限定 timeout 10min,如果超时(或者濒临超时),这个 PR 里可以将 10min 改为 15min

@jambow0320

Copy link
Copy Markdown
Contributor Author

@jambow0320 补个 gpu 测试吧,可以小点儿,主要测真实 GDN 在 CP=2 下,输出和梯度是否与 CP=1 一致。CP=1、CP=2 chunkwise、CP=2 all gather 数值一致。按照 @SigureMo 说的那样,在 H20 CI 上尽量 < 4 分钟,不要弄太大的计算。

@xiaoliang0601 可以的,我这边先调一下规模本地测试一下;
@SigureMo 以及我们的CI是不是后续也可以像sglang那样弄不同tag的CI,我主要顾虑到常态化地增加CI时间也不太好,应该是有涉及到GDN、CP之类的改动才应该跑相关的CI,这样才起到保护的作用

@SigureMo

Copy link
Copy Markdown
Member

应该是有涉及到GDN、CP之类的改动才应该跑相关的CI

精准测试其实是有风险的,一旦某块逻辑经常不跑,保护作用会失效的,根据之前的经验,如非必要,不引入这种机制

@jambow0320

Copy link
Copy Markdown
Contributor Author

了解~那我看看这块怎么设置测试更好

# ✅ Tests

- Compare real BF16 GDN CP=1 with CP=2 chunkwise and all_gather outputs,
  input gradients, and all parameter gradients using 128-wide heads and
  unequal packed sequences.
- Run the existing exact THD/SBHD layout checks in the same two-GPU spawn.
- Select one FLA autotune candidate only inside test workers, retaining
  real JIT kernels and NCCL while avoiding repeated candidate benchmarks.
- Update references to the consolidated GPU test filename.

## Validation

- Two H200s, empty compilation caches: 1 passed, 0 skipped; 60.224 seconds
  for the combined regression (61.740 seconds including pytest overhead).
- Python 3.11 CPU-only collection skips the GPU case as intended.
- pre-commit run --all-files --show-diff-on-failure passed.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

真实 GDN 前反向回归入口已恢复,原覆盖意见已在对应讨论更新。本轮仅测试变化,未发现新增问题;正文范围与验证版本的更新建议仍待处理。

当前 head 尚无 CI 结果。本机测试收集正常,但因无 CUDA 跳过,未验证 GPU/NCCL 数值或 H20 执行耗时。

Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.

@jambow0320

Copy link
Copy Markdown
Contributor Author

@xiaoliang0601 已补充 GDN 的 CP=1、CP=2 chunkwise/all_gather 输出与梯度对照并合并布局测试;通过固定 FLA 配置跳过 autotune 后,在h20上CI耗时49秒通过

@xiaoliang0601

Copy link
Copy Markdown
Contributor

我没问题了,cc @SigureMo 走合入流程吧~

@SigureMo
SigureMo merged commit 8a29ef2 into redai-studio:main Sep 23, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants