Add opt-in native two-node MI300X path for Kimi K3 AgentX / 新增 Kimi K3 AgentX 的 MI300X 双节点 native 路径 - #2401
Draft
edwingao28 wants to merge 16 commits into
Draft
Conversation
Add the exact TP8 x PP2 aggregate matrix and update the design with verified Slurm and node-local storage facts. 中文:新增精确的 TP8 x PP2 聚合式矩阵,并用已验证的 Slurm 与节点本地存储事实更新设计。 Signed-off-by: Wenyao Gao <wgao11@u.rochester.edu>
Serve the narrow two-rank TP8 x PP2 AMD topology while leaving the gfx942 AITER a8w4 switch caller-configurable. 中文:新增窄范围的双 rank TP8 x PP2 AMD 启动入口,并保留 gfx942 AITER a8w4 开关由调用方配置。 Signed-off-by: Wenyao Gao <wgao11@u.rochester.edu>
Verify gfx942 topology and complete pinned weights, then atomically import and validate the K3 image in the writable node-local squash tree. 中文:逐节点验证 gfx942 拓扑与完整的固定版本权重,并在可写的节点本地 squash 目录中原子导入和校验 K3 镜像。 Signed-off-by: Wenyao Gao <wgao11@u.rochester.edu>
Add fail-fast two-node allocation, per-rank lifecycle, bounded artifact handoff, and cleanup for success, failure, and signals without changing the default launcher path. 中文:新增失败即停的 MI300X 双节点分配、分 rank 生命周期、受限产物交接,以及成功、失败和信号路径清理,同时保持默认启动路径不变。 Signed-off-by: Wenyao Gao <wgao11@u.rochester.edu>
中文:补充 Kimi K3 MI300X 的 perf-changelog 条目。 Signed-off-by: Wenyao Gao <wgao11@u.rochester.edu>
The multi-node workflow template never sets HF_HUB_CACHE, so the trace corpus was re-downloaded into a throwaway container layer on every job. 中文:多节点 workflow 模板不设置 HF_HUB_CACHE,导致每个作业都把 trace 语料重新下载到临时容器层。 Signed-off-by: Wenyao Gao <wgao11@u.rochester.edu>
The launcher flattened every client failure to 1, so a failing replay was indistinguishable from a launcher error. Artifacts are still published first, then the client's own status is preserved. 中文:launcher 之前把所有 client 失败都压成 1,无法与 launcher 自身错误区分。现在仍先落盘产物,再原样透传 client 退出码。 Signed-off-by: Wenyao Gao <wgao11@u.rochester.edu>
The native launcher had its own copies of the squeue liveness poll, the server-log bundling and the workspace copy. runners/slurm_utils.sh already provides all three, so source it the way launch_gb200-nv.sh does. 中文:native launcher 自带 squeue 存活轮询、server 日志打包和 workspace 拷贝三份实现,runners/slurm_utils.sh 已有,按 launch_gb200-nv.sh 的方式 source 复用。 Signed-off-by: Wenyao Gao <wgao11@u.rochester.edu>
Every sibling launcher and recipe carries set -x; the new scripts had none. Tracing is scoped the way launch_mi355x-amds.sh scopes it, off around the multi-hour health poll and the command-array build. Also names the image in the changelog entry and points the three scripts at the operator runbook. 中文:同目录 launcher 和 recipe 都有 set -x,新脚本缺失。按 launch_mi355x-amds.sh 的做法做范围控制:健康轮询与命令数组构建处关闭。changelog 补上镜像名,三个脚本加 runbook 指引。 Signed-off-by: Wenyao Gao <wgao11@u.rochester.edu>
Adversarial review found three confirmed defects. scancel ran up to two minutes into cleanup, so a cancelled CI job leaked two exclusive nodes for the full 8h wall clock; only the scratch removal needs the allocation, so it now runs on a short deadline and log packaging moved after scancel. Killing the wrapper subshell left both vLLM ranks running, so the srun PID is recorded and killed directly. The client HF cache mounted the parent of the pinned checkpoint read-write, so it now gets its own subtree. Also validates the timing knobs, adds --container-writable to match the other AMD launchers, compares image builds across nodes, and stops the second path that masked the client exit code. 中文:对抗式 review 确认三处缺陷。scancel 最迟在清理开始约两分钟后才执行,CI 取消会让两个独占节点空占满 8 小时;只有 scratch 删除依赖 allocation,改为短超时并把日志打包移到 scancel 之后。kill wrapper 子 shell 无法停止 srun,改为记录并直接 kill srun PID。client HF 缓存把权重父目录以读写方式挂入,改为独立子目录。另外校验计时参数、补 --container-writable、跨节点比对镜像构建,并修掉第二处掩盖 client 退出码的路径。 Signed-off-by: Wenyao Gao <wgao11@u.rochester.edu>
The SIGTERM test asserted the exit code and scancel but never that the step died. The fake server step now records its PID; reverting the srun-PID kill makes this fail with "server step outlived cleanup". 中文:SIGTERM 测试此前只断言退出码和 scancel,没有验证 step 真的结束。fake server step 现在记录自身 PID;回退 srun-PID kill 会触发 "server step outlived cleanup" 失败。 Signed-off-by: Wenyao Gao <wgao11@u.rochester.edu>
The workflow names each test file, so a new module is silently absent from CI. Also triggers on the shell entrypoints the module drives, which live outside utils/matrix_logic. 中文:workflow 逐个文件列出测试,新模块不加进去就不会在 CI 跑。同时对该模块驱动的 shell 入口文件加触发路径,它们不在 utils/matrix_logic 下。 Signed-off-by: Wenyao Gao <wgao11@u.rochester.edu>
中文:把 changelog 条目的 pr-link 指向本 PR。 Signed-off-by: Wenyao Gao <wgao11@u.rochester.edu>
Folds five preflight rejections, the two rank cases and the two preflight-record rejections into parametrized tests, keeping the same 32 cases. Adds the docstrings AGENTS.md asks for, and restores the PEP 8 blank lines between top-level definitions. 中文:把五个 preflight 拒绝用例、两个 rank 用例和两个 preflight record 拒绝用例合并为参数化测试,用例数不变仍为 32。补上 AGENTS.md 要求的 docstring,并恢复顶层定义之间的 PEP 8 空行。 Signed-off-by: Wenyao Gao <wgao11@u.rochester.edu>
edwingao28
force-pushed
the
feat/kimik3-mi300x-native-multinode
branch
from
July 28, 2026 23:53
3a34ad9 to
0a0b5d7
Compare
Collaborator
|
@edwingao28 the AgentX/AIPerf harness has been updated, please merge origin/main into your branch and refresh your submission. Additional tuning may be necessary depending on the config. I apologize for any inconvenience. This is an automated message. |
Integrate the updated additive AIPerf warmup contract and preserve the Kimi K3 MI300X changelog entry at the append-only tail. 中文:合入最新 origin/main 的 AIPerf 逐 lane 追加预热协议,并按追加式约束将 Kimi K3 MI300X changelog 条目保留在文件末尾。 Signed-off-by: Wenyao Gao <wgao11@u.rochester.edu>
Integrate the RTX PRO 6000 and CollectiveX updates while preserving the Kimi K3 MI300X entry at the append-only changelog tail. 中文:合入最新的 RTX PRO 6000 与 CollectiveX 更新,并按追加式约束将 Kimi K3 MI300X changelog 条目保留在文件末尾。 Signed-off-by: Wenyao Gao <wgao11@u.rochester.edu>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Status: not yet validated on hardware
This is a draft. The MI300X end-to-end run is blocked on AITER gfx942 compatibility and has not happened yet. Everything below is validated on CPU only. Please do not read a green CI badge as "Kimi K3 runs on MI300X" — that gate comes later.
What is still outstanding, in order:
perf-changelog.yamlpr-linkto this PR's numberThe
full-sweep-fail-fastlabel is deliberately not applied yet. It would queue cluster runs against a dependency that is known to be broken. It goes on once the AITER probe is green.What this adds
An opt-in path that runs aggregated Kimi-K3 AgentX across two 8xMI300X nodes with plain vLLM. TP8 inside each node, PP2 across the node boundary, EP1, no DP attention. Rank 0 serves the endpoint, rank 1 runs headless.
The 1.5 TB checkpoint does not leave safe runtime headroom on one node's HBM, which is why this spans two nodes. PP2 across the boundary rather than a cross-node TP16, because TP16 would put a collective on every tensor-parallel op.
Four files are new:
runners/launch_mi300x-amds-native-multinode.sh— allocation, both server ranks, the AgentX client, artifact handoff, and cleanup on success, failure or signal.runners/mi300x_native_node_preflight.sh— runs on every allocated node./raidis node-local, so a model cached or an image imported on one node is invisible to the other. It never downloads the checkpoint; absent weights fail the job rather than being fixed inside a timed allocation.benchmarks/multi_node/agentic/kimik3_fp4_mi300x_vllm.sh— one vLLM rank. The caller supplies node count, rank and master address, so the same script works under Slurm or a manual two-terminal bring-up.utils/matrix_logic/test_kimik3_mi300x_native.py— the CPU gate, wired intotest-matrix-logic.yml.configs/amd-master.yamlgets one new key.runners/launch_mi300x-amds.shgets a four-line guard at the top. Nothing else in the single-node MI300X path changes, and it only diverges whenNATIVE_MULTINODE=1.AITER_SITUV2_A8W4is deliberately not set anywhere. gfx942 has no committed Kimi-K3 tuned MoE tables, so the a8w4 switch stays an operator input until an exact-shape test picks a mode.What the tests cover
32 tests, all CPU, no GPU and no cluster. They drive the real shell scripts with fake
salloc,srun,squeue,scancel,rocminfoandenrootonPATH.Three of them exist because a review found real bugs:
scancel. A cancelled CI job is killed well inside that window, so two exclusive nodes stayed allocated for the full eight-hour wall clock. Only the scratch removal actually needs the allocation, so it now has a short deadline and log packaging happens after the cancel.srunit was waiting on. That orphaned both vLLM ranks, which kept running and kept appending to the log the launcher was about to package and delete. There is now a test that fails with "server step outlived cleanup" if this regresses.One test is skipped on macOS because the pre-existing single-node launcher parses
sallocwith GNUgrep -P. It runs onubuntu-latest.Notes for reviewers
This touches
configs/amd-master.yaml, so it needs an AMD CODEOWNER. It also needs thefull-sweep-fail-fastlabel perCONTRIBUTING.md.中文说明
状态:尚未在硬件上验证
这是 draft。MI300X 端到端验证被 AITER gfx942 兼容性阻塞,还没有跑过。 下面所有内容只在 CPU 上验证过。CI 变绿不等于「Kimi K3 在 MI300X 上跑通了」。
剩余步骤按顺序:AITER gfx942 probe 通过 → 两个节点 staging 权重 → 对 rank 0 endpoint 发一个直接 vLLM 请求 → AgentX concurrency 1 → 完整 1/2/4/8 sweep → 把 changelog 里的
pull/XXX换成本 PR 编号。这个 PR 加了什么
一条 opt-in 路径,用 plain vLLM 在两台 8xMI300X 上跑 aggregated Kimi-K3 AgentX。节点内 TP8,跨节点 PP2,EP1,不开 DP attention。rank 0 提供 endpoint,rank 1 headless。
1.5 TB 权重在单节点 HBM 上没有安全余量,所以要跨两个节点。跨边界用 PP2 而不是 TP16,因为 TP16 会让每个 tensor-parallel op 都带一次 collective。
新增四个文件:launcher 负责 allocation、两个 server rank、AgentX client、产物回传,以及成功、失败和 signal 三种情况下的清理;preflight 在每个节点上各跑一次,因为
/raid是节点本地的,一个节点上缓存的模型另一个节点看不到,它不会下载权重,缺权重直接失败;rank entrypoint 与调度器无关,Slurm 和手工双终端都能用;测试是 CPU gate,已接进test-matrix-logic.yml。configs/amd-master.yaml加一个 key,runners/launch_mi300x-amds.sh顶部加四行 guard。单节点路径其他部分不变,只有NATIVE_MULTINODE=1时才会走新路径。AITER_SITUV2_A8W4到处都没有设。gfx942 上没有 Kimi-K3 的 tuned MoE 表,所以在有确切 shape 的测试之前,这个开关保持为操作者输入。测试覆盖
32 个测试,纯 CPU,不需要 GPU 和集群。用 fake 的
salloc、srun、squeue、scancel、rocminfo、enroot驱动真实的 shell 脚本。其中三个来自 review 发现的真实缺陷:清理时
scancel之前最多要等两分钟,CI 取消会在这之前就被 kill,导致两个独占节点空占满八小时;清理只 kill 了等待srun的 wrapper 子 shell,两个 vLLM rank 变成孤儿继续运行并继续写那个马上要被打包删除的日志;client 失败一律返回 1,无法与 launcher 自身错误区分。有一个测试在 macOS 上跳过,因为原有的单节点 launcher 用 GNU
grep -P解析salloc。在ubuntu-latest上会执行。给 reviewer 的说明
改动涉及
configs/amd-master.yaml,需要 AMD CODEOWNER review,并且按CONTRIBUTING.md需要full-sweep-fail-fastlabel。