Add Qwen3.5 FP4 B200 AgentX MTP - #2420
Conversation
Add the initial SGLang NEXTN AgentX recipe with golden synthetic acceptance and 256k traces.\n\n中文:新增 Qwen3.5 FP4 B200 AgentX MTP 初始配置,使用 SGLang NEXTN、黄金合成接受长度和 256k 轨迹数据集。
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
Append the Qwen3.5 FP4 B200 AgentX MTP benchmark entry.\n\n中文:追加 Qwen3.5 FP4 B200 AgentX MTP 基准测试触发记录。
Probe the allocated DGXC node for the complete Qwen3.5 NVFP4 checkpoint on /raid and fall back to shared Lustre when it is absent.\n\n中文:在分配到的 DGXC 节点上探测 /raid 中完整的 Qwen3.5 NVFP4 权重;若未暂存,则回退到共享 Lustre。
Add strict two-pool HiCache accounting, request cache reporting, session-sticky attention-DP routing, and AgentX request/graph limits for targeted B200 tuning.\n\n中文:为 B200 定向调优加入严格的双池 HiCache 容量核算、请求级缓存统计、会话粘性的注意力数据并行路由,以及 AgentX 请求数与 CUDA Graph 上限。
Move the configuration into the AgentX section, document the measured cache model, and add session-sticky DEP8 no-offload probes for the throughput frontier.\n\n中文:将配置移入 AgentX 区域,记录实测缓存模型,并为吞吐量前沿加入具备会话粘性路由的 DEP8 无卸载探测点。
中文:添加 B200 TP2 c16 HiCache 对照探针,并严格使用按 GPU 数量计算的 DRAM 上限。
中文:将 B200 TP2 HiCache 验证点设为独立的 c20,避免一次性重复启动无卸载配置。
Add the isolated TP4 c96 no-offload point after c64 remained healthy, so the HBM working-set cliff can be measured before adding matched HiCache coverage. 中文:在 c64 保持健康后新增独立的 TP4 c96 无卸载测试点,以便在添加匹配的 HiCache 覆盖前测量 HBM 工作集拐点。
Remove the DEP8 search arm after its c128 probe fragmented prefix reuse, completed no profiled requests in three minutes, and delivered less than half of TP4 c64 aggregate prefill throughput. 中文:DEP8 c128 测试出现前缀复用碎片化,三分钟内没有完成任何正式请求,聚合预填充吞吐量不足 TP4 c64 的一半,因此移除该搜索分支。
Account for SGLang target KV, Mamba, and native NEXTN draft host pools as H * 31/15 per rank, with page-alignment reserve, so projected use stays within runner-generated DRAM. 中文:按每个 rank 的 H * 31/15 计入 SGLang 目标 KV、Mamba 与原生 NEXTN 草稿主机缓存,并预留页对齐空间,确保预计用量不超过 runner 生成的 DRAM 上限。
Add matched no-offload and DRAM HiCache c72 points immediately after the measured c64 KV-pressure onset. 中文:在实测 c64 KV 压力拐点之后,添加匹配的 c72 无卸载与 DRAM HiCache 对照测试点。
中文:将最新 main 合并到 B200 提交分支,并将本 PR 的性能变更记录重新追加到文件末尾。
Simplify the B200 launcher script to TP-only serving after the measured DEP arm was removed from the search space. 中文:实测 DEP 分支已从搜索空间移除,因此将 B200 启动脚本简化为仅使用 TP 的服务路径。
Replace the dominated c24-c64 no-offload tail with dense c12-c22 points and matched HiCache coverage beginning at c16. 中文:移除性能受压的 c24-c64 无卸载尾部,改为密集采样 c12-c22,并从 c16 开始提供匹配的 HiCache 覆盖。
Densify TP4 around the measured c64-c72 transition and extend the HiCache curve beyond the HBM cliff.\n\n中文:细化 B200 Qwen TP4 在实测 c64-c72 缓存拐点附近的并发点,并将 HiCache 曲线扩展到 HBM 拐点之后。
Remove no-offload c22 after live profiling placed it past the TP2 cache cliff, while retaining the HiCache arm and c20 boundary.\n\n中文:实测表明无卸载 c22 已越过 TP2 缓存拐点,因此移除该点,同时保留 HiCache 配置与 c20 边界点。
中文:在缓存抖动前收紧 B200 TP4 并发范围。
中文:B200 TP2 在并发 14 后切换至 HiCache。
中文:裁剪性能受支配的 B200 TP4 并发 66 配置。
中文:使用 SGLang 默认单 tokenizer 启动路径,避免长时间 HiCache 初始化期间的多 tokenizer 共享内存竞态。
中文:移除已被 c64 严格支配的 B200 c68 无卸载点,并让 HiCache 从缓存悬崖边界开始接管。
Use six tokenizer workers for TP4 long-context replay while retaining SGLang default single-worker startup for TP2 HiCache. 中文:TP4 长上下文回放使用 6 个 tokenizer worker;TP2 HiCache 保持 SGLang 默认单 worker 启动路径,避免共享内存初始化竞态。
Keep the measured Pareto region and validate every node-local checkpoint shard before selecting NVMe weights.\n\n中文:保留实测帕累托区间,并在选择 NVMe 权重前校验节点本地检查点的全部分片。
|
Claude encountered an error after 23s —— View job I'll analyze this and get back to you. |
There was a problem hiding this comment.
LGTM — straightforward new benchmark recipe addition following the established pattern.
What was reviewed: the new AgentX bench script's HiCache DRAM sizing math (self-consistent floor/ceiling division, guarded against exceeding configured capacity), the new nvidia-master.yaml entry (validated it parses and generates correct matrix entries via generate_sweep_configs.py), and the shared runners/launch_b200-dgxc.sh changes (the new node-local NVMe checkpoint probe is scoped to qwen3.5+fp4 only and the reordered MODEL_PATH/MODEL export is behavior-preserving for all other model recipes).
Extended reasoning...
Overview
This PR adds a new single-node AgentX benchmark recipe for Qwen3.5-397B-A17B NVFP4 on B200 with SGLang native NEXTN MTP: a new bench script (benchmarks/single_node/agentic/qwen3.5_fp4_b200_sglang_mtp.sh), a new configs/nvidia-master.yaml entry with TP2/TP4 x no-offload/HiCache search spaces, a perf-changelog.yaml entry, and a shared-infra change in runners/launch_b200-dgxc.sh that adds an optional node-local /raid checkpoint probe for this specific model+precision combo.
Security risks
None. This is internal benchmarking infrastructure (Slurm job launch scripts and YAML sweep configs), not a production service. No user input, auth, or crypto is touched. The new Python heredoc in the launcher only reads local filesystem paths on the allocated compute node to validate checkpoint completeness.
Level of scrutiny
Low-to-moderate. Three of the four files are additive and scoped to the new recipe (bench script, yaml entry, changelog). The fourth, runners/launch_b200-dgxc.sh, is shared infra used by all B200 DGXC single-node recipes, which normally warrants closer scrutiny — but the new logic is gated behind MODEL_PREFIX == \"qwen3.5\" && PRECISION == \"fp4\"\", and the only change affecting other recipes is a reordering of the pre-existing export MODEL="$MODEL_PATH"line to after the Slurm job ID is resolved — $MODEL_PATH is not reassigned for any other model prefix in between, so this reorder is behavior-preserving. I verified the new yaml entry generates valid matrix rows viagenerate_sweep_configs.py, and both shell scripts pass bash -n` syntax checks. The HiCache DRAM sizing arithmetic in the bench script is domain-specific (based on Qwen3.5's layer/KV-head geometry) but is internally self-consistent (matching floor/ceiling divisions) and fails closed if capacity would be exceeded.
Other factors
No prior substantive review exists on this PR (only the automated recipe-reminder bot comment and my own initial "working" placeholder comment). The bug-hunting system found no issues, and my own review of the diff and cross-checks against benchmark_lib.sh (all referenced helper functions exist) turned up nothing additional. This follows the same structure as many prior model-recipe PRs in this repo (e.g. the DSv4, MiniMax-M3, Kimi-K3 entries in the same launcher file), so the pattern is well established.
Summary
Add Qwen3.5-397B-A17B NVFP4 AgentX coverage on B200 with SGLang native NEXTN MTP, golden synthetic acceptance, prefix caching, and the 256k trace dataset.