Add DeepSeek V4 Pro B200 workload#45
Conversation
Follow the official NVFP4 Blackwell recipe and retain the existing accuracy, BFCL, and serving benchmark coverage. Co-Authored-By: Codex <noreply@anthropic.com> Signed-off-by: khluu <khluu000@gmail.com>
Point TileLang at the CUDA 13.0 toolkit installed in the vLLM image so its NVCC compiler and CCCL headers remain version-matched.\n\nCo-Authored-By: Codex <noreply@anthropic.com> Signed-off-by: khluu <khluu000@gmail.com>
|
Buildkite #313 reached the DeepSeek V4 MHC kernel warmup after loading the 64-shard checkpoint, then failed consistently on all ranks while compiling TileLang's discovery order selected the pip |
|
Buildkite #318 has now crossed the exact #313 failure boundary. All TileLang MHC kernels, including |
Co-Authored-By: Codex <noreply@anthropic.com> Signed-off-by: khluu <khluu000@gmail.com>
|
Buildkite #318 validated the CUDA-toolkit fix, then exposed a separate topology issue: TP8 reduced DeepSeek V4's MLA dimensions to |
|
Replacement DP8 + EP validation is running in Buildkite #320 on |
Co-Authored-By: Codex <noreply@anthropic.com> Signed-off-by: khluu <khluu000@gmail.com>
|
Buildkite #320 reproduced the same zero-dimension error with TP=1 / DP=8, ruling out topology. The exact stack is |
|
Replacement validation is running in Buildkite #321 on |
|
Buildkite #321 passed in 34m00s on |
This PR was authored with assistance from Codex.
Summary
Recipe: https://recipes.vllm.ai/deepseek-ai/DeepSeek-V4-Pro
Local validation
lm_evalregistrybash -n lib/run.sh lib/server.sh lib/run_lm_eval.sh lib/run_vllm_bench.shpython3 .buildkite/test_generate_pipeline.py(6/6 passed)WORKLOADS=deepseek_v4_pro_b200pipeline generationgit diff --checkGPU validation
mhc_pre_big_fuse_broadcast_with_norm_tilelang7e9eb94setsCUDA_HOME=/usr/local/cuda, selecting the CUDA 13.0 compiler and headers installed together in the vLLM imagefa4_cutedsl_warmup()because DeepSeek V4 intentionally has no legacy MLAqk_nope_head_dimorv_head_dimeb441ccrestores the recipe-default TP8 + EP topology and setskernel_config.enable_jit_warmup=false, skipping only the incompatible generic FA4/MLA warmup