chore(data): add matched B200 SGLang 0.5.14 power - #1535
Conversation
|
Important Review skippedReview was skipped due to path filters ⛔ Files ignored due to path filters (24)
CodeRabbit blocks several paths by default. You can override this behavior by explicitly including those paths in the path filters. For example, including ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Comment |
Perf Parquet Diff Report
Compared
Per-File Row Diff PreviewShowing the first 3 rows per diff kind for each changed parquet file. Full exact CSVs are in aic-core/src/aiconfigurator_core/systems/data/b200_sxm/attention/sglang/0.5.14/context_attention_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,batch_size,isl,num_heads,num_key_value_heads,head_dim,beam_width,attn_dtype,kv_cache_dtype,step,window_size,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,context_attention,flashinfer,1,1,1,1,128,1,bfloat16,bfloat16,0,0,,305.366,,1000.0
SGLang,0.5.14,NVIDIA B200,context_attention,flashinfer,1,1,1,1,128,1,bfloat16,fp8,0,0,,300.529,,1000.0
SGLang,0.5.14,NVIDIA B200,context_attention,flashinfer,1,1,16,1,128,1,bfloat16,bfloat16,0,0,,385.902,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/attention/sglang/0.5.14/generation_attention_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,batch_size,isl,num_heads,num_key_value_heads,head_dim,beam_width,attn_dtype,kv_cache_dtype,step,window_size,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,generation_attention,flashinfer,1,1,1,1,128,1,bfloat16,bfloat16,1,0,,668.08,,1000.0
SGLang,0.5.14,NVIDIA B200,generation_attention,flashinfer,1,1,1,1,128,1,bfloat16,bfloat16,1023,0,,707.142,,1000.0
SGLang,0.5.14,NVIDIA B200,generation_attention,flashinfer,1,1,1,1,128,1,bfloat16,bfloat16,127,0,,543.522,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/comm/sglang/0.5.14/custom_allreduce_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,allreduce_dtype,num_gpus,message_size,backend,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,all_reduce,SGLang_CustomAllReduce_eager,bfloat16,2,1024,sglang_eager,,255.336,,1000.0
SGLang,0.5.14,NVIDIA B200,all_reduce,SGLang_CustomAllReduce_eager,bfloat16,2,1048576,sglang_eager,,331.051,,1000.0
SGLang,0.5.14,NVIDIA B200,all_reduce,SGLang_CustomAllReduce_eager,bfloat16,2,128,sglang_eager,,252.488,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/encoder_attention/sglang/0.5.14/encoder_attention_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,batch_size,isl,num_heads,head_dim,attn_dtype,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,encoder_attention,flash_attention_v4,1,1,1,64,bfloat16,,958.84,,1000.0
SGLang,0.5.14,NVIDIA B200,encoder_attention,flash_attention_v4,1,1,1,72,bfloat16,,992.717,,1000.0
SGLang,0.5.14,NVIDIA B200,encoder_attention,flash_attention_v4,1,1,10,64,bfloat16,,985.955,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/gemm/sglang/0.5.14/gemm_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,gemm_dtype,m,n,k,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,gemm,sglang_deepgemm_gemm_nt_f8f8bf16,fp8_block,1,1024,1024,,641.5605,,1000.0
SGLang,0.5.14,NVIDIA B200,gemm,sglang_deepgemm_gemm_nt_f8f8bf16,fp8_block,1,1024,10240,,581.8160000000001,,1000.0
SGLang,0.5.14,NVIDIA B200,gemm,sglang_deepgemm_gemm_nt_f8f8bf16,fp8_block,1,1024,12288,,830.154,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/linear_attention/sglang/0.5.14/gdn_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,phase,batch_size,seq_len,num_tokens,d_model,d_conv,num_k_heads,head_k_dim,num_v_heads,head_v_dim,model_name,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,gdn,causal_conv1d_fn,context,1,1,1,1024,4,16,128,16,128,Qwen/Qwen3.5-0.8B,,246.104,,1000.0
SGLang,0.5.14,NVIDIA B200,gdn,causal_conv1d_fn,context,1,1,1,1024,4,2,128,2,128,Qwen/Qwen3.5-0.8B,,250.571,,1000.0
SGLang,0.5.14,NVIDIA B200,gdn,causal_conv1d_fn,context,1,1,1,1024,4,4,128,4,128,Qwen/Qwen3.5-0.8B,,241.958,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/mhc/sglang/0.5.14/mhc_module_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,architecture,num_tokens,num_sites,hc_mult,hidden_size,sinkhorn_iters,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,post,sglang_tilelang_mhc_post,DeepseekV4ForCausalLM,1,2,4,4096,20,,244.082,,1000.0
SGLang,0.5.14,NVIDIA B200,post,sglang_tilelang_mhc_post,DeepseekV4ForCausalLM,1,2,4,7168,20,,243.281,,1000.0
SGLang,0.5.14,NVIDIA B200,post,sglang_tilelang_mhc_post,DeepseekV4ForCausalLM,1024,2,4,4096,20,,304.426,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/mla/sglang/0.5.14/context_mla_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,mla_dtype,kv_cache_dtype,num_heads,batch_size,isl,tp_size,step,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,mla_context,trtllm_mla,bfloat16,bfloat16,1,1,1,64,0,,256.178,,1000.0
SGLang,0.5.14,NVIDIA B200,mla_context,trtllm_mla,bfloat16,bfloat16,1,1,1024,64,0,,405.908,,1000.0
SGLang,0.5.14,NVIDIA B200,mla_context,trtllm_mla,bfloat16,bfloat16,1,1,10240,64,0,,700.2407999999999,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/mla/sglang/0.5.14/generation_mla_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,mla_dtype,kv_cache_dtype,num_heads,batch_size,isl,tp_size,step,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,mla_generation,trtllm_mla,bfloat16,bfloat16,1,1,1,64,1,,249.116,,1000.0
SGLang,0.5.14,NVIDIA B200,mla_generation,trtllm_mla,bfloat16,bfloat16,1,1,1,64,1023,,522.559,,1000.0
SGLang,0.5.14,NVIDIA B200,mla_generation,trtllm_mla,bfloat16,bfloat16,1,1,1,64,127,,346.071,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/mla_bmm/sglang/0.5.14/mla_bmm_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,bmm_dtype,num_tokens,num_heads,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,mla_gen_post,sglang_sgl_kernel_bmm_fp8,fp8,1,1,,243.905,,1000.0
SGLang,0.5.14,NVIDIA B200,mla_gen_post,sglang_sgl_kernel_bmm_fp8,fp8,1,12,,243.714,,1000.0
SGLang,0.5.14,NVIDIA B200,mla_gen_post,sglang_sgl_kernel_bmm_fp8,fp8,1,128,,243.329,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/moe/sglang/0.5.14/moe_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,moe_dtype,num_tokens,hidden_size,inter_size,topk,num_experts,moe_tp_size,moe_ep_size,distribution,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,moe,sglang_flashinfer_trtllm_moe,bfloat16,1,2048,512,8,256,1,1,balanced,,688.924,,1000.0
SGLang,0.5.14,NVIDIA B200,moe,sglang_flashinfer_trtllm_moe,bfloat16,1,2048,512,8,256,1,1,power_law_1.01,,490.709,,1000.0
SGLang,0.5.14,NVIDIA B200,moe,sglang_flashinfer_trtllm_moe,bfloat16,1,2048,512,8,256,1,1,power_law_1.2,,458.592,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/quantize/sglang/0.5.14/computescale_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,m,k,quant_dtype,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,compute_scale,sglang,1,1024,fp8,,264.515,,1000.0
SGLang,0.5.14,NVIDIA B200,compute_scale,sglang,1,10240,fp8,,256.113,,1000.0
SGLang,0.5.14,NVIDIA B200,compute_scale,sglang,1,12288,fp8,,264.819,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/quantize/sglang/0.5.14/scale_matrix_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,m,k,quant_dtype,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,scale_matrix,sglang,1,1024,fp8,,264.738,,1000.0
SGLang,0.5.14,NVIDIA B200,scale_matrix,sglang,1,10240,fp8,,256.129,,1000.0
SGLang,0.5.14,NVIDIA B200,scale_matrix,sglang,1,12288,fp8,,264.294,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/sparse_attention/sglang/0.5.14/dsa_context_module_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,model,architecture,mla_dtype,kv_cache_dtype,gemm_type,num_heads,batch_size,isl,tp_size,step,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,dsa_context_module,sglang_dsa_dense_mha_trtllm_ragged,deepseek-ai/DeepSeek-V3.2,DeepseekV32ForCausalLM,bfloat16,bfloat16,fp8_block,128,1,1024,1,0,,585.1595,,1000.0
SGLang,0.5.14,NVIDIA B200,dsa_context_module,sglang_dsa_dense_mha_trtllm_ragged,deepseek-ai/DeepSeek-V3.2,DeepseekV32ForCausalLM,bfloat16,bfloat16,fp8_block,128,1,1024,1,1,,551.1144,,1000.0
SGLang,0.5.14,NVIDIA B200,dsa_context_module,sglang_dsa_dense_mha_trtllm_ragged,deepseek-ai/DeepSeek-V3.2,DeepseekV32ForCausalLM,bfloat16,bfloat16,fp8_block,128,1,1024,1,1024,,494.4267,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/sparse_attention/sglang/0.5.14/dsa_generation_module_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,model,architecture,mla_dtype,kv_cache_dtype,gemm_type,num_heads,batch_size,isl,tp_size,step,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,dsa_generation_module,sglang_dsa_indexer_trtllm,deepseek-ai/DeepSeek-V3.2,DeepseekV32ForCausalLM,bfloat16,bfloat16,fp8_block,128,1,1,1,1,,327.728,,1000.0
SGLang,0.5.14,NVIDIA B200,dsa_generation_module,sglang_dsa_indexer_trtllm,deepseek-ai/DeepSeek-V3.2,DeepseekV32ForCausalLM,bfloat16,bfloat16,fp8_block,128,1,1,1,10000,,356.21,,1000.0
SGLang,0.5.14,NVIDIA B200,dsa_generation_module,sglang_dsa_indexer_trtllm,deepseek-ai/DeepSeek-V3.2,DeepseekV32ForCausalLM,bfloat16,bfloat16,fp8_block,128,1,1,1,1024,,304.919,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/sparse_attention/sglang/0.5.14/dsv4_csa_context_module_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,model,architecture,mla_dtype,kv_cache_dtype,gemm_type,num_heads,batch_size,isl,tp_size,step,compress_ratio,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,dsv4_csa_context_module,compressed_flashmla,deepseek-ai/DeepSeek-V4-Pro,DeepseekV4ForCausalLM,bfloat16,fp8_e4m3,bfloat16,128,1,1,1,0,4,,326.275,,1000.0
SGLang,0.5.14,NVIDIA B200,dsv4_csa_context_module,compressed_flashmla,deepseek-ai/DeepSeek-V4-Pro,DeepseekV4ForCausalLM,bfloat16,fp8_e4m3,bfloat16,128,1,1,1,1,4,,479.72166666666664,,1000.0
SGLang,0.5.14,NVIDIA B200,dsv4_csa_context_module,compressed_flashmla,deepseek-ai/DeepSeek-V4-Pro,DeepseekV4ForCausalLM,bfloat16,fp8_e4m3,bfloat16,128,1,1,1,10000,4,,463.895,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/sparse_attention/sglang/0.5.14/dsv4_csa_generation_module_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,model,architecture,mla_dtype,kv_cache_dtype,gemm_type,num_heads,batch_size,isl,tp_size,step,compress_ratio,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,dsv4_csa_generation_module,compressed_flashmla,deepseek-ai/DeepSeek-V4-Pro,DeepseekV4ForCausalLM,bfloat16,fp8_e4m3,bfloat16,128,1,1,1,1,4,,244.348,,1000.0
SGLang,0.5.14,NVIDIA B200,dsv4_csa_generation_module,compressed_flashmla,deepseek-ai/DeepSeek-V4-Pro,DeepseekV4ForCausalLM,bfloat16,fp8_e4m3,bfloat16,128,1,1,1,1024,4,,534.0245,,1000.0
SGLang,0.5.14,NVIDIA B200,dsv4_csa_generation_module,compressed_flashmla,deepseek-ai/DeepSeek-V4-Pro,DeepseekV4ForCausalLM,bfloat16,fp8_e4m3,bfloat16,128,1,1,1,10240,4,,471.00600000000003,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/sparse_attention/sglang/0.5.14/dsv4_csa_topk_calib_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,model,architecture,mla_dtype,kv_cache_dtype,gemm_type,num_heads,batch_size,isl,tp_size,step,compress_ratio,score_mode,power__base,power__head,power_limit__base,power_limit__head
SGLang,kernel-level,NVIDIA B200,dsv4_csa_topk_calib,topk_transform_v1,deepseek-ai/DeepSeek-V4-Pro,DeepseekV4ForCausalLM,fp8_e4m3,fp8_e4m3,fp8_block,128,1,1,1,0,4,v1_flat,,0.0,,1000.0
SGLang,kernel-level,NVIDIA B200,dsv4_csa_topk_calib,topk_transform_v1,deepseek-ai/DeepSeek-V4-Pro,DeepseekV4ForCausalLM,fp8_e4m3,fp8_e4m3,fp8_block,128,1,1,1,0,4,v1_top_last,,0.0,,1000.0
SGLang,kernel-level,NVIDIA B200,dsv4_csa_topk_calib,topk_transform_v1,deepseek-ai/DeepSeek-V4-Pro,DeepseekV4ForCausalLM,fp8_e4m3,fp8_e4m3,fp8_block,128,1,1,1,1,4,v1_flat,,0.0,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/sparse_attention/sglang/0.5.14/dsv4_hca_context_module_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,model,architecture,mla_dtype,kv_cache_dtype,gemm_type,num_heads,batch_size,isl,tp_size,step,compress_ratio,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,dsv4_hca_context_module,compressed_flashmla,deepseek-ai/DeepSeek-V4-Pro,DeepseekV4ForCausalLM,bfloat16,fp8_e4m3,bfloat16,128,1,1,1,0,128,,372.981,,1000.0
SGLang,0.5.14,NVIDIA B200,dsv4_hca_context_module,compressed_flashmla,deepseek-ai/DeepSeek-V4-Pro,DeepseekV4ForCausalLM,bfloat16,fp8_e4m3,bfloat16,128,1,1,1,1,128,,543.88725,,1000.0
SGLang,0.5.14,NVIDIA B200,dsv4_hca_context_module,compressed_flashmla,deepseek-ai/DeepSeek-V4-Pro,DeepseekV4ForCausalLM,bfloat16,fp8_e4m3,bfloat16,128,1,1,1,10000,128,,510.95425,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/sparse_attention/sglang/0.5.14/dsv4_hca_generation_module_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,model,architecture,mla_dtype,kv_cache_dtype,gemm_type,num_heads,batch_size,isl,tp_size,step,compress_ratio,power__base,power__head,power_limit__base,power_limit__head
SGLang,0.5.14,NVIDIA B200,dsv4_hca_generation_module,compressed_flashmla,deepseek-ai/DeepSeek-V4-Pro,DeepseekV4ForCausalLM,bfloat16,fp8_e4m3,bfloat16,128,1,1,1,1,128,,245.1135,,1000.0
SGLang,0.5.14,NVIDIA B200,dsv4_hca_generation_module,compressed_flashmla,deepseek-ai/DeepSeek-V4-Pro,DeepseekV4ForCausalLM,bfloat16,fp8_e4m3,bfloat16,128,1,1,1,1024,128,,279.623,,1000.0
SGLang,0.5.14,NVIDIA B200,dsv4_hca_generation_module,compressed_flashmla,deepseek-ai/DeepSeek-V4-Pro,DeepseekV4ForCausalLM,bfloat16,fp8_e4m3,bfloat16,128,1,1,1,10240,128,,297.8966666666667,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/sparse_attention/sglang/0.5.14/dsv4_paged_mqa_logits_module_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,model,architecture,mla_dtype,kv_cache_dtype,gemm_type,num_heads,batch_size,isl,tp_size,step,compress_ratio,power__base,power__head,power_limit__base,power_limit__head
SGLang,kernel-level,NVIDIA B200,dsv4_paged_mqa_logits_module,deep_gemm.fp8_paged_mqa_logits,sgl-project/DeepSeek-V4-Flash-FP8,DeepseekV4ForCausalLM,fp8_e4m3,fp8_e4m3,fp8_block,64,1,1,1,0,4,,243.179,,1000.0
SGLang,kernel-level,NVIDIA B200,dsv4_paged_mqa_logits_module,deep_gemm.fp8_paged_mqa_logits,sgl-project/DeepSeek-V4-Flash-FP8,DeepseekV4ForCausalLM,fp8_e4m3,fp8_e4m3,fp8_block,64,1,1,1,1,4,,244.755,,1000.0
SGLang,kernel-level,NVIDIA B200,dsv4_paged_mqa_logits_module,deep_gemm.fp8_paged_mqa_logits,sgl-project/DeepSeek-V4-Flash-FP8,DeepseekV4ForCausalLM,fp8_e4m3,fp8_e4m3,fp8_block,64,1,1,1,10000,4,,250.794,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/sparse_attention/sglang/0.5.14/glm5_dsa_attn_module_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,model,architecture,mla_dtype,kv_cache_dtype,gemm_type,num_heads,batch_size,isl,tp_size,step,compress_ratio,power__base,power__head,power_limit__base,power_limit__head
SGLang,kernel-level,NVIDIA B200,glm5_dsa_attn_module,flash_mla_sparse_fwd,nvidia/GLM-5.2-NVFP4,GlmMoeDsaForCausalLM,bfloat16,fp8_e4m3,fp8_block,64,1,1,1,1,1,,942.4306666666666,,1000.0
SGLang,kernel-level,NVIDIA B200,glm5_dsa_attn_module,flash_mla_sparse_fwd,nvidia/GLM-5.2-NVFP4,GlmMoeDsaForCausalLM,bfloat16,fp8_e4m3,fp8_block,64,1,1,1,1024,1,,272.5563333333333,,1000.0
SGLang,kernel-level,NVIDIA B200,glm5_dsa_attn_module,flash_mla_sparse_fwd,nvidia/GLM-5.2-NVFP4,GlmMoeDsaForCausalLM,bfloat16,fp8_e4m3,fp8_block,64,1,1,1,1048575,1,,277.932,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/sparse_attention/sglang/0.5.14/glm5_mqa_logits_module_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,model,architecture,mla_dtype,kv_cache_dtype,gemm_type,num_heads,batch_size,isl,tp_size,step,compress_ratio,power__base,power__head,power_limit__base,power_limit__head
SGLang,kernel-level,NVIDIA B200,glm5_mqa_logits_module,deep_gemm.fp8_mqa_logits,nvidia/GLM-5.2-NVFP4,GlmMoeDsaForCausalLM,fp8_e4m3,fp8_e4m3,fp8_block,64,1,1,1,1,1,,922.168,,1000.0
SGLang,kernel-level,NVIDIA B200,glm5_mqa_logits_module,deep_gemm.fp8_mqa_logits,nvidia/GLM-5.2-NVFP4,GlmMoeDsaForCausalLM,fp8_e4m3,fp8_e4m3,fp8_block,64,1,1,1,1024,1,,571.239,,1000.0
SGLang,kernel-level,NVIDIA B200,glm5_mqa_logits_module,deep_gemm.fp8_mqa_logits,nvidia/GLM-5.2-NVFP4,GlmMoeDsaForCausalLM,fp8_e4m3,fp8_e4m3,fp8_block,64,1,1,1,1048575,1,,280.77840000000003,,1000.0aic-core/src/aiconfigurator_core/systems/data/b200_sxm/sparse_attention/sglang/0.5.14/glm5_topk_module_perf.parquet
modified rows - full CSV: framework,version,device,op_name,kernel_source,model,architecture,mla_dtype,kv_cache_dtype,gemm_type,num_heads,batch_size,isl,tp_size,step,compress_ratio,score_mode,power__base,power__head,power_limit__base,power_limit__head
SGLang,kernel-level,NVIDIA B200,glm5_topk_module,fast_topk_transform_fused,nvidia/GLM-5.2-NVFP4,GlmMoeDsaForCausalLM,fp8_e4m3,fp8_e4m3,fp8_block,64,1,1,1,1,1,flat,,0.0,,1000.0
SGLang,kernel-level,NVIDIA B200,glm5_topk_module,fast_topk_transform_fused,nvidia/GLM-5.2-NVFP4,GlmMoeDsaForCausalLM,fp8_e4m3,fp8_e4m3,fp8_block,64,1,1,1,1,1,top_last,,0.0,,1000.0
SGLang,kernel-level,NVIDIA B200,glm5_topk_module,fast_topk_transform_fused,nvidia/GLM-5.2-NVFP4,GlmMoeDsaForCausalLM,fp8_e4m3,fp8_e4m3,fp8_block,64,1,1,1,1024,1,flat,,0.0,,1000.0Artifact Contents
|
Sanity Check Chart Generation Report📥 Download all sanity charts from workflow artifacts New perf data files were detected in this PR. Please use the link above to Below is a report of whether the chart generation was successful for each op. Chart Generation Report for system: b200_sxm, backend: sglang, backend_version: 0.5.14
|
1f47c48 to
130f081
Compare
4b85114 to
3dee700
Compare
1088503 to
6bc4f24
Compare
Signed-off-by: Kai Ma <kaim@nvidia.com>
Signed-off-by: kaim-eng <kaim@NV-8QHBYK4.localdomain>
6bc4f24 to
c8be175
Compare
|
Independent Arrow audit of the 24 files at c8be175 against the merge-base (2ed278a, the #1590 merge): every non-power column is value- and order-identical to stock, row counts unchanged, 1. 7,966 partial pairs (
They are exactly the stock rows with 2. The rebase onto current #1533 (b31d899) re-collected the SGLang 0.5.14 fp8_block DeepGEMM rows after moving the UE8M0 weight-scale pack out of the timed loop. Against this PR's stock: 29,526 rows have new latency on
|
Important
Dependency and zero-sentinel update
This pure-data PR depends on #1590. Merge #1590 first, then rebase/update this PR onto the resulting
mainbefore merging it.Latest signed-off data head:
6bc4f243cb5c7bf20a4e8e27d90819a41f186d93.The latest commit normalizes every unavailable
power/power_limitcell in the changed power-enabled tables to typed0.0. Row coverage, measured values, logical keys, and physical row order are unchanged.Summary
Pure-data PR constructed from the then-current
maindata snapshot (095f58a51c4ca8e61b66ec108d86f223f8d559ce). This revision rebuilds the B200 SGLang 0.5.14 power overlay from the third-pass campaign pinned to the same backend image used for the on-stock perf-only collection.float64power/power_limitonly through exact logical-key matches0.0/0.0sentinel pairs and excludes collection-only keysNo code, documentation, sidecar, or other artifact is included in this PR.
Exact-image provenance
cb43a2f412983fee709c087882ea3438ab5cb8f50c6d722ecbbbce39c477adceae2687eb7245d22clmsysorg/sglang:v0.5.14-cu130@sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3.sqsh) SHA-256:b3ecf5c8fcd521dcca67ceaa4c786d0b73b6adabdc5e08885c092b7c64c4551c49e384ce9d304648e9959666ecb8ce8cd98d0debBroad collection passes used
umb-b200-*8x-B200 allocations and parallelized independent single-GPU cases across the available GPUs. Final exact-key retries used isolated 1-GPU allocations where the operation did not require multiple GPUs; the custom-allreduce collection retained its required multi-GPU topology.The only approved key normalization is the one-to-one HCA model alias
sgl-project/DeepSeek-V4-Pro-FP8->deepseek-ai/DeepSeek-V4-Profor the two HCA tables. No performance field is relabeled or changed.Coverage for the packaged stock tables
power/power_limitpairs: 0Twenty-one files are byte-identical to the prior PR head. Three tables changed on
mainafter the original collection snapshot, so their power was rejoined by the complete logical key while preserving every current-main row, field, latency value, type, and physical row order:0.0pairsmoe_perf.parquetdsa_context_module_perf.parquetdsa_generation_module_perf.parquetNo power was inferred for new or changed current-main keys that were absent from the collection snapshot.
Overlay policy
The output is constructed from the stock tables, then power is joined by the complete logical key. Later isolated retries take precedence only for the same exact key. This makes the stock support surface authoritative: collected-only configurations are not inserted, and stock latency or configuration values cannot be overwritten.
Power-quality checks
power > 0andpower_limit > 0; unavailable pairs are exactly0.0/0.0Rebase and validation
maindata snapshot095f58a51c4ca8e61b66ec108d86f223f8d559ce; rebase onto the post-fix(collector): normalize missing power metrics to zero #1590mainis required before merge6bc4f243cb5c7bf20a4e8e27d90819a41f186d930.0sentinel pairs, 0 null or partial pairsThe PR is structurally independent of #1534. Targeted support-matrix validation with #1590 produced 118
PASS, 2HYBRID_PASS, and 22 unrelatedFAILrows, with zero null-access errors.