Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
27 commits
Select commit Hold shift + click to select a range
9726490
feat: add Qwen3.5 FP4 B300 AgentX MTP
cquil11 Jul 30, 2026
c8dd660
chore: trigger B300 AgentX sweep
cquil11 Jul 30, 2026
85521aa
perf: expand B300 AgentX preflight range
cquil11 Jul 30, 2026
ad6e2d3
chore: use SGLang v0.5.16 for Qwen B300
cquil11 Jul 30, 2026
a58dcef
chore: collect SGLang cache metrics on B300
cquil11 Jul 30, 2026
8ffac86
fix: remove unsupported SGLang tool-choice flag on B300
cquil11 Jul 30, 2026
87bf252
perf: prepare Qwen B300 cache and DEP tuning
cquil11 Jul 30, 2026
064b790
perf: add B300 HiCache cliff probes
cquil11 Jul 30, 2026
8570002
fix: enforce Qwen HiCache DRAM total
cquil11 Jul 30, 2026
d365cd3
chore: merge current main into B300 submission
cquil11 Jul 30, 2026
4064979
perf: densify B300 Qwen cache crossover
cquil11 Jul 30, 2026
05d97cf
perf: bracket B300 TP2 offload frontier
cquil11 Jul 30, 2026
0b5465c
perf: remove dominated B300 Qwen DEP topology
cquil11 Jul 30, 2026
e15c32d
perf: extend B300 TP2 HiCache tail
cquil11 Jul 30, 2026
3bb4e4d
perf: prune dominated B300 TP4 tail
cquil11 Jul 30, 2026
c5b8203
perf: cap B300 HiCache at valid working sets
cquil11 Jul 30, 2026
a3dac8d
perf: cap B300 TP2 before the cache cliff
cquil11 Jul 30, 2026
79dd0c7
perf: probe B300 TP4 HiCache at the cliff
cquil11 Jul 30, 2026
020a1aa
perf: prune the B300 TP2 no-offload cliff
cquil11 Jul 30, 2026
d2224b2
fix(agentx): refine B300 cache frontier
cquil11 Jul 30, 2026
89cd592
perf(agentx): densify B300 HiCache cliff
cquil11 Jul 30, 2026
c43766e
perf(agentx): densify B300 TP2 cache crossover
cquil11 Jul 30, 2026
e0b8e28
fix(agentx): scale Qwen tokenization by topology
cquil11 Jul 30, 2026
f8b5e3c
perf(agentx): densify B300 TP2 cache boundary
cquil11 Jul 30, 2026
023f543
perf(agentx): refine B300 cache frontier
cquil11 Jul 30, 2026
012e78f
fix(agentx): cap Qwen trace idle gaps
cquil11 Jul 30, 2026
c9187d7
chore(aiperf): pin merged trace idle cap branch
cquil11 Jul 30, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
174 changes: 174 additions & 0 deletions benchmarks/single_node/agentic/qwen3.5_fp4_b300_sglang_mtp.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,174 @@
#!/usr/bin/env bash

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 The PR title and description are English-only, but AGENTS.md line 7 mandates bilingual PR titles (<English title> / <中文标题>) and a Chinese translation section (e.g. ## 中文说明) in the body for every PR. / PR 标题和描述仅为英文,但 AGENTS.md 第 7 行要求每个 PR 的标题必须为双语格式(<English title> / <中文标题>),且正文需包含中文翻译部分(如 ## 中文说明)。请在标题追加中文标题,并在 Summary 后补充对应的中文说明段落。

Extended reasoning...

AGENTS.md explicitly and prominently states (line 7, bolded): "PR and GitHub-issue titles & descriptions must be bilingual — include a Simplified Chinese version in addition to English. Title format: <English title> / <中文标题>. In the PR/issue body, follow the English content with its Chinese translation (e.g. a ## 中文说明 section mirroring the summary...). This applies to every PR and every issue."

PR #2421's title is Add Qwen3.5 FP4 B300 AgentX MTP — English only, with no / <中文标题> suffix. The description body contains only an English ## Summary section:

Add Qwen3.5-397B-A17B NVFP4 AgentX coverage on B300 with SGLang native NEXTN MTP, golden synthetic acceptance, prefix caching, and the 256k trace dataset.

There is no ## 中文说明 section or any Chinese translation anywhere in the title or body, which is a direct, checked-in convention violation rather than a matter of subjective style.

Step-by-step verification:

  1. Open AGENTS.md at the repo root and read line 7 — it defines the exact required title format (<English title> / <中文标题>) and body structure (English content followed by a ## 中文说明 section).
  2. Compare against PR Add Qwen3.5 FP4 B300 AgentX MTP #2421's actual title: Add Qwen3.5 FP4 B300 AgentX MTP. There is no / separator and no Chinese characters anywhere in the title.
  3. Compare against the PR body: it contains exactly one section, ## Summary, entirely in English. No ## 中文说明 (or any Chinese-language section) is present.
  4. Conclusion: both the title-format rule and the body-translation rule from AGENTS.md line 7 are violated by this PR as submitted.

This convention exists so that the bilingual engineering audience described in the Code Conventions section (and referenced tooling like the CODEOWNER sign-off bot) can review PRs and issues without an English-only barrier; it is a checked-in, project-wide requirement rather than a stylistic nicety, so per review guidance it should be surfaced even though it does not affect any runtime or benchmark behavior in this recipe addition.

Impact: No functional or benchmark risk — the .sh script, configs/nvidia-master.yaml entry, and perf-changelog.yaml entry are all unaffected. This is purely a PR-metadata/process gap.

Fix: Update the PR title to Add Qwen3.5 FP4 B300 AgentX MTP / 新增 Qwen3.5 FP4 B300 AgentX MTP 基准测试, and append a ## 中文说明 section to the description body that mirrors the English ## Summary content in idiomatic technical Chinese (per the translation-quality bar and terminology table also given in AGENTS.md), keeping model names, hardware SKUs, framework names, and CLI/env-var identifiers in English.

set -euo pipefail
set -x

# AgentX trace replay for Qwen3.5-397B-A17B NVFP4 on B300 with SGLang
# native NEXTN MTP. Throughput uses the committed golden synthetic AL; evals
# retain real target-model verification.

source "$(dirname "$0")/../../benchmark_lib.sh"

check_env_vars \
MODEL TP CONC EP_SIZE KV_OFFLOADING \
TOTAL_CPU_DRAM_GB RESULT_DIR DURATION

SCHEDULER_RECV_INTERVAL=${SCHEDULER_RECV_INTERVAL:-10}

if [[ -n "${SLURM_JOB_ID:-}" ]]; then
echo "JOB $SLURM_JOB_ID running on ${SLURMD_NODENAME:-unknown}"
fi

if [[ -n "${MODEL_PATH:-}" ]]; then
if [[ ! -d "$MODEL_PATH" || -z "$(ls -A "$MODEL_PATH" 2>/dev/null)" ]]; then
hf download "$MODEL" --local-dir "$MODEL_PATH"
fi
else
hf download "$MODEL"
export MODEL_PATH="$MODEL"
fi
nvidia-smi

export WEKA_LOADER_OVERRIDE=semianalysis_cc_traces_weka_062126_256k
resolve_trace_source
install_agentic_deps

SERVER_LOG="$RESULT_DIR/server.log"
mkdir -p "$RESULT_DIR"

CACHE_ARGS=()
if require_agentic_kv_offload_backend hicache; then
REQUESTED_HICACHE_TOTAL_GB="${HICACHE_TOTAL_CPU_DRAM_GB:-$TOTAL_CPU_DRAM_GB}"
if [ "$REQUESTED_HICACHE_TOTAL_GB" -gt "$TOTAL_CPU_DRAM_GB" ]; then
echo "Error: requested HiCache pool ${REQUESTED_HICACHE_TOTAL_GB} GB exceeds configured capacity ${TOTAL_CPU_DRAM_GB} GB" >&2
exit 1
fi
TOTAL_CPU_DRAM_GB="$REQUESTED_HICACHE_TOTAL_GB"
# SGLang applies --hicache-size independently to Qwen's target KV and
# Mamba pools. Native NEXTN also creates a draft KV pool with the same
# slot count; its one attention layer adds 1/15 of the target KV bytes.
# Reserve 1 GB/rank for page alignment and enforce H * 31/15 per rank.
HICACHE_ALIGNMENT_RESERVE_GB=$TP
HICACHE_USABLE_TOTAL_GB=$((TOTAL_CPU_DRAM_GB - HICACHE_ALIGNMENT_RESERVE_GB))
if [ "$HICACHE_USABLE_TOTAL_GB" -lt 1 ]; then
echo "Error: insufficient DRAM after HiCache alignment reserve" >&2
exit 1
fi
MAX_HICACHE_SIZE_GB=$((HICACHE_USABLE_TOTAL_GB * 15 / TP / 31))
HICACHE_SIZE_GB="${HICACHE_SIZE_GB:-$MAX_HICACHE_SIZE_GB}"
if [ "$HICACHE_SIZE_GB" -lt 1 ] || [ "$HICACHE_SIZE_GB" -gt "$MAX_HICACHE_SIZE_GB" ]; then
echo "Error: HICACHE_SIZE_GB=$HICACHE_SIZE_GB outside 1..$MAX_HICACHE_SIZE_GB" >&2
exit 1
fi
PROJECTED_HICACHE_TOTAL_GB=$(((HICACHE_SIZE_GB * TP * 31 + 14) / 15 + HICACHE_ALIGNMENT_RESERVE_GB))
if [ "$PROJECTED_HICACHE_TOTAL_GB" -gt "$TOTAL_CPU_DRAM_GB" ]; then
echo "Error: projected HiCache use ${PROJECTED_HICACHE_TOTAL_GB} GB exceeds configured capacity ${TOTAL_CPU_DRAM_GB} GB" >&2
exit 1
fi
echo "HiCache CPU pools: ${HICACHE_SIZE_GB} GB target + Mamba + 1/15 draft per rank across TP=${TP}; projected node total ${PROJECTED_HICACHE_TOTAL_GB} GB <= ${TOTAL_CPU_DRAM_GB} GB"
CACHE_ARGS=(
--page-size 64
--enable-hierarchical-cache
--hicache-size "$HICACHE_SIZE_GB"
--hicache-io-backend kernel
--hicache-mem-layout page_first
--hicache-write-policy write_through_selective
)
fi

PARALLEL_ARGS=(
--tp "$TP"
--dp 1
--ep-size "$EP_SIZE"
)
STREAM_INTERVAL=50

# TP4 needs parallel tokenization to keep 256k AgentX warmups below the client
# request timeout. Keep TP2 on SGLang's single-worker default: multi-tokenizer
# startup races with the TP2 HiCache shared-memory initialization path.
TOKENIZER_ARGS=()
if [ "$TP" -ge 4 ]; then
TOKENIZER_ARGS=(--tokenizer-worker-num 6)
fi

# AgentX concurrency counts live session trees rather than individual HTTP
# requests. Leave room for subagent fan-out and avoid spending HBM on graphs
# above the batch sizes that remain useful for this long-context workload.
MAX_RUNNING_REQUESTS=$((2 * CONC))
CUDA_GRAPH_MAX_BS="$CONC"
[ "$CUDA_GRAPH_MAX_BS" -gt 64 ] && CUDA_GRAPH_MAX_BS=64

export TORCH_CUDA_ARCH_LIST="10.0"
export PYTHONNOUSERSITE=1
export NCCL_NVLS_ENABLE=1
export SGL_ENABLE_JIT_DEEPGEMM=false
export SGLANG_ENABLE_FLASHINFER_GEMM=true

if [ "${EVAL_ONLY:-false}" != "true" ]; then
export SGLANG_SIMULATE_ACC_LEN=3.39
export SGLANG_SIMULATE_ACC_METHOD=match-expected
export SGLANG_SIMULATE_ACC_TOKEN_MODE=real-draft-token
fi

SGLANG_CMD=(
python3 -m sglang.launch_server
--model-path "$MODEL_PATH"
--served-model-name "$MODEL"
--host 0.0.0.0
--port "$PORT"
--trust-remote-code
"${PARALLEL_ARGS[@]}"
--enable-symm-mem
--quantization modelopt_fp4
--fp4-gemm-backend flashinfer_cutlass
--kv-cache-dtype fp8_e4m3
--mamba-ssm-dtype bfloat16
--attention-backend trtllm_mha
--moe-runner-backend flashinfer_trtllm
--cuda-graph-max-bs "$CUDA_GRAPH_MAX_BS"
--max-running-requests "$MAX_RUNNING_REQUESTS"
--max-prefill-tokens 16384
--chunked-prefill-size 16384
--mem-fraction-static 0.80
--stream-interval "$STREAM_INTERVAL"
--scheduler-recv-interval "$SCHEDULER_RECV_INTERVAL"
"${TOKENIZER_ARGS[@]}"
--tokenizer-path "$MODEL"
--reasoning-parser qwen3
--tool-call-parser qwen3_coder
--speculative-algorithm NEXTN
--speculative-num-steps 3
--speculative-eagle-topk 1
--speculative-num-draft-tokens 4
--enable-metrics
--enable-cache-report
"${CACHE_ARGS[@]}"
)

printf '%q ' "${SGLANG_CMD[@]}" | tee "$RESULT_DIR/sglang_command.txt"
printf '\n' | tee -a "$RESULT_DIR/sglang_command.txt"
"${SGLANG_CMD[@]}" > "$SERVER_LOG" 2>&1 &
SERVER_PID=$!

capture_cache_metrics() {
{
echo "=== SGLang cache metrics snapshot $(date --iso-8601=seconds) ==="
curl -fsS "http://localhost:$PORT/metrics" 2>/dev/null \
| grep -E '^(sglang:(cache_hit_rate|cached_tokens_total|prompt_tokens_total|hicache_host_used_tokens|hicache_host_total_tokens|token_usage|num_requests_running|num_requests_waiting))' \
|| true
echo "============================================================"
} >> "$SERVER_LOG"
}

wait_for_server_ready --port "$PORT" --server-log "$SERVER_LOG" --server-pid "$SERVER_PID"

capture_cache_metrics
trap capture_cache_metrics EXIT

if [ "${EVAL_ONLY:-false}" = "true" ]; then
run_eval --port "$PORT"
else
build_replay_cmd "$RESULT_DIR"
REPLAY_CMD+=" --trace-idle-gap-cap-seconds 300"
REPLAY_CMD+=" --server-metrics http://localhost:$PORT/metrics"
run_agentic_replay_and_write_outputs "$RESULT_DIR"
fi
29 changes: 29 additions & 0 deletions configs/nvidia-master.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -7176,6 +7176,35 @@ qwen3.5-fp8-b300-sglang-agentic-hicache:
- { tp: 4, ep: 1, kv-offloading: none, conc-list: [1, 2, 4, 8, 16, 32] }
- { tp: 4, ep: 1, kv-offloading: dram, kv-offload-backend: { name: hicache }, conc-list: [16, 32, 48, 64] }

# Qwen3.5 target FP8 KV is 15,360 bytes/token before the native MTP layer.
# TP4 replicates the model's two KV heads and measured 30.5 KB/token node-wide;
# TP2 measured 15.3 KB/token. Against the 256k trace's 88.8k median and 148.8k
# p75 ISL, TP2 no-offload peaks at c32. HiCache overlaps from c32; probes put
# the throughput transition between c34 and c40, while host KV reaches 98% at
# c52 and c56 is the last useful stress point. c64+ warmups only deepen the
# queue and are excluded. TP4 no-offload remains healthy into the c56-72
# neighborhood, which is sampled every four requests around the cache cliff.
# TP4 HiCache is dominated there: c68 sustained 197k tok/s versus 224k for
# no-offload c64, while c76 fell to 158k with host KV full. Clean HiCache c94
# was also dominated by no-offload c56 (210k vs 234k tok/s, with 44 ms vs 7 ms
# median ITL), so the TP4 HiCache arm is excluded.
# DEP8 was rejected after fragmenting the trace across eight radix caches.
qwen3.5-fp4-b300-sglang-agentic-mtp:
image: lmsysorg/sglang:v0.5.16-cu130
model: nvidia/Qwen3.5-397B-A17B-NVFP4
model-prefix: qwen3.5
runner: cluster:b300-nv
precision: fp4
framework: sglang
multinode: false
scenarios:
agentic-coding:
- dram-utilization: 0.80
search-space:
- { tp: 4, ep: 1, spec-decoding: mtp, kv-offloading: none, conc-list: [1, 4, 8, 12, 16, 20, 24, 28, 32, 40, 48, 56, 60, 64, 68, 72] }
- { tp: 2, ep: 2, spec-decoding: mtp, kv-offloading: none, conc-list: [1, 4, 8, 12, 16, 20, 24, 28, 32] }
- { tp: 2, ep: 2, spec-decoding: mtp, kv-offloading: dram, kv-offload-backend: { name: hicache }, conc-list: [32, 34, 36, 38, 40, 44, 48, 52, 56] }

kimik2.5-fp4-b200-vllm-agentic-lmcache:
image: vllm/vllm-openai:v0.22.0
model: nvidia/Kimi-K2.5-NVFP4
Expand Down
9 changes: 9 additions & 0 deletions perf-changelog.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -5339,3 +5339,12 @@
- "Pass --use-chat-template to the benchmark, as required for EAGLE-style speculative decoding"
- "Validated on the node through the real launcher, every request successful: TEP4 concurrency 64 640/640 at 6895 tok/s (11% above the non-MTP arm's 6198 tok/s, mean TTFT 27.2s to 9.6s), and TEP4 concurrency 1 10/10 at 1233 tok/s (83% above the non-MTP arm's 672 tok/s, mean TPOT 12.6ms to 6.5ms)"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2312

- config-keys:
- qwen3.5-fp4-b300-sglang-agentic-mtp
description:
- "Add Qwen3.5-397B-A17B NVFP4 AgentX benchmark on B300 with SGLang native NEXTN MTP"
- "Use the 256k trace dataset and golden synthetic acceptance length 3.39"
- "Pin AIPerf trace idle-gap support and cap per-trace idle gaps at 300 seconds"
- "Image: lmsysorg/sglang:v0.5.16-cu130"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2421
Loading