Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
3735 commits
Select commit Hold shift + click to select a range
bedad85
[None][feat] AutoDeploy: Fix hardcoded configs (#14943)
taylor-yb-lee Jun 7, 2026
47666de
[#13718][feat] AutoDeploy MoE all-to-all: cache + runtime max-tokens …
greg-kwasniewski1 Jun 7, 2026
428cc3e
[TRTLLM-13177][doc] Add Nemotron 3 Ultra doc (#14964)
nv-guomingz Jun 7, 2026
dcd4e90
[#10710][feat] Make explicit CLI flags take precedence over --config …
marinayanov Jun 7, 2026
8be182d
[https://nvbugs/6260907][fix] unwaive test (#15058)
bo-nv Jun 8, 2026
b8d17d7
[None][chore] Increase GB200-4_GPUs-PyTorch shards (#14836)
tburt-nv Jun 8, 2026
71debd5
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 8, 2026
0e0ee27
[TRTLLM-12648][test] implement disagg cancellation canary thread (#15…
chienchunhung Jun 8, 2026
98a88f7
[TRTLLM-12507][feat] Cudagraph support for routed-expert MoE LoRA wit…
brb-nv Jun 8, 2026
86f33e6
[https://nvbugs/6245317][test] set Harmony tiktoken env for GPT-OSS d…
dongfengy Jun 8, 2026
b4d44d3
[https://nvbugs/6153955][test] unwaive GPT-OSS w4 DP4 CUTLASS (#14884)
dongfengy Jun 8, 2026
ca2bc5e
[None][perf] kv_cache_manager_v2: batch block-key SHA-256 hashing (#1…
lancelly Jun 8, 2026
2cad6db
[TRTLLM-13259][ci] Merge DGX_H100 DeepSeek and GptOss stages (#15035)
QiJune Jun 8, 2026
5fa68a4
[None][infra] Waive 11 failed cases for main in post-merge 2765 (#15080)
ZhanruiSunCh Jun 8, 2026
2632530
[None][infra] Waive 3 failed cases for main in post-merge 2765 (#15082)
ZhanruiSunCh Jun 8, 2026
7e49baa
[None][test] waive weekly qa ci failure cases (#15077)
crazydemo Jun 8, 2026
02f6b2f
[None][feat] AutoDeploy: propagate layer_type hint across pattern-mat…
greg-kwasniewski1 Jun 8, 2026
28dc25e
[None][test] Waive 15 failed cases for main in QA CI (#15056)
tensorrt-cicd Jun 8, 2026
2febb37
[None][infra] Waive 1 failed cases for main in pre-merge 41894 (#15089)
ZhanruiSunCh Jun 8, 2026
09c21b6
[TRTLLM-13262][ci] Move non-default-feature tests to post merge (#15038)
QiJune Jun 8, 2026
c93c63d
[None][feat] Enable disk cache config for KV cache v2 (#14845)
reasonsolo Jun 8, 2026
6dee167
[https://nvbugs/6185446][fix] Add warmup for trtllm-gen fmha JIT kern…
pengbowang-nv Jun 8, 2026
b14794c
[https://nvbugs/6162940][chore] Unwaive fixed test (#15078)
longlee0622 Jun 8, 2026
2bf4d3d
[None][perf] Support Gemma RMSNorm + interleaved mRoPE in fused_qk_no…
nv-guomingz Jun 8, 2026
9eaa468
[None][test] Half K25 Agg Multi Round to Solve Timeout Issue (#15083)
chenfeiz0326 Jun 8, 2026
9af8a16
[None][infra] Reduce Docker image layer count in release stage (#14972)
tburt-nv Jun 8, 2026
cb01607
[#14828][feat] AutoDeploy: support multi KV cache memory pool in trtl…
MrGeva Jun 8, 2026
15d06c0
[None][doc] Refine Nemotron Ultra doc (#15113)
nv-guomingz Jun 8, 2026
8036cde
[None][infra] Waive TestQwen3NextInstruct nvfp4 cases (#15086)
mzweilz Jun 8, 2026
1998324
[https://nvbugs/6248757][fix] Avoid running all reduce in aux stream …
tensorrt-cicd Jun 8, 2026
900d069
[https://nvbugs/6221483][fix] AutoDeploy: Fix Eagle metadata host syn…
govind-ramnarayan Jun 8, 2026
9827c21
[None][feat] add FLUX visual generation examples (#14987)
karljang Jun 8, 2026
b222246
[https://nvbugs/6261164][fix] In the kvcache insert transform (`_Inse…
tensorrt-cicd Jun 8, 2026
c1e9b00
[https://nvbugs/6211189][fix] Lower the reference to 46.5 (matching c…
tensorrt-cicd Jun 9, 2026
bfb4537
[None][refactor] split VisualGen pipeline and model configs (#14956)
bobboli Jun 9, 2026
5e3af40
[TRTLLM-11457][feat] Async Ulysses pipeline (Enabled for LTX-2 + WAN)…
luyiyun1021 Jun 9, 2026
09ebc59
[TRTLLM-11548][doc] Add Qwen3.5 deployment guide doc (#15111)
nv-guomingz Jun 9, 2026
a33dec7
[https://nvbugs/6181383][fix] Build inner text/vision/audio sub-confi…
tensorrt-cicd Jun 9, 2026
041ed83
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 9, 2026
2490441
[https://nvbugs/6273850][chore] waive TestQwen3_5_4B::test_bf16 for a…
tburt-nv Jun 9, 2026
64497e2
[None][doc] Add docs for AutoDeploy transforms (#15122)
bmarimuthu-nv Jun 9, 2026
9349fcc
[None][infra] Waive 4 failed cases for main in post-merge 2769 (#15140)
ZhanruiSunCh Jun 9, 2026
28845dd
[https://nvbugs/6227203][fix] Remove redundant TikTokenTokenizer shim…
tianyuxbear Jun 9, 2026
a197a5e
[None][fix] tunable_fp4_quantize: rename misnamed kwarg + add real SF…
luyiyun1021 Jun 9, 2026
e9402ab
[None][test] Fix gen_only missing prev_device_step_time race in perf …
tensorrt-cicd Jun 9, 2026
6bf3e49
[None][test] Fix disagg test result dir (#14864)
fredricz-20070104 Jun 9, 2026
a7e4a9b
[TRTLLM-13332][test] Remove TestLlama4ScoutInstruct tests (#15144)
QiJune Jun 9, 2026
6f7aea5
[https://nvbugs/6266705][fix] Gate FlashInfer GDN kernels to supporte…
nv-guomingz Jun 9, 2026
6254f3a
[https://nvbugs/6255037][fix] Count DSA indexer K-cache correctly as …
eopXD Jun 9, 2026
a90fd15
[https://nvbugs/6194812][test] Update llm_perf_core.yml to require a …
yufeiwu-nv Jun 9, 2026
34a94ee
[TRTLLMINF-112][infra] Reduce the waiting time between check node is …
EmmaQiaoCh Jun 9, 2026
b852703
[None][infra] Waive 1 failed cases for main in pre-merge 41821 (#15135)
ZhanruiSunCh Jun 9, 2026
178f4e6
[None][infra] CBTS Layer 3: pass test-db via Artifactory instead of e…
crazydemo Jun 9, 2026
45e2523
[TRTLLM-13264][feat] Add native bias epilogue to NVFP4 GEMM (#15053)
luyiyun1021 Jun 9, 2026
2ee96cf
[https://nvbugs/6278380][unwaive] unwaive ad cases (#15148)
crazydemo Jun 9, 2026
104b9d7
[https://nvbugs/6244474][fix] AutoDeploy: Remove llama perf test from…
MrGeva Jun 9, 2026
ba6ba1f
[https://nvbugs/6212252][fix] Select CUTLASS MoE backend on non-Black…
xxi-nv Jun 9, 2026
d620851
[TRTLLM-13302][feat] Register NVIDIA Wan2.2-T2V quantized checkpoints…
zhenhuaw-me Jun 9, 2026
484e6c9
[None][chore] add VisualGen team as the codeowner of the VisualGen At…
zhenhuaw-me Jun 9, 2026
487330e
[None][feat] Default on FlashInferTrtllmGenAttention (#14618)
yihwang-nv Jun 9, 2026
58fbfb9
[None][infra] Test DFW with BSL branch (#14597)
yuanjingx87 Jun 9, 2026
451dbb8
[TRTLLM-12214][perf] customMoeRoutingKernel: lower BLOCK_SIZE to 128,…
xwang233 Jun 9, 2026
f0ba8c7
[TRTLLM-12214][perf] DeepGemmFusedMoE: skip redundant data expand via…
xwang233 Jun 9, 2026
736dc22
[TRTLLM-12648][test] implement disagg cancellation load thread (#15124)
chienchunhung Jun 9, 2026
f1d39ea
[None][fix] Fix regression from SageAttention kernel: Use static sche…
xrq-phys Jun 9, 2026
680c6c4
[TRTLLM-12467][feat] EPD improvements (#13864)
venkywonka Jun 9, 2026
0edbbfe
[None][feat] Expose stored block-hash chain to KV cache connector (#1…
jthomson04 Jun 9, 2026
358505c
[#12805][fix] Fall back to local cache when loading tokenizer for gat…
1MrazorT1 Jun 9, 2026
3ddef66
[None][feat] Support partial RoPE fusion for Hopper kernels in XQA fo…
DomBrown Jun 9, 2026
b0216c6
[None][infra] Add nv-xtf, rahul-steiger-nv, tedzhouhk, tensorrt-cicd …
ZhanruiSunCh Jun 9, 2026
98393f3
[None][feat] Add Prometheus metrics for prompt cache, speculative dec…
vedularaghu Jun 9, 2026
f39a79c
[None][chore] Unwaive DSV32 helix tests (#14871)
brb-nv Jun 9, 2026
e0a909a
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 9, 2026
ddef2d0
[None][fix] unset UCX_TLS=tcp (#15008)
tburt-nv Jun 9, 2026
884520c
[None][feat] Port 13 AutoDeploy custom models to sharding IR + opt th…
greg-kwasniewski1 Jun 9, 2026
48d2b89
[None][chore] Make image paths absolute in blog22 (#15177)
brb-nv Jun 9, 2026
0f7e1db
Fix PyExecutor FPM iteration timing (#14922)
tedzhouhk Jun 9, 2026
ffcd8e6
[#13816][feat] AutoDeploy: Optimize gpt-oss-120b perf (#14202)
taylor-yb-lee Jun 9, 2026
9a7f76f
[None][fix] Register Multimodal Placeholders for Qwen3.5 MoE VLM Serv…
anurags25 Jun 9, 2026
edfc667
[None][feat] Weight trtllm-bench AR/AL averages by output length (#14…
zhaoyangwang-nvidia Jun 10, 2026
3b945f7
[TRTLLM-13052][feat] Enable TRTLLM moe backend for nemotron-h BF16 ck…
Wanli-Jiang Jun 10, 2026
8e40515
[None][fix] Fix and unwaive nemotron related bugs (#15085)
Wanli-Jiang Jun 10, 2026
9c6cb35
[https://nvbugs/6140226][test] Add DFlash coverage for Qwen3.5 MoE va…
yingguo-trt Jun 10, 2026
2763557
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 10, 2026
9c100bb
[None][test] temporarily waive Cosmos3 B200 failures (#15195)
bobboli Jun 10, 2026
5741389
[NVBUG-6241842][fix] DSA DSL atom-split: guard against MTP draft next…
limin2021 Jun 10, 2026
2148a3e
[#11423][feat] AutoDeploy: Basic Disagg Support (#14057)
govind-ramnarayan Jun 10, 2026
9bc4321
[https://nvbugs/6280060][fix] Scope disagg-ctx cache-transfer quorum …
tensorrt-cicd Jun 10, 2026
27b52b3
[None][test] Add e2e example tests for flux1/2, ltx2, wan_i2v, and co…
chang-l Jun 10, 2026
2878b30
[#12632][feat] Add pipeline cache support for AutoDeploy (#13729)
nvchenghaoz Jun 10, 2026
31e730a
[None][test] Add support for nemotron_3_ultra_550b_nvfp4 model in per…
yufeiwu-nv Jun 10, 2026
b206f68
[None][feat] Indexer TopK: single-block / multi-pass radix (#14268)
dcampora Jun 10, 2026
3b46728
[None][fix] Clear workspace in run_mla_generation to avoid potential …
yihwang-nv Jun 10, 2026
90cb7ff
[None][chore] Unwaive AutoDeploy accuracy tests (#14971)
bmarimuthu-nv Jun 10, 2026
2d196f7
[None][test] Increase kv_transfer_timeout_ms for b200 deepseek-r1 dis…
tensorrt-cicd Jun 10, 2026
62c6521
[None][feat] Enable MTP for Step-3.7 NVFP4 and port Step-3.7VL vision…
kaiyux Jun 10, 2026
9635f7d
[https://nvbugs/6266370][fix] Fix MAX_UTILIZATION reuse token budget …
brb-nv Jun 10, 2026
69f5add
[https://nvbugs/6272573][ci] Unwaive skipped test (#15118)
2ez4bz Jun 10, 2026
74d8a48
[https://nvbugs/6245279][fix] AutoDeploy: Unwaive accuracy tests (#15…
galagam Jun 10, 2026
6db3233
[TRTLLM-12491][feat] Align VisualGen serve request schema with Visual…
zhenhuaw-me Jun 10, 2026
dab3400
[None][test] Add MLA chunked-prefill SM dispatch regression coverage …
DhineshPonnarasan Jun 10, 2026
0be1447
[TRTLLM-12648][test] enable disagg cancellation stress test (#15174)
chienchunhung Jun 10, 2026
03ed843
[None][feat] Preserve cache_salt string in KV cache events (#13051)
jthomson04 Jun 10, 2026
309c764
[https://nvbugs/6104831][fix] Port dataTransceiver shared_ptr<LlmRequ…
chienchunhung Jun 10, 2026
7301075
[None][fix] Fix AutoDeploy transform docs generation (#15228)
bmarimuthu-nv Jun 10, 2026
bb74da1
[None][feat] Targeted warmup-waste cleanup (#14609)
dominicshanshan Jun 11, 2026
0d44f33
[None][fix] Remove TLLM_RUBIN_FEATURES (#15143)
yuxianq Jun 11, 2026
cab198d
[https://nvbugs/6108994][fix] add kv_transfer_timeout_ms to avoid tim…
bo-nv Jun 11, 2026
af5a22e
[TRTLLM-12657][infra] Fix periodic-junit in unittest pytest (#14075)
yiqingy0 Jun 11, 2026
5e3f012
[https://nvbugs/6143883][fix] Preserve ip:port for trtllm-serve visua…
JunyiXu-nv Jun 11, 2026
228829c
[TRTLLM-12958][feat] Enable gen-only spec dec (#14546)
bo-nv Jun 11, 2026
55bbcf4
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 11, 2026
016fb4c
[None][test] Remove 78 closed-bug waive entries for main (#15061)
tensorrt-cicd Jun 11, 2026
205920d
[https://nvbugs/6278399][fix] Add x86_64 path using CU_MEM_HANDLE_TYP…
tensorrt-cicd Jun 11, 2026
a622e30
[TRTLLM-11538][feat] Blackwell custom mask fmha support (#12958)
sunnyqgg Jun 11, 2026
01d8ccb
[None][infra] Waive 6 failed cases for main in post-merge 2773 (#15250)
ZhanruiSunCh Jun 11, 2026
d96c0df
[None][feat] Enhance CuteDSL NVF4 MOE (#15092)
liyuhannnnn Jun 11, 2026
6b7d8cf
[None][infra] Waive 3 failed cases for main in post-merge 2772 (#15253)
ZhanruiSunCh Jun 11, 2026
84b349f
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 11, 2026
835fd61
[None][test] Update K2.5 andGLM-5 into CI Perf Test (#14960)
chenfeiz0326 Jun 11, 2026
54dec4f
[None][feat] enable GQA and cross-attention for attn2d (#14961)
NVShreyas Jun 11, 2026
d3748a3
[#12230][fix] Add bounds checking in autotuner _find_nearest_profile …
mihai-chiorean Jun 11, 2026
80f18fe
[None][refactor] visual_gen Attention: drop redundant enable_ulysses …
luyiyun1021 Jun 11, 2026
9ab3501
[None][fix] Generalize FP8 checkpoint loading for Qwen3.5 (#15067)
amukkara Jun 11, 2026
19b5d0e
[#13858][fix] AutoDeploy fix the piecewise vlm issue (#14006)
nvchenghaoz Jun 11, 2026
be7117c
[TRTLLM-12507][feat] Cudagraph support for per-expert lora in Cutlass…
brb-nv Jun 11, 2026
c859650
[None][test] Remove stale perf sanity waives (#15269)
cascade812 Jun 11, 2026
81f7baf
[None][infra] Waive 8 failed cases for main in pre-merge 42699 (#15273)
ZhanruiSunCh Jun 11, 2026
d19b6a8
[None][fix] Install processor-output validation filter at module impo…
aswinvisva Jun 11, 2026
67e1097
[None][infra] Waive 10 failed cases for main in pre-merge 42753 (#15275)
ZhanruiSunCh Jun 11, 2026
eb5674b
[TRTLLM-12534][fix] Nemotron Nano - properly account for text prompts…
moraxu Jun 11, 2026
b586ebf
[None][doc] Fix stale --disable_xqa reference in legacy docs (#13395)
Erfandarzi Jun 11, 2026
1b360ee
[TRTLLM-11403][doc] Cache-DiT documentation (#15268)
o-stoner Jun 11, 2026
00ed78c
[#15022][fix] Guided decoding (xgrammar) + EAGLE-3 + draft_len_schedu…
chungen04 Jun 11, 2026
ccc0708
[TRTLLM-12154][test] Add Qwen3-32B FP8 disagg stress test (#14278)
brnguyen2 Jun 11, 2026
aef7d47
[TRTLLM-13141][feat] Add backend-agnostic SourceIdentity gate for wei…
chienchunhung Jun 12, 2026
ae9226e
[None][feat] Add PyTorch reset_prefix_cache API (#14970)
milesial Jun 12, 2026
2dd5c67
[None][fix] Stabilize Mamba replay state update (#14841)
sunnyqgg Jun 12, 2026
5a77356
[None][infra] Waive remaining AutoDeploy Disagg tests until fix lands…
govind-ramnarayan Jun 12, 2026
82ddf75
[None][test] Sunset the old disagg test cases for the qa side (#15290)
fredricz-20070104 Jun 12, 2026
02957e9
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 12, 2026
82d3811
[None][infra] Waive 1 failed cases for main in pre-merge 42836 (#15293)
ZhanruiSunCh Jun 12, 2026
2b3af6d
[None][fix] Fix max_context_length value for attention workspace sizi…
pengbowang-nv Jun 12, 2026
fb7a1d0
[TRTLLM-12038][feat] Add accuracy tests for nemotron-v3-ultra (#14808)
Wanli-Jiang Jun 12, 2026
8e2b7b2
[#14672][fix] AutoDeploy: Vendor OpenELMConfig locally to fix OpenELM…
plapagesse Jun 12, 2026
c323881
[https://nvbugs/6035425][fix] Fix KV cache host splitting logic (#14373)
mikeiovine Jun 12, 2026
85d5e6e
[None][refactor] Move KV cache manager V2 to separate file (#14680)
jiaganc Jun 12, 2026
f18d18d
[TRTLLM-12963][refactor] LTX-2 attention: drop dead k_pe parameter; r…
luyiyun1021 Jun 12, 2026
44550bc
[TRTLLM-10184][chore] Remove legacy XQA precompiled code path (#14941)
pengbowang-nv Jun 12, 2026
82ca2c5
[TRTLLM-35882][feat] cute dsl gvr-top multi-cta optimization (#15198)
limin2021 Jun 12, 2026
db7161b
[None][fix] Revert "Add PyTorch reset_prefix_cache API (#14970)" (#15…
xxi-nv Jun 12, 2026
b03b78f
Revert "[None][test] Add support for nemotron_3_ultra_550b_nvfp4 mode…
tburt-nv Jun 12, 2026
646464b
[https://nvbugs/6309375][test] AutoDeploy: Remove stale fallback test…
govind-ramnarayan Jun 12, 2026
19ae053
[None][fix] AutoDeploy: set enable_spec_decode on ADEngine for disagg…
Shixiaowei02 Jun 12, 2026
380d96a
[TRTLLM-12498][feat] Add support for beam search in disaggregated ser…
athena-nv Jun 12, 2026
be7e978
[None][chore] 2 more WAN multi-gpu tests (#15223)
NVShreyas Jun 12, 2026
57bb6ee
[TRTLLM-12721][feat] Add disagg transfer state consensus (#15139)
chienchunhung Jun 12, 2026
c2b7cd9
[None][infra] Waive 1 failed cases for main in pre-merge 43047 (#15326)
ZhanruiSunCh Jun 12, 2026
cd65070
[#12715][fix] disable NCCL_SYMMETRIC tactic on GB10 (DGX Spark) (#12902)
nv-lschneider Jun 13, 2026
706a91f
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 13, 2026
bb32597
[None][feat] AutoDeploy: Qwen3.5: Apply whielist based sharding and a…
taylor-yb-lee Jun 13, 2026
ec47baa
[https://nvbugs/6293015][fix] Add a delegating `@property def vocab_s…
tensorrt-cicd Jun 13, 2026
4e1776a
[TRTLLM-12842][feat] Maximal LLMAPI capture in usage telemetry (#14398)
venkywonka Jun 13, 2026
1283c6b
[TRTLLM-12427][perf] Qwen2.5/3/3.5-VL Performance Optimization (#11943)
yechank-nvidia Jun 13, 2026
4f46653
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 14, 2026
e6c9964
[TRTLLM-11408][test] Add e2e Tensor Parallel LPIPS tests for VisualGe…
yingguo-trt Jun 15, 2026
221a0e1
[None][infra] Waive 1 failed cases for main in pre-merge 43173 (#15358)
ZhanruiSunCh Jun 15, 2026
aa3236b
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 15, 2026
801cde1
[None][infra] Record CBTS decision to OpenSearch for CI-health monito…
crazydemo Jun 15, 2026
b1ee4ab
[None][feat] MNNVL Performance Optimization and FP8/NVFP4 Quant Fusio…
timlee0212 Jun 15, 2026
26ea499
[None][refactor] Remove TensorRT performance baseline and update to P…
yufeiwu-nv Jun 15, 2026
91a271b
[None][test] Waive 1 failed cases for main in QA CI (#15315)
tensorrt-cicd Jun 15, 2026
870f9b5
[https://nvbugs/6029882][fix] Fix attentionOp fp8 mla kvreuse workspa…
pengbowang-nv Jun 15, 2026
1c069d3
[None][infra] pin pytest and click workaround (#15357)
cascade812 Jun 15, 2026
20b6068
[None][feat] skip-softmax on SM120: TMA-load + sync-MMA warp-speciali…
dcampora Jun 15, 2026
130ae82
[None][fix] Fix beam search log_probs non-determinism with batch_size…
achartier Jun 15, 2026
0d4bab9
[None][fix] Forward secondary_offload_min_priority to KVCacheManager …
Saddss Jun 15, 2026
7cefb4a
[None][chore] Bump version to 1.3.0rc19 (#15188)
yuanjingx87 Jun 15, 2026
feca41c
[TRTLLMINF-103][feat] Keep SLURM timeouts non-retryable (#15183)
dpitman-nvda Jun 15, 2026
35c9704
[TRTLLM-12982][feat] support multi item scoring in LLM.encode (#14693)
ixlmar Jun 15, 2026
f451726
[https://nvbugs/6281014][fix] fix the repeated cute.compile and simpi…
JadoTu Jun 16, 2026
e171875
[None][chore] Integration tests for MoE lora & bugfixes (#15271)
brb-nv Jun 16, 2026
2ef2ea5
[TRTLLM-12339][feat] enable TRTLLM cross attention backend (#15345)
cascade812 Jun 16, 2026
d6967a1
[TRTLLM-12807][test] Guard thop attention kwarg aliases (#15335)
yuxianq Jun 16, 2026
f49d09f
[None][infra] Waive 21 failed cases for main in post-merge 2780 (#15373)
ZhanruiSunCh Jun 16, 2026
b7f3673
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 16, 2026
09449d4
[None][fix] pool-qualify KV cache transfer pending keys (#15272)
chienchunhung Jun 16, 2026
0ec3250
[None][refactor] Enhance pytest integration by updating test node gen…
yufeiwu-nv Jun 16, 2026
72ecb98
[None][test] Waive 1 failed cases for main in QA CI (#15377)
tensorrt-cicd Jun 16, 2026
0b0a03e
[https://nvbugs/312578][fix] split test_cache_transceiver_single_proc…
chuangz0 Jun 16, 2026
e45dda9
[None][infra] Update the new duration base on opensearch result (#15364)
EmmaQiaoCh Jun 16, 2026
f3b718a
[https://nvbugs/6245861][fix] Gate the two ID None-checks on `finish_…
tensorrt-cicd Jun 16, 2026
163be83
[https://nvbugs/6223556][fix] Propagate gen-first ctx usage via aux b…
reasonsolo Jun 16, 2026
9206812
[None][test] Fix Mamba hybrid transceiver helper (#15323)
chienchunhung Jun 16, 2026
81e57e0
[None][feat] Qwen3-VL: support per-request mm_processor_kwargs (#14702)
aswinvisva Jun 16, 2026
08fba40
[TRTLLM-12982][chore] NVTX-annotate logits processor (#15408)
ixlmar Jun 16, 2026
dfad249
[TRTLLM-12339][feat] Support T5 and BART in the PyTorch backend (#13919)
cascade812 Jun 16, 2026
275c172
[TRTLLM-13333][feat] Add prefetch_reuse_blocks and configurable prefe…
reasonsolo Jun 16, 2026
18cb08e
[None][feat] DSv4 prep: attention op plumbing (#15384)
lfr-0531 Jun 17, 2026
2b14cfb
[None][test] Waive 8 failed cases for main in post-merge (#15389)
tensorrt-cicd Jun 17, 2026
500ebf2
[#15182][fix] Fix embedding vocab mask for handling rejection samplin…
chungen04 Jun 17, 2026
c025b86
[None][test] Waive 1 failed cases for main in QA CI (#15320)
tensorrt-cicd Jun 17, 2026
9a081b8
[None][refactor] Refactor Skip Softmax Attention Interface (#14687)
bobboli Jun 17, 2026
fd5de7e
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 17, 2026
9b77c8e
[None][infra] Waive 1 failed cases for main in pre-merge 43656 (#15439)
ZhanruiSunCh Jun 17, 2026
29b228e
[None][infra] Waive 11 failed cases for main in post-merge 2782 (#15395)
ZhanruiSunCh Jun 17, 2026
7e24365
[https://nvbugs/6248837][fix] Densify trtllm-gen fmha warmup grid to …
pengbowang-nv Jun 17, 2026
0593968
[TRTLLM-13378][feat] Drop legacy --extra_visual_gen_options CLI alias…
zhenhuaw-me Jun 17, 2026
2772b99
[TRTLLM-12950][feat] Add MegaMoECuteDsl NVFP4 MoE backend (#14608)
xxi-nv Jun 17, 2026
5fe0a17
[None][perf] DSv4 prep: attention fusion custom ops (#15390)
lfr-0531 Jun 17, 2026
6e0c1c2
[TRTLLM-12669][refactor] Eagle3 sampling: auto-detect greedy fast-pat…
zhaoyangwang-nvidia Jun 17, 2026
0e74256
[TRTLLMINF-137][infra] Skip to create perf report when there is not p…
yiqingy0 Jun 17, 2026
071c287
[https://nvbugs/6270671][fix] Enable multi-block mode for XQA HMMA sp…
tensorrt-cicd Jun 17, 2026
40db402
[TRTLLMINF-113][infra] Add timeout protection to Setup/Initialize sta…
ZhanruiSunCh Jun 17, 2026
821165c
[None][infra] Waive 1 failed cases for main in pre-merge 43720 (#15449)
ZhanruiSunCh Jun 17, 2026
52cbeee
[None][infra] Waive 2 failed cases for main in post-merge 2785 (#15450)
ZhanruiSunCh Jun 17, 2026
a590a2d
[None][perf] executor: avoid deepcopy of prompt_token_ids on enqueue …
lancelly Jun 17, 2026
0ffa09f
[None][infra] Waive 1 failed cases for main in pre-merge 43712 (#15447)
ZhanruiSunCh Jun 17, 2026
d202244
[None][ci] tighten VisualGen CBTS routing (#15259)
zhenhuaw-me Jun 17, 2026
42a3e55
[None][fix] fix tinygemm barrier bug (#15338)
yweng0828 Jun 17, 2026
9e69568
[TRTLLM-12199][feat] WideEP FT: add EPGroupHealth thread-safe rank ma…
chienchunhung Jun 18, 2026
79ea125
[None][infra] Waive 18 failed cases for main in pre-merge 43878 (#15469)
ZhanruiSunCh Jun 18, 2026
08f4bb1
[None][fix] Fix stale sparse attention kwargs (#15460)
bobboli Jun 18, 2026
c390d2f
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 18, 2026
5c8a359
[None][test] Waive 1 failed cases for main in QA CI (#15411)
tensorrt-cicd Jun 18, 2026
1aa232a
[TRTLLM-12807][feat] Add multiple FMHA library support to TRTLLM atte…
yuxianq Jun 18, 2026
c25fa74
[None][infra] Waive 1 failed cases for main in pre-merge 43917 (#15478)
ZhanruiSunCh Jun 18, 2026
4a8b7af
[None][feat] Side-stream for MM encoder (#14322)
2ez4bz Jun 18, 2026
2a18bd4
[None][feat] BREAKING: Add MiniMax-M3 PyTorch backend bring-up with A…
WeiHaocheng Jun 19, 2026
4d44595
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 19, 2026
7060827
[https://nvbugs/6215678][fix] Point `--output-artifact-dir` at a uniq…
tensorrt-cicd Jun 19, 2026
d9041f8
[None][fix] fix CppMambaHybridCacheManager to handle dp dummy request…
bo-nv Jun 19, 2026
30c40dc
[None][test] Waive 5 failed cases for main in post-merge (#15392)
tensorrt-cicd Jun 19, 2026
3baa571
[None][test] Waive 9 failed cases for main in post-merge (#15391)
tensorrt-cicd Jun 19, 2026
b9b132b
[None][test] Waive 5 failed cases for main in QA CI (#15360)
tensorrt-cicd Jun 19, 2026
5e1a28f
[None][test] Waive 8 failed cases for main in QA CI (#15342)
tensorrt-cicd Jun 19, 2026
a76c818
[None][feat] Checkpointing variant of replay for MTP for mamba models…
hnover-nv Jun 19, 2026
3297cb9
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 20, 2026
d3d1b11
[None][test] Waive 23 failed cases for main in QA CI (#15337)
tensorrt-cicd Jun 20, 2026
53b392e
[None][test] Waive 3 failed cases for main in QA CI (#15319)
tensorrt-cicd Jun 20, 2026
6f9e32e
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 21, 2026
a8c5955
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 22, 2026
e47359c
[None][fix] AutoDeploy: Fixed wrong dist_backend AUTO detection when …
MrGeva Jun 22, 2026
416bdb2
[None][test] Waive 2 failed cases for main in QA CI (#15341)
tensorrt-cicd Jun 22, 2026
4d7bf0d
[TRTLLMINF-81][feat] Avoid failed runners on infra retry (#15237)
dpitman-nvda Jun 22, 2026
2e6abd1
[https://nvbugs/6179661][fix] Harden disagg cache transceiver teardow…
chienchunhung Jun 22, 2026
e1135bb
[https://nvbugs/6273846][test] gate GPT-OSS TRTLLM Gen MoE tests to S…
dongfengy Jun 22, 2026
eddaa3a
[None][fix] avoid type checking failures due to pip dependency resolu…
ixlmar Jun 22, 2026
f7dd7ec
[None][feat] VisualGen: async mp4 encode + fixed noise latent via env…
wu6u3tw Jun 22, 2026
6348278
Honor deterministic mode in PyTorch backend
cursoragent Jun 22, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
The diff you're trying to view is too large. We only load the first 3000 changed files.
80 changes: 80 additions & 0 deletions .claude/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,80 @@
# Custom Claude Code Skills & Agents for TensorRT-LLM

## Background: Skills & agents in Claude Code

Claude Code supports two extensibility mechanisms — **skills** and **agents** —
that let teams encode domain expertise into reusable, version-controlled
components.

**Skills** are markdown playbooks that Claude follows step-by-step when
triggered. They are invoked via `/slash-commands` (e.g. `/perf-analysis`) or
matched automatically from natural-language requests. Each skill lives in its
own directory under `.claude/skills/` and can bundle reference materials that
Claude reads during execution. See
[Custom slash commands](https://code.claude.com/docs/en/skills)
for details.

**Agents** (sub-agents) are specialist workers that Claude spawns in a separate
context to handle focused tasks. Each agent has its own system prompt, tool
access, and domain knowledge. Claude delegates to them when it determines a task
fits a specialist's scope, while you can also invoke agents directly. Agent
definitions live under `.claude/agents/`. See
[Custom sub-agents](https://code.claude.com/docs/en/sub-agents)
for details.

## How skills and agents are loaded

For users who are working with Claude Code under TensorRT-LLM project directory,
skills and agents are automatically discovered by Claude Code at startup — no
manual registration needed. Files placed in `.claude/skills/` and
`.claude/agents/` are picked up by convention.

To verify what's loaded, launch Claude Code under TensorRT-LLM project directory
and type `/skills` or `/agents` in the Claude Code prompt to see available
skills and sub-agents.

## How to use skills and agents

There are two ways to trigger skills and agents:

1. **Automatic dispatch** — just describe what you need in plain language
(e.g. "profile this workload", "compile TensorRT-LLM"). Claude Code will
match your request to the appropriate skill or delegate to the right
sub-agent automatically.

2. **Manual invoke** — type `/<skill-name>` (e.g. `/perf-analysis`,
`/trtllm-serve-config-guide`) to explicitly run a skill. For sub-agents, type
`@"<agent-name>" (agent)` (e.g. `@"exec-compile-specialist (agent)"`) to
delegate a task directly. This is useful when you know exactly which workflow you want.

In most cases, automatic dispatch is sufficient — you don't need to memorize
skill or agent names. Manual invoke is there for when you want precise control.

References:
* [Extend Claude with skills](https://code.claude.com/docs/en/skills)
* [Work with subagents](https://code.claude.com/docs/en/sub-agents#work-with-subagents)

## Naming convention

Every skill and agent name uses the format `<prefix>-<descriptive-name>`.
The prefix identifies the primary work area; the descriptive part should be
short and not repeat it.

| Prefix | Domain | Definition |
|---|---|---|
| `ad-` | AutoDeploy | Model onboarding, pipeline debugging, and execution for the AutoDeploy backend |
| `exec-` | Execution infra | Environment setup and job execution (compile, run, container) |
| `kernel-` | Kernel development | Kernel writing, generation, and kernel-specific transforms |
| `perf-` | Performance work | Profiling, analysis, and tuning above the kernel layer (kernel modifications belong under `kernel-`) |
| `trtllm-` | TRT-LLM project workflows | Project-specific workflows: codebase exploration, contribution, dependency upgrades, and serving configuration (static subsystem knowledge belongs in repo docs) |

Guidelines:

* If a skill doesn't fit any prefix, propose a new one and agree on its
boundary before using it.
* Use the prefix of the skill's **primary** domain, even if it orchestrates
across multiple domains.
* Agents follow the same convention.
* Good: `exec-local-compile`, `kernel-cuda-writing`, `perf-host-analysis`
* Bad: `exec-trtllm-compile`, `kernel-cuda-kernel-writing`,
`perf-trtllm-host-analysis`
30 changes: 30 additions & 0 deletions .claude/agent-tests/perf-test-sync/build_prompt.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
"""Promptfoo prompt builder.

Loads perf-test-sync.md, strips Claude Code specific sections, and appends the
user request.

Stripped sections:
- YAML frontmatter between leading `---` markers
- `# Persistent Agent Memory` section and everything after it (Claude Code
memory infrastructure, not relevant to prompt-quality evaluation)
"""

import os
import re

_SCRIPT_DIR = os.path.dirname(os.path.abspath(__file__))
_AGENT_MD = os.path.normpath(os.path.join(_SCRIPT_DIR, "..", "..", "agents", "perf-test-sync.md"))


def _load_agent_body() -> str:
with open(_AGENT_MD, "r", encoding="utf-8") as f:
text = f.read()
text = re.sub(r"\A---\n.*?\n---\n", "", text, count=1, flags=re.DOTALL)
text = text.split("# Persistent Agent Memory", 1)[0].rstrip()
return text


def build(context: dict) -> str:
user_prompt = context["vars"]["prompt"]
agent_body = _load_agent_body()
return f"{agent_body}\n\n## User request\n\n{user_prompt}\n"
Loading