forked from NVIDIA/TensorRT-LLM
-
Notifications
You must be signed in to change notification settings - Fork 1
Pull requests: nv-guomingz/TensorRT-LLM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[None][perf] Chat frontend fastpath: single template render + suffix token-id cache + to_thread offload
#7
opened Jul 8, 2026 by
lingjiew
Collaborator
Loading…
[None][fix] Align auto-filled max_tokens_in_buffer to tokens_per_block
#5
opened Jul 6, 2026 by
lingjiew
Collaborator
Loading…
[None][fix] Log and validate transceiver runtime selection for hybrid managers
#4
opened Jul 6, 2026 by
lingjiew
Collaborator
Loading…
[None][fix] Stop forcing Mixed hybrid manager for PYTHON transceiver runtime
#3
opened Jul 6, 2026 by
lingjiew
Collaborator
Loading…
[None][fix] Guard V2 mamba state-index tables against ADP dummy-request overflow
#2
opened Jul 6, 2026 by
lingjiew
Collaborator
Loading…
ProTip!
Find all pull requests that aren't related to any open issues with -linked:issue.