[TRTLLM-14022][feat] BREAKING: Remove python modules and tests for legacy TensorRT backend#15918
Conversation
dbe01c5 to
e6e3c0a
Compare
e6e3c0a to
6a4cdca
Compare
|
/bot run --disable-fail-fast |
📝 WalkthroughWalkthroughThis PR removes the TensorRT engine-building backend from TensorRT-LLM. It deletes builder, layers, graph-rewriting, and per-architecture model modules; removes ChangesTensorRT backend removal
Estimated code review effort: 4 (Complex) | ~75 minutes Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
|
PR_Github #57955 [ run ] triggered by Bot. Commit: |
|
/bot run --disable-fail-fast |
|
PR_Github #58959 [ run ] triggered by Bot. Commit: |
|
PR_Github #58959 [ run ] completed with state
|
chienchunhung
left a comment
There was a problem hiding this comment.
Thanks for checking the broader tree.
One item I'd like to flag: tests/unittest/others/test_leak.py calls APIs removed here, including Builder, net_guard, graph_rewriting, and LLaMAForCausalLM. Because those lookups occur inside the test body, collection succeeds, but running this test—or the documented full unit suite—fails at tllm.Builder(). Please remove the test if its legacy coverage is no longer applicable, or migrate it before merge. However, it seems like test_leak.py may not be exercised in regular CI; not sure why.
There may be other consumers as dormant examples, helpers, and legacy integration paths that are not currently selected by the checked-in CI configuration. This is not blocking this PR from merging, but would be great to track them in follow-up PRs.
9ad11ab to
f18278b
Compare
|
/bot run --only-multi-gpu-test --disable-fail-fast |
|
PR_Github #59083 [ run ] triggered by Bot. Commit: |
|
PR_Github #59083 [ run ] completed with state
|
Remove the TensorRT execution backend from the Python package so that `import tensorrt_llm` and the PyTorch backend run without the `tensorrt` package installed. The `tensorrt` pip dependency itself is dropped in the follow-up C++/packaging step; `requirements.txt` is unchanged here. - Delete the TensorRT graph/build/runtime machinery: builder, network, functional graph ops, layers/, plugin/, the legacy models/ implementations, the runtime/ session and ModelRunner* classes, _tensorrt_engine, TRT quant layers and tools, and the trtllm-build/refit/prune CLIs. - Keep import paths stable: partly-TensorRT modules (functional.py, models/modeling_utils.py, _common.py, runtime/, bench/build) are gutted in place so the TensorRT-free symbols used by the PyTorch backend stay put; no `_torch/` import churn. - Drop the TensorRT public API surface (Builder/BuildConfig/TrtLlmArgs/...); `LLM(backend="tensorrt")` now raises and `trtllm-serve --backend` accepts pytorch/_autodeploy only (api_stability serve-CLI reference updated). - Remove or de-TRT the tests that exercised the deleted code (disaggregated trt-backend tests, TRT bench/openai e2e params, mixed llmapi unittest suites) and delist them from test-db/qa/waives; shared test harnesses keep their signatures for the PyTorch twins. - Port test_gather_generation_logits_cuda_graph to the PyTorch accuracy suite; repoint examples/apps to the PyTorch LLM; regenerate ruff-legacy-baseline.json. Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>
f18278b to
42ce5e0
Compare
Test conclusionPre-merge
Post-merge
Test / GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU1-GEN1-NODE1-GPU4-Post-Merge-2 / test_e2e[disagg_upload-gen_only-gb200_gpt-oss-120b-fp4_1k1k_con64_ctx1_tp1_gen1_tp4_eplb0_mtp0_ccb-NIXL] – GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU1-GEN1-NODE1-GPU4-Post-Merge-2.perf.test_perf_sanity Test / GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU1-GEN1-NODE1-GPU4-Post-Merge-5 / test_e2e[disagg_upload-gen_only-gb200_gpt-oss-120b-fp4_8k1k_con4_ctx1_tp1_gen1_tp4_eplb0_mtp0_ccb-NIXL] – GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU1-GEN1-NODE1-GPU4-Post-Merge-5.perf.test_perf_sanity Test / GB200-8_GPUs-2_Nodes-PyTorch-PerfSanity-Node2-GPU8-Post-Merge-2 / test_e2e[aggr_upload-ctx_only-gb200_deepseek-r1-fp4_128k8k_con1_ctx1_pp8_gen1_tep8_eplb0_mtp3_ccb-NIXL] – GB200-8_GPUs-2_Nodes-PyTorch-PerfSanity-Node2-GPU8-Post-Merge-2.perf.test_perf_sanity Test / GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU2-GEN1-NODE2-GPU8-Post-Merge-5 / test_e2e[disagg_upload-gen_only-gb300_glm-5-fp4_8k1k_con1_ctx1_dep2_gen1_tep8_eplb0_mtp3_ccb-NIXL] – GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU2-GEN1-NODE2-GPU8-Post-Merge-5.perf.test_perf_sanity Test / GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1 / test_e2e[disagg_upload-e2e-gb300_deepseek-r1-fp4_128k8k_con256_ctx1_pp4_gen1_dep8_eplb0_mtp1_ccb-NIXL] – GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1.perf.test_perf_sanity Test / GB300-36_GPUs-9_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU2-GEN1-NODE8-GPU32-Post-Merge-2 / test_e2e[disagg_upload-gen_only-gb300_glm-5-fp4_8k1k_con512_ctx1_dep2_gen1_dep32_eplb0_mtp3_ccb-NIXL] – GB300-36_GPUs-9_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU2-GEN1-NODE8-GPU32-Post-Merge-2.perf.test_perf_sanity Test / GB200-4_GPUs-PyTorch-Post-Merge-1 / test_nvfp4_4gpus_online_eplb[moe_backend=TRTLLM] – GB200-4_GPUs-PyTorch-Post-Merge-1.accuracy.test_llm_api_pytorch.TestNemotronV3Super |
|
/bot skip --comment "full test analysis can be found in #15918 (comment) and we can regard it as passed" |
juney-nvidia
left a comment
There was a problem hiding this comment.
Approve due to that this PR has already received the approvals before and just get invalidated due to the changes of Code owner file.
|
PR_Github #59167 [ skip ] triggered by Bot. Commit: |
|
PR_Github #59167 [ skip ] completed with state |
Overview
This PR removes the legacy TensorRT execution backend from the
tensorrt_llmPythonpackage so that
import tensorrt_llmand the PyTorch backend run without thetensorrtpackage installed. It is step 1.5 of the staged TensorRT-backend removal(after the examples #15763, docs #15767, tests #15810, and Triton C++ backend #15907
removals). Dropping the
tensorrtpip dependency itself is deferred to the C++/packagingstep;
requirements.txtis unchanged here.held TensorRT code is gutted in place — the TensorRT parts are deleted and the
TensorRT-free symbols the PyTorch backend / VisualGen still need stay exactly where
they were, so consumer import paths under
_torch/**are unchanged and the diff staysremoval-shaped.
deletions break — TRT-only tests are deleted, mixed suites are de-TRT'd (PyTorch parts
kept, shared harness signatures unchanged), and all affected
test-db/qa/waivesentries are delisted (§ H–I).
import tensorrt_llmgreen withtensorrtblocked; AST import-graphscan over the surviving package = 0 imports of any deleted module or symbol; all
changed
.pybyte-compile; full test-list cross-reference (4k+ entries) = 0 danglingfiles/functions;
pre-commitfully green (incl.ruff-legacywith the regeneratedbaseline and the test-list validators).
_on_trt_backendremoval,_ModelFormatKindremoval), QiJune (plugin/relocation;quantize_by_modelopt.py/functional.pyfull removals deferred to follow-ups), brb-nv (LoRA sign-off + commentstyle), tburt-nv (coverage:
test_gather_generation_logits_cuda_graphported, resttracked as follow-ups) — see § K.
A. Gutted-in-place core (TensorRT stripped, TRT-free symbols kept)
functional.pyimport tensorrt as trt,graph_rewriting/networkimports,_common.default_net/default_trtnet/precision, and all graph-op machinery (DimRange, the graphTensorclass and every functional op). Kept backend-agnostic enums (RotaryScalingType,PositionEmbeddingType,AttentionMaskType,AllReduceStrategy,AllReduceFusionOp), param classes (AllReduceParams,MoEAllReduceParams), andRopeEmbeddingUtils. The only added line isfrom torch import Tensor, repointing the retainedOptional[Tensor]annotations off the deleted graphTensor.tensorrtgraph API and have no surviving consumer; the retained enums/params/RopeEmbeddingUtilsare TRT-free and still imported by_torch/ VisualGen at the same path. Full relocation (~55 consumer repoints) deferred per review (§ K).models/modeling_utils.pySpeculativeDecodingMode,QuantConfig,LayerQuantConfig,PretrainedConfig,Gemma2/Gemma3ConfigGroup; removed the TRT model base classes (PretrainedModel,DecoderModelForCausalLM), per-arch helpers, and engine-build/load helpers.quantization/functional.pypreprocess_weights_for_mixed_gemm; alltensorrt-based quant graph ops dropped._common.py_initno longer loads the TensorRT plugin lib (onlylibth_common.socustom ops + MPI init); deleteddefault_net/default_trtnet/set_network/precision, engine (de)serialization helpers, and thenetglobal._is_building/_BuildingFlagkept with a TODO — they are the Python half of theIS_BUILDINGPython↔C++ contract, removed together with the C++isBuilding()in the C++ decouple step.import tensorrt_llmmust work withouttensorrt; the graph-context machinery is TRT-only._common.py_initre-dlopenslibs/libtensorrt_llm.sowithctypes.RTLD_GLOBAL(non-Windows).libtensorrt_llm_nixl_wrapper.so,libtensorrt_llm_ucx_wrapper.so) are dlopen'ed at runtime (transferAgent.cpp:44,RTLD_LAZY) without aDT_NEEDEDonlibtensorrt_llm.so(would be circular) — they resolveTllmExceptionetc. from the process-global symbol table. Python extensions loadRTLD_LOCAL; onmainthe promotion was a hidden side effect of_load_plugin_lib()'sCDLL(libnvinfer_plugin_tensorrt_llm.so, RTLD_GLOBAL)(plugin lib NEEDS libtensorrt_llm.so). Deleting the plugin load broke everyCacheTransceiver/NIXL path (undefined symbol: _ZTI…TllmException, x86_64 + aarch64). Repro'd + fix validated locally via ctypes for both wrappers._utils.pyimport tensorrt as trtand all TRT dtype/version helpers (str_dtype_to_trt,trt_dtype_to_*,np_dtype_to_trt,torch_dtype_to_trt,dims_array,numpy_array,trt_version/trt_gte,is_same_dtype); prunedtrt.DataTypebranches fromTensorWrapper. AddedCUASSERT+ thecuda.bindings.runtime/cudartimport guard, relocated from the deletedruntime/generation.py.trttypes with no surviving consumer;CUASSERTis a TRT-free CUDA-error checker still needed by four_torchconsumers (§ D).bench/build/build.pytrtllm-bench buildengine command (build_command,apply_build_mode_settings) and itsBuildConfig/_tensorrt_engineimports; keptget_model_config/get_benchmark_engine_settings.DEFAULT_MAX_BATCH_SIZE/DEFAULT_MAX_NUM_TOKENSnow readTorchLlmArgs.model_fields[...].default(wasBuildConfig).trtllm-benchauto-tuning. Defaults track the args class so they can't drift.B. Package
__init__/ re-exports / runtime relocation / packaging__init__.pyfunctionalgraph API (Tensor,constant,net_guard),Builder/BuilderConfig/build/BuildConfig,Network,Parameter,Module,PluginBase,str_dtype_to_trt/torch_dtype_to_trt,TrtLlmArgs;from ._common import _initonly.import tensorrt_llmworks withouttensorrt.llmapi/__init__.pyTrtLlmArgs,BuildConfig, and thebuild_cacheexports (BuildCacheConfig) from imports +__all__.models/__init__.pyQuantConfig,LayerQuantConfig,PretrainedConfig,QuantAlgo,SpeculativeDecodingMode) from.modeling_utils;MODEL_MAP = {}kept forautomodelimportability.tensorrt_llm._torch.models.runtime/__init__.pySession/TensorInfo,GenerationSession,SamplingConfig,ModelRunner,ModelRunnerCpp,EncDecModelRunner,MultimodalModelRunner, KV cache manager v1, stopping/logits-processor lists);ModelConfignow imported from the new.model_config; keptPYTHON_BINDINGS+kv_cache_manager_v2.tensorrt_llm.runtime.ModelConfigimport path preserved.runtime/model_config.py(ADDED)ModelConfigdataclass +from_model_config_cpp, relocated ~verbatim from the deletedruntime/generation.py. One deviation:language_adapter_configtypedOptional[Any](the concrete type lived in the deletedlayers)._torchresource manager.serialization.pyTrtLlmArgs,plugin.plugin: PluginConfig, andbuilder: BuildConfigentries from the pickle allow-list.setup.pytrtllm-build,trtllm-prune,trtllm-refitconsole scripts.commands.{build,prune,refit}modules are deleted.C. LLM API (
llmapi/) + telemetryllmapi/llm.pyBaseLLM._get_llm_args_clselse-branch raisesValueError: Unknown backend … Supported backends are 'pytorch' and '_autodeploy'.(wasllm_args_cls = TrtLlmArgs); the entire_TrtLLMclass (workspace,save(), the TRT_build_model) deleted.LLM(backend="tensorrt"/"trt")is rejected at construction; the engine-execution path no longer exists.llmapi/llm.py_on_trt_backendproperty removed (not kept as aFalseshim): workspace init collapsed toself._workspace = None(deadargs.workspace+tempfiledropped), ctx-only guard collapsed toif is_ctx_only:.Falseshims.llmapi/llm.py_ModelInfo/_ModelRuntimeContextplumbing removed (never instantiated;runtime_contexttokenizer read was a dead branch);_validate_args_for_torch_backendnow validates against theTorchLlmArgsfield set, with the comment/docstring reworded to describe allowed args without naming the removed backend.TrtLlmArgsremoval + review wording feedback.llmapi/llm_args.pyTrtLlmArgsclass deleted (fields, validators,_load_config_from_engine/_ckpt); thebackendtelemetry categorical drops'tensorrt';update_llm_args_with_extra_dictstrips allbuild_confighandling; the whole_ModelFormatKindmachinery removed (enum,get_model_format,model_formatproperty + validator) since the format is always HF;TRT_LLMARGS_EXPLICIT_DOCSTRING/TRT_LLM_DOCSTRINGremoved outright.llmapi/llm_utils.pyModelLoader/CachedModelLoader(build pipeline, engine cache staging,get_engine_dir,_build_model, multi-gpu build tasks);CachedModelLoader.__call__is the pytorch/_autodeployHF-download + quant-config path only;BuildConfig/BuildCacheConfig/_ModelInfo/_ModelRuntimeContext/model_formatdropped from imports and__all__.llmapi/build_cache.py(DELETED)BuildCache/CachedStage/CacheRecordand alsoBuildCacheConfig+get_build_cache_config_from_env.enable_build_cachewas aTrtLlmArgs-only field, not inapi_stability). Originally gutted, fully deleted per review.usage/llmapi_config.pygolden_manifest()drops the"TrtLlmArgs": manifest_rows(TrtLlmArgs)entry + its import.usage/usage_lib.py_extract_architecture_class_namedocstring/comments drop the deleted_ModelFormatKind.TLLM_CKPTmentions (logic unchanged).D. Executor +
_torchconsumersexecutor/base_worker.pybuilder.{Engine, ConfigEncoder, EngineConfig}; removed the in-memory-Engineexecutor branch, the engine-config→ModelConfigpopulation, LoRA-plugin and prompt-adapter setup (attrs stayNone);engine: Union[Path, Engine]→Path;ModelConfignow from..runtime.Engine.executor/{executor,worker,ray_gpu_worker,rpc_worker}.pyfrom ..builder import Engine; narrowed theengineparams toPath.Engine-removal rewire._torch/pyexecutor/py_executor.py,_torch/disaggregation/native/transfer.py,_torch/disaggregation/native/bounce/{gather_scatter,impl}.pyfrom tensorrt_llm.runtime.generation import CUASSERT→from tensorrt_llm._utils import CUASSERT.runtime/generation.pyis deleted. The twobounce/files arrived onmainafter the original scan; missing them was the root cause of every disaggregation failure in the first CI run (workers died onModuleNotFoundError)._torch/distributed/allreduce_helper.py(ADDED)CustomAllReduceHelper(workspace/IPC-buffer size computation) +force_all_reduce_deterministic, relocated ~verbatim from the deletedplugin/plugin.py.plugin/package — relocate the two TRT-free survivors and delete the package._torch/distributed/ops.py,_torch/custom_ops/torch_custom_ops.py,tests/microbenchmarks/{all_reduce,minimax_all_reduce}.pyfrom tensorrt_llm.plugin.plugin import CustomAllReduceHelper→ the newallreduce_helperhome._torch/models/checkpoints/hf/qwen3_moe_weight_mapper.pymodels.modeling_utils.DecoderModelForCausalLMimport;modelannotatednn.Module.nn.Module._torch/auto_deploy/llm_args.pyBuildConfigimport and the vestigialbuild_configfield +ensure_no_build_configvalidator.TrtLlmArgs.build_config; without thisLLM(backend="_autodeploy")raisedImportError.E. CLI + benchmarking + api_stability
commands/serve.py--backend→click.Choice(["pytorch", "_autodeploy"]); alltensorrt/trtbranches,BuildConfig/TrtLlmArgs/_tensorrt_engineusage removed; theChoiceWithAliasclass deleted (its only purpose was thetrt→tensorrtvalue alias); the no-op"build_config": Nonedict entry removed (defensive.pop("build_config")guards kept). CLI defaults (max_batch_size/max_num_tokens/max_beam_width) readTorchLlmArgs.model_fields[...].default;max_seq_lendefault →None(same valueBuildConfigcarried).tests/unittest/api_stability/references/trtllm_serve_cli.yamlbackend: type:narrowedChoice(['_autodeploy', 'pytorch', 'tensorrt'])→Choice(['_autodeploy', 'pytorch']).commands/eval.py--backend→click.Choice(["pytorch"]);_tensorrt_engine/BuildConfig+ the tensorrt branch removed; defaults fromTorchLlmArgs.model_fields.commands/bench.pybuild_commandimport + registration removed.trtllm-bench buildsubcommand.bench/benchmark/__init__.py_tensorrt_engineimport dropped; defaultllm_cls→PyTorchLLM.bench/benchmark/utils/asynchronous.pyfrom .._tensorrt_engine import LLM→from tensorrt_llm import LLM.LLM.bench/benchmark/utils/general.pyALL_SUPPORTED_BACKENDSdrops"tensorrt".F. Logging, profiling, LoRA, misc consumers, evaluate
logger.pytrt_loggerproperty,_trt_loggerfield, theset_levelTRT branch, the module-levelimport tensorrt, and the TRT-severity column ofseverity_map(now[python_level]+ optional polygraphy);_tensorrt_enginemodule abbreviation dropped. Zerotensorrtreferences remain — no lazy import /TYPE_CHECKINGshim.profiler.check_gpt_mem_usage) is deleted (next row), so the earlier lazy-import gut was superseded by full deletion.profiler.pycheck_gpt_mem_usage(legacy TRT-engine memory estimator; no in-repo callers even onmain) + its now-unused_is_building/tracebackimports. Zerotensorrtreferences remain.tensorrt/trt_loggerdependencies intologger.py/profiler.py.lora_manager.pyload_hf_lora+load_nemo_lora(both mutate a TRTPretrainedModelin place) and theColumnLinear/split_matrix_tp/pad_vocab_sizeimports. The PyTorch chain (load_torch_lora→load_torch_hf_lora/load_torch_nemo_lora,LoraManagerruntime weight loading) is untouched.lora_helper.use_lora, invoked exclusively by deleted TRT model classes — verified 0 live callers.lora_helper.pyuse_lorawrapper;LoraConfig+ the helpers the_torchpath imports are untouched.quantization/quantize_by_modelopt.pycombine_medusa_weight; the mllama config workaround) + 3 now-unused imports. Full file removal deferred (§ K).models.medusa.weight/models.mllama.config.serve/openai_server.py_tensorrt_engineimport + theisinstance(_tensorrt_engine.LLM)health branch dropped;LLMfromllmapi.llm.disaggregated_params.pyimport tensorrt as trt(atensorrt_libspreload hack) dropped.__init__.py(next row).__init__.py_preload_tensorrt_libs(): atry: import tensorrt / except ImportError: passshim, run before the firsttensorrt_llm.bindingsimport.bindingsextension has aDT_NEEDEDonlibnvinfer.so.10until the Phase C decouple; onmainthe module-levelimport tensorrtin_utils.py/disaggregated_params.pypreloaded it from thetensorrt_libswheel. Removing every such import brokeimport tensorrt_llmin CI (ImportError: libnvinfer.so.10), where TRT exists only as the pip wheel and not on the system loader path (local validation passed because the dev container has TRT installed system-wide). The guard keeps import working when thetensorrtPython package is blocked/absent in envs with system TRT libs.evaluate/{lm_eval,mmlu,cnn_dailymail,covost2,json_mode_eval,longbench_v2}.pyfrom .._tensorrt_engine import LLM;Union[LLM, PyTorchLLM]hints narrowed toPyTorchLLM.examples/apps/{chat,fastapi_server}.pyfrom tensorrt_llm._tensorrt_engine import LLM→from tensorrt_llm import LLM;BuildConfigusage replaced by direct LLM kwargs (max_batch_size/max_input_len/max_num_tokens/max_beam_width).BaseLlmArgsfield defaults verified identical to the oldBuildConfigdefaults.G. Docstring / comment style (review)
All prose added during the gut was trimmed per review: module docstrings that did not exist
on
mainwere not (re)introduced, removal-narrative comments were dropped, and the addedcomments/docstrings avoid double-backticks and em-dashes where the surrounding file style is
plain (
lora_manager.py,models/__init__.py,models/modeling_utils.py,runtime/model_config.py,_common.py) — brb-nv's uniformity nit, applied diff-wide.H. Test-suite changes (this PR deletes the symbols, so it owns the breakage)
Deleted test files (9 of the 205 deletions)
tests/integration/defs/accuracy/test_llm_api.py_tensorrt_engineimport breakspytest --co); PyTorch twintest_llm_api_pytorch.pycarries the surviving coverage (§ K for the coverage-review disposition).tests/integration/defs/llmapi/test_llm_e2e.py_tensorrt_engine.LLM+BuildConfig).tests/integration/defs/disaggregated/test_configs/disagg_config{,_gen_only,_ctxtp2_gentp1}_trt_backend.yaml(3)backend: trtdisagg configs; their tests are deleted (below).tests/unittest/llmapi/apps/_test_openai_multi_chat.py,_test_openai_consistent_chat.py_tensorrt_engine.LLM+BuildConfigand serve with--backend trt; theirtest_e2e.pywrappers also deleted; neither scheduled anywhere.tests/unittest/tools/plugin_gen/(2)tools.plugin_gen.core(its tests were removed in #15810).tests/microbenchmarks/{build_time_benchmark,build_time_dashboard}.py+README.md(3)build_time_benchmark.pyimports the deletedBuildConfig/build/AutoModelForCausalLM+tensorrt;build_time_dashboard.pyis its only driver; the README documents only this benchmark. Executed bytest_e2e.py::test_build_time_benchmark_sanity(also removed, with itsl0_h100.ymlpost-merge-TRT entry — that stage is commented out in Jenkins, so this was latent, not an active failure).Integration defs
disaggregated/test_disaggregated.pytest_disaggregated_{single_gpu,multi_gpu,benchmark_gen_only}_trt_backend+ theirconfig_maprows. The single-gpu one ran unwaived in pre-merge CI and would turn red the moment the backend rejection landed.test_e2e.pytest_trtllm_bench_sanity/test_trtllm_bench_latency_sanity; dropped the TRT param fromtest_trtllm_bench_iteration_logand thetrt_backendparam fromtest_trtllm_bench_llmapi_launch; removed the[trt]param from the 5 openai/serve wrappers; removed the engine-build paths fromtrtllm_bench_prolog(now returns(model_path, dataset_path)) andBenchRunner(always--backend pytorch); rewrotetest_trtllm_bench_request_rate_and_concurrencyto pytorch (was--backend tensorrt, actively scheduled); droppedbuild --helpfromtest_trtllm_bench_help_sanity; deleted the two TRT app-test wrappers.llmapi/test_llm_api_qa.py+llmapi/_run_llmapi_llm.pytest_llm_args_type_tensorrt;test_llm_args_loggingkeeps only the pytorch half; the launcher script rewritten pytorch-only (noBuildConfig/engine save).accuracy/accuracy_core.py,tests/unittest/utils/util.pyaccuracy/test_llm_api_pytorch.pyTestLlama3_1_8BInstruct::test_gather_generation_logits_cuda_graph(PyTorch port of the removed TRT test:gather_generation_logits=True+cuda_graph_config=CudaGraphConfig(); unique cuda-graph×gather-logits coverage; review follow-up, tburt-nv).Unittest suites (de-TRT'd;
pytest --costays green)llmapi/test_llm.pyfast_build). All 15 cross-imported harness/fixture symbols kept with unchanged signatures (llm_test_harness,llm_get_stats*,llm_return_logprobs*,get_model_path,llama_model_path,prompts,cnn_dailymail_path,global_kvcache_config*, …) — the PyTorch twins (test_llm_pytorch.py,test_llm_multi_gpu_pytorch.py) and ~30apps/_test_*files import them and run unchanged.llmapi/test_llm_args.pyTorchLlmArgs; 2 backend-agnostic YAML tests converted toTorchLlmArgs;LoraConfigrepointed totensorrt_llm.lora_helper. Scheduled pre-merge.llmapi/test_grpc.pyLLM;fast_build=Truedropped ×2 (gRPC service coverage kept).llmapi/test_llm_quant.pyllmapi/test_llm_download.pyLLM; the deletedenable_build_cachearg dropped.llmapi/test_llm_telemetry.pyTestTelemetryTRTBackend+ the_tensorrt_engineimport deleted; PyTorch telemetry classes untouched.llmapi/test_llm_utils.pytest_ModelLoader/test_CachedModelLoader(engine building,_ModelFormatKind) deleted;test_LlmArgs_default_gpus_per_nodeconverted toTorchLlmArgs; explicit imports.llmapi/test_memory_profiling.pyBuildConfigimport (collection breaker in an active pre-merge block) replaced by explicit values;BaseLlmArgsdefaults verified identical (2048/None/1/8192).llmapi/run_llm.py,run_llm_with_postproc.pyllmapi/apps/_test_openai_{chat,completions,misc,reasoning,multi_gpu}.py,_test_trtllm_serve_top_logprobs.py,_test_llm_chat.py,_test_llm_server.pyif backend == "trt"branches/skips and TRT-only extra-options fixtures deleted; the twoapps.{chat,fastapi_server}tests use LLM kwargs instead ofBuildConfig.usage/test_llmapi_config_capture.pyquant_configis only aPrivateAttronTorchLlmArgs— no pydantic-field carrier remains).dynamo/test_imports.py("tensorrt_llm.llmapi", "BuildConfig")import-assert tuple removed.utils/test_logger.pytest_has_trt_loggerdeleted (asserted the deletedLogger.trt_logger).conftest.py_maybe_force_rayTRT-LLM-class guard collapsed (permanently false once_tensorrt_engineis gone).I. Test lists / CI metadata
test-db/l0_{a10,a100,a30,dgx_h100,dgx_h200,gh200,h100,l40s}.yml*_trt_backendentries (incl. the pre-merge pytorchl0_a10one), the[trt]/TRT-param e2e entries,test_trtllm_bench_sanity/latency_sanity,TestTelemetryTRTBackend, theaccuracy/test_llm_api.py+llmapi/test_llm_e2e.pyentries, and the medusa/multimodal entries in the still-active TRT-labelled blocks. Entries for de-TRT'd-but-kept files (e.g.test_llm_quant.py) stay.qa/llm_function_core.txt,qa/llm_function_l20.txttest_llm_api_qatensorrt entry + a pre-existing duplicate line); added the portedtest_gather_generation_logits_cuda_graphentry tollm_function_core.txt(was first added to the l20 list where the removed TRT entry lived, then moved per StanleySun639's review — L20 is a deprecated config slated for deletion).waives.txt[trt]openai waives,iteration_log[TRT-…], the orphaned multimodaltp:1-bs:1waive) — the duplicated-waives and AST validators gate this..test_durationsruff-legacy-baseline.jsonDead
backend: tensorrttest-db blocks whose Jenkins stages are commented out, and therelocation of backend-neutral tests out of TRT-labelled CI stages + stage retirement in
L0_Test.groovy, are deliberately not in this PR — tracked intrt-testlist-audit.md(Buckets 2–4) andplan.md.J. Deleted files (205) — by area
Every deleted file serves the TensorRT execution/build path and has no surviving
consumer (AST import-graph + deleted-symbol scan over the surviving package, re-run
after each rebase and after the test sweep).
models/**config.py/model.py/convert.py/weight.py,unet/**,enc_dec, multimodal encoders,generation_mixin,model_weights_loader) + package__init__s. Kept:models/__init__.py,models/modeling_utils.py(both gutted).runtime/**session.py,generation.py(ModelConfig+CUASSERTrelocated),model_runner{,_cpp}.py,enc_dec_model_runner.py,multimodal_model_runner.py,kv_cache_manager.py(v1),medusa_utils.py,redrafter_utils.py,memory_pools/**,processor_wrapper/**.layers/**tools/**plugin_gen/**),multimodal_builder.py,onnx_utils.py.builder.py,network.py,graph_rewriting.py,module.py,parameter.py,python_plugin.py,top_model_mixin.py.commands/**build.py,prune.py,refit.py.plugin/**plugin.py+__init__.py— survivors relocated to_torch/distributed/allreduce_helper.py(§ D).quantization/**layers.py,quantize.py.llmapi/build_cache.py_tensorrt_engine/LLMmirror.tools/plugin_gen/×2).K. Review-driven changes (chronological)
BuildConfigimport (ImportErroronbackend="_autodeploy"); thelogger.pylazy-TRT path (later superseded by full deletion of thetrt_loggerpath oncecheck_gpt_mem_usagewas found dead)._on_trt_backendremoved outright; the always-HF_ModelFormatKindmachinery removed. Hisnvidia-modeloptrequirements-drop suggestion rejected — modelopt is a core PyTorch-backend dependency (~9 non-quantize_by_modeloptmodules import it).plugin/package deleted with survivors relocated (done); fullquantize_by_modelopt.pyremoval deferred to the examples-coordination PR; fullfunctional.pyrelocation (~55_torchrepoints) deferred to a dedicated refactor.bounce/files' danglingruntime.generationimports caught and fixed — the root cause of every disaggregation failure in the first CI run (the run's 105 failures all mapped to this + the not-yet-pushed test sweep + the api_stability yaml; seetrt-testlist-audit.md§ CI-failure triage).test_llm_args.py,test_memory_profiling.py, the bench request-rate test) would break the moment the PR landed.load_torch_lorachain is untouched); comment-style nit applied diff-wide (§ G).test_gather_generation_logits_cuda_graphported (§ H);test_logprobsalready covered by pytorch unittest harnesses; W8A16 weight-only + fp8-rowwise are TRT in-flight-quantization feature gaps (the pytorch backend loads pre-quantized checkpoints only); aggregated Ulysses cp2 looks feasible but unvalidated — all tracked as follow-ups inplan.md.Summary by CodeRabbit
Description
Test Coverage
PR Checklist
Please review the following before submitting your PR:
PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.
PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.
Test cases are provided for new code paths (see test instructions)
If PR introduces API changes, an appropriate PR label is added - either
api-compatibleorapi-breaking. Forapi-breaking, includeBREAKINGin the PR title.Any new dependencies have been scanned for license and vulnerabilities
CODEOWNERS updated if ownership changes
Documentation updated as needed
Update tava architecture diagram if there is a significant design change in PR.
The reviewers assigned automatically/manually are appropriate for the PR.
Please check this after reviewing the above items as appropriate for this PR.
GitHub Bot Help
To see a list of available CI bot commands, please comment
/bot help.