Skip to content

Prototype: TensorRT execution provider (linux, CUDA 13.0) - #200

Draft
diegoferigo-rai wants to merge 58 commits into
conda-forge:mainfrom
diegoferigo-rai:diegoferigo/tensorrt-support
Draft

Prototype: TensorRT execution provider (linux, CUDA 13.0)#200
diegoferigo-rai wants to merge 58 commits into
conda-forge:mainfrom
diegoferigo-rai:diegoferigo/tensorrt-support

Conversation

@diegoferigo-rai

Copy link
Copy Markdown
Contributor

Draft / prototype — not meant to be merged. TensorRT packages come from a
local channel, so CI here cannot build it. Scoped to linux; posted as a
starting point to adapt. Relates to #199.

Approach

Optional TensorRT execution provider on the CUDA 13.0 linux builds
(linux-64 and linux-aarch64), following the pytorch-feedstock style of
encoding the accelerator in the build string. A dedicated tensorrt variant
axis (linux) gates it; the _tensorrt marker only appears on the TRT build:

Variant Build string
CUDA 13.0 + TRT cuda130_tensorrt_py310_h…_200
CUDA 13.0 (no TRT) cuda130_py310_h…_200
CUDA 12.9 cuda129_py310_h…_200
CPU cpu_py310_h…_0

The inert tensorrt: true combinations (CPU / CUDA 12.9) are skipped, so no
duplicate packages. The CI matrix is restricted to linux for the prototype.

Changes

  • recipe/recipe.yaml: tensorrt_enabled context gated to CUDA 13.0 via the
    tensorrt axis; _tensorrt appended to the build string only on the TRT build;
    TensorRT host deps libnvinfer-devel + libnvonnxparser-devel under
    if: tensorrt_enabled (run_exports pin the runtime libs); package_contents
    lib check for onnxruntime_providers_tensorrt; skip not linux.
  • recipe/build-cpp.nu: cmake flags onnxruntime_USE_TENSORRT=ON,
    onnxruntime_USE_TENSORRT_BUILTIN_PARSER=ON, onnxruntime_TENSORRT_HOME=$PREFIX.
    Builtin parser -> shared libnvonnxparser, nothing static.
  • recipe/conda_build_config.yaml: tensorrt axis (linux) + local
    channel_sources for the locally-built TensorRT packages.

To adapt

  • Replace the local channel_sources
    (file:///…/staged-recipes/build_artifacts,conda-forge) with the real TensorRT
    packages once they land on conda-forge.
  • The local channel only ships linux-64 artifacts; the CUDA 13.0
    linux-aarch64 TRT build won't solve until aarch64 packages are built.
  • Restore the non-linux platforms and decide how far to extend TensorRT.
  • Confirm the builtin (dynamic libnvonnxparser) vs. static parser choice.

cbourjau added 30 commits August 8, 2026 20:41
(cherry picked from commit bbb7162)
(cherry picked from commit 1729bf1)
(cherry picked from commit a603d7a)
(cherry picked from commit 013a296)
(cherry picked from commit b9e50c1)
(cherry picked from commit b4ddaab)
(cherry picked from commit 0ceeae9)
(cherry picked from commit 8178088)
(cherry picked from commit 0c0eb26)
[cherry picked from commit 256790e; recipe changes only, rerender output dropped]
[cherry picked from commit d3457f3; recipe changes only, rerender output dropped]
(cherry picked from commit 0f79f21)
(cherry picked from commit 2332847)
(cherry picked from commit 9724cd6)
(cherry picked from commit de3e786)
(cherry picked from commit 68ded51)
(cherry picked from commit 9417bea)
[cherry picked from commit 666a846; recipe changes only, rerender output dropped]
(cherry picked from commit 85eb668)
(cherry picked from commit af69e8d)
(cherry picked from commit 542e19a)
cbourjau and others added 26 commits August 8, 2026 20:42
(cherry picked from commit 953701e)
[cherry picked from commit add924b; recipe changes only, rerender output dropped]
(cherry picked from commit 9dcca6c)
(cherry picked from commit ae48510)
(cherry picked from commit dd1d48d)
(cherry picked from commit 72e183d)
(cherry picked from commit 9407828)
(cherry picked from commit 8f432b4)
(cherry picked from commit 66edfee)
<details><summary>Claude's draft</summary>

Adapt the cherry-picked rattler-build migration (conda-forge#171 by @cbourjau) to the
current 1.28.0 feedstock state:

- Bump version to 1.28.0 with the matching sha256 and add patches 0006/0007
  (onnx-less test guards introduced for 1.28.0).
- Re-enable the -novec variants that were skipped while debugging conda-forge#171; this
  required templating the staging output name in the `inherit:` keys.
- Port CoreML execution provider support for osx-arm64 (conda-forge#182):
  onnxruntime_USE_COREML=ON, nlohmann_json host dep, provider test,
  BSD-3-Clause license and vendored license files.
- Port Jetson Thor support (conda-forge#190): add 110-real to the CUDA 13.0 arch list.
- Port nvcc_threads=2 (Windows CUDA 13.0 OOM mitigation).
- Pin cmake <4 (vendored protobuf 3.21 does not build with CMake 4).
- Match main's 1.28.0 runtime deps: drop sympy, add __cuda run dep on the
  python output.
- Add onnxruntime_test --help entry-point test.

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume bceefe2e-0d08-4ad8-9b4c-3518466f1ccc
```
</details>
…026.08.08.17.43.08

Other tools:
- conda-build 26.7.0
- rattler-build 0.72.2
- rattler-build-conda-compat 1.4.19
<details><summary>Claude's draft</summary>

The top-level recipe name is a placeholder (onnxruntime-dummy), so the
conda-smithy linter needs an explicit extra.feedstock-name: onnxruntime.

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume bceefe2e-0d08-4ad8-9b4c-3518466f1ccc
```
</details>
<details><summary>Claude's draft</summary>

All CI jobs failed at cmake configure with "CMAKE_CXX_STANDARD must be at
least 20. It is 17." — 1.28.0 enforces C++20 via
onnxruntime_language_standard_versions.cmake.

The C++17 choice in conda-forge#171 was a workaround for C++20 module-scanning issues
at 1.24.4; upstream 1.28.0 now sets CMAKE_CXX_SCAN_FOR_MODULES=OFF itself
(cmake/CMakeLists.txt:21), so C++20 is safe. This also matches main's
build.py invocation which passed CMAKE_CXX_STANDARD=20.

run_cpp_test.bat stays at /std:c++17: consuming the public headers only
needs C++17 and main is green with that combination.

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume bceefe2e-0d08-4ad8-9b4c-3518466f1ccc
```
</details>
<details><summary>Claude's draft</summary>

Both Windows CUDA jobs failed at configure with "Could not find compiler
set in environment variable CC: cl.exe" while the Windows CPU job (which
skips the CUDA branch) proceeded fine.

Cause: on Windows, nushell exposes the search path as a *list* named
`Path`. The CUDA branch did

    $env.PATH = $"($build_lib_prefix)/bin;($env.PATH)"

which interpolates that list into a string ("[C:\..., C:\...]") under the
separate key `PATH`, shadowing the real `Path` in child processes — so
cmake could no longer find cl.exe. Prepend to the `Path` list instead.

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume bceefe2e-0d08-4ad8-9b4c-3518466f1ccc
```
</details>
<details><summary>Claude's draft</summary>

The Windows CUDA 12.9 job now compiles (PATH fix worked) but fails on every
.cu file targeting sm_100a/sm_120a: CCCL's clusterlaunchcontrol.h inline asm
uses long2, and long is 32-bit under MSVC, producing "asm operand type
size(4) does not match constraint 'l'".

main already carries the answer in bld.bat: SM 100+ is broken with CUDA 12.9
on Windows (fixed in 13.0), and SM 110 (Thor) is Linux-only. My 1.28.0
adaptation had wrongly applied the Linux arch lists to all platforms. Split
the arch lists per platform to match main exactly:

- win 12.9: 70..90-real (no 100/120)
- win 13.0: 75..100-real,120 (no 110)
- linux unchanged

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume bceefe2e-0d08-4ad8-9b4c-3518466f1ccc
```
</details>
<details><summary>Claude's draft</summary>

The osx_arm64 jobs failed compiling onnxruntime_autoep_test: the vendored
googletest's gtest-printers.h formats std::chrono::time_point via
std::format, whose floating-point path needs std::to_chars — gated by
libc++ availability to macOS 13.4+, while MACOSX_DEPLOYMENT_TARGET is 11.0.

main never hits this because its meta.yaml keeps gmock in host, so
onnxruntime's FetchContent FIND_PACKAGE_ARGS resolves conda-forge's
gtest/gmock instead of vendoring one. Restore that dep, scoped to osx —
Linux and Windows are green with the vendored copy, and the megabuild
prefers a minimal host env to keep the shared cmake cache stable.

Added to both the staging and python stages to keep their dependency
sets identical (cache-sharing requirement).

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume bceefe2e-0d08-4ad8-9b4c-3518466f1ccc
```
</details>
<details><summary>Claude's draft</summary>

The osx jobs now compile, pass all 10 ctest suites, and install — but fail
at packaging with rattler-build's "No license files were copied". That
error fires when ANY itemized license path matches nothing (the detailed
missing-path list goes to a log level the CI does not surface), and the
set of vendored _deps differs per platform (CoreML adds coremltools/fp16/
psimd on osx; deps provided from conda-forge, e.g. nlohmann_json/gtest on
osx, may not be vendored at all).

Replace the itemized _deps license list with globs:

    build-ci/Release/_deps/*/LICENSE*
    build-ci/Release/_deps/*/COPYING*

This ships the license of every dep actually vendored into the build on
each platform — including several that are compiled in but were previously
not shipped (mp11, date, dlpack, kleidiai, cutlass on cuda) — and cannot
break when the vendored set changes.

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume bceefe2e-0d08-4ad8-9b4c-3518466f1ccc
```
</details>
<details><summary>Claude's draft</summary>

Replace the license globs with an explicit list, conditioned on exactly the
variants that vendor each dependency. The per-variant sets were read from
the packaging logs of the passing glob-based CI run:

- base (all variants): LICENSE + abseil, date, dlpack, eigen (COPYING.MPL2
  only, EIGEN_MPL2_ONLY=ON), flatbuffers, gsl, onnx, protobuf,
  pytorch_cpuinfo, re2, safeint
- nlohmann_json: not osx (conda-forge host dep there, not vendored)
- googletest: only native non-CUDA builds (tests off ⇒ never fetched), and
  not osx (conda-forge gtest/gmock there)
- kleidiai: ARM targets (linux-aarch64, osx-arm64)
- microsoft_wil: win
- coremltools/fp16/psimd: osx (CoreML EP)

Verified with rattler-build --render-only that each variant's rendered list
matches what its CI packaging log actually copied. Not itemized: cutlass /
cudnn_frontend on CUDA builds (parity with main, which never shipped them);
can be added once a CUDA packaging log confirms the paths.

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume bceefe2e-0d08-4ad8-9b4c-3518466f1ccc
```
</details>
Wire up an optional TensorRT execution provider for the CUDA 13.0 linux
builds (linux-64 and linux-aarch64), following the pytorch-feedstock style
of encoding the accelerator in the build string.

- recipe.yaml: `tensorrt` variant axis gated to CUDA 13.0; `_tensorrt`
  appended to the build string only on the TRT build; TensorRT host deps
  (libnvinfer-devel, libnvonnxparser-devel) and a provider-lib package test.
- build-cpp.nu: enable onnxruntime_USE_TENSORRT and the builtin ONNX parser
  (dynamic libnvonnxparser, nothing static), TENSORRT_HOME=$PREFIX.
- conda_build_config.yaml: `tensorrt` axis (linux) + local channel_sources
  for the locally-built TensorRT packages.

Prototype only: TensorRT packages come from a local channel and CUDA 13.0
aarch64 artifacts are not built yet.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Regenerated .ci_support, pixi.toml and CI workflow for the TensorRT axis.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Skip non-linux platforms so the prototype only builds linux-64 and
linux-aarch64, where the TensorRT execution provider is enabled.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Regenerated CI configs after restricting the matrix to linux.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Move the onnxruntime_providers_tensorrt.so check from an explicit `files`
path to a `lib` entry, letting rattler-build expand it to the
platform-specific library filename.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@conda-forge-admin

conda-forge-admin commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Hi! This is the friendly automated conda-forge-linting service.

I just wanted to let you know that I linted all conda-recipes in your PR (recipe/recipe.yaml) and found it was in an excellent condition.

I do have some suggestions for making it better though...

For recipe/recipe.yaml:

  • ℹ️ output 0 output overrides versions pinned in the feedstock:
    ['- In section host: python 3.13.*']
    Requirement spec should not list version specifiers to respect conda-forge-pinning. If you need to force another version, please override the pin via conda_build_config.yaml.
  • ℹ️ License is not an SPDX identifier (or a custom LicenseRef) nor an SPDX license expression.

Documentation on acceptable licenses can be found conda-forge.org > Docs > Maintainer Documentation > Contributing packages > SPDX Identifiers and Expressions.

This message was generated by GitHub Actions workflow run https://github.com/conda-forge/conda-forge-webservices/actions/runs/31599157505. Examine the logs at this URL for more detail.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants