Wrapper for Triton C++ Client Libraries
This repository provides a header-only C++ wrapper around the official Triton C++ client libraries for making interaction with a Triton Inference Server easier in ROS 2 and standalone CMake projects.
The wrapper does not host neural networks itself. A Triton Inference Server with a compatible exported model repository must be available at runtime.
π Quick Start β’ π» Development β’ π Documentation
Important
This repository is part of OpenADS, the Open Automated Driving Systems project. OpenADS and its modules have been initiated and are currently being maintained by the Institute for Automotive Engineering (ika) at RWTH Aachen University.
- CMake 3.18 or newer
- A C++17 compiler
- Eigen3
- NVIDIA Triton C++ client libraries
- Optional: CUDA Toolkit for CUDA input shared memory
Set TRITON_CLIENT_DIR to the Triton client installation prefix when it is not
installed at /opt/tritonclient.
Clone and install the package:
git clone https://github.com/openads-project/triton_cpp.git
cmake -S triton_cpp -B build/triton_cpp -DROS_VERSION=0
cmake --build build/triton_cpp
cmake --install build/triton_cpp --prefix /path/to/prefixAfter installation, link your CMake target to the package:
find_package(triton_cpp 1.1 CONFIG REQUIRED)
target_link_libraries(my_target PRIVATE triton_cpp::triton_cpp)Alternatively, add this repository directly to your CMake project:
add_subdirectory(triton_cpp)
target_link_libraries(my_target PRIVATE triton_cpp)For ROS 2 workspaces, add this repository to the workspace and declare:
<depend>triton_cpp</depend>find_package(triton_cpp REQUIRED)
ament_target_dependencies(${TARGET_NAME} triton_cpp)// Include
#include <triton_cpp/triton_interface.hpp>
#include <memory>
#include <iostream>
// Connect to the server
std::unique_ptr<triton_cpp::TritonInterface> ti =
std::make_unique<triton_cpp::TritonInterface>("PBOD", "1", "127.0.0.1:8001", false);
// Query for model info (input & output names, shapes, and datatypes) as human-readable string
std::cout << ti->getModelInfo() << std::endl;
// Get the number of inputs and outputs programmatically
int n_inputs = ti->nInputs();
int n_outputs = ti->nOutputs();
// Create all data buffers for model input and output
// Note that, if the server doesn't know some output sizes and thus prints them as -1, you need to
// provide them as argument, like {{"reg_logits", {65536l, 7l}}}
ti->initInOutputs({});
// Get an interface to the model inputs. You need to provide the name and shape as arguments
// and the datatype as template parameter.
// The return type will be:
// - an Eigen::Map to an Eigen::Vector, if the shape has only one entry
// - an Eigen::Map to an Eigen::Matrix, if the shape has two entries
// - an Eigen::TensorMap to an Eigen::Tensor, if the shape has at least three entries
// If this does not correspond to the expected number of bytes, a std::invalid_argument will be thrown.
// If the number of bytes matches, the raw data will be interpreted as the requested type, no
// matter if the datatype and shape equal the expected ones. So you can easily perform
// reshapes, get rid of unnecessary batch dimensions, ...
// You can also get the raw pointer and bytesize by providing no shape, but that is strongly discouraged
auto points_xyz_map = ti->getInputTensor<float>("points_xyz", 200000, 3);
// fill with data (See Eigen documentation)
points_xyz_map(0,0) = 3.5; //...
// Infer
ti->infer();
// Get the results. The types and shapes work exactly as for input
auto class_logits = ti->getOutputTensor<float>("class_logits", 65536, 3);
std::cout << class_logits(0,0);
// Note that the Infer call takes the inputs that are currently stored inside the data buffers, and the
// output will be valid as long as no other inference is called. This means, that multi-threading
// is currently not supported!The full constructor signature is:
TritonInterface(const std::string& model_name,
const std::string& model_version,
const std::string& server_url,
bool shm,
bool variable_input_size = false,
bool retry_connection = false,
double client_timeout_s = 0.0,
bool cuda_input_shm = false)| Parameter | Type | Default | Description |
|---|---|---|---|
model_name |
string |
β | Name of the model to load from the Triton server |
model_version |
string |
β | Version of the model (e.g. "1") |
server_url |
string |
β | gRPC URL of the Triton server (e.g. "127.0.0.1:8001") |
shm |
bool |
β | Use host/system shared memory for Triton input and output transport (see Use host shared memory) |
variable_input_size |
bool |
false |
Allow input shapes to change between inference calls (see Variable input size) |
retry_connection |
bool |
false |
Retry connecting to the server until it becomes available (see Retry connection) |
client_timeout_s |
double |
0.0 |
Client-side inference timeout in seconds; 0.0 disables it (see Inference timeout) |
cuda_input_shm |
bool |
false |
Require Triton CUDA shared memory for input tensors. Construction fails with an informative error if it is unavailable. |
Transport combinations:
shm=false,cuda_input_shm=false: standard Triton transport for inputs and outputsshm=true,cuda_input_shm=false: system shared memory for inputs and outputsshm=false,cuda_input_shm=true: CUDA shared memory for inputs, standard Triton transport for outputsshm=true,cuda_input_shm=true: CUDA shared memory for inputs, system shared memory for outputsvariable_input_size=truecannot be combined with eithershm=trueorcuda_input_shm=true
By default, triton_cpp provides the input and output as a serialized Protobuf stream to the server. If the server and client run on the same machine, using shared memory is much more efficient.
-
If both are running locally, you can skip to step 3. If both are running in docker containers, you need to expose the shared memory between them. For this, start the server with the arguments
--ipc=shareable --shm-size=2gb, or indocker-compose.ymlipc: shareable shm_size: '2gb'
(Replace 2gb with a reasonable size for your application and machine)
-
Start the client container with the argument
--ipc=container:triton-triton-server-1, replacingtriton-triton-server-1with the name of the container running triton, or indocker-compose.ymlipc: "service:[triton-server]" # replace triton-server with the service's name that provides shm
-
Now in the code, change
std::unique_ptr<triton_cpp::TritonInterface> ti = std::make_unique<triton_cpp::TritonInterface>("PBOD", "1", "127.0.0.1:8001", false);to
std::unique_ptr<triton_cpp::TritonInterface> ti = std::make_unique<triton_cpp::TritonInterface>("PBOD", "1", "127.0.0.1:8001", true);and you're done.
-
To test whether SHM is working as expected, go to
/dev/shmin both containers and observe whether files containing data buffers are created while inference is running.
Shared-memory regions are registered with Triton using per-client unique names, so multiple independent
TritonInterface instances can use SHM safely at the same time, even when they point to the same Triton server.
When SHM is enabled, tensor offsets inside each shared region are aligned using max(type_alignment, 8) to avoid
misaligned accesses for mixed datatypes (for example FP32, INT64, and BOOL) in the same packed buffer.
If your input tensors are already on GPU memory, cuda_input_shm=true avoids copying them back to host memory before sending them to Triton.
std::unique_ptr<triton_cpp::TritonInterface> ti =
std::make_unique<triton_cpp::TritonInterface>(
"PBOD", "1", "127.0.0.1:8001", false, false, false, 0.0, true);Notes:
- this is for input tensors only; outputs still follow the normal path unless
shm=true - if CUDA shared memory is unsupported by the local client, Triton server, or
triton_cppbuild, construction fails with an informative error - host-mapped
getInputTensor(...)access is not available for CUDA-backed input buffers - for CUDA-backed inputs, use:
usesCudaInputSharedMemory()getInputTensorDevice(...)copyInputTensorToDevice(...)
Example:
std::unique_ptr<triton_cpp::TritonInterface> ti =
std::make_unique<triton_cpp::TritonInterface>(
"PBOD", "1", "127.0.0.1:8001", false, false, false, 0.0, true);
ti->initInOutputs({});
if (!ti->usesCudaInputSharedMemory()) {
throw std::runtime_error("CUDA input SHM was requested but is not enabled");
}
// Option 1: use CUDA code to write directly into Triton's CUDA-backed input buffer.
auto [points_device, points_bytes] = ti->getInputTensorDevice("points_xyz");
my_cuda_kernel<<<blocks, threads>>>(points_device, points_bytes);
// Option 2: copy from an existing host buffer into the CUDA-backed input.
std::vector<float> host_points(num_points * 3);
fill_points(host_points);
ti->copyInputTensorToDevice("points_xyz", host_points.data(), host_points.size() * sizeof(float));
ti->infer();getInputTensorDevice(...) and copyInputTensorToDevice(...) can be combined
in the same integration, depending on whether individual inputs already reside
on the GPU or originate in host memory.
By default (variable_input_size = false) input buffers are allocated once during initInOutputs() and reused on every infer() call, which is the most efficient mode.
Set variable_input_size = true when the number of elements in an input tensor changes between calls. In this mode the Triton client reallocates the input memory on every infer() call, so there is a small per-call overhead.
std::unique_ptr<triton_cpp::TritonInterface> ti =
std::make_unique<triton_cpp::TritonInterface>(
"behavior-planning", "1", "127.0.0.1:8001", false, true);Note:
variable_input_size = truecannot be combined withshm = trueorcuda_input_shm = true. If you try, theTritonInterfaceconstructor throwsstd::invalid_argument.
By default (retry_connection = false) the constructor throws std::runtime_error immediately if it cannot reach the server or fetch the model configuration.
Set retry_connection = true to have the constructor keep retrying (with a 1-second delay between attempts) until the server is reachable and the model is loaded. This is useful when the client node may start before the Triton server is fully ready (e.g. inside a Docker Compose stack).
std::unique_ptr<triton_cpp::TritonInterface> ti =
std::make_unique<triton_cpp::TritonInterface>(
"PBOD", "1", "127.0.0.1:8001", false, false, true);TritonInterface supports an optional client-side inference timeout via the constructor argument
client_timeout_s:
std::unique_ptr<triton_cpp::TritonInterface> ti =
std::make_unique<triton_cpp::TritonInterface>(
"PBOD", "1", "127.0.0.1:8001", false, false, false, 2.0);- Unit: seconds
client_timeout_s == 0.0: timeout disabledclient_timeout_s > 0.0: failed/blocked inference requests return with an error after the timeoutclient_timeout_s < 0.0: invalid (throws)
Implementation details are found in the Source Code Documentation.
The source code in this repository is licensed under Apache-2.0, see LICENSE. Container images provided by this repository may contain third-party software shipped with their own license terms.
In particular, the NVIDIA Triton C++ client libraries are licensed under the BSD 3-Clause License (license source).
Development and maintenance of this repository are supported by the following projects. We acknowledge the funding of the respective institutions.
| Project | Funding Institution | Grant Number |
|---|---|---|
| AIGGREGATE | πͺπΊ European Union | 101202457 |
| autotech.agil | π©πͺ Federal Ministry for Research, Technology and Space (BMFTR) | 1IS22088A |
Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Climate, Infrastructure and Environment Executive Agency (CINEA). Neither the European Union nor CINEA can be held responsible for them.
