Redline provides lightning-fast kernel dispatch for ROCm by recording a fixed kernel sequence once and replaying it as retained PM4 over public ROCr/HSA queues. Choose the integration surface that matches who owns graph capture and kernel loading in your application.
Current distribution status: Redline is usable from a source checkout.
redline-dispatchis not yet published on PyPI, the C SDK is not yet attached to GitHub Releases, and the Rust crates are not yet on crates.io. The commands below deliberately build the current checkout instead of referring to packages that do not exist yet.
| You have | Use | Current source status |
|---|---|---|
| A C or C++ engine that owns its kernels and buffers | redline-capi |
Real module load, graph or direct retained-PM4 build, kernarg patch, replay, and multiqueue APIs |
| A Python application | redline-dispatch Python module |
Graph authoring plus real retained-PM4 load/build/replay; direct Python PM4 build is currently GFX12-only |
An existing application that already calls hipGraph* |
redline-hipgraph |
LD_PRELOAD interposer for supported captures; unsupported operations fall through to HIP |
| A Rust application or engine | redline-dispatch |
Native graph, public-AQL, HIP multistream, and retained-PM4 APIs from a source/path dependency |
For engine integrations, start with the C ABI unless Rust is already part of your process. The explicit API makes kernel identity, kernargs, dependencies, and ownership visible; the preload path is for compatible applications that cannot change their graph call sites.
The commands in this guide are verified on x86-64 Linux with the ROCm Core SDK 7.14 TheRock layout:
export PATH=/opt/rocm/core/bin:/opt/rocm/core/lib/llvm/bin:$PATH
export ROCM_PATH=/opt/rocm/core
export HIP_PATH=/opt/rocm/coreYou need:
- ROCm Core SDK 7.14 or newer and a visible AMD RDNA GPU;
- Rust 1.85 or newer for source builds;
- a C compiler for the ABI-only smoke, and
hipccfor the real-GPU C smoke; - Python 3.9 or newer plus
maturinfor the Python module.
ROCR_VISIBLE_DEVICES selects the physical devices visible to both ROCr and
HIP. After filtering, Redline device ordinal 0 means the first visible GPU:
ROCR_VISIBLE_DEVICES=0 your_command| Path | Coverage |
|---|---|
| Public ROCr/AQL replay | Architecture-neutral; exercised on gfx1010, gfx1030, gfx1100, gfx1151, and gfx1201 |
| C/Rust retained direct PM4 | Family-specific GFX10/GFX11 and GFX12 encoders; a device-family mismatch fails closed |
| Automatic independent-queue policy | Q2 on gfx1100 and gfx12, Q4 on other measured gfx11 devices, Q1 on unswept gfx10 |
Python Gpu.build() |
GFX12 direct-PM4 encoder today; use the C/Rust paths for architecture-dispatched GFX10/GFX11 replay |
Legacy direct PM4 accepts zero-scratch HSA kernels whose implicit inputs are the kernarg pointer and optional private-segment buffer. Unsupported scratch, queue, dispatch, or flat-scratch contracts are rejected rather than replayed with guessed state.
Every integration ultimately supplies the same information.
- Exact code object and symbol. Load the HSACO bytes that contain the
kernel. Symbols commonly carry the
.kdsuffix; use the name present in the code-object metadata. - Launch geometry. Redline APIs use global work-item counts for
gridand local work-item counts forblock. - Packed kernarg segment. Pack pointers and scalars at the offsets declared by the kernel metadata. Device pointers must remain GPU-accessible through replay. Query the expected segment size instead of assuming a host struct layout.
- Dependencies and memory access. The graph API takes explicit resource
reads/writes and rejects unordered hazards. The lower-level PM4 builder takes
explicit dependency boundaries such as
rl_pm4_wait_rmw. - Object lifetimes. Keep the GPU binding, loaded module, code object, and every referenced allocation alive through the last replay. Destroy the retained IB before the module and GPU.
A verified Radiowave manifest is optional. Supplying one binds the manifest to the exact code-object bytes and allows Redline to use a narrower VMEM-only consumer boundary where inspection proves it safe. Missing or ambiguous certification retains the broader same-agent scalar/vector boundary.
The direct retained lifecycle is:
select GPU -> load HSACO -> pack kernargs -> record dispatches and waits
-> finalize once -> patch changed kernargs -> replay and wait
-> free IB -> free module -> free GPU
finalize consumes its builder. replay waits for completion before returning,
so patching a retained kernarg after one replay and before the next is
race-free. Do not mutate kernargs while a replay is in flight.
Use the C ABI when an engine already owns HIP allocations, code objects, and its fixed launch sequence.
cargo build --release -p redline-capiArtifacts:
crates/redline-capi/include/redline_dispatch.h
target/release/libredline_dispatch.so
target/release/libredline_dispatch.a
There is not yet a CMake package, pkg-config file, or install target. Point your build at the header and one of the emitted libraries.
cc crates/redline-capi/examples/smoke.c \
-I crates/redline-capi/include \
-L target/release \
-Wl,-rpath,"$PWD/target/release" \
-lredline_dispatch -lpthread -ldl -lm \
-o /tmp/redline-capi-smoke
/tmp/redline-capi-smokeExpected output includes:
abi=1 lanes=1
C-ABI smoke OK
This smoke validates graph recording, hazard compilation, plan fingerprinting,
and C ownership. Its launch_mock call does not submit GPU work.
The core low-level sequence is:
RlGpu *gpu = rl_gpu_new(0);
RlModule *module = NULL;
int rc = rl_gpu_load_module(gpu, hsaco_bytes, hsaco_len, &module);
long kernarg_size = rl_module_kernarg_size(module, "my_kernel.kd");
/* Pack exactly kernarg_size bytes using the kernel ABI. */
RlPm4Builder *builder = rl_pm4_builder_new(gpu);
rc = rl_pm4_dispatch(builder, module, "my_kernel.kd",
grid_x, grid_y, grid_z,
block_x, block_y, block_z,
dynamic_group_bytes, kernarg, kernarg_len);
RlPm4Ib *ib = NULL;
rc = rl_pm4_finalize(gpu, builder, &ib); /* consumes builder */
rc = rl_pm4_replay(ib); /* submit + completion wait */
rl_pm4_ib_free(ib);
rl_module_free(module);
rl_gpu_free(gpu);Check every return value against RL_OK; the header defines stable error
classes for null inputs, UTF-8, recording, compilation, replay, handles, and
certification. The complete correctness-gated example is
crates/redline-capi/examples/gpu_smoke.c.
To build and run that example with its included counter kernel:
export GPU_ARCH=gfx1201 # replace with the exact target reported by rocminfo
hipcc --genco --offload-arch="$GPU_ARCH" \
bench/floor_kernel_ctr.hip -o /tmp/redline-counter.co
hipcc -x hip crates/redline-capi/examples/gpu_smoke.c \
-I crates/redline-capi/include \
-L target/release \
-Wl,-rpath,"$PWD/target/release" \
-lredline_dispatch -lpthread -ldl -lm \
-o /tmp/redline-capi-gpu-smoke
ROCR_VISIBLE_DEVICES=0 \
/tmp/redline-capi-gpu-smoke 256 /tmp/redline-counter.coA successful run ends with:
real-GPU C-ABI gate: counter = 256 / 256 certified=no [PASS]
Pass a matching Radiowave JSON manifest as the fourth argument to exercise certified module loading.
After finalization, update only the scalars or pointers that changed:
uint32_t position = next_position;
rc = rl_pm4_ib_set_kernargs(
ib, dispatch_index, position_byte_offset,
(const uint8_t *)&position, sizeof(position));
rc = rl_pm4_replay(ib);The retained PM4 packet keeps the same kernarg address, so this does not rebuild
the IB. rl_pm4_finalize_multi and rl_pm4_replay_multi provide one retained IB
per independent public queue; memory used by separate lanes must be independent.
Use rl_gpu_pm4_queue_count(..., RlQueueAuto, independent_width) instead of
hard-coding a queue count.
The higher-level C graph path combines resource declarations with real replay:
record nodes with rl_graph_kernel_ex, instantiate with
rl_graph_instantiate, and launch with rl_graphexec_launch.
The Python module contains two layers:
Graph/GraphExecfor graph authoring, dependency validation, lane inspection, fingerprinting, and mock replay;Gpu/Module/Pm4Ibfor real module load, allocation, retained PM4 build, kernarg patching, and replay.
python3 -m venv .venv-redline
. .venv-redline/bin/activate
python -m pip install --upgrade pip maturin
maturin develop --release --manifest-path crates/redline-py/Cargo.toml
python -c 'import redline_dispatch as rl; print(rl.Graph, rl.Gpu)'To produce an installable wheel instead of installing into the active virtual environment:
maturin build --release --strip \
--manifest-path crates/redline-py/Cargo.tomlThis creates an abi3 wheel for Python 3.9 or newer. It is a local artifact;
pip install redline-dispatch will not work until the package is published.
import redline_dispatch as rl
graph = rl.Graph(mode="latency")
activations = graph.buffer("activations", 4096)
project = graph.kernel(
"project", (32, 1, 1), (64, 1, 1),
accesses=[
(activations, 0, 2048, False),
(activations, 2048, 2048, True),
],
)
graph.kernel(
"consume", (32, 1, 1), (64, 1, 1),
accesses=[(activations, 2048, 2048, False)],
deps=[project],
)
exec = graph.instantiate()
exec.launch_mock()
print(exec.lane_count, exec.fingerprint())import redline_dispatch as rl
gpu = rl.Gpu(0)
module = gpu.load_module(open("counter.co", "rb").read(), None)
counter = gpu.alloc(4)
kernarg = counter.address().to_bytes(8, "little")
dispatches = [
("ctr_k.kd", (1, 1, 1), (1, 1, 1), 0, kernarg, True)
for _ in range(256)
]
ib = gpu.build(module, dispatches)
ib.replay()
assert counter.read_u32(0) == 256A dispatch tuple is:
(symbol, grid, block, dynamic_group_bytes, packed_kernarg_bytes, serialize)
serialize=True inserts a safe RMW boundary before the next consumer. A
verified VMEM-only next consumer gets the narrower boundary; uncertified or
ambiguous code uses the generic fail-closed boundary.
Run the complete GPU example after compiling the counter code object as shown in the C section:
ROCR_VISIBLE_DEVICES=0 \
python crates/redline-py/examples/gpu_smoke.py \
256 /tmp/redline-counter.coFor retained decode, call
ib.set_kernargs(dispatch_index, bytes, byte_offset=...) between replays. See
crates/redline-py/examples/decode_kernargs.py.
Build the non-Python interposer:
cargo build --release -p redline-hipgraph
LD_PRELOAD="$PWD/target/release/libredline_hipgraph.so" /usr/bin/trueThen preload the same artifact into the HIP graph process:
LD_PRELOAD="$PWD/target/release/libredline_hipgraph.so" your_hip_appSupported module-loaded and static-fatbin graph captures can resolve to Redline's retained PM4 path. Unsupported graph operations fall through to the real HIP implementation; the preload library is not a complete HIP runtime replacement.
The optional Python control module requires the python feature and a symlink
to the same backing shared object so capture and interposition share one set of
Rust statics. Follow the exact build and lifecycle instructions in
crates/redline-hipgraph/README.md.
Until the crates are published, use a path dependency from a source checkout:
[dependencies]
redline-dispatch = { path = "../redline/crates/redline-dispatch" }The graph API uses a familiar create → add nodes → instantiate → replay shape, with explicit buffer accesses added so Redline can derive correct minimal fences. Run the architecture-neutral migration example:
cargo run --release -p redline-dispatch --example hipgraph_migrationThat example uses MockBackend to demonstrate graph construction and the
compiled replay plan. For a real public-ROCr retained replay using the counter
code object built above:
ROCR_VISIBLE_DEVICES=0 \
REDLINE_GRAPH_HSACO=/tmp/redline-counter.co \
REDLINE_GRAPH_SYMBOL=ctr_k \
cargo run --release -p redline-dispatch --example graph_launch_smokeA successful run prints:
OBSERVED=2 EXPECTED=2
Use the native APIs when you need custom artifact catalogs, resource bindings, HIP multistream fallback, public-AQL lowering, or direct access to the per-generation PM4 encoders.
Redline fails closed when it cannot prove a replay contract.
| Symptom | Check |
|---|---|
rl_gpu_new returns null or Gpu(0) raises |
Confirm ROCm ≥7.14, ROCR_VISIBLE_DEVICES, and the /opt/rocm/core environment |
| Module or symbol lookup fails | Compile for the exact GPU architecture and use the symbol spelling in HSACO metadata, commonly name.kd |
| Kernarg recording fails | Query rl_module_kernarg_size / module.kernarg_size() and pack the exact ABI offsets |
| PM4 compilation rejects the kernel | Check scratch, private-segment, implicit-user-SGPR, LDS, and device-family requirements |
| Output is stale or wrong | Declare graph resource accesses or insert the required low-level RMW dependency; keep device allocations live |
| Multiqueue stalls or aliases data | Use the automatic queue policy and ensure every lane has an independent memory footprint |
| Preload remains on HIP | The captured operation or kernel contract is unsupported; inspect with the interposer API rather than assuming PM4 replay |
C functions return RL_OK or a negative RL_ERR_* class. Rust APIs return typed
errors. Python converts those failures to exceptions. Do not suppress a
certification or compilation error and continue with a partially constructed
retained object.
One retained PM4 IB holds at most 1,048,575 dwords (20-bit PM4
INDIRECT_BUFFER size field). Finalize fails closed with
RL_ERR_COMPILE and a stderr diagnostic when a recording exceeds that
ceiling; split long graphs across multiple IBs rather than packing more
dispatches into one buffer.
Measured density on a gfx1201 card with a single-kernel counter shape
(crates/redline-capi/examples/gpu_smoke.c, N atomic-increment dispatches
in one retained IB):
REDLINE_PM4_FULL_STATE |
Measured dwords/dispatch | Implied ceiling | Largest observed PASS | Observed FAIL |
|---|---|---|---|---|
0 (default, SH-elided) |
18.0 | ~58,200 dispatches | 52,000 | 60,000 (1,080,028 dwords) |
1 (full state per dispatch) |
47.0 | ~22,300 dispatches | 20,000 | 26,000 (1,221,998 dwords) |
Densities are exact: the diagnostic reports the recorded dword count, so 60,000 elided dispatches measured 1,080,028 dwords and 26,000 full-state dispatches measured 1,221,998. Full state costs 2.61x the stream, so it cuts the dispatches that fit in one IB by roughly the same factor. These numbers are shape-dependent — a kernel with more kernarg words or a different dependency pattern shifts them — so treat the implied ceiling as guidance and the fail-closed diagnostic as the authority. Long decode graphs must be split across IBs.
ROCr/KFD enumerate GPUs in discovery order. rocm-smi sorts by PCI bus. The
same physical card therefore has different indices in different tools, and an
ordinal copied from the wrong tool can open the wrong device.
Measured on hipx (four agents):
| Physical device | BDF | ROCr index | rocm-smi | Anchor |
|---|---|---|---|---|
| gfx1100 | 0000:66:00.0 |
0 | uuid:GPU-43390a851e296ee5 |
|
| gfx1151 Strix Halo APU | 0000:bf:00.0 |
1 | GPU[3] | bdf:0000:bf:00.0 (no UUID) |
| gfx1010 | 0000:6e:00.0 |
2 | bdf:0000:6e:00.0 (no UUID) |
|
| gfx1030 | 0000:99:00.0 |
3 | uuid:GPU-c7ff6b154d0128bc |
On hipx, rocm-smi GPU[3] is ROCr index 1 — the integrated APU. A device reset
on that APU takes the whole host down. Two of the four GPUs report no UUID and
must be BDF-anchored. On hiptrx, four identical gfx1201 boards sit at
0000:03/c3/e3/13:00.0; any name:gfx1201 selector is ambiguous.
Pin with UUID when the device reports one, else PCI address (BDF). Never an
index. BDF is the join key across ROCr, HIP, sysfs, rocm-smi, and amdgpu
fault lines in dmesg — which is what lets a dmesg VM-fault line be attributed
to a specific benchmark row.
| Form | Example | Stability / refusal |
|---|---|---|
uuid: |
uuid:GPU-43390a851e296ee5 |
Stable when the device reports a real UUID (not GPU-XX) |
bdf: |
bdf:0000:66:00.0 or short bdf:66:00.0 (domain 0) |
Stable PCI address; preferred when UUID is absent |
slot: |
slot:3 |
Host PCI slot label; refused when the host has no slot labels |
name: |
name:gfx1100 |
Product/ISA name; refused when more than one agent matches |
index: |
index:1 |
Volatile ROCr discovery ordinal under the current visibility filter |
@alias |
@dev0 |
Expands through the host manifest; refused if undefined |
Unprefixed input is a hard error. Prefer uuid: / bdf: / @alias in engine
config; keep index: for throwaway local smoke only.
RlGpu *gpu = rl_gpu_open("uuid:GPU-43390a851e296ee5");
if (!gpu) {
/* stderr already has one redline: … diagnostic */
return 1;
}
char desc[256];
rl_gpu_describe(gpu, desc, sizeof(desc));
/* e.g. "gfx1100 0000:66:00.0 [uuid:GPU-43390a851e296ee5] card0 …" */
/* … load / build / replay … */
rl_gpu_free(gpu);rl_gpu_new(ordinal) still works but is volatile-by-index: the ordinal is
discovery order under the current ROCR_VISIBLE_DEVICES filter and is not
stable across tools or reboots. Both entry points enforce the host deny-list.
Aliases and the deny-list live in a TOML manifest. Format and real host entries:
docs/devices.toml.example. Minimal shape:
[host.hipx]
dev0 = "uuid:GPU-43390a851e296ee5"
deny = ["bdf:0000:bf:00.0"] # Strix Halo APU: device reset takes the host downDiscovery order (later overrides earlier for the same alias / deny list):
/etc/redline/devices.toml$XDG_CONFIG_HOME/redline/devices.toml(else~/.config/redline/devices.toml).redline/devices.tomlwalking up from cwd
The active section is [host.<hostname>], overridable with REDLINE_HOST.
deny is evaluated on the resolved device, so no selector form
(uuid: / bdf: / slot: / name: / index: / @alias) can bypass it. The
trailing # comment on a deny entry is quoted back in the error.
Both candidate sysfs signals were measured on hipx; both fail:
- the APU's KFD node reports
cpu_cores_count=0exactly like the discrete cards; - every node reports
heap_type=1because the APU's 96 GiB unified memory presents as public framebuffer.
There is therefore no integrated / is-dangerous bit in device identity. The deny-list is the only guard, and it is human-written.
Read-only inventory (no queues, no dispatch — safe on a busy machine):
cargo run -p redline-rocr --example device_listWorked example from hipx:
host hipx: 4 GPU agent(s)
[0] gfx1100 0000:66:00.0 [uuid:GPU-43390a851e296ee5] card0 kfd_node=1 rocr#0(volatile)
[1] gfx1151 0000:bf:00.0 [bdf:0000:bf:00.0] card1 kfd_node=2 rocr#1(volatile)
[2] gfx1010 0000:6e:00.0 [bdf:0000:6e:00.0] slot=1 card2 kfd_node=3 rocr#2(volatile)
[3] gfx1030 0000:99:00.0 [uuid:GPU-c7ff6b154d0128bc] slot=3 card3 kfd_node=4 rocr#3(volatile)
Copy the bracketed anchor (uuid:… or bdf:…) into engine config or the host
manifest. Do not copy the leading [N] ordinal.
ROCm/legacy-rocm-build#6529 contention A/B
Honest scope. The harness in
scripts/6529-contention-ab.sh is a
probability-shifting stress tool, not a deterministic reproducer. A clean
matrix cannot prove absence of the gfx1100 address-zero SQC fault. Use it
only to raise or lower confidence in the mid-IB CWSR/MCBP × SH-register
elision hypothesis described in
docs/investigations/2026-07-31-rocm-6529-address-zero-sqc-fault.md.
On a pinned ROCr device it sweeps:
REDLINE_PM4_FULL_STATE={0,1}(stateful SH elision vs full state per dispatch), and- a second independent HIP process
(
bench/contention_load.hip) on or off,
for R repetitions of hipfire-6409-bench (default filter:
serial_latency). Each cell records exit status, stderr, and any new
amdgpu / VM_L2_PROTECTION_FAULT / SQC / MES dmesg lines. Results land
under examples/hipfire-6409/results/<arch>/6529-contention-ab-<UTC>/.
ROCR_VISIBLE_DEVICES is required. HIP_VISIBLE_DEVICES is a CLR filter and
does not affect HSA enumeration, so an unpinned two-card host may bind the
wrong GPU.
The script refuses non-gfx11 agents unless --force is passed (local harness
smoke only). It never runs modprobe and never reboots.
Binary: hipfire-6409-bench
(examples/hipfire-6409/Cargo.toml
package name). Parsed in
examples/hipfire-6409/src/main.rs:
| Flag / env | Line(s) |
|---|---|
--out |
1021 |
--warmups |
1022 |
--samples |
1023 |
--filter |
1024 (substring match on row key; 210–213) |
--max-rows |
1025–1030 (truncates after filter; 216–217) |
--backends |
1094–1096 (subsets must include redline and vulkan; 1283–1284) |
--matrix |
1098–1102 |
HIPFIRE_BENCH_ARCH |
1007–1008 |
There is no start-row / end-row / row-index-range flag. Subset a matrix
with --filter and/or --max-rows only, or pass extra args after --.
These require a reboot or amdgpu module reload. The script only records
the live values of /sys/module/amdgpu/parameters/{cwsr_enable,mcbp}.
-
CWSR off vs on (primary preemption gate):
# kernel cmdline (persistent): amdgpu.cwsr_enable=0 # or modprobe.d: # /etc/modprobe.d/amdgpu-cwsr.conf options amdgpu cwsr_enable=0Then reboot (or unload all DRM clients and
modprobe -r amdgpu && modprobe amdgpu). Re-run the contention script. Repeat withcwsr_enable=1.Side effects:
cwsr_enable=0disables the GPU debugger (ROCgdb / wave-level debug) and can change scheduling for other compute workloads on the node. Treat it as a lab-only toggle. -
MCBP — a control, not a probe.
amdgpu.mcbp=-1(the default, and the value on the faulting host) leavesadev->gfx.mcbpfalse on any non-SR-IOV system:amdgpu_device_set_mcbponly forces the flag on formcbp=1or an SR-IOV VF (amdgpu_device.c:3632-3643). Settingmcbp=0therefore changes nothing on a bare-metal host. MCBP also governs the graphics ring mux andIB_FLAG_PREEMPT, not the KFD compute HQD/MES path that retained PM4 IBs execute on.# kernel cmdline — expected to be a no-op on bare metal; run it only to # confirm the null result, not as a hypothesis test: amdgpu.cwsr_enable=1 amdgpu.mcbp=0The informative second arm is contention, not MCBP: MES quantum switching (10 ms process / 1 ms gang quanta, programmed on every
add_queue_mes) can preempt a long retained IB with no eviction event at all, which is why the script sweeps contention on/off.
Interpretation sketch (still probabilistic): faults that track
cwsr_enable=1 under contention, and drop with REDLINE_PM4_FULL_STATE=1,
support the SH-elision x mid-IB preemption story. Absence of faults does not
clear the bug.
export PATH=/opt/rocm/core/bin:/opt/rocm/core/lib/llvm/bin:$PATH
export ROCM_PATH=/opt/rocm/core HIP_PATH=/opt/rocm/core
cd examples/hipfire-6409
HIPFIRE_BENCH_ARCH=gfx1100 cargo build --release --bin hipfire-6409-bench \
--target-dir target/gfx1100
cd ../..
ROCR_VISIBLE_DEVICES=0 scripts/6529-contention-ab.sh \
--bench examples/hipfire-6409/target/gfx1100/release/hipfire-6409-bench \
--arch gfx1100 \
--reps 5 \
--filter serial_latency \
--max-rows 32 \
--warmups 1 --samples 3- A source-built C shared/static library and public header.
- A locally built abi3 Python wheel.
- Source-built hipGraph preload and Rust APIs.
These are sufficient for developers with repository access. External developers need versioned artifacts and a public source/release location.
-
PyPI: publish
redline-dispatchabi3 wheels through.github/workflows/publish-pypi.ymlusing PyPI trusted publishing or a scoped token. Validate the uploaded wheel on a ROCm host before calling the release complete. -
C SDK: attach a versioned Linux x86-64 archive to the same GitHub Release:
redline-c-sdk-VERSION-linux-x86_64/ ├── include/redline_dispatch.h ├── lib/libredline_dispatch.so ├── lib/libredline_dispatch.a ├── LICENSE ├── NOTICE └── SHA256SUMS -
Later, if demanded: add CMake/pkg-config metadata and an apt/deb package. Publish Rust crates only after internal path dependencies have versioned crates.io metadata.
Do not document a PyPI install or C SDK download as available until those artifacts have been published and independently installed on a clean ROCm host.