Research snapshot: 2026-07-16 (Europe/Istanbul).
This is a factual orientation document. It separates observations made on this machine from upstream documentation, and separates PyTorch framework work from the still-missing AMD execution backend. The source-build recovery recorded below is complete; its reproducible evidence is in BUILD_REPORT.md.
The host's kernel-side NPU support is alive: the PCI device is bound to the
in-tree amdxdna driver, firmware loads, /dev/accel/accel0 exists, and the
user has access. On 2026-07-16, a locally built XRT 2.20.0 plus matching AMD
XDNA shim was installed and passed all three xrt-smi validate workloads on
the physical NPU. The Ubuntu kernel module was not replaced.
The current upstream main userspace stack also builds on Ubuntu 26.04, but
its XRT 2.26 validation fails with DRM_IOCTL_AMDXDNA_EXEC_CMD EINVAL
against Ubuntu's 7.0.0-27 in-tree driver. The observed mismatch is the command
buffer size contract: the Ubuntu driver rejects command BOs over 32 KiB, while
the newer userspace issued larger ones. The compatible 1.6/XRT 2.20 line
retains the older contract and is the locally proven baseline.
AMD's packaged Ryzen AI 1.7.1 Linux stack remains documented for Ubuntu 24.04 and Python 3.12, not Ubuntu 26.04. This is therefore a source-built, locally validated configuration rather than a vendor-certified Ubuntu 26.04 setup.
In parallel, PyTorch issue #189138 is legitimate upstream extensibility work, but it does not provide an AMD XDNA backend. It removes compile-stack device hardcoding so a separately implemented PrivateUse1/Inductor backend can plug in cleanly.
| Item | Observed value |
|---|---|
| Distribution | Ubuntu 26.04 LTS |
| Kernel | 7.0.0-27-generic |
| Processor | AMD Ryzen AI 9 HX 370 with Radeon 890M |
| NPU PCI address | 0000:67:00.1 |
| PCI ID | 1022:17f0, revision 0x10 |
| PCI description | Strix/Krackan/Strix Halo Neural Processing Unit |
| Kernel driver | amdxdna from Ubuntu's kernel tree |
| Driver initialization | amdxdna_accel_driver 0.7.0 |
| Firmware selected | amdnpu/17f0_10/npu_7.sbin |
| Device node | /dev/accel/accel0 |
| Access | user <user> has rw- ACL |
| Login memlock limit | unlimited soft and hard |
Evidence commands:
lsb_release -ds
uname -r
lspci -nnk -d 1022:17f0
journalctl -k -b | rg -i 'amdxdna|amdnpu'
ls -l /dev/accel/accel0
getfacl -cp /dev/accel/accel0
ulimit -lThe device node and ACL mean the immediate problem is not simple device permission denial. Do not replace the kernel module merely because XRT is failing.
- The system installation is XRT
2.20.0(xrt-base,xrt-base-dev, andxrt-npu) plusxrt_plugin-amdxdna2.20, all installed successfully. - It was built from
amd/xdna-driverbranch 1.6 atf0b7bd5fbc9d45892d050cb55701541350390b99, whose XRT submodule is021204355eeaa034ff69aae407ace2265adf047a. xrt-smi examinereports XRT2.20.0, NPU firmware1.1.2.64, andRyzenAI-npu4at0000:67:00.1.xrt-smi validatepasses GEMM, latency, and throughput.- The installed optional
pyxrtbinding targets CPython 3.14. It imports from/usr/bin/python33.14.4 but is not available in the Python 3.12 XRT environment; build it separately under that virtual environment if needed. - The old rootless XRT
2.21.75tree remains historical evidence only; do not add it toLD_LIBRARY_PATHahead of/opt/xilinx/xrt/lib.
The historical rootless failure and the former 8 MiB memlock limit are both
resolved. AMD's guidance about unlimited locked memory still applies; verify
the active shell with ulimit -l before runtime work.
The repositories were clean at research time:
| Checkout | Branch/state | Commit |
|---|---|---|
<workspace>/pytorch |
main...origin/main, clean |
4e0d62edff80bd68d67779f7f018fa8c1b133769 |
<workspace>/xdna-driver |
main...origin/main, clean |
ca5bce25951b1f8e4c01405d7ff1a033d5e8fd81 |
| XRT submodule | tag-near 202610.2.23.0_Canonical-142 |
b4bbf24c54b355b585d59f15264aba47d9aa54b9 |
The AMD build script at xdna-driver/build/build.sh exposes:
-nokmod Don't build or install the kernel module
Its CMake project maps this to SKIP_KMOD, which excludes the driver and
firmware packaging paths. That is the appropriate default on this host.
AMD Ryzen AI Software 1.7.1's official Linux page, last updated 2026-07-10, requires:
- Ubuntu 24.04 LTS;
- kernel 6.10 or newer;
- Python 3.12.x;
- 64 GiB RAM recommended.
Its downloadable driver bundle names only 24.04 DEBs: XRT base, base-dev, NPU, and the AMD XDNA XRT plugin. There are no Ubuntu 26.04 package names in that document. This confirms that using those binary packages directly on 26.04 is outside AMD's documented package matrix.
The upstream amd/xdna-driver repository has a broader source policy:
Ubuntu 22.04 or newer (and Arch) with kernel 6.10 or newer. It documents source
builds of the XRT submodule and the XDNA plugin. It mentions Ubuntu 25.04's
in-kernel driver but does not specifically certify Ubuntu 26.04. Source build is
therefore the supported-looking engineering route, but 26.04 remains an
uncertified combination that needs local validation.
Primary sources:
- AMD Ryzen AI 1.7.1 Linux installation
- AMD XDNA driver and XRT shim README
- AMD Ryzen AI Software 1.7 release
Verified on 2026-07-16:
| Component | Public source | Prebuilt distribution found |
|---|---|---|
XRT base (libxrt_core, xrt-smi) |
Xilinx/XRT and source release archives |
AMD's authenticated Ryzen AI bundle provides Ubuntu 24.04 DEBs; no official public Ubuntu 26.04 DEB was found |
AMD XDNA XRT shim (libxrt_driver_xdna) |
amd/xdna-driver |
AMD's Ubuntu 24.04 bundle provides the xrt_plugin...amdxdna.deb; the GitHub repository has no binary Releases assets |
| Ryzen AI higher-level compiler/runtime | AMD distributes ryzen_ai-1.7.1.tgz through its download portal |
Documented for Ubuntu 24.04 and Python 3.12; it does not replace the required XRT base/plugin packages |
The XRT 2.21.75 GitHub release
publishes a source tarball and Windows SDK, but no Linux DEB. The AMD XDNA
repository has a 2.21.75 source tag and no GitHub releases. At that tag it
pins XRT commit 4eb1f4392a012b4e6eca759762389c612537f7c7, which resolves to
XRT tag 2.21.75; its build script already supports -nokmod.
Therefore AMD's direction to compile locally on Ubuntu 26.04 is consistent with the available distribution channels. The current local clone already is the necessary source distribution and includes a matched XRT submodule; no second source download is required.
All four recovery gates below passed on 2026-07-16. They remain the required order for a clean rebuild or a future upgrade attempt.
The persistent configuration was installed and verified after reboot on
2026-07-16. The active shell, user@1000.service, per-user manager defaults,
and gnome-terminal-server.service all report unlimited/infinite soft and hard
locked-memory limits.
Installed configuration:
/etc/security/limits.d/99-amdxdna.confsets soft and hard memlock to unlimited for<user>;/etc/systemd/system/user@1000.service.d/99-amdxdna-memlock.confgives the per-user systemd manager an unlimited hard/soft limit;<workspace>/.config/systemd/user.confmakes unlimited memlock the default for services launched by that user manager.
The systemd unit and manager configuration parsed successfully and is active.
AMD's documented drop-in is equivalent to:
* soft memlock unlimited
* hard memlock unlimited
See POST_REBOOT_CHECK.md for the acceptance commands and the verified values.
The reviewed AMD dependency script was run with explicit authorization. Its raw installation log remains local; re-running it later is a system change and still requires explicit authorization.
The 1.6 compatibility checkout and its XRT submodule were built with all 24 logical CPUs. The exact commands were:
cd <workspace>/xdna-driver-ubuntu7-compat/xrt/build
./build.sh -npu -opt -j 24 -ccache -disable-werror
cd <workspace>/xdna-driver-ubuntu7-compat/build
./build.sh -release -nokmod -j 24XRT's final build passed all 70 configured tests (two intentionally skipped).
The shim package audit found no kernel module, DKMS source, or firmware
payload. The three Ubuntu 26.04 compatibility changes are preserved in
patches/ubuntu26/.
The validation order was:
- new login shows adequate
ulimit -l; xrt-smi examineenumerates the NPU;xrt-smi validatecompletes the applicable tests;- an XRT-level sample or AMD Ryzen AI quick test runs;
- only then begin an actual PyTorch PrivateUse1/XDNA adapter.
The final installed-stack result was 51.0 TOPS GEMM, 46.0 us average
latency, and 80547.0 op/s average throughput. This establishes XRT/NPU
userspace health, not PyTorch integration or model-level performance.
The umbrella RFC is open and triaged:
- #189138 — DeviceInterface registry support for PrivateUse1 in torch.compile
- #189135 — device identity/capability
- #189136 — Dynamo registry splicing
- #189137 — Inductor codegen registration interfaces
All four were open on 2026-07-16. The umbrella is labeled feature, large,
triaged, module: PrivateUse1, module: backend, module: dynamo,
module: inductor, and oncall: pt2. It is not labeled actionable.
The RFC intentionally splits the work into three independently reviewable slices:
- query device identity/capability through
DeviceInterfacerather than hardcoded lists; - populate Dynamo types and trace rules through registered interfaces;
- discover/configure Inductor codegen backends and obtain device-specific compile/runtime information through existing contracts.
The stated boundary is also important: the proposed work does not solve the AOTInductor C-shim ABI, which would need a separate proposal.
At the pinned local commit:
torch/_inductor/utils.pystill definesGPU_TYPES = ["cuda", "mps", "xpu", "mtia"];is_gpu()still tests membership in that list;get_gpu_type()still asserts at most one available listed accelerator;device_need_guard()still carries the hand-written MPS exception;torch/utils/_triton.py::has_triton()still iterates a hardcoded device map;torch/_dynamo/device_interface.pyhas the genericregister_interface_for_device(), but no dedicatedregister_interface_for_device_PrivateUse1primitive proposed by the RFC.
This confirms the RFC has not already landed under another form.
The available evidence does not show that maintainers rejected the technical direction.
PR #189367 attempted the RFC 1 identity/capability slice. It was open for about one hour on 2026-07-09, had no maintainer reviews, and was closed by its author. The visible blockers were:
- EasyCLA could not proceed because the commit email was not linked to the author's GitHub account;
- CI workflows awaited maintainer approval;
- the author closed the PR before review.
This was an author/workflow closure, not a recorded technical rejection.
PR #171969 was an earlier,
smaller PrivateUse1 has_triton() change. It received positive PrivateUse1
review feedback and the visible CI bot reported no failures. It later became
stale, was rebased, remained stale, and was automatically closed by the stale
bot on 2026-06-07 while still awaiting final review.
This is evidence that the area can be technically acceptable yet still fail to land if ownership, final review, and stale-label handling are not maintained.
The current Ultimate Guide to PyTorch Contributions
states that only PRs addressing issues labeled actionable will be considered;
PRs for non-actionable issues will be closed. PyTorch's local
CONTRIBUTING.md repeats the same requirement for
new contributors.
Therefore, the next contribution should not begin with another broad PR. First ask maintainers on the RFC which child issue/slice is actionable, what test fixture they want, and whether a replacement for #189367 is welcome.
PyTorch's official Accelerator Integration guide and the in-tree OpenReg reference cover:
- renaming/registering the PrivateUse1 device;
- allocators and storage;
- device hooks, guards, streams, events, and generators;
- operator registration and CPU fallback;
- AMP, profiler, distributed backend, and autoload integration;
- cross-repository CI for downstream accelerator repos.
OpenReg is deliberately a mechanism-verification backend, not a complete or high-performance backend.
Separately, a torch.compile custom backend accepts an FX GraphModule plus
example inputs and returns an equivalent callable. It can be registered with
torch._dynamo.register_backend or the torch_dynamo_backends Python entry
point. AOTAutograd can normalize the graph to core ATen ops and provide forward
and backward graphs.
OpenReg currently registers an openreg compiler backend that validates its FX
graph and returns gm.forward. Its tests prove graph capture, fake tensors,
guards, dynamic shapes, autocast, autograd, and device metadata. They do not
prove Inductor generates optimized XDNA code.
Current tests demonstrate manual runtime registration of an extension device's
scheduler, wrapper code generator, device overrides, and DeviceInterface.
Issue #189138 is about replacing the remaining assumptions and manual
monkeypatching with a coherent discovery/capability contract.
There is no ready AMD XDNA2 PyTorch PrivateUse1 backend in the checked-out AMD or PyTorch repositories. A real backend would still need, at minimum:
- an XRT-backed allocator, tensor storage, transfers, device API, guards, streams/events, and synchronization;
- the minimal PrivateUse1 operator set and/or an intentional fallback policy;
- graph metadata/fake implementations for compile paths;
- a compiler lowering from FX/core ATen to an execution format the Ryzen AI stack accepts, or a custom-op/partitioning bridge to AMD's supported model compiler/runtime;
- correctness, unsupported-op, caching, and synchronization tests;
- performance measurement that proves NPU execution rather than CPU fallback.
AMD's public Ryzen AI Linux flow is centered on compiled inference models and XRT/Vitis AI-related runtime components. The upstream XDNA README explicitly says the NPU is for inference, not ML training. Initial PyTorch/XDNA goals should therefore target bounded inference first.
After maintainer confirmation, take one registry/capability slice only:
- preserve all built-in device behavior exactly;
- add safe-default capability queries rather than another device-name list;
- exercise the behavior with OpenReg or the existing extension-device test, so no AMD hardware is needed in PyTorch CI;
- run the narrowly related Dynamo/Inductor tests plus
--openregand lint; - keep AMD-specific code entirely outside the PyTorch core patch.
This makes the PyTorch contribution independently useful and reviewable while the local XDNA runtime work proceeds on its own timeline.