Skip to content

Latest commit

 

History

History
380 lines (284 loc) · 16.5 KB

File metadata and controls

380 lines (284 loc) · 16.5 KB

Findings: PyTorch OOT compile support and AMD XDNA2

Research snapshot: 2026-07-16 (Europe/Istanbul).

This is a factual orientation document. It separates observations made on this machine from upstream documentation, and separates PyTorch framework work from the still-missing AMD execution backend. The source-build recovery recorded below is complete; its reproducible evidence is in BUILD_REPORT.md.

Executive finding

The host's kernel-side NPU support is alive: the PCI device is bound to the in-tree amdxdna driver, firmware loads, /dev/accel/accel0 exists, and the user has access. On 2026-07-16, a locally built XRT 2.20.0 plus matching AMD XDNA shim was installed and passed all three xrt-smi validate workloads on the physical NPU. The Ubuntu kernel module was not replaced.

The current upstream main userspace stack also builds on Ubuntu 26.04, but its XRT 2.26 validation fails with DRM_IOCTL_AMDXDNA_EXEC_CMD EINVAL against Ubuntu's 7.0.0-27 in-tree driver. The observed mismatch is the command buffer size contract: the Ubuntu driver rejects command BOs over 32 KiB, while the newer userspace issued larger ones. The compatible 1.6/XRT 2.20 line retains the older contract and is the locally proven baseline.

AMD's packaged Ryzen AI 1.7.1 Linux stack remains documented for Ubuntu 24.04 and Python 3.12, not Ubuntu 26.04. This is therefore a source-built, locally validated configuration rather than a vendor-certified Ubuntu 26.04 setup.

In parallel, PyTorch issue #189138 is legitimate upstream extensibility work, but it does not provide an AMD XDNA backend. It removes compile-stack device hardcoding so a separately implemented PrivateUse1/Inductor backend can plug in cleanly.

Verified host state

Operating system and NPU

Item Observed value
Distribution Ubuntu 26.04 LTS
Kernel 7.0.0-27-generic
Processor AMD Ryzen AI 9 HX 370 with Radeon 890M
NPU PCI address 0000:67:00.1
PCI ID 1022:17f0, revision 0x10
PCI description Strix/Krackan/Strix Halo Neural Processing Unit
Kernel driver amdxdna from Ubuntu's kernel tree
Driver initialization amdxdna_accel_driver 0.7.0
Firmware selected amdnpu/17f0_10/npu_7.sbin
Device node /dev/accel/accel0
Access user <user> has rw- ACL
Login memlock limit unlimited soft and hard

Evidence commands:

lsb_release -ds
uname -r
lspci -nnk -d 1022:17f0
journalctl -k -b | rg -i 'amdxdna|amdnpu'
ls -l /dev/accel/accel0
getfacl -cp /dev/accel/accel0
ulimit -l

The device node and ACL mean the immediate problem is not simple device permission denial. Do not replace the kernel module merely because XRT is failing.

XRT state

  • The system installation is XRT 2.20.0 (xrt-base, xrt-base-dev, and xrt-npu) plus xrt_plugin-amdxdna 2.20, all installed successfully.
  • It was built from amd/xdna-driver branch 1.6 at f0b7bd5fbc9d45892d050cb55701541350390b99, whose XRT submodule is 021204355eeaa034ff69aae407ace2265adf047a.
  • xrt-smi examine reports XRT 2.20.0, NPU firmware 1.1.2.64, and RyzenAI-npu4 at 0000:67:00.1. xrt-smi validate passes GEMM, latency, and throughput.
  • The installed optional pyxrt binding targets CPython 3.14. It imports from /usr/bin/python3 3.14.4 but is not available in the Python 3.12 XRT environment; build it separately under that virtual environment if needed.
  • The old rootless XRT 2.21.75 tree remains historical evidence only; do not add it to LD_LIBRARY_PATH ahead of /opt/xilinx/xrt/lib.

The historical rootless failure and the former 8 MiB memlock limit are both resolved. AMD's guidance about unlimited locked memory still applies; verify the active shell with ulimit -l before runtime work.

Local source snapshots

The repositories were clean at research time:

Checkout Branch/state Commit
<workspace>/pytorch main...origin/main, clean 4e0d62edff80bd68d67779f7f018fa8c1b133769
<workspace>/xdna-driver main...origin/main, clean ca5bce25951b1f8e4c01405d7ff1a033d5e8fd81
XRT submodule tag-near 202610.2.23.0_Canonical-142 b4bbf24c54b355b585d59f15264aba47d9aa54b9

The AMD build script at xdna-driver/build/build.sh exposes:

-nokmod  Don't build or install the kernel module

Its CMake project maps this to SKIP_KMOD, which excludes the driver and firmware packaging paths. That is the appropriate default on this host.

AMD's current supported status

AMD Ryzen AI Software 1.7.1's official Linux page, last updated 2026-07-10, requires:

  • Ubuntu 24.04 LTS;
  • kernel 6.10 or newer;
  • Python 3.12.x;
  • 64 GiB RAM recommended.

Its downloadable driver bundle names only 24.04 DEBs: XRT base, base-dev, NPU, and the AMD XDNA XRT plugin. There are no Ubuntu 26.04 package names in that document. This confirms that using those binary packages directly on 26.04 is outside AMD's documented package matrix.

The upstream amd/xdna-driver repository has a broader source policy: Ubuntu 22.04 or newer (and Arch) with kernel 6.10 or newer. It documents source builds of the XRT submodule and the XDNA plugin. It mentions Ubuntu 25.04's in-kernel driver but does not specifically certify Ubuntu 26.04. Source build is therefore the supported-looking engineering route, but 26.04 remains an uncertified combination that needs local validation.

Primary sources:

Userspace distribution availability

Verified on 2026-07-16:

Component Public source Prebuilt distribution found
XRT base (libxrt_core, xrt-smi) Xilinx/XRT and source release archives AMD's authenticated Ryzen AI bundle provides Ubuntu 24.04 DEBs; no official public Ubuntu 26.04 DEB was found
AMD XDNA XRT shim (libxrt_driver_xdna) amd/xdna-driver AMD's Ubuntu 24.04 bundle provides the xrt_plugin...amdxdna.deb; the GitHub repository has no binary Releases assets
Ryzen AI higher-level compiler/runtime AMD distributes ryzen_ai-1.7.1.tgz through its download portal Documented for Ubuntu 24.04 and Python 3.12; it does not replace the required XRT base/plugin packages

The XRT 2.21.75 GitHub release publishes a source tarball and Windows SDK, but no Linux DEB. The AMD XDNA repository has a 2.21.75 source tag and no GitHub releases. At that tag it pins XRT commit 4eb1f4392a012b4e6eca759762389c612537f7c7, which resolves to XRT tag 2.21.75; its build script already supports -nokmod.

Therefore AMD's direction to compile locally on Ubuntu 26.04 is consistent with the available distribution channels. The current local clone already is the necessary source distribution and includes a matched XRT submodule; no second source download is required.

Recommended XRT recovery sequence

All four recovery gates below passed on 2026-07-16. They remain the required order for a clean rebuild or a future upgrade attempt.

Gate 1: persistent memlock — passed

The persistent configuration was installed and verified after reboot on 2026-07-16. The active shell, user@1000.service, per-user manager defaults, and gnome-terminal-server.service all report unlimited/infinite soft and hard locked-memory limits.

Installed configuration:

  • /etc/security/limits.d/99-amdxdna.conf sets soft and hard memlock to unlimited for <user>;
  • /etc/systemd/system/user@1000.service.d/99-amdxdna-memlock.conf gives the per-user systemd manager an unlimited hard/soft limit;
  • <workspace>/.config/systemd/user.conf makes unlimited memlock the default for services launched by that user manager.

The systemd unit and manager configuration parsed successfully and is active.

AMD's documented drop-in is equivalent to:

* soft memlock unlimited
* hard memlock unlimited

See POST_REBOOT_CHECK.md for the acceptance commands and the verified values.

Gate 2: source dependencies — passed

The reviewed AMD dependency script was run with explicit authorization. Its raw installation log remains local; re-running it later is a system change and still requires explicit authorization.

Gate 3: build matched XRT and shim — passed

The 1.6 compatibility checkout and its XRT submodule were built with all 24 logical CPUs. The exact commands were:

cd <workspace>/xdna-driver-ubuntu7-compat/xrt/build
./build.sh -npu -opt -j 24 -ccache -disable-werror

cd <workspace>/xdna-driver-ubuntu7-compat/build
./build.sh -release -nokmod -j 24

XRT's final build passed all 70 configured tests (two intentionally skipped). The shim package audit found no kernel module, DKMS source, or firmware payload. The three Ubuntu 26.04 compatibility changes are preserved in patches/ubuntu26/.

Gate 4: validate bottom-up — passed

The validation order was:

  1. new login shows adequate ulimit -l;
  2. xrt-smi examine enumerates the NPU;
  3. xrt-smi validate completes the applicable tests;
  4. an XRT-level sample or AMD Ryzen AI quick test runs;
  5. only then begin an actual PyTorch PrivateUse1/XDNA adapter.

The final installed-stack result was 51.0 TOPS GEMM, 46.0 us average latency, and 80547.0 op/s average throughput. This establishes XRT/NPU userspace health, not PyTorch integration or model-level performance.

PyTorch issue #189138 status

The umbrella RFC is open and triaged:

All four were open on 2026-07-16. The umbrella is labeled feature, large, triaged, module: PrivateUse1, module: backend, module: dynamo, module: inductor, and oncall: pt2. It is not labeled actionable.

The RFC intentionally splits the work into three independently reviewable slices:

  1. query device identity/capability through DeviceInterface rather than hardcoded lists;
  2. populate Dynamo types and trace rules through registered interfaces;
  3. discover/configure Inductor codegen backends and obtain device-specific compile/runtime information through existing contracts.

The stated boundary is also important: the proposed work does not solve the AOTInductor C-shim ABI, which would need a separate proposal.

Current main still exhibits the problem

At the pinned local commit:

  • torch/_inductor/utils.py still defines GPU_TYPES = ["cuda", "mps", "xpu", "mtia"];
  • is_gpu() still tests membership in that list;
  • get_gpu_type() still asserts at most one available listed accelerator;
  • device_need_guard() still carries the hand-written MPS exception;
  • torch/utils/_triton.py::has_triton() still iterates a hardcoded device map;
  • torch/_dynamo/device_interface.py has the generic register_interface_for_device(), but no dedicated register_interface_for_device_PrivateUse1 primitive proposed by the RFC.

This confirms the RFC has not already landed under another form.

Why related PRs closed

The available evidence does not show that maintainers rejected the technical direction.

PR #189367

PR #189367 attempted the RFC 1 identity/capability slice. It was open for about one hour on 2026-07-09, had no maintainer reviews, and was closed by its author. The visible blockers were:

  • EasyCLA could not proceed because the commit email was not linked to the author's GitHub account;
  • CI workflows awaited maintainer approval;
  • the author closed the PR before review.

This was an author/workflow closure, not a recorded technical rejection.

PR #171969

PR #171969 was an earlier, smaller PrivateUse1 has_triton() change. It received positive PrivateUse1 review feedback and the visible CI bot reported no failures. It later became stale, was rebased, remained stale, and was automatically closed by the stale bot on 2026-06-07 while still awaiting final review.

This is evidence that the area can be technically acceptable yet still fail to land if ownership, final review, and stale-label handling are not maintained.

The new procedural gate

The current Ultimate Guide to PyTorch Contributions states that only PRs addressing issues labeled actionable will be considered; PRs for non-actionable issues will be closed. PyTorch's local CONTRIBUTING.md repeats the same requirement for new contributors.

Therefore, the next contribution should not begin with another broad PR. First ask maintainers on the RFC which child issue/slice is actionable, what test fixture they want, and whether a replacement for #189367 is welcome.

What PyTorch already supports OOT

PrivateUse1/OpenReg eager/runtime path

PyTorch's official Accelerator Integration guide and the in-tree OpenReg reference cover:

  • renaming/registering the PrivateUse1 device;
  • allocators and storage;
  • device hooks, guards, streams, events, and generators;
  • operator registration and CPU fallback;
  • AMP, profiler, distributed backend, and autoload integration;
  • cross-repository CI for downstream accelerator repos.

OpenReg is deliberately a mechanism-verification backend, not a complete or high-performance backend.

torch.compile custom compiler backend path

Separately, a torch.compile custom backend accepts an FX GraphModule plus example inputs and returns an equivalent callable. It can be registered with torch._dynamo.register_backend or the torch_dynamo_backends Python entry point. AOTAutograd can normalize the graph to core ATen ops and provide forward and backward graphs.

OpenReg currently registers an openreg compiler backend that validates its FX graph and returns gm.forward. Its tests prove graph capture, fake tensors, guards, dynamic shapes, autocast, autograd, and device metadata. They do not prove Inductor generates optimized XDNA code.

Inductor extension path

Current tests demonstrate manual runtime registration of an extension device's scheduler, wrapper code generator, device overrides, and DeviceInterface. Issue #189138 is about replacing the remaining assumptions and manual monkeypatching with a coherent discovery/capability contract.

What is still missing for this NPU

There is no ready AMD XDNA2 PyTorch PrivateUse1 backend in the checked-out AMD or PyTorch repositories. A real backend would still need, at minimum:

  • an XRT-backed allocator, tensor storage, transfers, device API, guards, streams/events, and synchronization;
  • the minimal PrivateUse1 operator set and/or an intentional fallback policy;
  • graph metadata/fake implementations for compile paths;
  • a compiler lowering from FX/core ATen to an execution format the Ryzen AI stack accepts, or a custom-op/partitioning bridge to AMD's supported model compiler/runtime;
  • correctness, unsupported-op, caching, and synchronization tests;
  • performance measurement that proves NPU execution rather than CPU fallback.

AMD's public Ryzen AI Linux flow is centered on compiled inference models and XRT/Vitis AI-related runtime components. The upstream XDNA README explicitly says the NPU is for inference, not ML training. Initial PyTorch/XDNA goals should therefore target bounded inference first.

Recommended first upstream contribution

After maintainer confirmation, take one registry/capability slice only:

  1. preserve all built-in device behavior exactly;
  2. add safe-default capability queries rather than another device-name list;
  3. exercise the behavior with OpenReg or the existing extension-device test, so no AMD hardware is needed in PyTorch CI;
  4. run the narrowly related Dynamo/Inductor tests plus --openreg and lint;
  5. keep AMD-specific code entirely outside the PyTorch core patch.

This makes the PyTorch contribution independently useful and reviewable while the local XDNA runtime work proceeds on its own timeline.