torch-xdna is an independently maintained experimental downstream backend. Contributions must preserve its narrow, evidence-backed support boundary. A passing CPU emulation, FakeTensor trace, or generated wrapper is not evidence of physical XDNA execution.
Every change should name one ownership class:
- downstream torch-xdna: PrivateUse1 runtime and operators, XRT BO ownership, XDNA Inductor adapters, cache and artifact contracts, packaging, tests, CI, acceptance tools, and project documentation;
- generic upstream PyTorch: a device-independent problem in an existing PyTorch extension contract, demonstrated without AMD-specific code and suitable for OpenReg or another hardware-free fixture; or
- vendor/system: XRT, the XDNA driver, firmware, MLIR-AIE, Peano, kernel or distribution packaging, and machine configuration.
Do not fix an upstream or vendor problem by hiding it in the downstream runtime, and do not put XDNA-specific behavior into PyTorch core. Upstream PyTorch work is frozen until the maintainer explicitly reopens it. Never open, comment on, tag, or push an upstream issue or pull request from a downstream task without direct authorization and the applicable upstream contribution gate.
Read AGENTS.md, SECURITY.md, docs/DEVELOPER_PREVIEW.md, and the nearest design and result note before editing. Start from repository facts:
git status --short --branch
git rev-parse HEAD
git diff --statUse a dedicated feature branch and isolated worktree. Preserve published
history and annotated milestone tags. Do not develop directly on main, move
a tag, rewrite a published branch, or create v0.1.0a1 without explicit
maintainer approval.
Product support is exactly CPython >=3.12,<3.13. Use an isolated uv-managed
3.12 environment and the pinned PyTorch/XRT/compiler inputs described in the
developer-preview contract. Research results from other Python versions do
not expand package metadata.
Build the native extension against the intended installed PyTorch ABI through
the repository's PEP 517/editable workflow with build isolation disabled.
Set CC, CXX, XILINX_XRT, and a bounded MAX_JOBS explicitly. Do not
reuse an extension from another interpreter, PyTorch build, or environment.
Do not install or replace system packages, the driver, firmware, XRT, or
system configuration as part of a repository change.
- Keep the supported graph, dtype, layout, device, and synchronization claims as small as the physical evidence.
- Use PyTorch's existing PrivateUse1, DeviceInterface, and Inductor extension registries. Do not create a parallel registration architecture or monkeypatch private global tables.
- Do not add CPU, NumPy, eager, dispatcher, Meta, or FakeTensor user-arithmetic fallback. Meta and FakeTensor are graph-analysis tools only.
- Reject unsupported behavior deterministically before accepting a generated output or launching the device. Tests must compare allocation, cache, and XRT sequence state where applicable.
- Treat streams, events, autograd, training, dynamic shapes, general Triton,
AOTInductor, new operators, and new dtypes as separate vertical slices. The
Triton-XDNA product contracts are exactly the pinned Strategy B bodies for
static contiguous BF16
[1024]tensor add and the two independently receipt-bound BF16 matrix products[128,64] @ [64,128] -> [128,128]and[64,64] @ [64,128] -> [64,128]; none is a reusable general Triton, matmul, or linear API. - Include every code-generation, ABI, runtime, compiler, artifact, and capability input in the cache identity before allowing reuse.
- Validate the nine-file IRON set against its strict manifest before native initialization. Validate each separate Strategy B direct ELF against its own compiler lock, receipt, ownership, mode, size, hash, final ABI, and kernel identity before enabling its graph. Never bundle generated XCLBIN, ELF, instruction, wheel, or compiler-cache payloads.
- Preserve the known-good in-tree
amdxdnadriver and XRT 2.20 runtime unless a separate, explicitly authorized system task changes that baseline.
Run the smallest relevant test first, then the repository gates appropriate to
the diff. Hardware-free work normally includes focused pytest, the complete
default suite, Ruff formatting and lint, strict type checking, shell syntax,
ShellCheck, packaging checks, and git diff --check. Use exact commands from
the current workflow and acceptance scripts rather than copying stale counts
from an older result note.
Physical tests are opt-in. Run them only after non-hardware gates pass and only with the reviewed external artifacts and validated runtime. A physical claim must record, at minimum:
- output residency on
xdna:0before explicit copy-back; - bit-exact output against an independent oracle;
- XRT sequence advancement and
ERT_CMD_STATE_COMPLETED; - a positive launch-to-completion interval;
- expected XCLBIN UUID and kernel identity for XCLBIN-backed kernels, or the exact ELF digest, receipt, final ABI, and kernel identity for a direct-ELF kernel; and
- evidence that CPU, NumPy, dispatcher, eager, Meta, and FakeTensor did not perform user arithmetic.
A timing sample is correctness evidence, not a performance result. A missing device is an explicit hardware-unavailable outcome, not a pass. A self-hosted runner labeled as XDNA2 must fail when its runtime, device, or artifact contract is invalid.
The focused Strategy B receipt and the older exhaustive INT32 receipt are exact-checkpoint evidence. Neither transfers to a later consolidated native, compiler, wrapper, or package identity. Changes to those identities require a targeted physical smoke; the complete alpha stack requires one final object-independent qualification before a release gate can be reviewed.
Keep raw logs, build trees, environments, wheels, device binaries, caches, and machine-specific evidence outside the repository. Commit only reviewed, sanitized summaries without usernames, absolute paths, raw BDF assumptions, or secrets.
Prefer small commits that can be independently reverted and validated. Before each private push:
git status --short
git diff --check
git diff --stat
git diff
git ls-files --others --exclude-standardAlso scan the complete range for credentials, passwords, tokens, private paths, binaries, artifacts, raw logs, build outputs, caches, and virtual environments. Do not delete unrelated user state or validated external artifacts during cleanup.
Commit messages must explain:
- the exact problem;
- ownership classification;
- the implemented boundary;
- literal test commands and actual results;
- physical evidence when applicable;
- known limitations; and
- AI-assisted development disclosure.
Do not squash or rewrite already validated milestone commits. Update Issue #1 only after the corresponding acceptance criteria and sanitized evidence have been reviewed. A workflow file without a successful intended run does not complete a CI gate.
Follow SECURITY.md. Never commit credentials, private keys, cookies, package-manager authentication, passwords, raw environment dumps, or unreviewed external binaries. Treat artifact directories and compiler caches as untrusted mutable input; fail closed on symlinks, ownership violations, unexpected inventory, or digest mismatch.
There is no top-level project LICENSE at this checkpoint. Do not infer
redistribution rights for repository source or generated artifacts from source
availability; licensing is an explicit maintainer/legal release blocker.
AI tools may assist development, but the human maintainer remains responsible for understanding and approving every code, test, claim, commit, issue, and release action. Disclose material AI assistance in commits and downstream pull requests. Remove generated verbosity and verify every factual claim against the repository or recorded evidence.
For any future upstream PyTorch communication, also obey PyTorch's current
AI_POLICY.md: do not paste raw or lightly reviewed model output; visibly
contain disclosed AI-generated material and accompany it with the human
contributor's own explanation and judgment. Downstream authorization never
authorizes an upstream write.
Issue #1 is the source of truth for v0.1.0a1. A gate may be checked only
after its stated evidence exists on the reviewed commit. Do not imply AMD
affiliation, production readiness, general XDNA/PyTorch support, Windows
support, artifact redistribution permission, or unsupported compiler/runtime
features. Do not create a SemVer tag or GitHub release without the maintainer's
explicit final approval, even when every technical gate appears to pass.