Skip to content

Latest commit

 

History

78 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Phoenix SDR-DSP

License: Apache 2.0 Target: AMD Phoenix NPU1 Host: Windows 11 Pro Validation: 34/34 mixed-backend PASS Release: v1.0.0 Post-Quantum Cryptography Xilinx XRT Xilinx MLIR-AIE Compiler: LLVM Peano I/Q: 7.46 Msps CI

Windows-native SDR/DSP and finite-field engineering corpus for AMD Ryzen AI Phoenix silicon (XDNA1 / AIE2)

Built on Xilinx XRT, Xilinx MLIR-AIE / IRON, and LLVM Peano.

Third-party source and test-vector provenance is recorded in THIRD_PARTY_NOTICES.md.

7.46 Msps of real I/Q on a 10 TOPS AMD laptop NPU. 29.8 MB/s in · 59.7 MB/s in+out · ~92% NPU. No discrete GPU. No FPGA.

The 34-entry regression matrix completed 34/34 PASS on 2026-08-17 in 126.29 seconds. Its accurate backend accounting is 29 direct-hardware entries, four host/NPU composer entries (M32e, M33d, M33e Sign, and M33e Verify), and one intentional CPU reference entry (M12). M33a and M33b are native, fail-closed silicon gates; the higher-level composers dispatch those primitives from Python and are not fully device-resident. See the M33 validation record, v1.0.0 validation errata, and PQC status summary. New-user path: clone the repository, then run py .\install; a successful install automatically invokes the canonical regression runner.

InstallArchitectureDirectory StructureValidation MatrixI/Q ThroughputEngineering Issues & FixesQuickstartReferencesCreditsDocumentationPublication readiness


Installation

A new Windows 11 machine with a Phoenix / Hawk Point NPU only needs a clone of this repository. The extensionless, stdlib-only install launcher wraps the official Xilinx / AMD native-Windows stack:

Component Pin Upstream
Xilinx XRT (Xilinx Runtime) Windows SDK 2.21.75 host DMA / pyxrt
Xilinx MLIR-AIE + IRON wheel v1.4.1 + source 3ca0193 AIE dialect, iron.Runtime, aiecc
LLVM Peano (llvm-aie) 21.0.0.2026080301+c9c5ecb7 AIE2 clang++
AMD NPU driver / xrt-smi 32.0.20102.3930 already on the laptop

Do not install mlir_aie from the rolling latest-wheels-4 channel. An untagged checkout of pin 3ca0193 would otherwise resolve to an older series (observed: 1.3.4). The installer downloads the published v1.4.1 cp313 wheel into a local wheelhouse and passes --wheelhouse to official iron_setup.py.

New-user steps

Clone the repository, open PowerShell in that clone, and run the one setup command:

py .\install

The launcher uses its internal installer implementation, installs the pinned XRT / MLIR-AIE / IRON stack and PQC reference packages, then invokes the canonical run_all_silicon_tests.py after a successful full installation. py is the Windows Python launcher and binds to system CPython. The runner re-execs third_party\mlir-aie\ironenv\Scripts\python.exe (where Xilinx IRON installed numpy / mlir_aie / pyxrt) and sets PEANO_INSTALL_DIR from that same checkout. Do not activate ironenv or manually run pip install for the normal one-command setup flow.

To run an individual example or test after installation, activate the checkout-local environment from the repository root:

.\scripts\activate_ironenv.ps1

The activation helper is drive-independent. It selects this clone's third_party\mlir-aie\ironenv, sets PEANO_INSTALL_DIR to this clone's Peano installation, and relies on the installer-validated checkout-local XRT Python binding. A new PowerShell window must activate the environment again before using generic python commands. The whole-suite command does not require manual activation because its runner performs the same interpreter and Peano selection automatically.

For non-install maintenance, the launcher forwards --check-only, --download-only, and --self-test without invoking the canonical regression or dispatching kernels. --check-only and --download-only can call xrt-smi examine as a prerequisite probe; --self-test uses only local temporary files:

py .\install --check-only

Prerequisites

  • AMD Ryzen 7040 / 8040 APU (Ryzen 9 7940HS or similar Phoenix/Hawk Point silicon).
  • Windows 11 Pro 22H2 / 23H2 / 25H2 with the AMD NPU driver enabled.
  • CPython 3.13, Git, CMake, and Visual Studio 2022/18 with the C++ and Clang/LLVM workloads (llvm-objcopy is required for the Peano wheel fixup). Official path: mlir-aie 1.4.1 buildHostWinNative.

Verified clean clone

On 2026-08-15 a wipe-and-clone of main on a second Windows volume ran the two commands above and reported 16/16 PASS in 95.91 s (cold xclbin compile) on a Ryzen 9 7940HS Phoenix NPU1. A cached re-run on the development tree was 17.46 s. After native M33 runner integration, the 34-entry development tree completed 34/34 PASS in 126.29 s on 2026-08-17. That is a mixed-backend regression result, not 34 fully device-resident workloads: 29 entries dispatch directly to hardware, four are host/NPU composers, and M12 is an intentional CPU reference. The current boundary is documented in docs/M33_SILICON_VALIDATION_20260817.md.

Longer Windows walkthrough: docs/SETUP_WINDOWS.md. Pin rationale: docs/M2_TOOLCHAIN_PIN.md. v1.0.0 PQC summary: docs/PQC_COMPLETE_V1.md.

1. System & Hardware Architecture

The Phoenix SDR-DSP project provides a native Windows 11 execution and validation path for SDR processing and finite-field lattice-cryptography experiments on AMD Ryzen 7040/8040 series processors.

  • Target APU: AMD Ryzen 9 7940HS (8 Cores / 16 Threads @ 4.0–5.2 GHz)
  • NPU Silicon: AMD XDNA1 / 1st Gen Ryzen AI (npu1), up to 10 TOPS (INT8 on Phoenix 7040)
    • Tile Array: 20-tile 4×5 AI Engine 2 (AIE2) array; row/column orientation follows the platform documentation and is not inferred here
    • Vector Architecture: 512-bit SIMD registers supporting 64-lane bfloat16, 32-lane int16, and 16-lane cint16
    • Local Memory: 64 KB local data memory per tile (four 16 KB banks)
  • Host Operating System: Windows 11 Pro 25H2
  • Compilation Toolchain: LLVM Peano clang++ (--target=aie2-none-unknown-elf)
  • Runtime Environment: Xilinx MLIR-AIE / IRON Python eDSL JIT + native Windows Xilinx XRT (xrt_core.dll / CachedXRTHostRuntime)

2. Repository Structure

phoenix-sdr-dsp/
├── include/
│   └── sdr_dsp/
│       ├── sdr_dsp_common.hpp      # Vector types, lane constants, Q15 definitions
│       ├── fir_filter.hpp           # 64-lane vectorized FIR filtering
│       ├── complex_mixer.hpp        # Complex NCO & I/Q frequency shifter
│       ├── power_detector.hpp       # I^2 + Q^2 energy / RSSI detector
│       ├── modular_arithmetic.hpp   # Barrett & Montgomery modular reduction mod q=3329
│       └── ntt_butterfly.hpp        # Cooley-Tukey & Gentleman-Sande butterflies
├── kernels/
│   └── fft_stockham_f32.cc          # Radix-4 Stockham FFT (adapted from AMD FFT_R4_AIE, Apache-2.0)
├── tests/
│   ├── m3_saxpy/                    # Milestone 3:  Single-Core SAXPY Vector Operation (bfloat16)
│   ├── m5_fir/                      # Milestone 5:  8-Tap Vectorized Low-Pass FIR Filter
│   ├── m6_mixer/                    # Milestone 6:  Complex Mixer / NCO Frequency Downconverter
│   ├── m7_power/                    # Milestone 7:  Power / RSSI Energy Detector
│   ├── m8_pipeline/                 # Milestone 8:  Streaming Multi-Stage Fused Demodulator Pipeline
│   ├── m9_parallel/                 # Milestone 9:  4-Column Parallel FIR Filter (Hardware Scaling)
│   ├── m9b_parallel_pipeline/       # Milestone 9b: 4-Column Parallel Multi-Stage Demodulator Pipeline
│   ├── m10_modular/                 # Milestone 10: Modular Arithmetic & Barrett Reduction
│   ├── m11_butterfly/               # Milestone 11: Radix-2 NTT Butterfly Kernel
│   ├── m12_ntt_ref/                 # Milestone 12: CPU NTT/INTT Reference & Constant Generator
│   ├── m13_ntt16/                   # Milestone 13: 16-Point Vectorized NPU NTT (64 Batches)
│   ├── m14_ntt256/                  # Milestone 14: 256-Point Vectorized NPU NTT (4 Batches)
│   ├── m15_polymul/                 # Milestone 15: NPU INTT & Cyclic Polynomial Multiplication
│   ├── m15b_negacyclic/             # Milestone 15b: Negacyclic Polynomial Multiplication (Kyber / ML-KEM ring)
│   ├── m32_mlkem/                   # Milestone 32: FIPS 203 ML-KEM (planned; not in the 16-suite)
│   ├── m16_fft_ref/                 # Milestone 16: CPU DFT/FFT Reference (three implementations, CI)
│   ├── m17_radix2_fft/              # Milestone 17: 64-Point NPU Radix-4 Stockham FFT + IFFT
│   ├── m17p_fft_parallel/           # Milestone 17p: 4-Column Parallel FFT Channelizer
│   └── npu_visible/                 # Demo: 4-column I/Q throughput (not in the 16-suite)
├── scripts/                         # Windows environment audit, bootstrap, and activation scripts
├── docs/                            # Milestones, mathematics, ROADMAP, Windows setup, toolchain pin
├── requirements/                    # Pinned toolchain versions
├── toolchain.yaml                   # Machine-readable pinned stack (silicon-verified components)
├── install                          # Windows clean-clone launcher: py .\install
├── install.py                       # Internal implementation used by install
├── run_all_silicon_tests.py         # Automated Master Regression Suite
├── CITATION.cff                     # Citation metadata (validated with cffconvert)
├── LICENSE                          # Apache License 2.0
├── CONTRIBUTING.md                  # Contribution Guidelines
└── README.md                        # Master Project Documentation

3. Validation Matrix

Hardware-backed rows below execute on physical Phoenix NPU silicon (npu1) and compare against an independent mathematical reference. CPU and reference-only rows are labeled explicitly.

Release-maintenance validation

For a clean-drive, host-safe audit that never accesses the NPU, run PowerShell 7 as a normal user:

pwsh -File .\scripts\validate_clean_clone.ps1 `
    -InstallHostDependencies

The explicit switch installs and verifies pinned numpy==2.5.2, then writes one timestamped report beneath ignored release-evidence/. Without it, a missing or mismatched NumPy version causes an actionable refusal. The separate -RunSilicon switch runs the unchanged canonical runner only after the host checks pass and must be used only on an approved Phoenix test host. Publication scope, evidence limits, and retention requirements are in Publication readiness and the journal reproducibility checklist.

Milestone Component / DSP Primitive Target Array Workload / Dimensions Silicon Status Verification Result
M3 Single-Core SAXPY Vector Operation Tile (0,2) 4096 bfloat16 elements PASS Bit-Exact ($0.0$ error)
M5 8-Tap Vectorized Low-Pass FIR Tile (0,2) 4096 elements PASS $L_\infty \le 0.007812$
M6 Complex Mixer / NCO Downconverter Tile (0,2) 2048 I/Q pairs PASS $L_\infty \le 0.007812$
M7 Power / RSSI Energy Detector Tile (0,2) 2048 I/Q $\to$ 2048 P PASS $L_\infty \le 0.015625$
M8 Multi-Stage Fused Demodulator Tile (0,2) RF I/Q $\to$ Mix $\to$ FIR $\to$ Pwr PASS Zero stack allocation
M9 4-Column Parallel FIR Scaling 4 Columns (0..3,2) 4096 samples (1024/core) PASS 4-Core Parallel Lockstep
M9b 4-Column Parallel Multi-Stage Pipeline 4 Columns (0..3,2) 2048 I/Q burst / column PASS 2400.71 µs / burst, 0.85 MSamples/sec
M10 Modular Arithmetic & Barrett Reduction Tile (0,2) 1024 pairs mod $q=3329$ PASS Bit-Exact Match
M11 Radix-2 NTT Butterfly Kernel Tile (0,2) 1024 CT butterflies mod $q$ PASS Bit-Exact Match
M12 NTT Constant & Reference Engine CPU Reference $N=16, 256$, $\omega^N \equiv 1$ PASS Bit-Exact Match
M13 16-Point Vectorized NPU NTT Tile (0,2) 64 parallel frames (1024 elems) PASS Bit-Exact Match
M14 256-Point Vectorized NPU NTT Tile (0,2) 4 parallel frames (1024 elems) PASS Bit-Exact Match
M15 NPU INTT & Cyclic Polynomial Multiplication Tile (0,2) $C(x) = A(x) \times B(x) \pmod{x^{256}-1}$ PASS Bit-Exact Match
M15b Negacyclic Polynomial Multiplication (Kyber / FIPS 203 ring) Tile (0,2) $C(x) = A(x) \times B(x) \pmod{x^{256}+1}$ PASS Bit-Exact Match
M16 CPU DFT/FFT Reference (three implementations) CPU Reference (CI) $N \in {8..1024}$ PASS $\le 10^{-13}$ vs NumPy fft.fft
M17 64-Point NPU Radix-4 Stockham FFT + IFFT Tile (0,2) 64-point complex bfloat16 PASS FFT SNR 138.79 dB, IFFT round-trip 135.11 dB
M17p 4-Column Parallel FFT Channelizer 4 Columns (0..3,2) 64 parallel 64-point frames PASS 1,993 FFTs/sec, 0.51 MB/s I/Q
M19 8-Tap Complex FIR (complex taps × complex I/Q) Tile (0,2) 4096 complex samples PASS Bit-Exact vs CPU reference
M20 Fused Polyphase Decimator (M=4) + Interpolator (L=4) Tile (0,2) 4096 complex samples PASS Bit-Exact vs CPU reference
M21 Fused Digital Down-Converter (DDC) Tile (0,2) Complex NCO at −f_s/8 + Kaiser LPF + decim-by-4 PASS Bit-Exact vs CPU reference
M22 Fused Digital Up-Converter (DUC) Tile (0,2) Interp-L=4 + Kaiser×L LPF + complex NCO at +f_s/8 PASS Bit-Exact vs CPU reference
M23 Fused Polyphase Channelizer (M-path) Tile (0,2) M=8 commutator + M-path FIR + 8-point matmul-DFT PASS Bit-Exact vs CPU reference
M24 Fused Barker-13 Matched-Filter Correlator Tile (0,2) Reversed-tap FIR pair on I and Q, L=13 PASS Bit-Exact vs CPU reference
M25 Fused BPSK / QPSK Receiver Tile (0,2) Gardner TED + linear interp + NCO derotate + Costas PASS Receiver-theoretic gates (‹π/8 residual)
M26 Fused QAM-16 Receiver + Soft-Decision Demapping Tile (0,2) M25 core + Gray slicer + DD phase detector + max-log LLR PASS LLR consistency ≥ 0.75 (LSBs) / ≥ 0.85 (MSBs)
M27 Fused OFDM Loopback Tile (0,2) FFT + CP + pilots + LS pilot estimates + linear interpolation + zero-forcing equalization HARDWARE PASS Reuses M17 radix-4 Stockham FFT
M32b Post-Quantum Cryptography — FIPS 203 ML-KEM NTT Tile (0,2) Algorithms 9–12, Z_3329, pq-crystals ζ-table PASS Bit-Exact vs kyber-py 1.0.1
M32c Post-Quantum Cryptography — FIPS 202 Keccak / SHA-3 / SHAKE + samplers Tile (0,2) Keccak-f[1600] permutation, 5 dispatch modes PASS Bit-Exact vs FIPS 202 test vectors
M32d Post-Quantum Cryptography — FIPS 203 K-PKE Component Tile (0,2) Algorithms 13–15 (K-PKE.KeyGen / Encrypt / Decrypt) PASS Bit-Exact vs kyber-py K-PKE
M32e Post-Quantum Cryptography — FIPS 203 ML-KEM internal-interface composer Host + tile Algorithms 16–18, ML-KEM-512 HARDWARE SMOKE PASS 60 host KATs plus 3 silicon vectors per operation
M33a Post-Quantum Cryptography — FIPS 204 ML-DSA NTT Tile (0,2) NTT / INTT / basemul / reduce, Z_8380417 SILICON PASS 420 / 420; m33a:silicon
M33b Post-Quantum Cryptography — FIPS 204 Rounding & Hint Tile (0,2) Power2Round / Decompose / MakeHint / UseHint / CheckNorm SILICON PASS 700 / 700; m33b:silicon
M33d Post-Quantum Cryptography — FIPS 204 ML-DSA.KeyGen composer Host + tile Algorithm 6, ML-DSA-{44, 65, 87} HYBRID PASS 75 / 75; native M33a/M33b primitives
M33e Post-Quantum Cryptography — FIPS 204 ML-DSA.{Sign_internal, Verify_internal} composer Host + tile Algorithms 7 and 8, ML-DSA-{44, 65, 87} HYBRID PASS 180 / 180; native M33a/M33b primitives

Post-Quantum Cryptography track (M32 + M33, v1.0.0)

The v1.0.0 tree contains FIPS-aligned ML-KEM and ML-DSA experiments with different validation boundaries:

  • FIPS 203 ML-KEM — M32b, M32c, and M32d dispatch directly to the NPU. M32e exercises the ML-KEM-512 internal deterministic interfaces (Algorithms 16–18) in a host/NPU composition with 60 host KATs and a nine-vector silicon smoke gate; it does not establish public Algorithms 19–21 coverage. Reference oracle: kyber-py 1.0.1.
  • FIPS 204 ML-DSA — M33a and M33b are native, fail-closed silicon gates. M33d/e are host-orchestrated composers that dispatch those polynomial primitives to the NPU while SHAKE, sampling, packing, accumulation, and control remain host-side. The recorded gates are M33a 420/420, M33b 700/700, M33d 75/75, M33e Sign 90/90, and M33e Verify 90/90. Reference oracle: dilithium-py 1.4.0.

Full v1.0.0 release summary: docs/PQC_COMPLETE_V1.md. Per-milestone design notes live at docs/M32b_DESIGN.md, docs/M32c_DESIGN.md, docs/M32d_DESIGN.md, docs/M32e_DESIGN.md, docs/M33a_DESIGN.md, docs/M33b_DESIGN.md, docs/M33d_DESIGN.md, docs/M33e_DESIGN.md. kyber-py and dilithium-py are version-pinned; pytest is required but not yet locked, so the installation is not fully dependency-reproducible. All SHAKE / SHA-3 primitives come from the CPython hashlib standard library, so no separate SHAKE / Keccak wheel is required.

I/Q throughput demo (not in the 34-invocation suite)

Host-visible 4-column streamed complex mixer in tests/npu_visible/. Measured 2026-08-15 on a Ryzen 9 7940HS Phoenix NPU1 (10 TOPS). First buffer matches the M6 complex-multiply reference ($L_\infty = 0.007812$). Kernel vectorization is deferred.

Metric 1-column 8 KB loop 4-column stream
IQ in 3.85 MB/s 29.84 MB/s
IQ out 3.85 MB/s 29.84 MB/s
IQ in+out 7.70 MB/s 59.68 MB/s
Complex rate 0.963 Msps 7.459 Msps
Task Manager NPU ~53% ~92%
python tests\npu_visible\test_iq_throughput.py

4. Engineering Challenges & Technical Solutions

During development on native Windows 11 with the AMD IRON/AIE2 toolchain, several architecture-specific hurdles were identified and resolved:

1. XRTTensor Host Buffer Alignment & Type Constraints

  • Issue: Passing 16-bit integers directly to XRTTensor failed with TypeError: Cannot cast array data from dtype('int16') to dtype('uint32') according to the rule 'same_kind'.
  • Root Cause: AMD XRT DMA host buffers enforce 32-bit word alignment.
  • Solution: Packed adjacent 16-bit operands ($I/Q$ sample pairs or modular $(A, B)$ polynomials) into native uint32 arrays on the host, unpacking them into SIMD vector registers inside the AIE2 kernel.

2. AIE2 Tile Local Data Memory Overflow (64 KB Bank Budget)

  • Issue: Kernel compilation aborted with [aiecc] error: 'aie.tile' op allocated buffers exceeded available memory.
  • Root Cause: AIE2 tile data memory is strictly 64 KB (divided into four 16 KB banks). Allocating 16 KB double-buffered ping-pong ObjectFIFOs for both input and output exceeded available SRAM.
  • Solution: Right-sized burst buffer lengths to 1024 elements (4 KB per buffer), allowing input/output ping-pong buffering ($16\text{ KB}$ total) while reserving remaining banks for kernel stack and precomputed twiddle LUTs.

3. Peano Header Resolution Across Cache Directories

  • Issue: Peano failed with fatal error: 'sdr_dsp/...' file not found during JIT compilation.
  • Root Cause: IRON generates temporary compilation units in user cache directories (%USERPROFILE%\.npu\cache\...), breaking relative C++ include paths.
  • Solution: Implemented self-contained kernels or programmatically forwarded absolute include paths via include_dirs=[cxx_header_path(), str(include_sdr_dir)].

4. Decimation-in-Time NTT Twiddle Table Stride Indexing

  • Issue: 16-Point and 256-Point NTT stage-2/stage-3 butterflies exhibited non-trivial bin mismatches against direct $O(N^2)$ DFT.
  • Root Cause: Radix-2 Decimation-in-Time (DIT) butterflies require twiddle powers $\omega^{j \cdot (N / 2^s)}$ at stage $s$. Using flat twiddle indexing introduced phase errors.
  • Solution: Derived programmatic stage stride step indexing (W[j * (N >> stage)]), achieving bit-exact match ($0$ error) across all transform sizes.

5. M17 FFT Stage-1 Butterfly Inversion

  • Issue: The initial radix-2 M17 FFT produced valid magnitude but wrong bin ordering. Every stage after the first had a systematic butterfly-index inversion, yielding SNR $< 12$ dB against numpy.fft.fft.
  • Root Cause: The Stockham auto-sort schedule pairs indices $(k, k + m/2)$ where $m$ is the current subtransform size. The initial kernel applied the twiddle to the wrong lane of each pair.
  • Solution: Rewrote as radix-4 Stockham (twiddle applied per quadruplet, autonomous permutation between stages), reaching 138.79 dB forward SNR vs NumPy — better than double-precision floor for a 64-point transform.

6. IFFT Without Separate Device Code

  • Issue: Shipping a full inverse-FFT kernel would double the memory + build footprint of M17.
  • Solution: Applied the identity IFFT(Y) = conj(FFT(conj(Y))) / N in the host driver. The M17 forward kernel is reused as-is; only conjugate-and-scale runs on the host. Silicon result: 135.11 dB round-trip SNR on random complex vectors.

7. Upstream mlir-aie v1.4.1 iron.Runtime API Break

  • Issue: After moving to upstream mlir-aie v1.4.1 (pin commit 3ca0193, 2026-08-14), the full silicon sweep failed 14/16 milestones with Runtime.__init__() missing 1 required positional argument: 'seq_fn'.
  • Root Cause: Upstream deprecated the Runtime() context-manager pattern in favor of Runtime(seq_fn, fn_args=[...]), and moved worker enrollment and task-group management out of the Runtime object into Program(..., workers=[...]) and per-sequence TaskGroup objects.
  • Solution: Migrated all 12 iron-based tests in one sweep. Single-worker kernels use the new Runtime(seq_fn, [...]) signature; multi-worker channelizers use TaskGroup() inside the sequence body with tg.finish() at the end and endpoint-native prod_ep.fill(buf, tap=tap, group=tg) for per-column DMA. Full detail in docs/ROADMAP.md.

8. AIE2 Program-Memory Overflow on Unrolled ObjectFifo Loops

  • Issue: A 64-iteration for around acquire/mix/release in the I/Q throughput worker failed aiecc at CDO load with _XAie_LoadProgMemSection(): Overflow of program memory / XAie_LoadElf failed.
  • Root Cause: AIE2 core program memory is 16 KB. Static ObjectFifo lowering unrolled the frame loop into the ELF .text section.
  • Solution: One acquire/release per worker, Worker(while_true=True, dynamic_objfifo_lowering=True), and a 64-frame host TAP. IRON streams the tokens; the core binary stays compact.

5. Quickstart & Silicon Verification

The Installation command already runs the canonical regression after a successful install. To rerun it later from the clone:

py .\run_all_silicon_tests.py

Expected output (mlir-aie v1.4.1 pin 3ca0193, cached xclbin):

======================================================================
                     REGRESSION EXECUTION SUMMARY
======================================================================
 [ PASS ] Milestone 3: Single-Core SAXPY Vector Operation                        (1.11s)
 [ PASS ] Milestone 5: 8-Tap Vectorized Low-Pass FIR Filter                      (1.16s)
 [ PASS ] Milestone 6: Complex Mixer / NCO Frequency Downconverter               (1.03s)
 [ PASS ] Milestone 7: Vectorized Power / RSSI Energy Detector                   (1.07s)
 [ PASS ] Milestone 8: Streaming Multi-Stage Fused Demodulator Pipeline          (1.08s)
 [ PASS ] Milestone 9: 4-Column Parallel FIR Filter (Hardware Scaling)           (1.04s)
 [ PASS ] Milestone 9b: 4-Column Parallel Multi-Stage Demodulator Pipeline       (1.18s)
 [ PASS ] Milestone 10: Modular Arithmetic & Barrett Reduction (mod 3329)        (1.08s)
 [ PASS ] Milestone 11: Radix-2 NTT Butterfly Kernel (mod 3329)                  (1.03s)
 [ PASS ] Milestone 12: CPU NTT/INTT Reference & Constant Generator              (0.28s)
 [ PASS ] Milestone 13: 16-Point Vectorized NPU NTT (64 Batches)                 (0.62s)
 [ PASS ] Milestone 14: 256-Point Vectorized NPU NTT (4 Batches)                 (1.22s)
 [ PASS ] Milestone 15: NPU INTT & Cyclic Polynomial Multiplication              (1.18s)
 [ PASS ] Milestone 15b: NPU Negacyclic Polynomial Multiplication (Kyber ring)   (0.63s)
 [ PASS ] Milestone 17: 64-Point Radix-4 Stockham FFT + IFFT (NPU1)              (1.07s)
 [ PASS ] Milestone 17p: 4-Column Parallel 64-Point FFT Channelizer              (2.68s)
----------------------------------------------------------------------
 Total Tests Run: 16 | Passed: 16 | Failed: 0
 Total Elapsed Time: 17.46 seconds

Optional: I/Q throughput

Not part of run_all_silicon_tests.py. After the suite, or instead of it:

python tests\npu_visible\test_iq_throughput.py

Expected on Phoenix NPU1: first-buffer $L_\infty = 0.007812$, then ~7.5 Msps / ~30 MB/s I/Q in over a 5 s window. See I/Q Throughput.


6. References & Upstream Projects


7. Credits & Acknowledgments

  • Lead Architect & Maintainer: Midhat Nashar (@midhatn)
  • AI Architecture & Engineering Partner: Perplexity AI (Senior AMD XDNA / AIE & DSP Copilot)
  • Upstream toolchain (Advanced Micro Devices, Inc. — formerly Xilinx): mlir-aie, llvm-aie (Peano), XRT, xdna-driver, and the FFT_R4_AIE radix-4 Stockham FFT reference (Apache-2.0), which kernels/fft_stockham_f32.cc is adapted from.
  • Academic foundations: J. W. Cooley & J. W. Tukey (1965) for the radix-2 FFT that seeds M16/M17; P. Barrett (1986) for the modular-reduction method underlying M10–M15b.
  • Community reference: hal-lab-u-tokyo/ntt-aie NTT-on-AIE reference implementation.
  • License: Apache License 2.0 for original project work. File-level exceptions and third-party notices remain authoritative; see LICENSE, LICENSE_HISTORY.md, NOTICE, and THIRD_PARTY_NOTICES.md. Immutable upstream anchors and local SHA-256 identities are recorded in THIRD_PARTY_PROVENANCE.md. Citation metadata is provided in CITATION.cff and .zenodo.json.

About

Research-grade DSP, SDR, transforms, and the retained M32/M33 PQC foundation on AMD Phoenix XDNA1 using MLIR-AIE, IRON, Peano, and XRT.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages