Skip to content

Latest commit

 

History

History
286 lines (213 loc) · 49.8 KB

File metadata and controls

286 lines (213 loc) · 49.8 KB

Phoenix SDR-DSP Roadmap

This roadmap records repository-local milestones and historical planning notes. The previously cited external “Phoenix SDR-DSP Master Prompt” is not tracked in this repository, so it is not used as a reproducibility dependency. Every technical claim in this document is intended to be citable to a primary source; the References section at the bottom collects all sources with canonical URLs.

Project positioning

phoenix-sdr-dsp is a bit-accurate, silicon-validated NPU DSP kernel library for the first-generation AMD XDNA1 architecture as implemented in the Ryzen 7040-series "Phoenix" APUs. The XDNA1 NPU is a spatial dataflow array of AI Engine (AIE2) tiles, organized as a 4×5 grid of compute tiles plus memory tiles and shim DMAs, per AMD's official architecture description and the upstream Linux amdxdna kernel driver documentation. Each AIE2 tile is a VLIW SIMD vector processor with 512-bit vector datapath and local L1 program/data memory (AMD XDNA overview; IEEE Micro 2024 "AMD XDNA NPU in Ryzen AI Processors").

The software toolchain used across all shipped milestones is the open-source MLIR-AIE / IRON stack (Xilinx/mlir-aie GitHub; IRON documentation v1.4.1), which provides close-to-metal Python bindings that compile to MLIR and then to AI Engine core binaries via the Peano LLVM backend (Peano/llvm-aie GitHub; Peano announcement on LLVM Discourse, 2024; AMD IRON tutorial PDF, MICRO 2024). Device dispatch uses the Xilinx Runtime (XRT) with the XDNA shim.

The SDR-integration track (LimeSDR, ring buffers, real-time streaming) is deferred until suitable SDR hardware is available for validation. The pure-DSP track (transforms, filters, modular arithmetic, NTT/FFT) continues without RF-hardware dependency.

Post-v1.0.0 development plan

This document is the v1.0.0 closure log. The forward-looking roadmap — RF-live SDR demonstrations, CPU-baseline benchmarks, and community integration — lives at ROADMAP_v1.1.0.md.

Status legend

  • ✅ Shipped — implemented and present in the repository. The row must state whether its validation is hardware-backed or reference-only; presence in run_all_silicon_tests.py alone is not proof of NPU execution.
  • 🧪 Shipped, unintegrated — silicon-validated but not part of the regression runner; scheduled for renumbering and integration.
  • 🚧 Next up — no hardware dependency, actively planned.
  • 🔒 Deferred — SDR hardware — blocked pending acquisition of a supported SDR device.
  • ⏸️ Deferred — depends on prior deferred milestone — blocked transitively.
  • 💡 Optional / research — post-v1.0.

Foundational milestones (M0–M2)

M# Focus Status Artifact
M0 Windows environment audit ✅ scripts/windows/windows_audit.ps1, audit/
M1 Native Windows architecture decision ✅ docs/M1_ARCHITECTURE_DECISION.md
M2 Pinned Windows toolchain ✅ docs/M2_TOOLCHAIN_PIN.md, toolchain.yaml

The native-Windows execution path is used because MLIR-AIE and Peano ship first-class native-Windows wheels (MLIR-AIE 1.2 release notes, Phoronix) and because the amdxdna NPU driver binds to the Windows host — WSL2 cannot directly access the NPU device (Linux amdxdna documentation). See docs/M1_ARCHITECTURE_DECISION.md for the full rejection rationale.

Track 1 — NPU DSP kernels (historical v0.4.0 suite: 16/16)

Vector & scalar primitives (canonical §16 M3–M6)

M# Focus Status Notes
M3 Native Windows MLIR-AIE pass-through example ✅ Shipped as SAXPY (tests/m3_saxpy/). SAXPY (y = a·x + y) is a strict superset of the pass-through demo the master prompt calls for; SAXPY exercises vector multiply-add and validates the compile → device-load → buffer → kernel → verify path in one step.
M5 Native Windows NPU vector-copy kernel ✅ Shipped as 8-tap FIR (tests/m5_fir/) — vector-copy is a degenerate FIR (tap = identity). Divergence documented; no separate vector-copy milestone.
M6 Vector arithmetic + complex I/Q primitives ✅ Shipped as complex mixer/NCO (tests/m6_mixer/). Covers complex multiply-add, the core primitive for downstream mixers, correlators, and phase rotators.

DSP kernels (extensions beyond §16 numbering)

M# Focus Status Notes
M7-ext Power / RSSI energy detector ✅ tests/m7_power/. Not in §16 sequence; useful for spectrum monitoring, squelch, and correlation-based preamble detection.
M8-ext Fused DSP pipeline (mixer + FIR + power) ✅ tests/m8_pipeline/. Multi-stage kernel fusion pattern, no SDR yet.
M9-ext 4-column parallel FIR ✅ tests/m9_parallel/. Multi-column parallelization exercises the XDNA1 4×5 tile grid geometry (Linux amdxdna docs); pattern reusable for future kernels.
M10-ext 4-column parallel multi-stage demodulator pipeline ✅ tests/m9b_parallel_pipeline/. In run_all_silicon_tests.py as Milestone 9b.
demo-iq 4-column streamed I/Q throughput 🧪 tests/npu_visible/. Host-visible IRON+DMA mixer measured 2026-08-15 on Ryzen 9 7940HS: 7.459 Msps, 29.84 MB/s I/Q in, ~92% Task Manager NPU, first-buffer $L_\infty = 0.007812$. Phoenix NPU is 10 TOPS (INT8). Not in run_all_silicon_tests.py. Kernel vectorization deferred.

Modular arithmetic & NTT track (canonical §16 M10–M15)

The NTT track implements the Number-Theoretic Transform, a finite-field analogue of the Discrete Fourier Transform that runs bit-exactly over integers modulo a prime (Emergent Mind NTT survey; Ingonyama ICICLE NTT documentation). The prime modulus q = 3329 and polynomial degree N = 256 are the CRYSTALS-Kyber / ML-KEM parameters used by NIST's post-quantum key-encapsulation standard (arXiv 2601.17806, Algorithm-Targeted NTT Hardware Acceleration, 2026; PLOS ONE, Area-time efficient pipelined NTT for CRYSTALS-Kyber).

M# Focus Status Artifact
M10 Modular arithmetic (Barrett reduction, q = 3329) ✅ tests/m10_modular/. Uses the Barrett-reduction algorithm from PLOS ONE 2025.
M11 Radix-2 NTT butterfly ✅ tests/m11_butterfly/. Cooley-Tukey butterfly generalized to finite field, form X = U + T, Y = U − T where T = V · ω mod q (Emergent Mind NTT topic).
M12 CPU NTT/INTT reference ✅ tests/m12_ntt_ref/. Bit-exact software reference the NPU implementation is verified against.
M13 16-point NPU NTT ✅ tests/m13_ntt16/
M14 256-point vectorized NPU NTT ✅ tests/m14_ntt256/. Matches Kyber N = 256.
M15 NPU INTT + cyclic polynomial multiplication ✅ tests/m15_polymul/
M15+ Negacyclic polynomial multiplication (Z_q[x]/(x^N + 1)) ✅ tests/m15b_negacyclic/test_negacyclic_m16.py. Kyber / ML-KEM ring per Isabelle/AFP. Ported 2026-08-15 to the same iron.Runtime sequence-function API as M15. Schoolbook O(N²) kernel, bit-exact on Phoenix NPU1. Filename still says m16; keep until a dedicated renumber. The FIPS 203 KEM itself is M32, not this row.

FFT track (canonical §16 M16–M18)

The Fast Fourier Transform in radix-2 form is the Cooley–Tukey algorithm (1965), which reduces the DFT operation count from O(N²) to O(N log N) by recursive decomposition into even/odd subsequences (Rice University FFT tutorial; Brian McFee, Digital Signals Theory §8.2).

M# Focus Status Notes
M16 CPU DFT/FFT reference ✅ tests/m16_fft_ref/test_fft_reference_m16.py. Three independent implementations cross-validated: direct O(N²) DFT via twiddle matrix, recursive radix-2 Cooley-Tukey 1965 FFT, iterative in-place radix-2 FFT with bit-reversed permutation (dataflow proxy for the M17 NPU kernel). All match NumPy fft.fft to double-precision round-off (~1e-13 relative). Tests: impulse, DC constant, pure tone, random complex, x = IFFT(FFT(x)) round-trip, Parseval energy conservation. Sizes N ∈ {8, 16, 32, 64, 128, 256, 512, 1024}. Runs on Ubuntu in CI in ~0.3 s.
M17 NPU FFT/IFFT ✅ Shipped as 64-point radix-4 Stockham FFT (tests/m17_radix2_fft/test_fft_m17_v3.py) delivering O(N log N) complexity per Cooley-Tukey 1965 and the Stockham auto-sort variant that avoids bit-reversed permutation. Forward FFT SNR 138.79 dB vs numpy.fft.fft. Round-trip IFFT via conj(FFT(conj(Y)))/N reuses the forward kernel with no separate IFFT device code, RMS SNR 135.11 dB. Supersedes the earlier direct-DFT prototype in tests/m17_fft_dft/ (retained pending removal). Wired into run_all_silicon_tests.py.
M17-parallel 4-column parallel NPU FFT ✅ tests/m17p_fft_parallel/. Parallel channelizer running the M17 radix-4 Stockham kernel across all four AIE2 tile columns of the Phoenix NPU1 grid (Linux amdxdna docs). Batch throughput 1,993 FFTs/sec on 64 parallel 64-point frames. Wired into run_all_silicon_tests.py.
M17-butterfly NPU FFT via radix-2/radix-4 Cooley-Tukey butterflies ✅ Delivered by the M17 radix-4 Stockham kernel above. Row retained for §16 traceability.
M18 Streaming FFT spectrum analyzer connected to SDR 🔒 Requires SDR hardware.

Filtering & resampling (canonical §16 M19–M23)

M# Focus Status Notes
M19 Complex FIR filter ✅ Real-valued 8-tap FIR shipped as tests/m5_fir/. Complex-valued (complex taps × complex I/Q input) 8-tap variant shipped as tests/m19_complex_fir/ and wired as the 17th silicon regression entry. Design: docs/M19_DESIGN.md.
M20 Polyphase decimation & interpolation ✅ Fused decim-M=4 + interp-L=4 kernel on one AIE2 core shipped as tests/m20_polyphase/, 16-tap Kaiser-window prototype LPF (Kaiser 1974) with taps *= L interpolator scaling matching scipy.signal.resample_poly and GNU Radio pfb. Polyphase decomposition per Vaidyanathan 1993 ch. 4 and Harris 2004 ch. 6. Silicon PASS at atol=0.01 on random I/Q. Wired as the 18th silicon regression entry. Design: docs/M20_DESIGN.md.
M21 Digital downconverter (DDC) ✅ Fused DDC (negative-exponent complex NCO with f_LO = +f_s/8 + 16-tap Kaiser LPF + decim-by-M=4) on one AIE2 core, shipped as tests/m21_ddc/. 8-sample cordic-free LO LUT per Analog Devices MT-085 "Fundamentals of DDS". Signal chain follows Harris 2004 ch. 8 "The Digital Down-Converter"; block topology matches the GNU Radio Frequency Xlating FIR Filter reference. LPF reuses the M20 16-tap Kaiser prototype (Kaiser 1974). Silicon PASS at max err 0.003906 (atol=0.01) on random I/Q (seed 789); on-carrier tone at +f_s/8 decodes to mag=1.0000/phase=0.0000, image rejection 55.8 dB at -f_s/8. Wired as the 19th silicon regression entry. Design: docs/M21_DESIGN.md.
M22 Digital upconverter (DUC) ✅ Fused DUC (zero-stuff interp by L=4 + 16-tap Kaiser×L LPF + complex NCO at f_c = +f_s/8) on one AIE2 core, shipped as tests/m22_duc/. Zero-stuff-and-filter interpolation follows the polyphase commutator identity Vaidyanathan 1993 Eq. 4.3.13 and Harris 2004 ch. 7; DUC signal chain per Harris 2004 §8.4 "The Digital Up-Converter"; block topology matches the GNU Radio Frequency Xlating FIR Filter run with negative decimation. Interpolator tap scaling taps *= L follows the scipy.signal.resample_poly convention. 8-sample cordic-free LO LUT (sign-flipped from M21) per Analog Devices MT-085. LPF prototype reuses the M20 16-tap Kaiser design (Kaiser 1974). Silicon PASS at max err 0.007812 (atol=0.01) on random I/Q (seed 792); DC baseband upconverts to +f_s/8 at mag 0.9976 (FFT peak at bin 192), baseband tone at -f_bb/8 lands at +3f_s/32 (FFT peak at bin 144). Wired as the 20th silicon regression entry. Design: docs/M22_DESIGN.md.
M23 Channelizer & filter bank ✅ Fused polyphase channelizer (input commutator + M=8-path 8-tap Kaiser FIR + 8-point matmul-DFT) on one AIE2 core, shipped as tests/m23_channelizer/. M-path analysis-bank topology per Harris 2004 ch. 6 §6.3 Fig. 6.8; polyphase commutator identity per Vaidyanathan 1993 §4.3, Eq. 4.3.13. Natural sample-to-branch order (p = q) matches the GNU Radio pfb_channelizer_ccf and NVIDIA MatX channelize_poly conventions. 64-tap Kaiser prototype (β ≈ 5.653, cutoff π/M, 60 dB stop-band) per Kaiser 1974 via scipy.signal.firwin with scale=True — sum(h) = 0.99977, exact even symmetry. 8-point DFT uses fully-embedded twiddles (matmul-style, same pattern as M17p parallel_fft64_kernel.cc). Silicon PASS at max err 0.003906 (atol=0.02) on random I/Q (seed 793); DC→ch0 iso 66.2 dB, tone→ch3 iso 66.2 dB, two-tone→ch1+ch5 iso 64.5 dB. Sandbox transliteration is bit-exact to host reference (0/4096 slots differ). Wired as the 21st silicon regression entry. Design: docs/M23_DESIGN.md.

Modulation & synchronization (canonical §16 M24–M27, partially SDR-blocked)

M# Focus Status Notes
M24 Correlation, preamble detection, packet sync ✅ Fused Barker-13 matched-filter correlator on one AIE2 core, shipped as tests/m24_correlator/. Matched-filter theory follows Proakis & Salehi 5e §5.1.5 and Massey 1972; correlation-as-reversed-FIR identity per Oppenheim & Schafer 3e §2.6.2; block topology matches the GNU Radio Correlation Estimator and liquid-dsp detector_cccf. Barker-13 preamble (+1,+1,+1,+1,+1,-1,-1,+1,+1,-1,+1,-1,+1) per Barker 1953 and Wikipedia "Barker code"; PSL = 1 (
M25 BPSK / QPSK receiver pipeline ✅ Fused BPSK/QPSK receiver on one AIE2 core, shipped as tests/m25_psk_rx/. Signal chain: on-tile NCO derotator (e^{-jθ[k]}, open-coded 7th-order Taylor sin/cos with π/2 fold since Peano NOCPP has no libc <math.h>), linear fractional interpolator with fractional delay μ ∈ [0, 1), Gardner 1986 mid-symbol timing-error detector e_τ[k] = (x[k] - x[k-2]) · x[k-1], an order-2 or order-4 Costas phase detector, and two second-order PI loop filters with Rondeau 2011 gains. Order-2 detector e_φ = z_I · z_Q for BPSK per Costas 1956, wirelesspi Costas, and GNU Radio Costas wiki; order-4 decision-directed cross form e_φ = z_I · sgn(z_Q) - z_Q · sgn(z_I) for QPSK per US Patent 4344178A and GNU Radio costas_loop_cc phase_detector_4. Single templated psk_rx_body<ORDER> C++ body with two @iron.jit entry points (psk_rx_bpsk/psk_rx_qpsk); I/O bfloat16, internal math float32. Silicon PASS on Phoenix NPU1: BPSK order-2 seed 795 gate (a) max err 0.003906, gate (b) |z| median 0.9961, RMS Costas error 0.0000; QPSK order-4 seed 796 gate (a) max err 0.003906, gate (b) |z| median 0.9975, RMS residual phase 0.2841 rad (< π/8 = 0.3927 per the canonical lock criterion of NASA JPL TDA Progress Report 42-130). Sandbox transliteration is bit-exact to host reference on both orders (0/1024 slots differ). Four bring-up incidents documented in docs/M25_DESIGN.md §4b: (1) Peano NOCPP has no libc <math.h> → open-coded Taylor sin/cos + π/2 fold; (2) scalar (x>=0)?1:-1 miscompiles → IEEE-754 sign-bit reinterpret via union{float;uint32_t}; (3) Peano -O2 folds the union form into llvm.copysign which AIE2 rejects as unable to legalize G_FCOPYSIGN → integer-OR into 0x3F800000 with a volatile uint32_t intermediate to defeat the pattern-matcher; (4) Costas + Gardner is a closed-feedback dynamical system, so CPU vs AIE2 float32 rounding integrates to different steady-state equilibria after ~1/BW_φ symbols — not fixable in kernel per NASA JPL TDA 42-130, Kuznetsov et al 2018 arXiv:1810.00071, and Analog Devices "Practical Costas Loops"; PASS gate revised to three receiver-theoretic checks (acquisition atol, steady-state |z| median + RMS phase residual under π/8, diagnostic-only first-divergence slot). Wired as the 23rd silicon regression entry. Design: docs/M25_DESIGN.md.
M26 QAM receiver pipeline ✅ Fused QAM-16 receiver on one AIE2 core, shipped as tests/m26_qam_rx/. Extends the M25 receiver core (on-tile NCO derotator with Taylor sin/cos + π/2 fold, linear fractional interpolator, Gardner 1986 mid-symbol TED, Rondeau-tuned PI loop filters) with three new blocks: a Gray-labelled QAM-16 hard-decision slicer on the unit-average-energy {±1, ±3}/√10 constellation (Proakis & Salehi 5e §4.3.1; Rice 2e §5.3), a decision-directed order-M phase detector e_φ = z_I · â_Q - z_Q · â_I (Godard 1980; Barry-Lee-Messerschmitt 3e §8.5), and a max-log soft-output demapper emitting 4 LLRs per symbol via the axis-separable closed form LLR(b_MSB) ≈ 4·z_axis, LLR(b_LSB) ≈ 4·(2 - |z_axis|) (Tosato & Bisaglia 2002; Alvarado & Fabregas 2009). Single @iron.jit entry point qam16_rx with first three-argument kernel signature in the suite (in_iq, out_iq, out_llr); I/O bfloat16, internal math float32. Loop bandwidths narrowed to BW_φ = 2π/200 (half of M25) to keep the loop inside the DD detector's linear region for QAM-16's 2.24×-smaller phase margin per Rice §7.4.4. Silicon PASS gates match M25 discipline: (a) acquisition first-32-symbol max err < 0.10, (b1) steady-state median distance of |z| to nearest QAM-16 magnitude class {0.4472, 1.0, 1.3416} < 0.15, (b2) steady-state RMS(z − QAM16_slice(z)) < 0.10 at unit-average energy (the 2D constellation-error metric replaces M25's "residual angle mod π/2" because DD-QAM16 lacks π/2 rotational symmetry per Barry-Lee-Messerschmitt 3e §8.5.3 and Rice 2e §7.4.4), (c) hard-decision SER vs reference (min over QAM-16's 4-fold rotational symmetry group) — diagnostic only, not asserted per M26_DESIGN.md Amendment #1 because two independent DD + Gardner timing loops (silicon float32-SIMD vs CPU float32-serial) drift apart by 1+ symbols over the burst even when both are individually locked to a valid QAM-16 grid (M25 incident #4 pattern applied to the timing integrator; the rotation-invariant printout observed [1.0, 0.7188, 0.7344, 0.9922] on seed 826, ruling out phase ambiguity and confirming timing drift as the root cause per Gardner 1986, Barry-Lee-Messerschmitt 3e §8.5.4, and NASA JPL TDA 42-130), (d) LLR/hard consistency ≥ 0.85 for MSB bits (b3, b1) and ≥ 0.75 for LSB bits (b2, b0). Silicon PASS on Phoenix NPU1 seed 826 (2026-08-15): gate (a) max_err 0.0039, gate (b1) magnitude-class median 0.0020, gate (b2) RMS(z − slice(z)) 0.0027, gate (d) LLR MSB b3 = b1 = 1.000 and LLR LSB b2 = b0 = 1.000. Sandbox transliteration bit-exact to host reference on both hard-symbol and LLR buffers (0/1024 hardSym slots differ, 0/2048 LLR slots differ, tools/m26_kernel_transliteration_check.py on seeds 826 and 827). Inherits all four M25 bring-up mitigations verbatim (open-coded Taylor sin/cos + π/2 fold since Peano NOCPP has no libc <math.h>, dead-zone sgn_bit with volatile uint32_t OR into 0x3F800000 to defeat -O2 llvm.copysign folding, receiver-theoretic PASS gates per NASA JPL TDA 42-130, Kuznetsov et al 2018 arXiv:1810.00071, and Analog Devices "Practical Costas Loops"). Two M26-specific test-side bring-up incidents documented in docs/M26_DESIGN.md §4b: (1) initial gate (b2) borrowed M25's "residual angle mod π/2" metric which is invalid for QAM-16 because DD-QAM16 lacks π/2 cost-function symmetry — replaced with RMS(z − QAM16_slice(z)); (2) initial gate (c) asserted sample-wise SER which is unreachable because two independent DD + Gardner timing integrators drift apart by 1+ symbols — gate (c) reduced to diagnostic-only per Amendment #1. Wired as the 24th silicon regression entry. Design: docs/M26_DESIGN.md.
M27 OFDM: FFT + CP + pilots + channel estimation + equalization ✅ Fused OFDM loopback on the Phoenix NPU, shipped as tests/m27_ofdm/. The implemented receive path uses LS pilot estimates, linear interpolation across data subcarriers, and one-tap zero-forcing equalization. It does not implement LMMSE. Design: docs/M27_DESIGN.md. Wired as the 25th hardware-backed regression entry.

Track 2 — SDR integration (🔒 blocked on hardware)

All milestones below require a working SDR device. Deferred until hardware is available.

M# Focus Master prompt intent
M4 SDR enumeration + Windows streaming test (no NPU) §16: LimeSDR specifically; generalized here to any SDR via the ISdrDevice abstraction (§7 of the master prompt).
M7-canonical Continuous SDR → host ring buffer §9 real-time streaming architecture (double/triple buffering).
M8-canonical Host buffer → NPU → host streaming bridge §9 pipelined architecture.
M9-canonical Real-time pass-through of SDR I/Q through NPU §16 integration checkpoint.
M18 Streaming FFT spectrum analyzer (uses M17 + SDR) §16 M18.
M28 Beamforming & multi-channel processing Requires multi-channel SDR.
M31 Unified native Windows SDR-DSP API Final integration.

Hardware plan when unblocked: Master prompt §7 specifies LimeSDR as the primary reference device. ISdrDevice abstraction is designed to accept additional backends. Realistic candidate devices when hardware is acquired: LimeSDR (primary, per master prompt), RTL-SDR (cheap RX validation), HackRF, PlutoSDR, USRP B-series, BladeRF. API choice (native LimeSuite / UHD / hackrf / rtlsdr vs SoapySDR) is deferred to the point of hardware acquisition, following the master prompt's §7 constraint ("do not choose between Lime Suite and SoapySDR without testing").

Track 3 — Advanced / research (💡 optional)

M# Focus Notes
M29 Adaptive filtering & interference suppression Can be prototyped on synthetic data; real validation needs SDR.
M30 Optional learned AI blocks for SDR Master prompt marks explicitly optional. Depends on Track 1 primitives + SDR integration.

Track 4 — FIPS 203 ML-KEM (✅ Post-Quantum Cryptography, v1.0.0)

Extra milestone after the shipped M10–M15b stack. Numbered M32 so it does not collide with master-prompt §16 (M0–M31 are SDR/DSP). Design entry point: docs/M32_FIPS203_MLKEM.md; per-sub-milestone: docs/M32b_DESIGN.md, docs/M32c_DESIGN.md, docs/M32d_DESIGN.md, docs/M32e_DESIGN.md. Test tree: tests/m32_mlkem/.

FIPS 203 (13 August 2024, DOI 10.6028/NIST.FIPS.203) specifies ML-KEM, derived from round-3 CRYSTALS-Kyber (§1.1). Ring R_q = Z_3329[X]/(X^{256}+1) is already the M15b ring. The KEM product is Algorithms 9–12 (NTT), not M15b schoolbook. M32b, M32c, and M32d are hardware-backed; M32e covers ML-KEM-512 with 60 host KATs and a 9-vector hardware smoke set. ML-KEM-768 and ML-KEM-1024 are not claimed as implemented or validated.

M# Focus Status Notes
M32 FIPS 203 ML-KEM (Post-Quantum Cryptography) ✅ ML-KEM-512 scope only. M32e runs 60 host KATs and 9 hardware smoke vectors. Hashes from FIPS 202. Reference oracle: kyber-py.
M32b NPU NTT-domain negacyclic product ✅ FIPS 203 Algorithms 9–12 on Phoenix NPU1. Silicon-dispatched NTT kernel over Z_3329 with pq-crystals ζ-table (256-th root of unity) matching the pq-crystals reference ntt.c. Replaces M15b schoolbook for the KEM path. 26th silicon regression entry. Design: docs/M32b_DESIGN.md.
M32c Keccak-f[1600] + FIPS 202 SHA-3 / SHAKE + samplers ✅ FIPS 202 permutation-based hash and extendable-output functions plus FIPS 203 Algorithms 7–8 (SampleNTT / SamplePolyCBD) on Phoenix NPU1. Kernel dispatches SHAKE128 / SHAKE256 / SHA3-256 / SHA3-512 in five modes over one Keccak-f permutation core. 27th silicon regression entry. Design: docs/M32c_DESIGN.md.
M32d K-PKE component ✅ FIPS 203 Algorithms 13–15. Silicon-dispatched K-PKE KeyGen / Encrypt / Decrypt orchestrated on top of M32b + M32c. Not approved standalone per FIPS 203 §3.3; wrapped by M32e for the approved KEM. 28th silicon regression entry. Design: docs/M32d_DESIGN.md.
M32e ML-KEM internal-interface composer ✅ ML-KEM-512 only: 60 host KATs plus 9 hardware smoke vectors (3 each for internal key generation, encapsulation, and decapsulation). The host composes FIPS 203 Algorithms 16–18 calls to M32b/c/d through a SiliconBackend seam; public Algorithms 19–21 are not claimed. Design history: docs/M32e_DESIGN.md.

Track 5 — FIPS 204 ML-DSA (native primitives; hybrid composers)

FIPS 204 (13 August 2024, DOI 10.6028/NIST.FIPS.204) specifies ML-DSA, derived from round-3 CRYSTALS-Dilithium. Ring R_q = Z_q[X]/(X^{256}+1) with q = 8380417. M33a/M33b use native, fail-closed runners. M33d/e compose those primitives from Python and remain hybrid rather than fully device-resident.

M# Focus Status Notes
M33 FIPS 204 ML-DSA (Post-Quantum Cryptography) ✅ Native M33a/M33b polynomial primitives plus host/NPU ML-DSA-44/65/87 composition. Reference oracle: dilithium-py.
M33a ML-DSA NTT / INTT / basemul / reduce ✅ Native Phoenix silicon, 420/420 PASS. Design: docs/M33a_DESIGN.md.
M33b ML-DSA rounding / hint / norm ✅ Native Phoenix silicon, 700/700 PASS. Design: docs/M33b_DESIGN.md.
M33c SHAKE / Keccak 🚧 Current M33 composers use host hashlib; streamed or reusable device SHAKE is required for fully device-resident ML-DSA.
M33d ML-DSA.KeyGen composer ✅ Host/NPU composition for ML-DSA-44/65/87, 75/75 PASS using native M33a/M33b. Design: docs/M33d_DESIGN.md.
M33e ML-DSA.Sign_internal + Verify_internal composer ✅ Host/NPU composition, Sign 90/90 and Verify 90/90, using native M33a/M33b. Design: docs/M33e_DESIGN.md.

Corrected v1.0.0 validation boundary

The runner contains 34 invocations: 29 direct-hardware, four host/NPU composers, and one intentional CPU reference (M12). It completed 34/34 PASS in 126.29 seconds on 2026-08-17. M32e is limited to ML-KEM-512; M33d/e are not fully device-resident.

The stack (Xilinx XRT, Xilinx MLIR-AIE v1.4.1 + pin 3ca0193, LLVM Peano 21.0.0.2026080301+c9c5ecb7) is unchanged from v0.4.0; the PQC track adds only two Python-level reference packages (kyber-py, dilithium-py) plus pytest inside the existing ironenv. All SHAKE / SHA-3 primitives come from the CPython hashlib standard library, so no separate SHAKE / Keccak wheel is required. Since v1.0.0, install.py auto-installs these packages into ironenv.

Completed — v0.4.0 silicon + new-user install (2026-08-15)

The DSP / NTT / FFT kernel library is closed at 16/16 PASS. A wipe-and-clone of main ran install.py then run_all_silicon_tests.py and passed in 95.91 s (cold xclbin). Cached re-run on the development tree is 17.46 s. Stack: Xilinx XRT 2.21.75, Xilinx MLIR-AIE / IRON v1.4.1, LLVM Peano 21.0.0.2026080301+c9c5ecb7. New-user path is documented on the landing page Installation section.

Items 1–5 below are historical. The live next NTT item is M32a.

  1. v0.2.1 polish — completed. Dependabot, CI badge, v0.4.0 tag.
  2. Directory renumbering pass — completed in v0.2.1. Scheme B directories renamed to §16 canonical names. Blob SHAs preserved so git log --follow still tracks each file.
  3. M16 CPU FFT reference — completed. tests/m16_fft_ref/test_fft_reference_m16.py; CI cpu-reference-tests on Ubuntu.
  4. M17-butterfly — completed. Radix-4 Stockham FFT at 138.79 dB forward / 135.11 dB round-trip SNR.
  5. M15b negacyclic port to iron.Runtime — completed 2026-08-15. Schoolbook kernel, bit-exact. Closes the 16-suite.
  6. M32 FIPS 203 ML-KEM — historical planning entry. M32 is now represented in the current matrix, with M32e limited to ML-KEM-512 internal interfaces (Algorithms 16–18), not a public Algorithms 19–21 claim. Historical plan: docs/M32_FIPS203_MLKEM.md.
  7. M19 complex FIR — DSP-track alternative: extend the shipped real-valued FIR to complex taps × complex I/Q.
  8. M20 polyphase — completed. Fused decim + interp with scipy-convention tap scaling on top of M19.
  9. M21 DDC — completed. Fused complex-NCO + Kaiser-LPF + decim-by-4 on one AIE2 core; reuses M20's LPF prototype and adds an 8-sample cordic-free LO LUT.
  10. M22 DUC — completed. Mathematical mirror of M21: fused zero-stuff interp-by-L=4 (polyphase, taps *= L) + 16-tap Kaiser LPF (reused from M20) + complex NCO at +f_s/8 on one AIE2 core.
  11. M23 channelizer — completed. Fused M=8 polyphase analysis bank (input commutator + M-path 8-tap Kaiser FIR + 8-point matmul-DFT with fully-embedded twiddles) on one AIE2 core; closes the DSP-track filtering & resampling block (M19–M23).

Toolchain events

2026-08-14 — Upstream mlir-aie v1.4.1 runtime API break (pinned at commit 3ca0193)

Upstream release mlir-aie v1.4.1 (2026-08-11) reshaped the aie.iron.Runtime surface incompatibly with the pre-v1.4.1 tests. The Phoenix SDR-DSP regression is pinned at commit 3ca0193 — "Retain executable per kernel handle to fix run_chain use-after-free" (2026-08-14, v1.4.1 + 13 commits) because that commit additionally contains the run_chain executable-lifetime fix required by the parallel-DMA milestones (M9, M9b, M17p). The four API changes introduced at v1.4.1 are:

  • The context-manager form rt = Runtime(); with rt.sequence(...) as (...): was removed. Runtime.__init__ now requires a seq_fn: Callable positional argument, verified against python/iron/runtime/runtime.py at that revision.
  • rt.start(worker) was replaced by passing workers directly to Program(device, rt, workers=[...]).
  • rt.task_group() / rt.finish_task_group(tg) were replaced by a per-sequence TaskGroup constructed inside the sequence body with tg.finish().
  • rt.fill(fifo.prod(), buf, tap, task_group=tg) was replaced by endpoint-native prod_ep.fill(buf, tap=tap, group=tg), with the endpoints passed as fn_args to Runtime(seq_fn, fn_args=[...]). The canonical multi-worker + TaskGroup shape is documented in the upstream runtime-sequence test suite and the single-core matmul example programming_examples/getting_started/03_matrix_multiplication_single_core/matrix_multiplication_single_core.py.

An initial silicon sweep after the pull failed 14 of 16 milestones with an identical Runtime.__init__() missing 1 required positional argument: 'seq_fn' traceback. All 12 iron-based tests were ported to the new API in a single sweep:

  • Single-worker: M3 SAXPY (new tests/m3_saxpy/, first canonical port using upstream saxpy.cc), M5 FIR, M6 mixer, M7 power, M8 fused pipeline, M10 modular arithmetic, M11 NTT butterfly, M13 16-point NTT, M14 256-point NTT, M15 cyclic polymul.
  • Multi-worker with TaskGroup + per-column taps: M9 4-column FIR, M9b 4-column demod pipeline (2-input), M17-parallel 4-column FFT channelizer.

Post-migration silicon sweep: 15 / 16 PASS on Phoenix NPU1 (AIE2, Win11). Only tests/m15b_negacyclic/ still used the low-level aie.dialects + runtime_sequence + XRTTensor API rather than iron.Runtime.

The iron ports of M3–M15 / M17 / M17p landed on feat/m17-radix2-fft-npu, fast-forwarded to main, and pushed as commit 1ec80c8.

2026-08-15 — M15b iron.Runtime port closes the suite

M15b was rewritten to the M15 host shape: @iron.jit, ExternalFunction, two input ObjectFifos plus one output, Runtime(seq_fn), Program(..., workers=[...]), and uint32 XRTTensor buffers (the v1.4.1 host tensor rejects int32 same_kind copies). The schoolbook kernel and Barrett constants (MU = 20165, shift 26) are unchanged. Ring is the Kyber / ML-KEM ring Z_3329[x]/(x^256+1). Silicon result: bit-exact vs negacyclic_polymul_ref, seed 42. Full suite 16 / 16 PASS in 17.46 s on Phoenix NPU1.

2026-08-15 — install.py + clean-clone 16/16

install.py is on main. A real-user wipe-and-clone downloaded the published mlir_aie 1.4.1 cp313 wheel (not rolling latest-wheels-4), put VS llvm-objcopy on PATH for the Peano Windows fixup, and retried shallow fetch over HTTP/1.1. run_all_silicon_tests.py re-execs checkout ironenv because py binds to system CPython. Clean-clone suite: 16/16 PASS in 95.91 s (cold xclbin) on Phoenix NPU1.

Divergences from master prompt §16 — honest disclosure

  • Repo's early M-numbers (M3, M5, M6) map to §16 concepts but with different exact scope. M3 SAXPY covers §16 M3 pass-through with extra work; M5 FIR is not §16 M5 (vector-copy); M6 mixer is a subset of §16 M6 vector-arithmetic. Numbering is preserved for git history stability; canonical mapping documented in this table.
  • Repo M7, M8, M9 are DSP-track additions, not the §16 SDR-integration milestones. This roadmap suffixes them -ext and reserves the canonical §16 M7/M8/M9 numbering for Track 2.
  • Two m10_, m11_, m12_ test directories exist (one Scheme A modular/NTT, one Scheme B parallel-pipeline/FFT). Step 2 of the "Immediate next steps" resolves this via renumbering.
  • FFT was originally shipped as direct O(N²) DFT. As of the M17 v3 radix-4 Stockham port the shipped version is O(N log N) and aligns with §16 intent; the earlier direct-DFT prototype in tests/m17_fft_dft/ is retained pending removal. Prior divergence documented for history.
  • SDR-integration milestones (M4, M7-canonical, M8-canonical, M9-canonical, M18, M28, M31) are all unshipped — the master prompt's original ordering has them preceding the NTT track. In this repo they are deferred to Track 2 pending hardware.

References

Hardware architecture

Toolchain

Cooley-Tukey FFT

Number-Theoretic Transform (Kyber / ML-KEM)

Project-internal references

  • Repository-local milestone and validation policy: this document and MILESTONES_AND_MATHEMATICS.md
  • Shipped milestone details: MILESTONES_AND_MATHEMATICS.md
  • Engineering rules: master prompt §13
  • Response format for new milestones: master prompt §20