| 2026-06-16 | Animated splash-screen preview on the NV3007 LCD (`splash-test` feature + `make splash-test-hw`) | Ported the three `assets/splash-1{6,7,8}-*.standalone.html` revisions (hyperspace / horizon / nebula) — which share one identical `<canvas>` animation registry — to a no_std Rust renderer in `secure/src/ui/splash_test.rs`. Renders into a 1-bit packed **landscape** framebuffer (428×142, 7.6 KB — a full RGB565 native frame would be 119 KB vs the 192 KB secure SRAM budget) then blits via the SAME landscape→native transpose the trusted-UI text path uses (`ui::lcd` `FLIP=(true,false)` → native `(nx,ny)=(141-ly, lx)`). f32 math via a new **optional** `micromath` dep (pure-Rust no_std approximations; `num-traits` off; only compiled under `splash-test`). `main()` short-circuits into `ui::splash_test::run()` after `ui::init()` brings the panel up + SysTick starts (so `timeout::now()` is the animation clock); cycles the 3 revisions ~12 s each forever. One intentional divergence: integer-LCG PRNG (the browser's f64 multiply loses precision before masking) → star coordinates differ but field density/character match. **Verified:** clean `cargo check` for thumbv8m (zero diagnostics in the new module), non-splash builds unaffected, `WORDMARK` byte-identical to source, and a 4-agent adversarial fidelity workflow (per-animation + scaffolding/orientation reviewers, each finding independently verified) returned **0 confirmed divergences**. **On-silicon confirmed (all 3 render correctly), then an optimization pass landed for smoothness:** (1) blit rewritten from 428 per-row `set_window` calls/frame to ONE full-frame `set_window` + a single continuous chunked RAMWR stream via a new `lcd_nv3007::write_pixels_with(n, closure)` primitive (mirrors the proven `write_pixels_solid` framing; scan-order verified pixel-identical against the HW-validated `blit_glyph` multi-row fill) — makes all 3 animations SPI-bound (~24 ms/121 KB-at-40 MHz floor ⇒ ~40 fps for hyperspace/horizon); (2) nebula per-pixel `powf(n,2.2)` → 256-entry LUT and per-pixel `sqrt` for `carve` eliminated via a squared-distance branch that skips the carved-out centre AND the saturated body (`sqrt` only in the 104..172 annulus) — the dominant nebula cost; (3) per-revision avg-FPS logged for bench readout. A 2-agent workflow verified both changes behavior-preserving (`safe:true`, no bugs; one optional `get` bounds-guard applied for `set`/`get` symmetry). **On-silicon FPS readout then exposed the REAL bottleneck:** hyperspace 16 fps / horizon 6 fps / nebula ~0.8 fps (1.2 s/frame) — far worse than any blit/compute estimate, because the firmware is otherwise all-integer so **the Cortex-M33 FPU was never enabled** and the default `thumbv8m.main-none-eabi` (soft-float ABI, no FPU feature) compiled every `sin`/`cos`/`sqrt`/`powf` to soft-float emulation (~24 µs/transcendental). **Fix: build `splash-test-hw` with `-C target-feature=+fp-armv8d16sp` (the M33 FPU, soft-float ABI KEPT so the CMSE veneer / NS interface is byte-unchanged — "softfp") + `enable_fpu()` flips `CPACR` CP10/CP11 at the top of `run()` before the first VFP op.** Binary-verified: 698 hardware VFP instructions (incl. `vmla.f32` in the noise/bilinear hot loops), zero f32 soft-float in the render bodies (the only residual is `fmodf` for the float `%` in hyperspace drift — inherent, no VFP modulo exists, negligible). This is the dominant smoothness win (nebula compute ~30× faster). **Post-FPU FPS (20/17/10) + a new compute-vs-blit DWT split** (logged per-revision) then localised the remaining bottleneck precisely: blit = **constant 48.6 ms for all three** (vs `fill_screen`'s 24 ms floor for the same byte count) while compute was tiny (hyperspace 1.4 ms) — i.e. the blit was **CPU-bound on the per-pixel transpose+bit-unpack**, not SPI-bound; also bumped `micromath` + splash hot fns to `opt-level=3` / `#[optimize(speed)]`. Restructured the framebuffer to **native scan order** (`set` does the `FLIP=(true,false)` transpose once per *lit* pixel: `k = lx*FRAME_WIDTH + (141-ly)`) so the blit became a branch-light **byte-wise expand** (verified pixel-identical to the prior workflow-blessed blit by an exhaustive host program over all 60 776 pixels — perfect bijection, 0 mismatches, no OOB; binary shows 0 multiplies in the blit body). **But the next on-silicon readout showed the blit UNCHANGED at exactly 48.6 ms** — disproving the CPU-bound hypothesis. Root cause: `spi_hw::init` runs the `ui-lcd` SPI at **÷8 = 20 MHz** (chosen for the *trusted UI* because the dev board's LD2 LED on PE13=SCK shimmers at 40 MHz), and 121 552 B × 8 / 20 MHz = exactly 48.6 ms — the blit was **SPI-clock-bound**, not CPU-bound, which is why the closure work was free (hidden in the FIFO-wait shadow). **Fix: the `splash-test` build alone runs SPI at ÷4 = 40 MHz** (mutually-exclusive cfg from the 20 MHz trusted-UI default; binary-confirmed `CFG1=0x10000007`) → blit ~halves to ~24 ms. The native-layout byte-wise blit, while not the bottleneck at 20 MHz, is what keeps the polled FIFO fed at 40 MHz (the heavier old closure would have starved). Net expected: hyperspace ~38 / horizon ~30 / nebula ~14 fps. On-silicon confirmed: hyperspace **38** / horizon **30** / nebula **14** fps, blit now a constant 24.3 ms (the 40 MHz floor) — hyperspace/horizon maxed for a full-frame repaint, **nebula compute-bound at 43.8 ms**. Took the safe compute lever first: **coarsened the nebula noise grid step 4→6** (`GW/GH = div_ceil(STEP)+1` so the bilinear's +1 neighbours stay in-grid even when STEP doesn't divide LW — host-verified bounds-safe across all 60 776 pixels, 0 OOB; **2.1× fewer `noise()` calls/frame**, 1825 vs 3888; also now renders the full 142 rows vs the old 2-row bottom skip). Scoped DMA fully (GPDMA1 secure base `0x5002_0000`, SPI1_TX req 7, `CFG1.TXDMAEN` bit 15; 16-bit SPI items make `TSIZE=60776` fit one transaction; CMSIS lacks GPDMA bitfields but the SVD has them) — but at 40 MHz DMA-overlap helps ONLY nebula (hyperspace/horizon are transfer-bound), and it's a from-scratch blind GPDMA+linked-list driver needing on-silicon debug iterations, so it's deferred in favour of the safe compute lever; DMA's real payoff is a future 80 MHz "all-three-faster" push (signal-integrity risk on the LD2-on-SCK dev board). Next on-silicon readout will show how much the grid coarsening recovered (implicitly the field-vs-pixel split) before deciding further compute cuts (step→8 / octaves 3→2) vs the DMA+80 MHz project. |
0 commit comments