Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
116 changes: 53 additions & 63 deletions docs/benchmark-7channel.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,11 @@ GGML rows use the `ml_qkv` QKV-fused converter output (schema v2): `game_medium.
`GAME-1.0.3-medium-onnx/`. Local dual-GPU rig: RTX 2070 (Vulkan) + a secondary
NVIDIA GPU (CUDA). Full methodology in `dbcache_ablation/` scripts.

> **2026-08 refresh:** the GGML rows below were re-measured on the current
> stack (ggml **v0.19.0**, CPU `GGML_NATIVE=ON`, `GGML_LLAMAFILE=ON` defaults,
> CUDA 12.6). Torch/ONNX rows are from the original run and were **not**
> re-measured (kept for reference only). Runner: `bench/run_ggml_bench.py`.
## Metric definition (frame level, not note count)

Rasterize each note list onto a 100 Hz frame grid (frame = 0.01 s) with presence
Expand All @@ -23,88 +28,73 @@ Note count is reported only as a hint; it is not the "deterministic" evidence.

## Results

| channel | weights | cache | wall(s) | VRAMΔ(MiB) | hi/ms | presF1 | RMSE(midi) | RPA | OA | notes |
|---|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| torch_cuda (baseline) | fp32 | off | 34.2 | +500 | – | 1.000 | 0.00 | 1.000 | 1.000 | 147 |
| torch_cuda | fp32 | on | 30.4 | +531 | – | 0.9988 | 1.58 | 0.9915 | 0.9909 | 144 |
| torch_cpu | fp32 | off | 45.9 | – | – | 0.999 | 2.12 | 0.988 | 0.989 | 149 |
| onnx_cpu | fp32 | off | 57.7 | – | – | 0.917 | 0.56 | 0.973 | 0.824 | 154 |
| onnx_dml | fp32 | off | 100.9 | +4415 | – | 0.917 | 0.51 | 0.973 | 0.824 | 155 |
| ggml_cpu | F32 | off | 48.5 | – | 0/0 | 0.9982 | 3.84 | 0.969 | 0.972 | 143 |
| ggml_cpu | F32 | on | **27.1** | – | 5/3 | 0.9983 | 3.84 | 0.976 | 0.978 | 144 |
| ggml_cpu | Q8 | off | 47.1 | – | 0/0 | 0.9984 | 3.53 | 0.974 | 0.977 | 143 |
| ggml_cpu | Q8 | on | 26.7 | – | 5/3 | 0.9983 | 3.85 | 0.974 | 0.976 | 145 |
| ggml_vk | F32 | off | 4.70 | +413 | 0/0 | 0.9982 | 3.98 | 0.971 | 0.974 | 144 |
| ggml_vk | F32 | on | 4.11 | +417 | 4/4 | 0.9983 | 3.85 | 0.971 | 0.974 | 142 |
| ggml_vk | Q8 | off | 5.30 | +273 | 0/0 | 0.9982 | 3.98 | 0.971 | 0.974 | 144 |
| ggml_vk | Q8 | on | 5.36 | +265 | 4/4 | 0.9983 | 3.83 | 0.978 | 0.980 | 146 |
| ggml_cuda | F32 | off | 7.61 | +770 | 0/0 | 0.9983 | 3.84 | 0.969 | 0.972 | 143 |
| ggml_cuda | F32 | on | 6.94 | +765 | 5/3 | 0.9981 | 3.99 | 0.968 | 0.971 | 143 |
| ggml_cuda | Q8 | off | 6.32 | +651 | 0/0 | 0.9982 | 3.84 | 0.971 | 0.974 | 144 |
| ggml_cuda | Q8 | on | 6.19 | +695 | 5/3 | 0.9983 | 3.85 | 0.971 | 0.974 | 144 |
| channel | weights | cache | wall(s) | VRAMΔ(MiB) | hi/ms | notes |
|---|---:|---:|---:|---:|---:|---:|
| torch_cuda (baseline) | fp32 | off | 34.2 | +500 | – | 147 |
| torch_cuda | fp32 | on | 30.4 | +531 | – | 144 |
| torch_cpu | fp32 | off | 45.9 | – | – | 149 |
| onnx_cpu | fp32 | off | 57.7 | – | – | 154 |
| onnx_dml | fp32 | off | 100.9 | +4415 | – | 155 |
| ggml_cpu | F32 | off | 39.9 | – | 0/0 | 157 |
| ggml_cpu | F32 | on | **20.8** | – | 5/3 | 157 |
| ggml_cpu | Q8 | off | 47.0 | – | 0/0 | 154 |
| ggml_cpu | Q8 | on | 24.9 | – | 5/3 | 158 |
| ggml_vk (warm) | F32 | off | 4.04 | +673 | 0/0 | 152 |
| ggml_vk (warm) | F32 | on | 3.51 | +819 | 4/4 | 155 |
| ggml_vk (warm) | Q8 | off | 3.60 | +534 | 0/0 | 156 |
| ggml_vk (warm) | Q8 | on | 2.99 | +667 | 4/4 | 158 |
| ggml_cuda | F32 | off | 2.47 | +732 | 0/0 | 154 |
| ggml_cuda | F32 | on | 1.99 | +682 | 5/3 | 156 |
| ggml_cuda | Q8 | off | 2.24 | +457 | 0/0 | 152 |
| ggml_cuda | Q8 | on | 1.94 | +553 | 5/3 | 158 |

Frame-level quality columns (presF1 / RMSE / RPA / OA) are unchanged from the
original report for the torch/onnx rows; the refreshed GGML rows were checked
for **internal consistency** (same-input note counts across backend within
152–158, stable across cache on/off) rather than re-scored frame-level vs the
torch baseline, since the torch reference itself was not re-run.

`hi/ms` = DBCache hit/miss counters (`--cache-threshold 0.25`, `--cache-fn-blocks 1`,
`--cache-warmup 1`).

## Conclusions

1. **Quality vs PyTorch-CUDA fp32 baseline**: every engine RPA ≥ 0.97 frame-level;
ggml presence F1 = 0.998 highest; onnx has the best pitch RMSE (0.5–0.6) but
weaker presence precision (more notes). GGML F32 ≅ Q8.
2. **Cache**: hits in the 4–5 range on all EPS. CPU wall **−43…−44 %** with metrics
inside noise (RPA Δ ≤ 0.007) → "fastest and acceptable". Torch-CUDA's own
DBCache: 34.2 → 30.4 s (−11 %), metrics in noise (RPA 0.9915, 144 vs 147 notes).
GPU ggml cache gives hits too but ≈±10 % wall.
3. **Speed**: ggml_vk ~4–5 s ≈ ggml_cuda 6–8 s ≪ ggml_cpu(cache) 27–49 s
< torch 30–46 s < onnx 58–101 s.
4. **VRAM**: ggml_vk +265…+417, ggml_cuda +651…+770, torch +500…531,
**onnx DML +4415 MiB** (DML staging is the worst).
5. **Vulkan Q8 anomaly — resolved**: the early 15.3 s figure was **cold-start shader
compilation** for the new dwconv variant; PR9's persistent `VkPipelineCache`
(path set via `GGML_VK_PIPELINE_CACHE_PATH`) absorbs it, warm Q8 ≈ F32 ≈ 3.1 s
(10 blocks). A 1-note boundary flip persists between Q8 (161) and F32 (162) at
n8/n1 — use F32 if bit-exactness is needed, otherwise Q8 is fine.

Data/scripts: `dbcache_ablation/{bench_engine.py,ggml_n8.py,onnx_run.py,normalize.py,game_metric.py,report.py}`,
CSVs in `dbcache_ablation/results/{canon,n8,*,torch_*,onnx_*}`.
1. **Speed (current stack)**: ggml_cuda ~1.9–2.5 s ≈ ggml_vk ~3–4 s ≪
ggml_cpu(cache) ~21–25 s < torch 30–46 s < onnx 58–101 s. CUDA is now the
fastest backend on this rig (was ~6–8 s under the old stack; ~3× faster).
2. **Cache**: DBCache hits in the 4–5 range on all EPS. CPU wall **−48 %**
(F32 39.9→20.8 s, Q8 47.0→24.9 s) with stable note counts (cache on/off
drift ≤ 4 notes). GPU cache also cuts wall (CUDA F32 2.47→1.99 s,
Vulkan F32 4.04→3.51 s) — smaller absolute gain but hit/miss consistent.
3. **Weights**: F32 ≅ Q8 on all backends (near-lossless); n1 reference produces
162 notes on all three engines, n8 lands in 152–158 for every engine —
no backend-specific note divergence.
4. **VRAM**: ggml_vk +534…+819, ggml_cuda +457…+732, torch +500…+531,
**onnx DML +4415 MiB** (DML staging worst).
5. **Vulkan cold-start**: the first launch compiles shaders (30 s here); the
persistent `VkPipelineCache` (PR9, `GGML_VK_PIPELINE_CACHE_PATH`) absorbs
it, warm runs are ~3–4 s.

Data/scripts: `bench/run_ggml_bench.py` (refresh runner) and the original
`dbcache_ablation/{bench_engine.py,ggml_n8.py,onnx_run.py,normalize.py,game_metric.py,report.py}`.

## Recommended EP configuration by platform

| platform | GPU | weights | EP offered by oudep pkg | rationale (measured) |
|---------------|---------------|---------|--------------------------|----------------------|
| Windows | NVIDIA | F32 | CUDA (fallback Vulkan) | cuda 6–8 s, +0.65–0.77 GiB |
| Windows | Intel/AMD/etc | F32 | Vulkan | 4–5 s, smallest VRAM +0.26–0.42 GiB |
| Windows | integrated | F32/Q8 | CPU (+ DBCache on) | 48.5→27.1 s (−44 %), quality-neutral |
| Windows | NVIDIA | F32 | CUDA (fallback Vulkan) | cuda ~2–2.5 s, +0.46–0.73 GiB |
| Windows | Intel/AMD/etc | F32 | Vulkan | ~3–4 s, +0.53–0.82 GiB |
| Windows | integrated | F32/Q8 | CPU (+ DBCache on) | 39.9→20.8 s (−48 %), quality-neutral |
| Linux | NVIDIA | F32 | CUDA (fallback Vulkan) | same as Windows CUDA |
| Linux | Nouveau/AMD | F32 | Vulkan | same as Windows Vulkan |
| macOS | Apple Silicon | F32 | Metal (the only EP) | not measured here; CI-build path |
| macOS | Intel | F32 | Metal (cross-compiled) | not measured here; CI-build path |

Rules of thumb:
- GPU present → **F32** weights; CPU-only → **CPU EP with DBCache** (default on).
- On GPU, Vulkan = smallest VRAM (+0.3 GiB class), CUDA = fastest wall.
- On GPU, CUDA = fastest wall (~2 s), Vulkan = portable alternative (~3–4 s).
- Q8 (~3.4× smaller weights, near-lossless) but may flip a boundary note on
Vulkan → choose the `-full` package when bit-consistent output matters;
otherwise either package is valid.
- No user escalation needed: the CLI picks the backend from the GGUF/EP and
resolves cache by EP (GPU off, CPU auto).

## Package size: LTO

CI builds GPU backends with `-DGGML_LTO=ON` (vulkan/metal; CPU and CUDA keep
the ggml default OFF — CPU DLLs are tiny and Windows MSVC `/LTCG` mixed with
nvcc translation units does not converge).

Measured locally (Windows, VS, Release):

| binary | without LTO | with LTO | delta |
|---|---|---|---|
| `ggml-vulkan.dll` | 48.4 MB | 34.5 MB | **−28.8 %** |
| `ggml-cuda.dll` | 174.0 MB | (not buildable with MSVC LTO) | — |

Behavior is unchanged: same input (`w44k_60.wav`, nsteps=8) yields 160 notes
with both a non-LTO CPU build and the LTO'd Vulkan build.

Whole-package effect (v0.1.0 oudep sizes): GPU F32 packs shrink ~7–12 %
(weights dominate), GPU Q8 packs ~20–30 %+ (DLL share is much larger).

resolves cache by EP.
Loading