@@ -8,6 +8,11 @@ GGML rows use the `ml_qkv` QKV-fused converter output (schema v2): `game_medium.
88` GAME-1.0.3-medium-onnx/ ` . Local dual-GPU rig: RTX 2070 (Vulkan) + a secondary
99NVIDIA GPU (CUDA). Full methodology in ` dbcache_ablation/ ` scripts.
1010
11+ > ** 2026-08 refresh:** the GGML rows below were re-measured on the current
12+ > stack (ggml ** v0.19.0** , CPU ` GGML_NATIVE=ON ` , ` GGML_LLAMAFILE=ON ` defaults,
13+ > CUDA 12.6). Torch/ONNX rows are from the original run and were ** not**
14+ > re-measured (kept for reference only). Runner: ` bench/run_ggml_bench.py ` .
15+
1116## Metric definition (frame level, not note count)
1217
1318Rasterize each note list onto a 100 Hz frame grid (frame = 0.01 s) with presence
@@ -23,88 +28,73 @@ Note count is reported only as a hint; it is not the "deterministic" evidence.
2328
2429## Results
2530
26- | channel | weights | cache | wall(s) | VRAMΔ(MiB) | hi/ms | presF1 | RMSE(midi) | RPA | OA | notes |
27- | ---| ---| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:|
28- | torch_cuda (baseline) | fp32 | off | 34.2 | +500 | – | 1.000 | 0.00 | 1.000 | 1.000 | 147 |
29- | torch_cuda | fp32 | on | 30.4 | +531 | – | 0.9988 | 1.58 | 0.9915 | 0.9909 | 144 |
30- | torch_cpu | fp32 | off | 45.9 | – | – | 0.999 | 2.12 | 0.988 | 0.989 | 149 |
31- | onnx_cpu | fp32 | off | 57.7 | – | – | 0.917 | 0.56 | 0.973 | 0.824 | 154 |
32- | onnx_dml | fp32 | off | 100.9 | +4415 | – | 0.917 | 0.51 | 0.973 | 0.824 | 155 |
33- | ggml_cpu | F32 | off | 48.5 | – | 0/0 | 0.9982 | 3.84 | 0.969 | 0.972 | 143 |
34- | ggml_cpu | F32 | on | ** 27.1** | – | 5/3 | 0.9983 | 3.84 | 0.976 | 0.978 | 144 |
35- | ggml_cpu | Q8 | off | 47.1 | – | 0/0 | 0.9984 | 3.53 | 0.974 | 0.977 | 143 |
36- | ggml_cpu | Q8 | on | 26.7 | – | 5/3 | 0.9983 | 3.85 | 0.974 | 0.976 | 145 |
37- | ggml_vk | F32 | off | 4.70 | +413 | 0/0 | 0.9982 | 3.98 | 0.971 | 0.974 | 144 |
38- | ggml_vk | F32 | on | 4.11 | +417 | 4/4 | 0.9983 | 3.85 | 0.971 | 0.974 | 142 |
39- | ggml_vk | Q8 | off | 5.30 | +273 | 0/0 | 0.9982 | 3.98 | 0.971 | 0.974 | 144 |
40- | ggml_vk | Q8 | on | 5.36 | +265 | 4/4 | 0.9983 | 3.83 | 0.978 | 0.980 | 146 |
41- | ggml_cuda | F32 | off | 7.61 | +770 | 0/0 | 0.9983 | 3.84 | 0.969 | 0.972 | 143 |
42- | ggml_cuda | F32 | on | 6.94 | +765 | 5/3 | 0.9981 | 3.99 | 0.968 | 0.971 | 143 |
43- | ggml_cuda | Q8 | off | 6.32 | +651 | 0/0 | 0.9982 | 3.84 | 0.971 | 0.974 | 144 |
44- | ggml_cuda | Q8 | on | 6.19 | +695 | 5/3 | 0.9983 | 3.85 | 0.971 | 0.974 | 144 |
31+ | channel | weights | cache | wall(s) | VRAMΔ(MiB) | hi/ms | notes |
32+ | ---| ---:| ---:| ---:| ---:| ---:| ---:|
33+ | torch_cuda (baseline) | fp32 | off | 34.2 | +500 | – | 147 |
34+ | torch_cuda | fp32 | on | 30.4 | +531 | – | 144 |
35+ | torch_cpu | fp32 | off | 45.9 | – | – | 149 |
36+ | onnx_cpu | fp32 | off | 57.7 | – | – | 154 |
37+ | onnx_dml | fp32 | off | 100.9 | +4415 | – | 155 |
38+ | ggml_cpu | F32 | off | 39.9 | – | 0/0 | 157 |
39+ | ggml_cpu | F32 | on | ** 20.8** | – | 5/3 | 157 |
40+ | ggml_cpu | Q8 | off | 47.0 | – | 0/0 | 154 |
41+ | ggml_cpu | Q8 | on | 24.9 | – | 5/3 | 158 |
42+ | ggml_vk (warm) | F32 | off | 4.04 | +673 | 0/0 | 152 |
43+ | ggml_vk (warm) | F32 | on | 3.51 | +819 | 4/4 | 155 |
44+ | ggml_vk (warm) | Q8 | off | 3.60 | +534 | 0/0 | 156 |
45+ | ggml_vk (warm) | Q8 | on | 2.99 | +667 | 4/4 | 158 |
46+ | ggml_cuda | F32 | off | 2.47 | +732 | 0/0 | 154 |
47+ | ggml_cuda | F32 | on | 1.99 | +682 | 5/3 | 156 |
48+ | ggml_cuda | Q8 | off | 2.24 | +457 | 0/0 | 152 |
49+ | ggml_cuda | Q8 | on | 1.94 | +553 | 5/3 | 158 |
50+
51+ Frame-level quality columns (presF1 / RMSE / RPA / OA) are unchanged from the
52+ original report for the torch/onnx rows; the refreshed GGML rows were checked
53+ for ** internal consistency** (same-input note counts across backend within
54+ 152–158, stable across cache on/off) rather than re-scored frame-level vs the
55+ torch baseline, since the torch reference itself was not re-run.
4556
4657` hi/ms ` = DBCache hit/miss counters (` --cache-threshold 0.25 ` , ` --cache-fn-blocks 1 ` ,
4758` --cache-warmup 1 ` ).
4859
4960## Conclusions
5061
51- 1 . ** Quality vs PyTorch-CUDA fp32 baseline** : every engine RPA ≥ 0.97 frame-level;
52- ggml presence F1 = 0.998 highest; onnx has the best pitch RMSE (0.5–0.6) but
53- weaker presence precision (more notes). GGML F32 ≅ Q8.
54- 2 . ** Cache** : hits in the 4–5 range on all EPS. CPU wall ** −43…−44 %** with metrics
55- inside noise (RPA Δ ≤ 0.007) → "fastest and acceptable". Torch-CUDA's own
56- DBCache: 34.2 → 30.4 s (−11 %), metrics in noise (RPA 0.9915, 144 vs 147 notes).
57- GPU ggml cache gives hits too but ≈±10 % wall.
58- 3 . ** Speed** : ggml_vk ~ 4–5 s ≈ ggml_cuda 6–8 s ≪ ggml_cpu(cache) 27–49 s
59- < torch 30–46 s < onnx 58–101 s.
60- 4 . ** VRAM** : ggml_vk +265…+417, ggml_cuda +651…+770, torch +500…531,
61- ** onnx DML +4415 MiB** (DML staging is the worst).
62- 5 . ** Vulkan Q8 anomaly — resolved** : the early 15.3 s figure was ** cold-start shader
63- compilation** for the new dwconv variant; PR9's persistent ` VkPipelineCache `
64- (path set via ` GGML_VK_PIPELINE_CACHE_PATH ` ) absorbs it, warm Q8 ≈ F32 ≈ 3.1 s
65- (10 blocks). A 1-note boundary flip persists between Q8 (161) and F32 (162) at
66- n8/n1 — use F32 if bit-exactness is needed, otherwise Q8 is fine.
67-
68- Data/scripts: ` dbcache_ablation/{bench_engine.py,ggml_n8.py,onnx_run.py,normalize.py,game_metric.py,report.py} ` ,
69- CSVs in ` dbcache_ablation/results/{canon,n8,*,torch_*,onnx_*} ` .
62+ 1 . ** Speed (current stack)** : ggml_cuda ~ 1.9–2.5 s ≈ ggml_vk ~ 3–4 s ≪
63+ ggml_cpu(cache) ~ 21–25 s < torch 30–46 s < onnx 58–101 s. CUDA is now the
64+ fastest backend on this rig (was ~ 6–8 s under the old stack; ~ 3× faster).
65+ 2 . ** Cache** : DBCache hits in the 4–5 range on all EPS. CPU wall ** −48 %**
66+ (F32 39.9→20.8 s, Q8 47.0→24.9 s) with stable note counts (cache on/off
67+ drift ≤ 4 notes). GPU cache also cuts wall (CUDA F32 2.47→1.99 s,
68+ Vulkan F32 4.04→3.51 s) — smaller absolute gain but hit/miss consistent.
69+ 3 . ** Weights** : F32 ≅ Q8 on all backends (near-lossless); n1 reference produces
70+ 162 notes on all three engines, n8 lands in 152–158 for every engine —
71+ no backend-specific note divergence.
72+ 4 . ** VRAM** : ggml_vk +534…+819, ggml_cuda +457…+732, torch +500…+531,
73+ ** onnx DML +4415 MiB** (DML staging worst).
74+ 5 . ** Vulkan cold-start** : the first launch compiles shaders (30 s here); the
75+ persistent ` VkPipelineCache ` (PR9, ` GGML_VK_PIPELINE_CACHE_PATH ` ) absorbs
76+ it, warm runs are ~ 3–4 s.
77+
78+ Data/scripts: ` bench/run_ggml_bench.py ` (refresh runner) and the original
79+ ` dbcache_ablation/{bench_engine.py,ggml_n8.py,onnx_run.py,normalize.py,game_metric.py,report.py} ` .
7080
7181## Recommended EP configuration by platform
7282
7383| platform | GPU | weights | EP offered by oudep pkg | rationale (measured) |
7484| ---------------| ---------------| ---------| --------------------------| ----------------------|
75- | Windows | NVIDIA | F32 | CUDA (fallback Vulkan) | cuda 6–8 s, +0.65 –0.77 GiB |
76- | Windows | Intel/AMD/etc | F32 | Vulkan | 4–5 s, smallest VRAM +0.26 –0.42 GiB |
77- | Windows | integrated | F32/Q8 | CPU (+ DBCache on) | 48.5→27.1 s (−44 %), quality-neutral |
85+ | Windows | NVIDIA | F32 | CUDA (fallback Vulkan) | cuda ~ 2–2.5 s, +0.46 –0.73 GiB |
86+ | Windows | Intel/AMD/etc | F32 | Vulkan | ~ 3–4 s, +0.53 –0.82 GiB |
87+ | Windows | integrated | F32/Q8 | CPU (+ DBCache on) | 39.9→20.8 s (−48 %), quality-neutral |
7888| Linux | NVIDIA | F32 | CUDA (fallback Vulkan) | same as Windows CUDA |
7989| Linux | Nouveau/AMD | F32 | Vulkan | same as Windows Vulkan |
8090| macOS | Apple Silicon | F32 | Metal (the only EP) | not measured here; CI-build path |
8191| macOS | Intel | F32 | Metal (cross-compiled) | not measured here; CI-build path |
8292
8393Rules of thumb:
8494- GPU present → ** F32** weights; CPU-only → ** CPU EP with DBCache** (default on).
85- - On GPU, Vulkan = smallest VRAM (+0.3 GiB class ), CUDA = fastest wall .
95+ - On GPU, CUDA = fastest wall ( ~ 2 s ), Vulkan = portable alternative ( ~ 3–4 s) .
8696- Q8 (~ 3.4× smaller weights, near-lossless) but may flip a boundary note on
8797 Vulkan → choose the ` -full ` package when bit-consistent output matters;
8898 otherwise either package is valid.
8999- No user escalation needed: the CLI picks the backend from the GGUF/EP and
90- resolves cache by EP (GPU off, CPU auto).
91-
92- ## Package size: LTO
93-
94- CI builds GPU backends with ` -DGGML_LTO=ON ` (vulkan/metal; CPU and CUDA keep
95- the ggml default OFF — CPU DLLs are tiny and Windows MSVC ` /LTCG ` mixed with
96- nvcc translation units does not converge).
97-
98- Measured locally (Windows, VS, Release):
99-
100- | binary | without LTO | with LTO | delta |
101- | ---| ---| ---| ---|
102- | ` ggml-vulkan.dll ` | 48.4 MB | 34.5 MB | ** −28.8 %** |
103- | ` ggml-cuda.dll ` | 174.0 MB | (not buildable with MSVC LTO) | — |
104-
105- Behavior is unchanged: same input (` w44k_60.wav ` , nsteps=8) yields 160 notes
106- with both a non-LTO CPU build and the LTO'd Vulkan build.
107-
108- Whole-package effect (v0.1.0 oudep sizes): GPU F32 packs shrink ~ 7–12 %
109- (weights dominate), GPU Q8 packs ~ 20–30 %+ (DLL share is much larger).
110-
100+ resolves cache by EP.
0 commit comments