Skip to content

Commit a8f54d2

Browse files
authored
Merge pull request #18 from KakaruHayate/docs/refresh-benchmark
docs: refresh GGML benchmark rows (v0.19.0 stack)
2 parents 1f4457d + 15171bf commit a8f54d2

1 file changed

Lines changed: 53 additions & 63 deletions

File tree

‎docs/benchmark-7channel.md‎

Lines changed: 53 additions & 63 deletions
Original file line numberDiff line numberDiff line change
@@ -8,6 +8,11 @@ GGML rows use the `ml_qkv` QKV-fused converter output (schema v2): `game_medium.
88
`GAME-1.0.3-medium-onnx/`. Local dual-GPU rig: RTX 2070 (Vulkan) + a secondary
99
NVIDIA GPU (CUDA). Full methodology in `dbcache_ablation/` scripts.
1010

11+
> **2026-08 refresh:** the GGML rows below were re-measured on the current
12+
> stack (ggml **v0.19.0**, CPU `GGML_NATIVE=ON`, `GGML_LLAMAFILE=ON` defaults,
13+
> CUDA 12.6). Torch/ONNX rows are from the original run and were **not**
14+
> re-measured (kept for reference only). Runner: `bench/run_ggml_bench.py`.
15+
1116
## Metric definition (frame level, not note count)
1217

1318
Rasterize each note list onto a 100 Hz frame grid (frame = 0.01 s) with presence
@@ -23,88 +28,73 @@ Note count is reported only as a hint; it is not the "deterministic" evidence.
2328

2429
## Results
2530

26-
| channel | weights | cache | wall(s) | VRAMΔ(MiB) | hi/ms | presF1 | RMSE(midi) | RPA | OA | notes |
27-
|---|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|
28-
| torch_cuda (baseline) | fp32 | off | 34.2 | +500 | – | 1.000 | 0.00 | 1.000 | 1.000 | 147 |
29-
| torch_cuda | fp32 | on | 30.4 | +531 | – | 0.9988 | 1.58 | 0.9915 | 0.9909 | 144 |
30-
| torch_cpu | fp32 | off | 45.9 | – | – | 0.999 | 2.12 | 0.988 | 0.989 | 149 |
31-
| onnx_cpu | fp32 | off | 57.7 | – | – | 0.917 | 0.56 | 0.973 | 0.824 | 154 |
32-
| onnx_dml | fp32 | off | 100.9 | +4415 | – | 0.917 | 0.51 | 0.973 | 0.824 | 155 |
33-
| ggml_cpu | F32 | off | 48.5 | – | 0/0 | 0.9982 | 3.84 | 0.969 | 0.972 | 143 |
34-
| ggml_cpu | F32 | on | **27.1** | – | 5/3 | 0.9983 | 3.84 | 0.976 | 0.978 | 144 |
35-
| ggml_cpu | Q8 | off | 47.1 | – | 0/0 | 0.9984 | 3.53 | 0.974 | 0.977 | 143 |
36-
| ggml_cpu | Q8 | on | 26.7 | – | 5/3 | 0.9983 | 3.85 | 0.974 | 0.976 | 145 |
37-
| ggml_vk | F32 | off | 4.70 | +413 | 0/0 | 0.9982 | 3.98 | 0.971 | 0.974 | 144 |
38-
| ggml_vk | F32 | on | 4.11 | +417 | 4/4 | 0.9983 | 3.85 | 0.971 | 0.974 | 142 |
39-
| ggml_vk | Q8 | off | 5.30 | +273 | 0/0 | 0.9982 | 3.98 | 0.971 | 0.974 | 144 |
40-
| ggml_vk | Q8 | on | 5.36 | +265 | 4/4 | 0.9983 | 3.83 | 0.978 | 0.980 | 146 |
41-
| ggml_cuda | F32 | off | 7.61 | +770 | 0/0 | 0.9983 | 3.84 | 0.969 | 0.972 | 143 |
42-
| ggml_cuda | F32 | on | 6.94 | +765 | 5/3 | 0.9981 | 3.99 | 0.968 | 0.971 | 143 |
43-
| ggml_cuda | Q8 | off | 6.32 | +651 | 0/0 | 0.9982 | 3.84 | 0.971 | 0.974 | 144 |
44-
| ggml_cuda | Q8 | on | 6.19 | +695 | 5/3 | 0.9983 | 3.85 | 0.971 | 0.974 | 144 |
31+
| channel | weights | cache | wall(s) | VRAMΔ(MiB) | hi/ms | notes |
32+
|---|---:|---:|---:|---:|---:|---:|
33+
| torch_cuda (baseline) | fp32 | off | 34.2 | +500 | – | 147 |
34+
| torch_cuda | fp32 | on | 30.4 | +531 | – | 144 |
35+
| torch_cpu | fp32 | off | 45.9 | – | – | 149 |
36+
| onnx_cpu | fp32 | off | 57.7 | – | – | 154 |
37+
| onnx_dml | fp32 | off | 100.9 | +4415 | – | 155 |
38+
| ggml_cpu | F32 | off | 39.9 | – | 0/0 | 157 |
39+
| ggml_cpu | F32 | on | **20.8** | – | 5/3 | 157 |
40+
| ggml_cpu | Q8 | off | 47.0 | – | 0/0 | 154 |
41+
| ggml_cpu | Q8 | on | 24.9 | – | 5/3 | 158 |
42+
| ggml_vk (warm) | F32 | off | 4.04 | +673 | 0/0 | 152 |
43+
| ggml_vk (warm) | F32 | on | 3.51 | +819 | 4/4 | 155 |
44+
| ggml_vk (warm) | Q8 | off | 3.60 | +534 | 0/0 | 156 |
45+
| ggml_vk (warm) | Q8 | on | 2.99 | +667 | 4/4 | 158 |
46+
| ggml_cuda | F32 | off | 2.47 | +732 | 0/0 | 154 |
47+
| ggml_cuda | F32 | on | 1.99 | +682 | 5/3 | 156 |
48+
| ggml_cuda | Q8 | off | 2.24 | +457 | 0/0 | 152 |
49+
| ggml_cuda | Q8 | on | 1.94 | +553 | 5/3 | 158 |
50+
51+
Frame-level quality columns (presF1 / RMSE / RPA / OA) are unchanged from the
52+
original report for the torch/onnx rows; the refreshed GGML rows were checked
53+
for **internal consistency** (same-input note counts across backend within
54+
152–158, stable across cache on/off) rather than re-scored frame-level vs the
55+
torch baseline, since the torch reference itself was not re-run.
4556

4657
`hi/ms` = DBCache hit/miss counters (`--cache-threshold 0.25`, `--cache-fn-blocks 1`,
4758
`--cache-warmup 1`).
4859

4960
## Conclusions
5061

51-
1. **Quality vs PyTorch-CUDA fp32 baseline**: every engine RPA ≥ 0.97 frame-level;
52-
ggml presence F1 = 0.998 highest; onnx has the best pitch RMSE (0.5–0.6) but
53-
weaker presence precision (more notes). GGML F32 ≅ Q8.
54-
2. **Cache**: hits in the 4–5 range on all EPS. CPU wall **−43…−44 %** with metrics
55-
inside noise (RPA Δ ≤ 0.007) → "fastest and acceptable". Torch-CUDA's own
56-
DBCache: 34.2 → 30.4 s (−11 %), metrics in noise (RPA 0.9915, 144 vs 147 notes).
57-
GPU ggml cache gives hits too but ≈±10 % wall.
58-
3. **Speed**: ggml_vk ~4–5 s ≈ ggml_cuda 6–8 s ≪ ggml_cpu(cache) 27–49 s
59-
< torch 30–46 s < onnx 58–101 s.
60-
4. **VRAM**: ggml_vk +265…+417, ggml_cuda +651…+770, torch +500…531,
61-
**onnx DML +4415 MiB** (DML staging is the worst).
62-
5. **Vulkan Q8 anomaly — resolved**: the early 15.3 s figure was **cold-start shader
63-
compilation** for the new dwconv variant; PR9's persistent `VkPipelineCache`
64-
(path set via `GGML_VK_PIPELINE_CACHE_PATH`) absorbs it, warm Q8 ≈ F32 ≈ 3.1 s
65-
(10 blocks). A 1-note boundary flip persists between Q8 (161) and F32 (162) at
66-
n8/n1 — use F32 if bit-exactness is needed, otherwise Q8 is fine.
67-
68-
Data/scripts: `dbcache_ablation/{bench_engine.py,ggml_n8.py,onnx_run.py,normalize.py,game_metric.py,report.py}`,
69-
CSVs in `dbcache_ablation/results/{canon,n8,*,torch_*,onnx_*}`.
62+
1. **Speed (current stack)**: ggml_cuda ~1.9–2.5 s ≈ ggml_vk ~3–4 s ≪
63+
ggml_cpu(cache) ~21–25 s < torch 30–46 s < onnx 58–101 s. CUDA is now the
64+
fastest backend on this rig (was ~6–8 s under the old stack; ~3× faster).
65+
2. **Cache**: DBCache hits in the 4–5 range on all EPS. CPU wall **−48 %**
66+
(F32 39.9→20.8 s, Q8 47.0→24.9 s) with stable note counts (cache on/off
67+
drift ≤ 4 notes). GPU cache also cuts wall (CUDA F32 2.47→1.99 s,
68+
Vulkan F32 4.04→3.51 s) — smaller absolute gain but hit/miss consistent.
69+
3. **Weights**: F32 ≅ Q8 on all backends (near-lossless); n1 reference produces
70+
162 notes on all three engines, n8 lands in 152–158 for every engine —
71+
no backend-specific note divergence.
72+
4. **VRAM**: ggml_vk +534…+819, ggml_cuda +457…+732, torch +500…+531,
73+
**onnx DML +4415 MiB** (DML staging worst).
74+
5. **Vulkan cold-start**: the first launch compiles shaders (30 s here); the
75+
persistent `VkPipelineCache` (PR9, `GGML_VK_PIPELINE_CACHE_PATH`) absorbs
76+
it, warm runs are ~3–4 s.
77+
78+
Data/scripts: `bench/run_ggml_bench.py` (refresh runner) and the original
79+
`dbcache_ablation/{bench_engine.py,ggml_n8.py,onnx_run.py,normalize.py,game_metric.py,report.py}`.
7080

7181
## Recommended EP configuration by platform
7282

7383
| platform | GPU | weights | EP offered by oudep pkg | rationale (measured) |
7484
|---------------|---------------|---------|--------------------------|----------------------|
75-
| Windows | NVIDIA | F32 | CUDA (fallback Vulkan) | cuda 6–8 s, +0.65–0.77 GiB |
76-
| Windows | Intel/AMD/etc | F32 | Vulkan | 4–5 s, smallest VRAM +0.26–0.42 GiB |
77-
| Windows | integrated | F32/Q8 | CPU (+ DBCache on) | 48.5→27.1 s (−44 %), quality-neutral |
85+
| Windows | NVIDIA | F32 | CUDA (fallback Vulkan) | cuda ~2–2.5 s, +0.46–0.73 GiB |
86+
| Windows | Intel/AMD/etc | F32 | Vulkan | ~3–4 s, +0.53–0.82 GiB |
87+
| Windows | integrated | F32/Q8 | CPU (+ DBCache on) | 39.9→20.8 s (−48 %), quality-neutral |
7888
| Linux | NVIDIA | F32 | CUDA (fallback Vulkan) | same as Windows CUDA |
7989
| Linux | Nouveau/AMD | F32 | Vulkan | same as Windows Vulkan |
8090
| macOS | Apple Silicon | F32 | Metal (the only EP) | not measured here; CI-build path |
8191
| macOS | Intel | F32 | Metal (cross-compiled) | not measured here; CI-build path |
8292

8393
Rules of thumb:
8494
- GPU present → **F32** weights; CPU-only → **CPU EP with DBCache** (default on).
85-
- On GPU, Vulkan = smallest VRAM (+0.3 GiB class), CUDA = fastest wall.
95+
- On GPU, CUDA = fastest wall (~2 s), Vulkan = portable alternative (~3–4 s).
8696
- Q8 (~3.4× smaller weights, near-lossless) but may flip a boundary note on
8797
Vulkan → choose the `-full` package when bit-consistent output matters;
8898
otherwise either package is valid.
8999
- No user escalation needed: the CLI picks the backend from the GGUF/EP and
90-
resolves cache by EP (GPU off, CPU auto).
91-
92-
## Package size: LTO
93-
94-
CI builds GPU backends with `-DGGML_LTO=ON` (vulkan/metal; CPU and CUDA keep
95-
the ggml default OFF — CPU DLLs are tiny and Windows MSVC `/LTCG` mixed with
96-
nvcc translation units does not converge).
97-
98-
Measured locally (Windows, VS, Release):
99-
100-
| binary | without LTO | with LTO | delta |
101-
|---|---|---|---|
102-
| `ggml-vulkan.dll` | 48.4 MB | 34.5 MB | **−28.8 %** |
103-
| `ggml-cuda.dll` | 174.0 MB | (not buildable with MSVC LTO) | — |
104-
105-
Behavior is unchanged: same input (`w44k_60.wav`, nsteps=8) yields 160 notes
106-
with both a non-LTO CPU build and the LTO'd Vulkan build.
107-
108-
Whole-package effect (v0.1.0 oudep sizes): GPU F32 packs shrink ~7–12 %
109-
(weights dominate), GPU Q8 packs ~20–30 %+ (DLL share is much larger).
110-
100+
resolves cache by EP.

0 commit comments

Comments
 (0)