Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

spiwar-llm-bench

Benchmarking and finding the best configurations for various MoE and small models.

Configs here should run comfortably with room to spare for other tasks if you have at least a 16GB GPU, but these would also work with lower VRAM/RAM configuration (down to a 12GB 3060), since they're built deliberately to leave headroom.

Tested on: RTX 4070 Ti SUPER (16 GB) · Ryzen 7 5800X (8C/16T) · 32 GB RAM · Windows · llama.cpp official CUDA builds.

Recommended configuration

Shared settings across all presets: flash attention on, q8_0 K/V cache, ubatch 2048 (except E4B, below). Set threads / threads-batch to your CPU's physical core count; it's a per-machine setting, not a tuning knob (thread count measured inside the noise here: results/threads.md).

Model Quant (HF) n-cpu-moe ubatch MTP spec p-min Decode @ ~87k Peak VRAM
Qwen3.6-35B-A3B UD-IQ4_NL 30 2048 n-max 2 0.4 ~38–47 t/s (content-dependent) ~12.9 GiB
Gemma 4 26B A4B UD-Q4_K_XL 22 2048 n-max 2 0.65 ~43–46 t/s ~11.5 GiB
Ornith 1.5 35B A3B i1 IQ4_XS 30 2048 n-max 2 0.25 ~50 t/s ~12.7 GiB
Gemma 4 E4B UD-Q4_K_XL n/a (fits in VRAM) 512 (default) n-max 2 0.0 (default) ~97–100 t/s ~8.7 GiB

Model names link to that model's preset block in configs/models_config.ini.

  • On a 12 GB card, raise n-cpu-moe by a few layers; each additional CPU-offloaded expert layer frees ~350–420 MiB (§13, §15), same knobs otherwise.
  • The settings these benchmarks landed on, with context: results/final-settings.md (§10, §14).

ubatch-size

  • On MoEs with hybrid inference, raising ubatch-size increased prefill for slightly more VRAM use and slightly lower token generation.
  • On models that fit completely in VRAM, causes a loss in prefill and token generation.

Full data: results/ubatch-and-batch-shape.md.

MTP testing

MTP (multi-token prediction) heads let llama.cpp draft several tokens per step and verify them in one batch.

  • Gains are variable per model: +30–36% decode on Qwen/Gemma-26B (§5), +19% on Ornith with the updated head at n-max 2 / p-min 0.25 (§18), +60% on the full-VRAM E4B (§12).
  • The knobs are model-specific: Qwen n-max 2 / p-min 0.4, Gemma-26B n-max 2 / p-min 0.65, E4B p-min 0 (default), Ornith n-max 2 / p-min 0.25. There is no universal value.
  • With a good MTP head, you can achieve better speed at the same VRAM usage by offloading more layers to RAM

Full grids: results/speculative-decoding-mtp.md.

Tested, no meaningful improvement

Parameters that didn't produce a measurable speedup on this machine, recorded so nobody re-tests them.

Parameter Result Data
--no-host Up to 9× slower prefill, −10–19% decode. Actively harmful results/threads.md (§9)
Backend sampling (-bs) No effect (0.99×) results/backend-sampling.md (§4)
Thread count (8 vs 12 vs 16) Inside noise; paired interleaved testing closed the question results/threads.md (§3, §11)
Decoupled threads-batch (-tb) No useful decoupling; -tb above -t costs decode results/threads.md (§9, §11)
ik_llama.cpp fork Wins small-prompt CPU-side prefill (+28–60%), loses at real workload shape (−24% prefill, −23% decode at an 11.5k prompt). Stayed on mainline results/llamacpp-vs-ik-fork.md (§20)

Additional info

  • Methodology (seed policy, paired interleaved A/B testing, invalidated-run catalog, measurement rules): docs/methodology.md
  • Full results log: every number in this README with its context, grouped by category: results/
  • Benchmark harnesses (--demo self-checks, no GPU needed for the check): scripts/
  • Context depth tax: ~−18–21% decode at 80k depth vs fresh context, both big models (§2). Benchmark at the depth you actually run. Data: results/context-depth.md.

About

Benchmarking and tuning llama.cpp on a 16GB GPU

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages