Benchmarking and finding the best configurations for various MoE and small models.
Configs here should run comfortably with room to spare for other tasks if you have at least a 16GB GPU, but these would also work with lower VRAM/RAM configuration (down to a 12GB 3060), since they're built deliberately to leave headroom.
Tested on: RTX 4070 Ti SUPER (16 GB) · Ryzen 7 5800X (8C/16T) · 32 GB RAM · Windows · llama.cpp official CUDA builds.
Shared settings across all presets: flash attention on, q8_0 K/V cache, ubatch 2048 (except E4B, below). Set threads / threads-batch to your CPU's physical core count; it's a per-machine setting, not a tuning knob (thread count measured inside the noise here: results/threads.md).
| Model | Quant (HF) | n-cpu-moe | ubatch | MTP spec | p-min | Decode @ ~87k | Peak VRAM |
|---|---|---|---|---|---|---|---|
| Qwen3.6-35B-A3B | UD-IQ4_NL | 30 | 2048 | n-max 2 | 0.4 | ~38–47 t/s (content-dependent) | ~12.9 GiB |
| Gemma 4 26B A4B | UD-Q4_K_XL | 22 | 2048 | n-max 2 | 0.65 | ~43–46 t/s | ~11.5 GiB |
| Ornith 1.5 35B A3B | i1 IQ4_XS | 30 | 2048 | n-max 2 | 0.25 | ~50 t/s | ~12.7 GiB |
| Gemma 4 E4B | UD-Q4_K_XL | n/a (fits in VRAM) | 512 (default) | n-max 2 | 0.0 (default) | ~97–100 t/s | ~8.7 GiB |
Model names link to that model's preset block in configs/models_config.ini.
- On a 12 GB card, raise
n-cpu-moeby a few layers; each additional CPU-offloaded expert layer frees ~350–420 MiB (§13, §15), same knobs otherwise. - The settings these benchmarks landed on, with context:
results/final-settings.md(§10, §14).
- On MoEs with hybrid inference, raising ubatch-size increased prefill for slightly more VRAM use and slightly lower token generation.
- On models that fit completely in VRAM, causes a loss in prefill and token generation.
Full data: results/ubatch-and-batch-shape.md.
MTP (multi-token prediction) heads let llama.cpp draft several tokens per step and verify them in one batch.
- Gains are variable per model: +30–36% decode on Qwen/Gemma-26B (§5), +19% on Ornith with the updated head at
n-max 2 / p-min 0.25(§18), +60% on the full-VRAM E4B (§12). - The knobs are model-specific: Qwen
n-max 2 / p-min 0.4, Gemma-26Bn-max 2 / p-min 0.65, E4Bp-min 0(default), Ornithn-max 2 / p-min 0.25. There is no universal value. - With a good MTP head, you can achieve better speed at the same VRAM usage by offloading more layers to RAM
Full grids: results/speculative-decoding-mtp.md.
Parameters that didn't produce a measurable speedup on this machine, recorded so nobody re-tests them.
| Parameter | Result | Data |
|---|---|---|
--no-host |
Up to 9× slower prefill, −10–19% decode. Actively harmful | results/threads.md (§9) |
Backend sampling (-bs) |
No effect (0.99×) | results/backend-sampling.md (§4) |
| Thread count (8 vs 12 vs 16) | Inside noise; paired interleaved testing closed the question | results/threads.md (§3, §11) |
Decoupled threads-batch (-tb) |
No useful decoupling; -tb above -t costs decode |
results/threads.md (§9, §11) |
| ik_llama.cpp fork | Wins small-prompt CPU-side prefill (+28–60%), loses at real workload shape (−24% prefill, −23% decode at an 11.5k prompt). Stayed on mainline | results/llamacpp-vs-ik-fork.md (§20) |
- Methodology (seed policy, paired interleaved A/B testing, invalidated-run catalog, measurement rules):
docs/methodology.md - Full results log: every number in this README with its context, grouped by category:
results/ - Benchmark harnesses (
--demoself-checks, no GPU needed for the check):scripts/ - Context depth tax: ~−18–21% decode at 80k depth vs fresh context, both big models (§2). Benchmark at the depth you actually run. Data:
results/context-depth.md.