Run a 52GB BF16 dense model bit-exactly on an 8GB GPU.
TCPP streams a large checkpoint layer-by-layer from NVMe through a small GPU and produces greedy outputs identical, bit for bit, to an in-VRAM eager forward of the same model. No approximation ever enters the authority path: quantization, drafters and rankers may predict, only the exact forward decides.
Résumé FR — TCPP exécute un modèle dense de 52GB BF16 (Qwen3.8-27B, 64 couches hybrides GDN/attention) sur un GPU de 8GB, avec une sortie greedy bit-exacte et déterministe inter-process. Vérification par lots à 172 positions/s ; génération ~1 token / sweep (~30s sur NVMe). Le but du projet : rendre utilisable — et ouvrir — l'inférence exacte de gros modèles denses sur matériel grand public.
pip install -r requirements.txt
# 1. sanity check (expects the exact argmax of the reference prompt)
python -m tcpp sanity --model /path/to/Qwen3.8-27B
# 2. batch verification: argmax for every position of a K-token context
python -m tcpp bench --model /path/to/Qwen3.8-27B --k 2048 4096 8192
# 3. exact greedy generation (slow: one full layer sweep per token)
python -m tcpp generate --model /path/to/Qwen3.8-27B \
--prompt "La tour Eiffel se trouve à Paris et" --max-new 8The model directory is any local HF snapshot in safetensors format
(zero-copy: TCPP builds a page index over the original shards, no repacking).
Supported architecture family in v0.1: qwen3_5 hybrid (Qwen3.5 / Qwen3.8
dense models, e.g. 9B / 27B).
| Component | Minimum (tested) |
|---|---|
| GPU VRAM | 8GB (RTX 3070 laptop, sm_86) |
| System RAM | 32GB (auto-budgeted resident layers; more = faster warm passes) |
| Disk | NVMe SSD, ~1.6GiB/s sustained, model size + headroom |
| OS | Linux (preadv, fadvise, /proc/meminfo) |
| Workload | Result |
|---|---|
| Batch verify K=2048 | 33.0s (62 pos/s) |
| Batch verify K=4096 | 37.5s (109 pos/s) |
| Batch verify K=8192 | 47.5s (172 pos/s) |
| Greedy generation | ~1 token / sweep (~30s) — I/O-bound, see below |
| Bit-exactness | query-chunked attention hash-identical to eager (K=2048/4096) |
| Determinism | 3 fresh processes → identical SHA256 of every layer's weights, activations and argmax |
- The attention is the eager algorithm replicated op-for-op (fp32 softmax → bf16 cast → matmul), only chunked along the query axis so the [T, T] matrix never materializes. Demonstrated hash-identical to eager at reference sizes; SDPA/flash kernels are NOT used (they change argmax on near-ties — measured 6/10 prompt divergences).
- Weights are read from the original BF16 safetensors; no dequantization.
- The runtime is deterministic across processes (validated by hashing every
intermediate tensor in 3 fresh processes — see
harness/determinism_gate.py).
Every generated token requires one full 52GB sweep through the memory hierarchy. At NVMe speed (~1.6GiB/s) the sweep floor is ~27-33s. The GPU itself could do ~400 tok/s compute-only if the weights were resident. The project's accompanying preprint quantifies why speculative decoding (EAGLE-3, DFlash, trees, pipeline speculation) cannot compensate: sequential acceptance is ~2-9 tokens per sweep for every drafter family tested, so the ceiling on this substrate is ~0.3-0.5 tok/s. The wall is the memory hierarchy, not the compute and not the drafter. Give this code 64GB of fast GPU memory and it is no longer the bottleneck.
Best-fit uses today: batch scoring/verification (172 pos/s), exactness audits, research on out-of-core execution. Interactive chat needs a resident substrate or a different model class (e.g. MoE with hot experts in VRAM).
tcpp/ the runtime (page index, pipelined loader, exact attention)
harness/ measurement tooling
determinism_gate.py 3-process bit-exactness gate (weights+activations hashes)
golden_collect.py self-contained (hidden states + argmax) collection
depth_probe.py depth→argmax-accuracy curve (zero-training probe)
paper/ the preprint (markdown source)
results/ measured results shipped with the repo
Reproduce the exactness claims:
# 1. inter-process determinism (3 fresh processes, compare all hashes)
python harness/determinism_gate.py --model /path/to/model
# 2. bit-exact chunked attention vs eager (same process, A/B toggle)
# (see paper §4 and harness/ source for the protocol)- Single model family (qwen3_5 dense); porting = adapting
bind_layer+ the attention-mask cadence for other architectures. - Generation has no KV cache: each token re-streams all layers (the exact sweep is the point; an incremental KV decoder changes the wall only by a small factor — see preprint §6).
- Linux only (preadv/fadvise).
sudo -nused once at init fordrop_caches(optional,--no-drop-cachesto skip). - The C++ pipeline (20.2s/sweep, 4 zstd workers) is not yet wired into the
Python package — see
results/for its measured numbers.
@misc{tcpp2026,
title = {TCPP: Exact Out-of-Core Inference for Large Dense LLMs},
author = {Garrot, Cyril},
year = {2026},
note = {Preprint; software},
url = {https://github.com/cgarrot/tcpp}
}MIT — see LICENSE.