Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

TCPP — Exact out-of-core inference for large dense LLMs

Run a 52GB BF16 dense model bit-exactly on an 8GB GPU.

TCPP streams a large checkpoint layer-by-layer from NVMe through a small GPU and produces greedy outputs identical, bit for bit, to an in-VRAM eager forward of the same model. No approximation ever enters the authority path: quantization, drafters and rankers may predict, only the exact forward decides.

Résumé FR — TCPP exécute un modèle dense de 52GB BF16 (Qwen3.8-27B, 64 couches hybrides GDN/attention) sur un GPU de 8GB, avec une sortie greedy bit-exacte et déterministe inter-process. Vérification par lots à 172 positions/s ; génération ~1 token / sweep (~30s sur NVMe). Le but du projet : rendre utilisable — et ouvrir — l'inférence exacte de gros modèles denses sur matériel grand public.

Quickstart

pip install -r requirements.txt

# 1. sanity check (expects the exact argmax of the reference prompt)
python -m tcpp sanity --model /path/to/Qwen3.8-27B

# 2. batch verification: argmax for every position of a K-token context
python -m tcpp bench --model /path/to/Qwen3.8-27B --k 2048 4096 8192

# 3. exact greedy generation (slow: one full layer sweep per token)
python -m tcpp generate --model /path/to/Qwen3.8-27B \
    --prompt "La tour Eiffel se trouve à Paris et" --max-new 8

The model directory is any local HF snapshot in safetensors format (zero-copy: TCPP builds a page index over the original shards, no repacking). Supported architecture family in v0.1: qwen3_5 hybrid (Qwen3.5 / Qwen3.8 dense models, e.g. 9B / 27B).

Hardware requirements

Component Minimum (tested)
GPU VRAM 8GB (RTX 3070 laptop, sm_86)
System RAM 32GB (auto-budgeted resident layers; more = faster warm passes)
Disk NVMe SSD, ~1.6GiB/s sustained, model size + headroom
OS Linux (preadv, fadvise, /proc/meminfo)

Performance (measured, Qwen3.8-27B BF16, RTX 3070 8GB, NVMe)

Workload Result
Batch verify K=2048 33.0s (62 pos/s)
Batch verify K=4096 37.5s (109 pos/s)
Batch verify K=8192 47.5s (172 pos/s)
Greedy generation ~1 token / sweep (~30s) — I/O-bound, see below
Bit-exactness query-chunked attention hash-identical to eager (K=2048/4096)
Determinism 3 fresh processes → identical SHA256 of every layer's weights, activations and argmax

What "exact" means here

  • The attention is the eager algorithm replicated op-for-op (fp32 softmax → bf16 cast → matmul), only chunked along the query axis so the [T, T] matrix never materializes. Demonstrated hash-identical to eager at reference sizes; SDPA/flash kernels are NOT used (they change argmax on near-ties — measured 6/10 prompt divergences).
  • Weights are read from the original BF16 safetensors; no dequantization.
  • The runtime is deterministic across processes (validated by hashing every intermediate tensor in 3 fresh processes — see harness/determinism_gate.py).

Why generation is ~30s/token (and why that is the honest number)

Every generated token requires one full 52GB sweep through the memory hierarchy. At NVMe speed (~1.6GiB/s) the sweep floor is ~27-33s. The GPU itself could do ~400 tok/s compute-only if the weights were resident. The project's accompanying preprint quantifies why speculative decoding (EAGLE-3, DFlash, trees, pipeline speculation) cannot compensate: sequential acceptance is ~2-9 tokens per sweep for every drafter family tested, so the ceiling on this substrate is ~0.3-0.5 tok/s. The wall is the memory hierarchy, not the compute and not the drafter. Give this code 64GB of fast GPU memory and it is no longer the bottleneck.

Best-fit uses today: batch scoring/verification (172 pos/s), exactness audits, research on out-of-core execution. Interactive chat needs a resident substrate or a different model class (e.g. MoE with hot experts in VRAM).

Repository layout

tcpp/            the runtime (page index, pipelined loader, exact attention)
harness/         measurement tooling
  determinism_gate.py   3-process bit-exactness gate (weights+activations hashes)
  golden_collect.py     self-contained (hidden states + argmax) collection
  depth_probe.py        depth→argmax-accuracy curve (zero-training probe)
paper/           the preprint (markdown source)
results/         measured results shipped with the repo

Validation tooling

Reproduce the exactness claims:

# 1. inter-process determinism (3 fresh processes, compare all hashes)
python harness/determinism_gate.py --model /path/to/model

# 2. bit-exact chunked attention vs eager (same process, A/B toggle)
#    (see paper §4 and harness/ source for the protocol)

Limitations (v0.1)

  • Single model family (qwen3_5 dense); porting = adapting bind_layer + the attention-mask cadence for other architectures.
  • Generation has no KV cache: each token re-streams all layers (the exact sweep is the point; an incremental KV decoder changes the wall only by a small factor — see preprint §6).
  • Linux only (preadv/fadvise). sudo -n used once at init for drop_caches (optional, --no-drop-caches to skip).
  • The C++ pipeline (20.2s/sweep, 4 zstd workers) is not yet wired into the Python package — see results/ for its measured numbers.

Citation

@misc{tcpp2026,
  title  = {TCPP: Exact Out-of-Core Inference for Large Dense LLMs},
  author = {Garrot, Cyril},
  year   = {2026},
  note   = {Preprint; software},
  url    = {https://github.com/cgarrot/tcpp}
}

License

MIT — see LICENSE.

About

Exact bit-for-bit out-of-core inference for large dense LLMs — run a 52GB BF16 27B model on an 8GB GPU

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages