WHAT: BigSmall decides how to run a model on your machine — which fidelity, which placement, at what speed — and tells you in one sentence before doing it.
WHY you care: Running local models normally means decisions: quant level, VRAM budget, offload settings, context size. Get one wrong and the model crashes, crawls, or silently degrades. The autopilot makes those calls from a measured hardware profile, applies one rule — highest fidelity that still runs at usable speed — and is built so it cannot downgrade quality without telling you. You never learn a settings vocabulary; you read one sentence.
One-time hardware profile (~10 seconds, saved to ~/.bigsmall/profile.json):
bigsmall profileprofile: C:\Users\you\.bigsmall\profile.json
cpu: 8 logical cores | ram: 30.04 GB total, 21.32 GB free at probe
disk: 0.657 GB/s sequential (FILE_FLAG_NO_BUFFERING)
gpu: NVIDIA RTX A4500 | 19.5 GB usable [UNVERIFIED (GPU busy)]
kernels: bf16 decode 484.4 MB/s, int4 decode 853.7 MB/s (CPU, measured)
The decision, without execution:
bigsmall plan model.bsdRunning perfect mode (bit-exact, receipt verified) in CPU RAM at full CPU speed while the GPU is busy — fast mode (lossy INT4) available with --mode fast.
The decision, executed:
bigsmall run model.bsdRunning perfect mode (bit-exact, receipt verified) in CPU RAM at full CPU speed while the GPU is busy — fast mode (lossy INT4) available with --mode fast.
loaded 290 tensors into host RAM in 27.0s [mode=perfect] (bit-exact receipt honoured)
Overrides and escape hatches:
bigsmall plan model.bsd --mode fast # force a fidelity; the downgrade is announced
bigsmall plan model.bsd --min-tok-s 0.1 # change the usable-speed floor (default 1.0 tok/s)
bigsmall plan model.bsd --json # the full decision: placements, fit math, alternatives
bigsmall run model.bsd --stream # force the layer-streaming executorForcing a downgrade prints the announcement first, always:
NOTE: running 'fast' — lossy INT4 groupwise (g=128); NOT the original weights — below this file's best fidelity ('perfect': bit-exact original bf16 (verified)). Why: you asked for it (--mode).
Running fast mode (lossy INT4) in CPU RAM at full CPU speed while the GPU is busy.
How the pick works, in order: read what the file actually offers (modes + verification receipts from the container manifest) → compute what fits where from your profile (GPU resident, GPU compressed-resident, streamed from RAM, streamed from disk, CPU resident) → take the highest fidelity whose estimated speed clears the floor. If nothing clears the floor, it picks the best capacity option and puts the honest speed in the sentence rather than refusing.
NUMBERS (measured on the reference machine — 8-core CPU, RTX A4500, NVMe; v4 measurement campaign, 2026-06)
| What | Measured |
|---|---|
| Profile probe time | ~10 s, one-time (256 MB unbuffered disk sample + 8–24 MB decode benches) |
| Speed-estimate honesty, fast mode | promised 0.864 passes/s vs 0.806 measured steady (×1.07) |
| Speed-estimate honesty, perfect mode | promised 0.041 vs 0.048 measured steady (×0.86 — conservative) |
| RAM promise vs reality (streamed perfect) | promised 1.24 GB, measured peak 1.01 GB (within promise, 19% headroom) |
| Honesty law coverage | every profile × file-shape × GPU-busy × speed-floor combination swept in the test suite; an unannounced downgrade anywhere is a suite failure |
The estimator's rule is written into its source: the formula follows the measurement, never the reverse. When a kernel rate is re-measured, the promise moves to match reality.
- Speed estimates are decode/IO-bound bounds labeled with
~; resident placements say "full speed" / "full CPU speed" — there is no invented FLOPs model. - A busy GPU is never probed with CUDA, so its fields are static lookups marked
UNVERIFIEDuntil a quiet-GPU probe refreshes them. The planner treats a busy GPU as off-limits entirely. bigsmall runloads into host RAM (or streams); it is not an inference server. Wire the loaded weights into your own stack, or useserve-streamfor generation.- The planner's
.bsspeed estimate assumes the file's tensors use the parallel-decode codec; older.bsfiles coded with the serial codec decode slower than the estimate.bigsmall transcodere-encodes them (this gap is recorded, not hidden — see the knobs report).