Five situations, what BigSmall concretely does in each, and the one command to start with. No jargon; every claim links to a page with the measured numbers.
Your card says 12, 16, or 24 GB and the models you want say more. BigSmall helps three ways: compressed models are ~34% smaller to download and store; the autopilot computes what actually fits on your card instead of you guessing; and when a model doesn't fit, the streaming executor runs it anyway from disk — slower, with the slowness stated in numbers before you commit.
bigsmall profile # once
bigsmall plan model.bs # "what would happen?" — one sentence
bigsmall run model.bs # do itYou never pick a quant level or an offload setting. If the honest answer is "this runs at 0.05 tokens/second," the sentence says so and you decide. A real receipt of the extreme case: a 14.2 GB model ran bit-exact in ~2.5 GB of working memory (streaming, capacity).
Checkpoints are sacred: a rerun must use exactly the bits you saved. BigSmall compresses every checkpoint ~34% with per-tensor md5 verification — verify re-proves bit-exactness any time, so "did compression touch my weights?" has a checkable answer, forever.
bigsmall compress run42/checkpoint-9000/ -o ckpt9000.bs
bigsmall verify ckpt9000.bsTwo more habits that pay: bigsmall xray tells you whether a checkpoint looks trained before you burn a day on a silently-corrupted load (integrity); and one honest warning from our own measurements — adjacent checkpoints from the same run do not delta-compress (they are bit-decorrelated; we measured it), so store each standalone. Fine-tune-vs-base is a different story — see the next persona.
Your users download your model thousands of times; 34% off is 34% off every download, and bit-exactness means zero re-validation — the model you QA'd is byte-for-byte the model they run. If you ship a fine-tune of a public base, ship the delta: users who have the base only download what changed (34–50% of full size for ≥7B instruct-class tunes, measured table).
bigsmall compress our-finetune/ --delta-from public-base/ -o patch.bsFor "try it now" UX, ship one .bsd — the Ferrell Duo: one file, two models, the fast one and the real one — instead of two artifacts. The same file loads as a fast INT4 preview (reading ~21% of the bytes) or as the bit-exact original, and the toolchain labels the INT4 mode lossy every single time, so nobody mistakes the preview for the product (dual-fidelity).
Twenty fine-tunes of the same base is twenty nearly-identical 14 GB folders. Store the base once, store each tune as a patch against it:
bigsmall compress tune-v7/ --delta-from base/ -o tune-v7.bs
bigsmall apply base/ tune-v7.bs -o tune-v7-restored/Patch size is pair-dependent — measured from under 1% (best ≥7B SFT pairs) to ~61% (small-model full tunes) — and since 3.15 the engine measures both codings per tensor and never produces a delta bigger than standalone, warning you up front when a pair is in the doesn't-pay regime. The full measured table, including the cases where delta is the wrong tool: delta-compression.md.
Your releases are already fp8 — half of bf16 — and conventional wisdom says that's the floor. Measured: a real fp8 release (Qwen3-0.6B-FP8) losslessly compressed to 0.829 of its fp8 bytes, all 507 tensors bit-exact, with the codec sitting at the measured entropy floor of the format (fp8). Both fp8 kinds (e4m3/e5m2) are native, in weights and in the KV-cache codec.
bigsmall compress our-fp8-release/ -o release.bs
bigsmall xray our-fp8-release/ # fp8-aware forensics, untrained-trap includedThe capacity consequence on a 20 GB card: ~21B fp8 params resident versus ~13B bf16 — fp8 plus lossless coding stacks (capacity).
Three behaviours you can rely on, because the test suite enforces them, not a style guide (integrity):
- Anything below bit-exact announces itself before it happens — never silently.
- Speed and memory promises are computed from measurements on your machine, and the streaming executor prints its actual usage next to its promise.
- Failures explain themselves in plain words ("needs ~X GB more than you have; closest option: …") and exit cleanly — no stack traces.
Start here: quickstart.md.