Skip to content

Latest commit

 

History

History
80 lines (49 loc) · 4.84 KB

File metadata and controls

80 lines (49 loc) · 4.84 KB

How BigSmall helps — find yourself on this page

Five situations, what BigSmall concretely does in each, and the one command to start with. No jargon; every claim links to a page with the measured numbers.


"I have a gaming GPU and I want to run bigger models"

Your card says 12, 16, or 24 GB and the models you want say more. BigSmall helps three ways: compressed models are ~34% smaller to download and store; the autopilot computes what actually fits on your card instead of you guessing; and when a model doesn't fit, the streaming executor runs it anyway from disk — slower, with the slowness stated in numbers before you commit.

bigsmall profile          # once
bigsmall plan model.bs    # "what would happen?" — one sentence
bigsmall run model.bs     # do it

You never pick a quant level or an offload setting. If the honest answer is "this runs at 0.05 tokens/second," the sentence says so and you decide. A real receipt of the extreme case: a 14.2 GB model ran bit-exact in ~2.5 GB of working memory (streaming, capacity).


"I'm a researcher with a disk full of checkpoints"

Checkpoints are sacred: a rerun must use exactly the bits you saved. BigSmall compresses every checkpoint ~34% with per-tensor md5 verification — verify re-proves bit-exactness any time, so "did compression touch my weights?" has a checkable answer, forever.

bigsmall compress run42/checkpoint-9000/ -o ckpt9000.bs
bigsmall verify ckpt9000.bs

Two more habits that pay: bigsmall xray tells you whether a checkpoint looks trained before you burn a day on a silently-corrupted load (integrity); and one honest warning from our own measurements — adjacent checkpoints from the same run do not delta-compress (they are bit-decorrelated; we measured it), so store each standalone. Fine-tune-vs-base is a different story — see the next persona.


"We ship models to users"

Your users download your model thousands of times; 34% off is 34% off every download, and bit-exactness means zero re-validation — the model you QA'd is byte-for-byte the model they run. If you ship a fine-tune of a public base, ship the delta: users who have the base only download what changed (34–50% of full size for ≥7B instruct-class tunes, measured table).

bigsmall compress our-finetune/ --delta-from public-base/ -o patch.bs

For "try it now" UX, ship one .bsd — the Ferrell Duo: one file, two models, the fast one and the real one — instead of two artifacts. The same file loads as a fast INT4 preview (reading ~21% of the bytes) or as the bit-exact original, and the toolchain labels the INT4 mode lossy every single time, so nobody mistakes the preview for the product (dual-fidelity).


"I fine-tune, and I keep every version"

Twenty fine-tunes of the same base is twenty nearly-identical 14 GB folders. Store the base once, store each tune as a patch against it:

bigsmall compress tune-v7/ --delta-from base/ -o tune-v7.bs
bigsmall apply base/ tune-v7.bs -o tune-v7-restored/

Patch size is pair-dependent — measured from under 1% (best ≥7B SFT pairs) to ~61% (small-model full tunes) — and since 3.15 the engine measures both codings per tensor and never produces a delta bigger than standalone, warning you up front when a pair is in the doesn't-pay regime. The full measured table, including the cases where delta is the wrong tool: delta-compression.md.


"Our lab works in fp8 now"

Your releases are already fp8 — half of bf16 — and conventional wisdom says that's the floor. Measured: a real fp8 release (Qwen3-0.6B-FP8) losslessly compressed to 0.829 of its fp8 bytes, all 507 tensors bit-exact, with the codec sitting at the measured entropy floor of the format (fp8). Both fp8 kinds (e4m3/e5m2) are native, in weights and in the KV-cache codec.

bigsmall compress our-fp8-release/ -o release.bs
bigsmall xray our-fp8-release/     # fp8-aware forensics, untrained-trap included

The capacity consequence on a 20 GB card: ~21B fp8 params resident versus ~13B bf16 — fp8 plus lossless coding stacks (capacity).


Whoever you are: the honesty rules hold

Three behaviours you can rely on, because the test suite enforces them, not a style guide (integrity):

  1. Anything below bit-exact announces itself before it happens — never silently.
  2. Speed and memory promises are computed from measurements on your machine, and the streaming executor prints its actual usage next to its promise.
  3. Failures explain themselves in plain words ("needs ~X GB more than you have; closest option: …") and exit cleanly — no stack traces.

Start here: quickstart.md.