Skip to content

Latest commit

 

History

History
79 lines (57 loc) · 4.48 KB

File metadata and controls

79 lines (57 loc) · 4.48 KB

Autopilot — profile, plan, run

WHAT: BigSmall decides how to run a model on your machine — which fidelity, which placement, at what speed — and tells you in one sentence before doing it.

WHY you care: Running local models normally means decisions: quant level, VRAM budget, offload settings, context size. Get one wrong and the model crashes, crawls, or silently degrades. The autopilot makes those calls from a measured hardware profile, applies one rule — highest fidelity that still runs at usable speed — and is built so it cannot downgrade quality without telling you. You never learn a settings vocabulary; you read one sentence.

HOW

One-time hardware profile (~10 seconds, saved to ~/.bigsmall/profile.json):

bigsmall profile
profile: C:\Users\you\.bigsmall\profile.json
  cpu: 8 logical cores | ram: 30.04 GB total, 21.32 GB free at probe
  disk: 0.657 GB/s sequential (FILE_FLAG_NO_BUFFERING)
  gpu: NVIDIA RTX A4500 | 19.5 GB usable [UNVERIFIED (GPU busy)]
  kernels: bf16 decode 484.4 MB/s, int4 decode 853.7 MB/s (CPU, measured)

The decision, without execution:

bigsmall plan model.bsd
Running perfect mode (bit-exact, receipt verified) in CPU RAM at full CPU speed while the GPU is busy — fast mode (lossy INT4) available with --mode fast.

The decision, executed:

bigsmall run model.bsd
Running perfect mode (bit-exact, receipt verified) in CPU RAM at full CPU speed while the GPU is busy — fast mode (lossy INT4) available with --mode fast.
loaded 290 tensors into host RAM in 27.0s [mode=perfect] (bit-exact receipt honoured)

Overrides and escape hatches:

bigsmall plan model.bsd --mode fast       # force a fidelity; the downgrade is announced
bigsmall plan model.bsd --min-tok-s 0.1   # change the usable-speed floor (default 1.0 tok/s)
bigsmall plan model.bsd --json            # the full decision: placements, fit math, alternatives
bigsmall run  model.bsd --stream          # force the layer-streaming executor

Forcing a downgrade prints the announcement first, always:

NOTE: running 'fast' — lossy INT4 groupwise (g=128); NOT the original weights — below this file's best fidelity ('perfect': bit-exact original bf16 (verified)). Why: you asked for it (--mode).
Running fast mode (lossy INT4) in CPU RAM at full CPU speed while the GPU is busy.

How the pick works, in order: read what the file actually offers (modes + verification receipts from the container manifest) → compute what fits where from your profile (GPU resident, GPU compressed-resident, streamed from RAM, streamed from disk, CPU resident) → take the highest fidelity whose estimated speed clears the floor. If nothing clears the floor, it picks the best capacity option and puts the honest speed in the sentence rather than refusing.

NUMBERS (measured on the reference machine — 8-core CPU, RTX A4500, NVMe; v4 measurement campaign, 2026-06)

What Measured
Profile probe time ~10 s, one-time (256 MB unbuffered disk sample + 8–24 MB decode benches)
Speed-estimate honesty, fast mode promised 0.864 passes/s vs 0.806 measured steady (×1.07)
Speed-estimate honesty, perfect mode promised 0.041 vs 0.048 measured steady (×0.86 — conservative)
RAM promise vs reality (streamed perfect) promised 1.24 GB, measured peak 1.01 GB (within promise, 19% headroom)
Honesty law coverage every profile × file-shape × GPU-busy × speed-floor combination swept in the test suite; an unannounced downgrade anywhere is a suite failure

The estimator's rule is written into its source: the formula follows the measurement, never the reverse. When a kernel rate is re-measured, the promise moves to match reality.

LIMITS, stated plainly

  • Speed estimates are decode/IO-bound bounds labeled with ~; resident placements say "full speed" / "full CPU speed" — there is no invented FLOPs model.
  • A busy GPU is never probed with CUDA, so its fields are static lookups marked UNVERIFIED until a quiet-GPU probe refreshes them. The planner treats a busy GPU as off-limits entirely.
  • bigsmall run loads into host RAM (or streams); it is not an inference server. Wire the loaded weights into your own stack, or use serve-stream for generation.
  • The planner's .bs speed estimate assumes the file's tensors use the parallel-decode codec; older .bs files coded with the serial codec decode slower than the estimate. bigsmall transcode re-encodes them (this gap is recorded, not hidden — see the knobs report).