One warm-up run does not warm the query up. The first timed run is
systematically the slowest, and #749's --print-runs makes it visible on its
first real use.
SF10, 84090b4, --runs 3 --print-runs, series in execution order:
runs IS1 spread 1.77x 0.05ms 0.03ms 0.03ms
runs IS2 spread 1.68x 0.11ms 0.07ms 0.08ms
runs IS6 spread 1.88x 0.06ms 0.03ms 0.03ms
runs IS7 spread 1.42x 0.17ms 0.12ms 0.13ms
runs IC4 spread 1.38x 27.5ms 21.9ms 19.9ms
runs IC7 spread 1.58x 0.41ms 0.28ms 0.26ms
runs IC11 spread 1.12x 54.5ms 50.1ms 48.7ms
Roughly 15 of 21 queries have run 1 as their slowest, and on the short reads
it is 1.4–1.9×. run_benchmark does one warm-up before the timed runs, so this
is a second warm-up effect the first one did not absorb.
Why it matters
Every spread statistic in this repo is inflated by it. #716 argued IC2 is
uniquely unstable partly from spread; the adjacency-tier comparison (#746) read
per-query differences of 1.3–1.9× off 3-run medians and two of them later
reversed. A systematic first-run penalty of up to 1.9× on short reads is a
plausible contributor to both, and nobody could see it because the series was
discarded.
It also biases min (never from run 1) and max (usually run 1), so any
max/min spread — which is what analyse_noise.py and CH-REGRESS reason
about — carries it.
The median of three is the least affected, which is why the published table has
held up. But the published table is not the only thing people read.
What to do
Two options, and I have deliberately not picked one because it shifts published
numbers:
- Warm up more than once. Two or three warm-ups, discarded. Simple, and it
makes --runs 3 mean what it says.
- Discard the first timed run when
runs > 3, reporting the rest. Keeps
the warm-up cheap and is what most benchmark harnesses do.
Either changes every number slightly, so it wants the same treatment as #746:
decide, then re-run the three-engine comparison on that basis rather than let
the default shift under a published table.
What is not in question is that the current output describes a distribution
that includes a warm-up artefact and calls it variance.
One warm-up run does not warm the query up. The first timed run is
systematically the slowest, and #749's
--print-runsmakes it visible on itsfirst real use.
SF10,
84090b4,--runs 3 --print-runs, series in execution order:Roughly 15 of 21 queries have run 1 as their slowest, and on the short reads
it is 1.4–1.9×.
run_benchmarkdoes one warm-up before the timed runs, so thisis a second warm-up effect the first one did not absorb.
Why it matters
Every spread statistic in this repo is inflated by it. #716 argued IC2 is
uniquely unstable partly from spread; the adjacency-tier comparison (#746) read
per-query differences of 1.3–1.9× off 3-run medians and two of them later
reversed. A systematic first-run penalty of up to 1.9× on short reads is a
plausible contributor to both, and nobody could see it because the series was
discarded.
It also biases min (never from run 1) and max (usually run 1), so any
max/minspread — which is whatanalyse_noise.pyand CH-REGRESS reasonabout — carries it.
The median of three is the least affected, which is why the published table has
held up. But the published table is not the only thing people read.
What to do
Two options, and I have deliberately not picked one because it shifts published
numbers:
makes
--runs 3mean what it says.runs > 3, reporting the rest. Keepsthe warm-up cheap and is what most benchmark harnesses do.
Either changes every number slightly, so it wants the same treatment as #746:
decide, then re-run the three-engine comparison on that basis rather than let
the default shift under a published table.
What is not in question is that the current output describes a distribution
that includes a warm-up artefact and calls it variance.