Skip to content

One warm-up run does not warm up: the first timed run is the slowest in ~15 of 21 queries #752

Description

@sandeepkunkunuru

One warm-up run does not warm the query up. The first timed run is
systematically the slowest, and #749's --print-runs makes it visible on its
first real use.

SF10, 84090b4, --runs 3 --print-runs, series in execution order:

runs IS1    spread  1.77x  0.05ms 0.03ms 0.03ms
runs IS2    spread  1.68x  0.11ms 0.07ms 0.08ms
runs IS6    spread  1.88x  0.06ms 0.03ms 0.03ms
runs IS7    spread  1.42x  0.17ms 0.12ms 0.13ms
runs IC4    spread  1.38x  27.5ms 21.9ms 19.9ms
runs IC7    spread  1.58x  0.41ms 0.28ms 0.26ms
runs IC11   spread  1.12x  54.5ms 50.1ms 48.7ms

Roughly 15 of 21 queries have run 1 as their slowest, and on the short reads
it is 1.4–1.9×. run_benchmark does one warm-up before the timed runs, so this
is a second warm-up effect the first one did not absorb.

Why it matters

Every spread statistic in this repo is inflated by it. #716 argued IC2 is
uniquely unstable partly from spread; the adjacency-tier comparison (#746) read
per-query differences of 1.3–1.9× off 3-run medians and two of them later
reversed. A systematic first-run penalty of up to 1.9× on short reads is a
plausible contributor to both, and nobody could see it because the series was
discarded.

It also biases min (never from run 1) and max (usually run 1), so any
max/min spread — which is what analyse_noise.py and CH-REGRESS reason
about — carries it.

The median of three is the least affected, which is why the published table has
held up. But the published table is not the only thing people read.

What to do

Two options, and I have deliberately not picked one because it shifts published
numbers:

  1. Warm up more than once. Two or three warm-ups, discarded. Simple, and it
    makes --runs 3 mean what it says.
  2. Discard the first timed run when runs > 3, reporting the rest. Keeps
    the warm-up cheap and is what most benchmark harnesses do.

Either changes every number slightly, so it wants the same treatment as #746:
decide, then re-run the three-engine comparison on that basis rather than let
the default shift under a published table.

What is not in question is that the current output describes a distribution
that includes a warm-up artefact and calls it variance.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions