Claypool

Index

html | docx | pdf | txt | md

v1.0, diff: [prev | all], License

Run Length - Choosing an Experimental Run Length

The Question

In an experiment that measures a system over time — frame rate, throughput, latency — each run produces one number, such as the mean fps. If the runs are too short, that number is noisy, and repeating the same experiment gives a different answer. If the runs are too long, time is wasted that could have gone into more runs or more conditions. The question - how long a run is "long enough" and how to figure this out?

Using Standard Deviation to Choose Run Length

This technique finds the shortest run length whose result is (about) as repeatable as a much longer run.

  1. Pick the metric. This is the per-run number you will report, such as mean f/s, median f/s or 1%-low f/s. Run the analysis separately for each metric, because tail metrics like 1%-low need longer runs than the mean does.

  2. Pick candidate lengths. Space them roughly geometrically, for example 5, 10, 15, 20, 30, 45, 60, 90, 120, 180, 240 and 300 s (i.e., step size 5, then 10, then 15, then 30, then 60).

  3. Remove warm-up. Drop "warm-up" time (e.g., cache fill, page-faults, clock ramp-up) from the front of the run before measuring. Run length means the measured part only.

  4. Repeat runs at each length. Use N independent runs per length. Aim for 20–30, but at least 10. You can get this cheaply by recording N long runs and truncating each one to the first L seconds after warm-up.

  5. Compute the std dev of the metric across runs at each length.

  6. Plot std dev against run length. The curve typically falls steeply at first and then flattens.

  7. Find the knee. This is the shortest length at which the curve has effectively reached its plateau. One concrete rule is the shortest L whose std dev is within about 25% of the plateau, where the plateau is the mean std dev of the longest few lengths. Another is the first L whose confidence band overlaps the plateau.

  8. Pick a run length just past the knee. To be "safe" and make sure it is long enough, pick a length that is past the knee and in the "flat" part of the curve. This assumes it is better to waste a bit of time (i.e., with longer runs) than it is to get the wrong answer (i.e., from shorter runs). A simple rule is the first candidate length at least 2× the knee.

Why the Curve Flattens

The variance of a run's mean has two parts:

Var(run mean)  ≈  σ²_between  +  σ²_within · τ / L

Once the first term is small compared with the second, a longer run does not give more value - the only way to gain more precision then is more runs.

Special Case: if the curve never flattens and keeps falling like 1/√L, the between-run variance is negligible. In that case, choose the length that gives the precision needed (see the alternatives below in the example) instead of looking for a knee.

Example: Game Frame Rate

Assume this setup:

A single run is depicted in Figure 1.

One run of frame rate over time

The std dev of the per-run mean f/s across the 30 runs, at each length. The table data is below (knee in italics, chosen length in bold).

Length (s) 5 10 15 20 30 45 60 90 120 180 240 300
Stddev (f/s) 1.55 1.20 1.01 1.10 0.99 0.67 0.73 0.64 0.68 0.54 0.59 0.54

A graph of the data is in Figure 2, with the shaded band is a bootstrap 90% interval on each std dev.

Std dev of mean f/s vs. run length

From Figure 2:

Practical notes