Your Benchmarks Are Wrong
I ran my powerix benchmark twice, back to back, same binary, same machine, same flags. The first run said pow_hierarchical<uint32> took 38 ns. The second said 44 ns. That is a 15% difference on a function I was trying to optimize by 10%.
I had no idea which number was right. Probably neither.
The Problem with Running Things Once
Google Benchmark is excellent at what it does: it handles iteration counting, warm-up, and precise time measurement within a single process run. But one run gives you one sample. And one sample tells you almost nothing about the true performance of your code.
The variance between runs comes from everywhere: turbo boost, thermal throttling, OS scheduler moving your process between cores, cache state, and on hybrid CPUs (Alder Lake+) you might land on a P-core or an E-core.
You can pass --benchmark_repetitions=10 to Google Benchmark, but that repeats within the same process. It does not capture the inter-process variance — the stuff that changes between two invocations of the binary.
So you end up doing what I did: run it, look at the numbers, run it again, squint, decide it is "probably fine." That is not engineering.
What I Built
meta-benchmark is a Python wrapper around Google Benchmark. It runs your benchmark binary multiple times and tells you when the results have statistically converged. If some cases are still noisy, it re-runs only those cases, not the entire suite.
The algorithm:
- Run all benchmarks for at least 5 iterations (
--min-meta-reps) - After each run, compute a 95% confidence interval for every case
- If the relative CI half-width (absolute CI divided by the mean) is below 3%, that case is stable. Stop re-running it.
- Re-run only the unstable cases
- Stop when everything converges or you hit the maximum (default: 30)
Why Student's t, Not the Normal Distribution
With 5 samples, you do not have enough data to trust the normal approximation. The sample standard deviation itself is uncertain, and the t-distribution accounts for that with heavier tails.
def rel_ci95_half(self) -> float: n = len(self.samples_ns) if n < 2: return float("inf") critical = T_CRITICAL_95.get(n - 1, Z_CRITICAL_95) half = critical * (self.stddev / math.sqrt(n)) return half / self.mean
At n=5 (df=4), the t-critical value is 2.776 — 42% larger than the z=1.96 you would use with a normal. Using z=1.96 would make your confidence intervals too narrow, declaring results stable when they are not. The critical values are hardcoded for degrees of freedom 1 through 30 (no scipy dependency) and fall back to z=1.96 beyond that.
The Smart Part: Selective Re-running
After the minimum runs, the tool constructs a regex filter for Google Benchmark. If you have 35 cases and 29 are still unstable, the next run only executes those 29.
This is where most of the time savings come from. A full suite might take 10 seconds. Re-running 5 stubborn cases takes 1 second.
Real Results: Powerix on This Machine
I ran meta-benchmark on the powerix integer benchmark suite (35 cases) on my Intel 14-core at 5.2 GHz. Two runs: one without core pinning, one pinned to core 0.
Without pinning — after 15 meta-runs, only 4 of 35 cases converged to 3% CI. Average CI: 14%. The worst offender was std_pow<uint32> at 25.4%.
Pinned to core 0 — 6 of 35 converged. Average CI: 6.3%.
| Metric | No pinning | Pinned core 0 |
|---|---|---|
| Stable cases | 4/35 | 6/35 |
| Average CI | 14.0% | 6.3% |
| Worst CI | 25.4% | 14.0% |
Pinning cut the average CI by more than half. The cases where it helped most:
| Case | CI (no pin) | CI (pinned) |
|---|---|---|
pow_c_raw<double> | 22.6% | 2.8% |
hierarchical<uint16> | 19.6% | 2.9% |
std_pow<uint32> | 25.4% | 9.4% |
But here is the interesting part: a few cases were actually more stable without pinning. pow_cached_unordered_pair<uint32> went from 1.8% CI unpinned to 9.0% pinned. Thermal throttling on a single core under sustained load can add variance that scheduling across cores would average out.
What This Tells You
Even with core pinning and 15 repetitions, only 17% of cases converged to 3% CI. If you ran any of these benchmarks once and reported the number, there is a good chance it was off by 10-25%. And if you compared two implementations and found a "5% improvement" — that might just be noise.
Going Deeper: Inner vs Outer Repetitions
Google Benchmark has --benchmark_repetitions=N which repeats within a single process. meta-benchmark launches N separate processes. Are these equivalent? I tested three strategies, all using 30 total samples per case:
Inner reps only (strategy A) have the widest CIs. All 30 samples share the same process — same thermal state, same cache, same scheduler decision. They capture micro-noise but miss the inter-process variance that actually matters. Pure outer reps (B) are better because each sample comes from a fresh process invocation. The mixed approach (C) wins: inner reps average out micro-noise within each invocation, while outer reps capture the real inter-process variance.
The conclusion is clear: if you only do in-process repetitions, you are measuring the wrong variance.
The Convergence Curve
I ran 20 meta-repetitions without selective filtering (re-running all 35 cases every time) and tracked how stability evolves:
Stable cases rise quickly as more data comes in, with most cases converging within 5-8 reps. The average CI drops sharply in the first few reps then levels off. The worst CI follows a similar pattern — a few stubborn cases take longer to settle.
On a well-powered machine with stable thermals, convergence is monotonic and predictable. But I also ran this experiment on battery power, and the results were dramatically different: stable cases would peak around rep 7, then drop back down as thermal throttling introduced systematic drift. The CI would widen as the statistics correctly captured this drift. The lesson: your power source affects your benchmarks. Always benchmark on AC power with stable thermals.
Selective vs Non-Selective Re-running
meta-benchmark only re-runs unstable cases. Does this filtering actually help?
Selective re-running takes roughly half the wall-clock time — it skips the cases that already converged. On a stable machine, both approaches reach full convergence; selective just gets there faster. The time savings grow with suite size: if you have 200 benchmarks and 180 converge after 5 runs, selective re-running avoids re-executing those 180 cases for the remaining iterations.
In my tests, non-selective re-running actually produces slightly tighter CIs because it collects more data points per case. But the extra precision is marginal — both approaches agree on which implementations are faster. The selective approach is the pragmatic choice for day-to-day optimization work.
Per-Case Convergence Trajectories
To understand how individual cases behave over many repetitions, I plotted the running mean and 95% CI band for 6 cases:
The fast integer kernels (pow_binary, pow_ultra_fast, hierarchical) converge tightly — their CI bands narrow quickly and the running mean stabilizes. The cached implementations and std_pow take longer to settle, with wider bands reflecting higher intrinsic variance.
What I find most useful about these trajectories: you can see when a case has converged. The running mean stops moving and the CI band stops shrinking. That is exactly what meta-benchmark's convergence criterion detects automatically.
Integration with BenchDiff
meta-benchmark pairs with BenchDiff for detecting performance regressions. The workflow: run meta-benchmark on your baseline, make changes, rebuild, run meta-benchmark again, then compare the two JSON outputs with BenchDiff. Because both measurements have converged CIs, BenchDiff can distinguish real regressions from noise.
What I Learned
- A single run is a sample, not a measurement. You need multiple samples and convergence checks before reporting anything.
- Inner reps are not outer reps. In-process repetitions capture different variance than separate invocations. The best strategy is a mix: moderate inner reps to smooth micro-noise, moderate outer reps to capture inter-process effects.
- Your environment matters more than your code. Battery vs AC power, core pinning, thermal state — these change your results by 10-30%. Always benchmark on AC power with stable thermals.
- Selective filtering saves time without sacrificing accuracy. It skips converged cases and focuses effort on the noisy ones. For a 35-case suite, it halves the wall-clock time.
- Core pinning helps on average but some cases behave worse when pinned. Test both ways and pick what works for your workload.
The code is on GitHub: meta-benchmark.