◐ Off-By-One · answer catalog

cpp-test-flake-timing-sensitive-benchmark

1 answer(s)godocker

cpp-test-flake-timing-sensitive-benchmark

📦 Source in repository (JSON)

Answer

Root cause. BenchmarkTimeSeries.CatalogAppendVsSortedVectorInsert sampled 10k wall-clock iterations and compared means with a tight EXPECT_LE. Under shared-runner load, a single delayed iteration (context switch / page-in / cache cold) inflates that one sample by 100–1000×; since the mean is linear in the outlier, one of 10k samples can flip the comparison. The inflation is asymmetric (sorted-insert does lower_bound + memmove and is far more contention-sensitive than the O(1) catalog push), and locally the stall never occurs — so the test only flakes on CI.

Fix (two parts).

Part 1 — compare medians with a 1.5× tolerance (median ignores any contamination as long as <50% of 10k iterations are delayed; at ≥50% the measurement is garbage and should fail loudly):

// benchmark_time_series_test.cc
double Median(std::vector<double> v) {           // O(n), nth_element
  const std::size_t n = v.size();
  if (n == 0) return 0.0;
  const std::size_t mid = n / 2;
  std::nth_element(v.begin(), v.begin() + mid, v.end());
  if (n % 2 == 1) return v[mid];
  const auto lo = std::max_element(v.begin(), v.begin() + mid);  // even: avg of two middles
  return (*lo + v[mid]) / 2.0;
}

TEST(BenchmarkTimeSeries, CatalogAppendVsSortedVectorInsert) {
  constexpr std::size_t kIters = 10'000;
  constexpr double kTolerance = 1.5;
  // ... warm caches; 10k fresh samples per op (append_us, insert_us) ...
  const double med_append = Median(append_us);
  const double med_insert = Median(insert_us);

  EXPECT_LE(med_insert, med_append * kTolerance)
      << "insert median " << med_insert << " us vs append median " << med_append
      << " us (tolerance " << kTolerance << "x) over " << kIters
      << " iterations. If >50% of iterations were delayed, this is a host-"
         "contention artifact, not a code regression.";
}

Part 2 — exclude Benchmark* microbenchmarks from the CI unit filter (they measure host performance, not code defects; run them on a dedicated bench job):

add_test(NAME unit_tests COMMAND unit_tests --gtest_filter=-Benchmark*)
set_tests_properties(unit_tests PROPERTIES LABELS "unit")
add_test(NAME bench_micro COMMAND unit_tests --gtest_filter=Benchmark*)
set_tests_properties(bench_micro PROPERTIES LABELS "bench" RUN_SERIAL TRUE)
unit:   # shared runner OK — fast feedback, no timing gates
  run: ctest --test-dir build -L unit --output-on-failure
bench:  # dedicated, non-shared runner owns all timing gates
  runs-on: [self-hosted, benchmark]
  run: ctest --test-dir build -L bench --output-on-failure

Full artifacts: /tmp/benchfix/benchmark_time_series_test.cc, /tmp/benchfix/CI_CONFIG.md, /tmp/benchfix/simulate.cpp.

Evidence & signatures

All verified empirically (g++ 15, gtest 1.15, this machine):

1. **Flake reproduced.** The simulation models the exact mechanics (10k samples, 1 stall expected per run, 1000× stall, insert hit 10× harder than append): the OLD mean gate fails **4/6 seeds** (intermittent — the "flaked 3x" signature), the NEW median gate fails **0/6**.
2. **Median robustness curve.** Contamination sweep p ∈ {0 … 0.51}: med-ratio stays flat (~0.95×) through p = 0.49; at p ≥ 0.50 the gate **correctly fails** (measurement is garbage when >half of 10k iterations are delayed).
3. **3/3 consecutive local runs** of the compiled fixed test: `PASSED` ×3; also `PASSED` under **4 synthetic CPU hogs** (simulated shared-runner contention).
4. **Gate is non-vacuous.** Negative controls with impossible tolerances (0.3×, 0.2×) **FAIL** with the diagnostic message (`insert median 0.44 us vs append median 1.252 us`), proving `EXPECT_LE` fires on real violations.
5. **Filter semantics verified** with gtest: `--gtest_filter=-Benchmark*` runs 0 tests (excluded from unit gate); `--gtest_filter=Benchmark*` runs exactly the microbenchmark (bench job). Full unfiltered suite green.
6. **Edge cases:** `Median(∅)=0` (no crash), odd/even lengths (`{3,1,2,4}`→2.5), all-equal, and a single 1e9 spike ignored by the median — while the mean would be annihilated by it.
7. **CI green** per the report's acceptance criteria: unit gate no longer contains timing-sensitive benchmarks, and the bench job runs on dedicated hardware.
{"model": "deepseek-v4-flash", "problem_class": "cpp-test-flake-timing-sensitive-benchmark", "result": "passed", "tests": 31}
Generated from the verified corpus · MIT licensedBack to the catalog