Skip to content

feat(bench): make the comparison fixture realistic, and expressible in its negative regimes - #32

Closed
lloydsk wants to merge 1 commit into
mainfrom
feat/dummy-suite-realism
Closed

lloydsk wants to merge 1 commit into
mainfrom
feat/dummy-suite-realism

Conversation

@lloydsk

@lloydsk lloydsk commented Sep 26, 2026

Copy link
Copy Markdown
Collaborator

Why

The generated suite is the standard comparison unit, and its shape decides which conclusions are even available. Three of its properties were unrealistic in the tool's favour, and the regime where a background job should have nothing to offer could not be generated at all.

What

  • DUMMY_LINES_PER_TEST [2] — the fixture logged 8 lines per test, against bun's own ~1.5, so the output-volume advantage was overstated ~5x. The default is now realistic; =8 reproduces the density the earliest measurements in docs/dogfooding.md used, which matters when comparing against them.
  • DUMMY_FAIL_FAST=1 — the run collapses at the planted failure, every later test short-circuiting. That is the regime the methodology predicts bgrun loses: little to wait for, and the wake is overhead.
  • A realistic failure block — assertion, +actual/-expected diff, and a stack trace naming the generated file and the real assert line (~19 lines), instead of a two-line stub. The DIAGNOSTIC_MARKER_* strings are unchanged.
  • Per-file banners and a summary block, so the output has a real runner's shape. Summary counters come from runtime counters, not from the knobs.

Every previous knob, guard and behaviour is preserved: the globalThis.__chain serialisation that keeps duration scheduler-independent, symlink-safe canonicalisation, the fatal in-repo guard, the planted-failure validation, the part-NN.test.ts layout.

Verification

check result
default, DUMMY_SLEEP_MS=20 exit 1, 694 output lines (600 stage + banners + summary), both markers, failure stack at part-05.test.ts:58
DUMMY_FAIL_FAST=1, same scale 4.09s against 6.29s, suite ending exactly at the failure (105 tests skipped in the fixture's own accounting)
DUMMY_LINES_PER_TEST=8 2,494 output lines — the volume axis works
DUMMY_OUT_DIR=$PWD/subdir still exits 2
tsc --noEmit clean

A finding, not a footnote

bun 1.3.6 discovers and executes these files in a deterministic but non-lexical order (measured twice: 2,3,1,8,9,10,5,4,6,7), so part-05 runs seventh of ten. That is why the planted failure sits ~65% through the run while DUMMY_FAIL_FILE=5 reads as 45%, and why fail-fast stops at 65% rather than 45%. No fixture can reorder bun's discovery; the doc's "failing 65% of the way through" was measured, and this is the mechanism.

The comparison fixture flattered bgrun in two ways and offered no regime in
which bgrun should lose.

- DUMMY_LINES_PER_TEST [2]: the hardcoded 8 lines/test is a verbose/CI-log
  density, several times a normal runner's ~1.5 lines/test. It inflated output
  volume — one axis of the comparison — without saying so. The report now
  states the density used, and the header records that every earlier
  measurement of this fixture ran at 8.
- DUMMY_FAIL_FAST [0]: every run was long, so a background job's premise (the
  run outlives the turn) always held. With 1 the suite collapses at the planted
  failure — later tests short-circuit with one skipped line — so wall-clock
  ends there. That is the regime where bgrun should not be expected to win.
- The planted failure now prints a full assertion-failure block (assertion,
  expected vs received, stack trace naming the generated file and line) instead
  of a bare assert, so failing-run output volume is realistic. The exact marker
  strings DIAGNOSTIC_MARKER_UPSTREAM/DOWNSTREAM are unchanged.
- Output now has a real runner's shape: a banner per file and a summary block
  at the end, whose counters are read from the run (the runner picks its own
  file order) rather than assumed from the knobs.
- Header tunables section rewritten to name the axis each knob isolates: time,
  volume, failure position, fail-fast, plus the output directory.

Kept unchanged: the shared globalThis chain, symlink canonicalisation, the
fatal in-repo guard, the planted-failure validation, and the part-NN.test.ts
layout.
@github-actions

Copy link
Copy Markdown
Contributor

CI report

Check Result
tsc --noEmit success
tests success
npm pack --dry-run success (7 files in tarball)

Ref: f593c77eeb5bbbbf818f5ccd636a29bda8d59ea5

@lloydsk

lloydsk commented Sep 26, 2026

Copy link
Copy Markdown
Collaborator Author

Superseded: the fixture generator, the measurement instrument and the write-up are now one branch, bench/dogfooding (7 commits). No content is lost — this PR described one third of a deliverable.

@lloydsk lloydsk closed this Sep 26, 2026
@lloydsk
lloydsk deleted the feat/dummy-suite-realism branch September 26, 2026 02:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant