Conversation
The comparison fixture flattered bgrun in two ways and offered no regime in which bgrun should lose. - DUMMY_LINES_PER_TEST [2]: the hardcoded 8 lines/test is a verbose/CI-log density, several times a normal runner's ~1.5 lines/test. It inflated output volume — one axis of the comparison — without saying so. The report now states the density used, and the header records that every earlier measurement of this fixture ran at 8. - DUMMY_FAIL_FAST [0]: every run was long, so a background job's premise (the run outlives the turn) always held. With 1 the suite collapses at the planted failure — later tests short-circuit with one skipped line — so wall-clock ends there. That is the regime where bgrun should not be expected to win. - The planted failure now prints a full assertion-failure block (assertion, expected vs received, stack trace naming the generated file and line) instead of a bare assert, so failing-run output volume is realistic. The exact marker strings DIAGNOSTIC_MARKER_UPSTREAM/DOWNSTREAM are unchanged. - Output now has a real runner's shape: a banner per file and a summary block at the end, whose counters are read from the run (the runner picks its own file order) rather than assumed from the knobs. - Header tunables section rewritten to name the axis each knob isolates: time, volume, failure position, fail-fast, plus the output directory. Kept unchanged: the shared globalThis chain, symlink canonicalisation, the fatal in-repo guard, the planted-failure validation, and the part-NN.test.ts layout.
Contributor
CI report
Ref: |
Collaborator
Author
|
Superseded: the fixture generator, the measurement instrument and the write-up are now one branch, |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
The generated suite is the standard comparison unit, and its shape decides which conclusions are even available. Three of its properties were unrealistic in the tool's favour, and the regime where a background job should have nothing to offer could not be generated at all.
What
DUMMY_LINES_PER_TEST[2] — the fixture logged 8 lines per test, against bun's own ~1.5, so the output-volume advantage was overstated ~5x. The default is now realistic;=8reproduces the density the earliest measurements indocs/dogfooding.mdused, which matters when comparing against them.DUMMY_FAIL_FAST=1— the run collapses at the planted failure, every later test short-circuiting. That is the regime the methodology predicts bgrun loses: little to wait for, and the wake is overhead.+actual/-expecteddiff, and a stack trace naming the generated file and the real assert line (~19 lines), instead of a two-line stub. TheDIAGNOSTIC_MARKER_*strings are unchanged.Every previous knob, guard and behaviour is preserved: the
globalThis.__chainserialisation that keeps duration scheduler-independent, symlink-safe canonicalisation, the fatal in-repo guard, the planted-failure validation, thepart-NN.test.tslayout.Verification
DUMMY_SLEEP_MS=20part-05.test.ts:58DUMMY_FAIL_FAST=1, same scaleDUMMY_LINES_PER_TEST=8DUMMY_OUT_DIR=$PWD/subdirtsc --noEmitA finding, not a footnote
bun 1.3.6 discovers and executes these files in a deterministic but non-lexical order (measured twice:
2,3,1,8,9,10,5,4,6,7), sopart-05runs seventh of ten. That is why the planted failure sits ~65% through the run whileDUMMY_FAIL_FILE=5reads as 45%, and why fail-fast stops at 65% rather than 45%. No fixture can reorder bun's discovery; the doc's "failing 65% of the way through" was measured, and this is the mechanism.