Skip to content

docs(dogfooding): measure foreground bash against bgrun, with the fixtures and method - #30

Closed
lloydsk wants to merge 7 commits into
mainfrom
docs/dogfooding-benchmarks
Closed

lloydsk wants to merge 7 commits into
mainfrom
docs/dogfooding-benchmarks

Conversation

@lloydsk

@lloydsk lloydsk commented Sep 26, 2026

Copy link
Copy Markdown
Collaborator

No description provided.

…tures and method

Adds the measured comparison to the dogfooding doc: how a foreground `bash` run and a
`bgrun` job compare on the same task, in live sessions, with per-run numbers and the
fixtures they come from.

- Method: three fixtures (the repo's own suite; a generated long, noisy, failing
  suite; and a short suite failing early), an identical prompt per arm except the
  mechanism clause, live sessions (a one-shot `pi -p` exits before the wake lands),
  and metrics read from the session transcript — executions, blocked-on-output,
  wall, context chars, and whether the diagnostic reached context.
- The mechanism-determined differences are stable (bgrun never blocks the session and
  runs the suite once); the agent-determined ones are not. On the long fixture,
  vanilla ran the suite twice for 359.9s of blocked time and 10,016 chars, against
  bgrun's single execution, 0.0s, and 8,049 chars — 28% of which was one pre-wake
  poll the guidance now discourages.
- It also records where the naive path converges on bgrun's technique by hand
  (redirecting a long run to a file, then grepping it), so the doc does not overclaim.

Fixture: scripts/make-dummy-suite.ts (bun run bench:make-suite).
@github-actions

github-actions Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

CI report

Check Result
tsc --noEmit success
tests success
npm pack --dry-run success (7 files in tarball)

Ref: 59e4430b50e1282e6ac7b59871d209d7d3c92d16

…grun does not help

Three additions, prompted by the obvious hazard: a fixture we can tune is a
fixture we can tune in the tool's favour.

- **How these numbers are produced** — the protocol as one table (arms, prompt
  template, model, live sessions, n=3 reported as median and range, metric
  definitions, and the command that re-derives every table from the transcripts),
  plus the `vanilla-hinted` control arm: the agent hand-rolling redirect-and-grep,
  which is the strongest baseline in the set.
- **Pre-registered expectations** — nine cells with the predicted winner written
  down before the runs, including the ones where bgrun is expected to lose (R1
  green, R2 tail-shaped failure, R5 fail-fast) and the one where our own story may
  be wrong (D1: if the advantage is reliability-driven it should grow with a weaker
  model). C1-C4 are dialogues — overlap, three jobs at once, a mid-run restart,
  asking where a run is — because no generated fixture can express them.
- **The fixture's design and why each property is under test** — the shared chain
  that makes duration scheduler-independent, the 65% failure position, the
  realistic failure block, and the knobs for the negative cells. Records that the
  earlier long-suite run used the verbose 8-lines-per-test density while the
  realistic default is 2, and that the difference is an axis rather than an
  oversight.
- **Where bgrun helps, and where it does not** — the strongest claim the data
  supports is the removal of the worst case, not a better best case (vanilla's
  `tail -100` run beat every bgrun run on cost). Names the papercuts the measuring
  found: the 28% pre-wake poll (#27 fixed), window-relative bggrep line numbers
  (#25 open), and re-paid log reads before delta `bgtail` (#4).
…he context

The standard protocol's first two vanilla cells, n=3 each, with the runs kept
rather than the aggregate:

- R1 (green suite, one-shot, same prompt and command): 4,894 / 5,018 / 21,787
  context chars. Two sessions bounded the output, one read the whole 241-line run
  — a 4.4x swing from the agent's choice of window alone. bgrun has no equivalent
  mode, which makes this a claim about variance rather than about the median.
- R2 (red suite, failure 78-90 lines from the end): executions 1, 3, 1; blocked
  21.8s, 66.2s, 21.9s; context 3,641 / 23,034 / 4,719. Across all six vanilla runs
  of this command the median is one execution and ~4.7 KB and the worst is four
  executions and 66.2s blocked.

The earlier three-run sample happened to be a worse draw, so the doc now carries
both samples and says so, instead of the rounder number. Also worth recording: in
all six runs the tail did reach the diagnostic — this fixture punishes a bad window
with another run, not with a missed failure.
Half the matrix cannot be scripted — the bgrun wake needs a live session — so the
protocol needs an executable run sheet: fresh session, prompt verbatim, then no
nudges (the wake is the mechanism under test), then profile the session directory.

Also names the two temptations this method has to resist: re-running a cell whose
number looks wrong (sampling on the outcome), and comparing a live bgrun cell with
a one-shot vanilla cell without recording the difference in the table.
…e row needs re-running

Records what the realism pass *verified* rather than what it intended:

- 694 output lines at the standard density, 2,494 at the verbose one.
- The 65% failure position now has a mechanism instead of a measurement: bun 1.3.6
  discovers these files in a non-lexical order (2,3,1,8,9,10,5,4,6,7), so part-05
  runs seventh of ten while DUMMY_FAIL_FILE=5 reads as 45%. Fail-fast therefore
  stops at 65%, not the 45% a lexical order would give.
- The failure block is ~19 lines, with a real +actual/-expected diff and a stack
  naming the generated file and line.
- The long-suite row below is relabelled as the verbose-density variant: its
  numbers came from ~2,760 lines of output, where the standard fixture emits 694.
  The failure's distance from the end *is* the mechanism, so those numbers cannot
  sit next to R1/R2's unqualified. The pre-registration table now says so at the
  cell, and the section says the row is pending re-measurement.
Port the throwaway arm profilers into a committed, dependency-free CLI so
the numbers in docs/dogfooding.md can be re-derived, not just trusted.

`bun scripts/measure-sessions.ts <dir-or-glob...> [--csv]` reports, per
session: wall (first->last message), executions (foreground suite runs plus
bgrun handoffs), blocked-on-output (sum of foreground suite waits), idle
(wall - blocked), the agent-chosen sleep/wait total, context chars, the
per-call (wait, size) ledger with the exact commands invoked, and whether
the LAST assistant text carried the upstream/downstream diagnostic markers.
Then a median+range row per arm where at least two sessions exist -- never a
single aggregate that hides the spread. Unknown entry types and junk lines
are skipped, so a crashed or partial transcript still measures.

Cross-checked against /tmp/arm_profile.py: span, calls, executions, blocked
and context agree exactly on all 17 captured sessions. The tool's idle is a
deliberate redefinition (wall - blocked); the profilers' "agent-chosen idle"
is preserved as its own sleep_s column so both readings stay available.
The comparison fixture flattered bgrun in two ways and offered no regime in
which bgrun should lose.

- DUMMY_LINES_PER_TEST [2]: the hardcoded 8 lines/test is a verbose/CI-log
  density, several times a normal runner's ~1.5 lines/test. It inflated output
  volume — one axis of the comparison — without saying so. The report now
  states the density used, and the header records that every earlier
  measurement of this fixture ran at 8.
- DUMMY_FAIL_FAST [0]: every run was long, so a background job's premise (the
  run outlives the turn) always held. With 1 the suite collapses at the planted
  failure — later tests short-circuit with one skipped line — so wall-clock
  ends there. That is the regime where bgrun should not be expected to win.
- The planted failure now prints a full assertion-failure block (assertion,
  expected vs received, stack trace naming the generated file and line) instead
  of a bare assert, so failing-run output volume is realistic. The exact marker
  strings DIAGNOSTIC_MARKER_UPSTREAM/DOWNSTREAM are unchanged.
- Output now has a real runner's shape: a banner per file and a summary block
  at the end, whose counters are read from the run (the runner picks its own
  file order) rather than assumed from the knobs.
- Header tunables section rewritten to name the axis each knob isolates: time,
  volume, failure position, fail-fast, plus the output directory.

Kept unchanged: the shared globalThis chain, symlink canonicalisation, the
fatal in-repo guard, the planted-failure validation, and the part-NN.test.ts
layout.
@lloydsk

lloydsk commented Sep 26, 2026

Copy link
Copy Markdown
Collaborator Author

Superseded: the fixture generator, the measurement instrument and the write-up are now one branch, bench/dogfooding (7 commits). No content is lost — this PR described one third of a deliverable.

@lloydsk lloydsk closed this Sep 26, 2026
@lloydsk
lloydsk deleted the docs/dogfooding-benchmarks branch September 26, 2026 02:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant