Conversation
…tures and method Adds the measured comparison to the dogfooding doc: how a foreground `bash` run and a `bgrun` job compare on the same task, in live sessions, with per-run numbers and the fixtures they come from. - Method: three fixtures (the repo's own suite; a generated long, noisy, failing suite; and a short suite failing early), an identical prompt per arm except the mechanism clause, live sessions (a one-shot `pi -p` exits before the wake lands), and metrics read from the session transcript — executions, blocked-on-output, wall, context chars, and whether the diagnostic reached context. - The mechanism-determined differences are stable (bgrun never blocks the session and runs the suite once); the agent-determined ones are not. On the long fixture, vanilla ran the suite twice for 359.9s of blocked time and 10,016 chars, against bgrun's single execution, 0.0s, and 8,049 chars — 28% of which was one pre-wake poll the guidance now discourages. - It also records where the naive path converges on bgrun's technique by hand (redirecting a long run to a file, then grepping it), so the doc does not overclaim. Fixture: scripts/make-dummy-suite.ts (bun run bench:make-suite).
Contributor
CI report
Ref: |
…grun does not help Three additions, prompted by the obvious hazard: a fixture we can tune is a fixture we can tune in the tool's favour. - **How these numbers are produced** — the protocol as one table (arms, prompt template, model, live sessions, n=3 reported as median and range, metric definitions, and the command that re-derives every table from the transcripts), plus the `vanilla-hinted` control arm: the agent hand-rolling redirect-and-grep, which is the strongest baseline in the set. - **Pre-registered expectations** — nine cells with the predicted winner written down before the runs, including the ones where bgrun is expected to lose (R1 green, R2 tail-shaped failure, R5 fail-fast) and the one where our own story may be wrong (D1: if the advantage is reliability-driven it should grow with a weaker model). C1-C4 are dialogues — overlap, three jobs at once, a mid-run restart, asking where a run is — because no generated fixture can express them. - **The fixture's design and why each property is under test** — the shared chain that makes duration scheduler-independent, the 65% failure position, the realistic failure block, and the knobs for the negative cells. Records that the earlier long-suite run used the verbose 8-lines-per-test density while the realistic default is 2, and that the difference is an axis rather than an oversight. - **Where bgrun helps, and where it does not** — the strongest claim the data supports is the removal of the worst case, not a better best case (vanilla's `tail -100` run beat every bgrun run on cost). Names the papercuts the measuring found: the 28% pre-wake poll (#27 fixed), window-relative bggrep line numbers (#25 open), and re-paid log reads before delta `bgtail` (#4).
…he context The standard protocol's first two vanilla cells, n=3 each, with the runs kept rather than the aggregate: - R1 (green suite, one-shot, same prompt and command): 4,894 / 5,018 / 21,787 context chars. Two sessions bounded the output, one read the whole 241-line run — a 4.4x swing from the agent's choice of window alone. bgrun has no equivalent mode, which makes this a claim about variance rather than about the median. - R2 (red suite, failure 78-90 lines from the end): executions 1, 3, 1; blocked 21.8s, 66.2s, 21.9s; context 3,641 / 23,034 / 4,719. Across all six vanilla runs of this command the median is one execution and ~4.7 KB and the worst is four executions and 66.2s blocked. The earlier three-run sample happened to be a worse draw, so the doc now carries both samples and says so, instead of the rounder number. Also worth recording: in all six runs the tail did reach the diagnostic — this fixture punishes a bad window with another run, not with a missed failure.
Half the matrix cannot be scripted — the bgrun wake needs a live session — so the protocol needs an executable run sheet: fresh session, prompt verbatim, then no nudges (the wake is the mechanism under test), then profile the session directory. Also names the two temptations this method has to resist: re-running a cell whose number looks wrong (sampling on the outcome), and comparing a live bgrun cell with a one-shot vanilla cell without recording the difference in the table.
…e row needs re-running Records what the realism pass *verified* rather than what it intended: - 694 output lines at the standard density, 2,494 at the verbose one. - The 65% failure position now has a mechanism instead of a measurement: bun 1.3.6 discovers these files in a non-lexical order (2,3,1,8,9,10,5,4,6,7), so part-05 runs seventh of ten while DUMMY_FAIL_FILE=5 reads as 45%. Fail-fast therefore stops at 65%, not the 45% a lexical order would give. - The failure block is ~19 lines, with a real +actual/-expected diff and a stack naming the generated file and line. - The long-suite row below is relabelled as the verbose-density variant: its numbers came from ~2,760 lines of output, where the standard fixture emits 694. The failure's distance from the end *is* the mechanism, so those numbers cannot sit next to R1/R2's unqualified. The pre-registration table now says so at the cell, and the section says the row is pending re-measurement.
Port the throwaway arm profilers into a committed, dependency-free CLI so the numbers in docs/dogfooding.md can be re-derived, not just trusted. `bun scripts/measure-sessions.ts <dir-or-glob...> [--csv]` reports, per session: wall (first->last message), executions (foreground suite runs plus bgrun handoffs), blocked-on-output (sum of foreground suite waits), idle (wall - blocked), the agent-chosen sleep/wait total, context chars, the per-call (wait, size) ledger with the exact commands invoked, and whether the LAST assistant text carried the upstream/downstream diagnostic markers. Then a median+range row per arm where at least two sessions exist -- never a single aggregate that hides the spread. Unknown entry types and junk lines are skipped, so a crashed or partial transcript still measures. Cross-checked against /tmp/arm_profile.py: span, calls, executions, blocked and context agree exactly on all 17 captured sessions. The tool's idle is a deliberate redefinition (wall - blocked); the profilers' "agent-chosen idle" is preserved as its own sleep_s column so both readings stay available.
The comparison fixture flattered bgrun in two ways and offered no regime in which bgrun should lose. - DUMMY_LINES_PER_TEST [2]: the hardcoded 8 lines/test is a verbose/CI-log density, several times a normal runner's ~1.5 lines/test. It inflated output volume — one axis of the comparison — without saying so. The report now states the density used, and the header records that every earlier measurement of this fixture ran at 8. - DUMMY_FAIL_FAST [0]: every run was long, so a background job's premise (the run outlives the turn) always held. With 1 the suite collapses at the planted failure — later tests short-circuit with one skipped line — so wall-clock ends there. That is the regime where bgrun should not be expected to win. - The planted failure now prints a full assertion-failure block (assertion, expected vs received, stack trace naming the generated file and line) instead of a bare assert, so failing-run output volume is realistic. The exact marker strings DIAGNOSTIC_MARKER_UPSTREAM/DOWNSTREAM are unchanged. - Output now has a real runner's shape: a banner per file and a summary block at the end, whose counters are read from the run (the runner picks its own file order) rather than assumed from the knobs. - Header tunables section rewritten to name the axis each knob isolates: time, volume, failure position, fail-fast, plus the output directory. Kept unchanged: the shared globalThis chain, symlink canonicalisation, the fatal in-repo guard, the planted-failure validation, and the part-NN.test.ts layout.
Collaborator
Author
|
Superseded: the fixture generator, the measurement instrument and the write-up are now one branch, |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.