Conversation
Port the throwaway arm profilers into a committed, dependency-free CLI so the numbers in docs/dogfooding.md can be re-derived, not just trusted. `bun scripts/measure-sessions.ts <dir-or-glob...> [--csv]` reports, per session: wall (first->last message), executions (foreground suite runs plus bgrun handoffs), blocked-on-output (sum of foreground suite waits), idle (wall - blocked), the agent-chosen sleep/wait total, context chars, the per-call (wait, size) ledger with the exact commands invoked, and whether the LAST assistant text carried the upstream/downstream diagnostic markers. Then a median+range row per arm where at least two sessions exist -- never a single aggregate that hides the spread. Unknown entry types and junk lines are skipped, so a crashed or partial transcript still measures. Cross-checked against /tmp/arm_profile.py: span, calls, executions, blocked and context agree exactly on all 17 captured sessions. The tool's idle is a deliberate redefinition (wall - blocked); the profilers' "agent-chosen idle" is preserved as its own sleep_s column so both readings stay available.
Contributor
CI report
Ref: |
Collaborator
Author
|
Superseded: the fixture generator, the measurement instrument and the write-up are now one branch, |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Every number in
docs/dogfooding.mdwas produced by two throwaway Python profilers in/tmp. A benchmark whose measuring instrument is uncommitted and unversioned is not reproducible, and the method section of that doc now tells readers to re-derive the tables — this is the tool it points at.What
bun scripts/measure-sessions.ts <dir-or-glob…> [--csv]— one row per session plus medians/ranges per arm, dependency-free Bun/TypeScript.Per session: label (from the directory), inferred arm (
bgrunwhen the transcript hands off to a job, elsevanilla), executions (foreground suite runs and bgrun handoffs, reported separately), blocked-on-output, wall, context characters, per-call(wait, size, command), and whether the failure diagnostic reached context (DIAGNOSTIC_MARKER_UPSTREAM/_DOWNSTREAM) — for a red fixture, that is whether the detail arrived at all.Two definitional notes, both deliberate:
idle_siswall − blocked(time in the session spent not waiting on the suite), while the profilers' reading — agent-chosen sleeps/waits — is kept separately assleep_s, so the old numbers stay cross-checkable instead of quietly changing meaning.executionscounts execution, not mention:cat long-job.shis not a run, and abun testthat only appears in prose is not either.Verification
sleep_sequals the profiler's agent-chosen idle on all 17.bun test scripts/measure-sessions.test.ts— 10 tests, synthetic transcripts for a blocking run, a multi-call session, a no-execution session, and bgrun detection.tsc --noEmitclean.Note on CI
package.json'stestscript becomesbun test extension/index.test.ts scripts/measure-sessions.test.ts, so the instrument is covered by the same job as the extension. Parsing is what turns transcripts into the doc's numbers, so it gets tests; a silent parser regression would rewrite the benchmark.