Skip to content

feat(bench): commit the session profiler every measurement came from - #31

Closed
lloydsk wants to merge 1 commit into
mainfrom
measure-sessions
Closed

lloydsk wants to merge 1 commit into
mainfrom
measure-sessions

Conversation

@lloydsk

@lloydsk lloydsk commented Sep 26, 2026

Copy link
Copy Markdown
Collaborator

Why

Every number in docs/dogfooding.md was produced by two throwaway Python profilers in /tmp. A benchmark whose measuring instrument is uncommitted and unversioned is not reproducible, and the method section of that doc now tells readers to re-derive the tables — this is the tool it points at.

What

bun scripts/measure-sessions.ts <dir-or-glob…> [--csv] — one row per session plus medians/ranges per arm, dependency-free Bun/TypeScript.

Per session: label (from the directory), inferred arm (bgrun when the transcript hands off to a job, else vanilla), executions (foreground suite runs and bgrun handoffs, reported separately), blocked-on-output, wall, context characters, per-call (wait, size, command), and whether the failure diagnostic reached context (DIAGNOSTIC_MARKER_UPSTREAM / _DOWNSTREAM) — for a red fixture, that is whether the detail arrived at all.

Two definitional notes, both deliberate:

  • idle_s is wall − blocked (time in the session spent not waiting on the suite), while the profilers' reading — agent-chosen sleeps/waits — is kept separately as sleep_s, so the old numbers stay cross-checkable instead of quietly changing meaning.
  • executions counts execution, not mention: cat long-job.sh is not a run, and a bun test that only appears in prose is not either.

Verification

  • Cross-checked field-by-field against the Python profiler over 17 sessions: 0 mismatches on span, calls, executions, blocked-on-output and context chars; sleep_s equals the profiler's agent-chosen idle on all 17.
  • bun test scripts/measure-sessions.test.ts — 10 tests, synthetic transcripts for a blocking run, a multi-call session, a no-execution session, and bgrun detection.
  • tsc --noEmit clean.

Note on CI

package.json's test script becomes bun test extension/index.test.ts scripts/measure-sessions.test.ts, so the instrument is covered by the same job as the extension. Parsing is what turns transcripts into the doc's numbers, so it gets tests; a silent parser regression would rewrite the benchmark.

Port the throwaway arm profilers into a committed, dependency-free CLI so
the numbers in docs/dogfooding.md can be re-derived, not just trusted.

`bun scripts/measure-sessions.ts <dir-or-glob...> [--csv]` reports, per
session: wall (first->last message), executions (foreground suite runs plus
bgrun handoffs), blocked-on-output (sum of foreground suite waits), idle
(wall - blocked), the agent-chosen sleep/wait total, context chars, the
per-call (wait, size) ledger with the exact commands invoked, and whether
the LAST assistant text carried the upstream/downstream diagnostic markers.
Then a median+range row per arm where at least two sessions exist -- never a
single aggregate that hides the spread. Unknown entry types and junk lines
are skipped, so a crashed or partial transcript still measures.

Cross-checked against /tmp/arm_profile.py: span, calls, executions, blocked
and context agree exactly on all 17 captured sessions. The tool's idle is a
deliberate redefinition (wall - blocked); the profilers' "agent-chosen idle"
is preserved as its own sleep_s column so both readings stay available.
@github-actions

Copy link
Copy Markdown
Contributor

CI report

Check Result
tsc --noEmit success
tests success
npm pack --dry-run success (7 files in tarball)

Ref: 2c68c5b4d282ae8daea1984884a33cf7c06bd394

@lloydsk

lloydsk commented Sep 26, 2026

Copy link
Copy Markdown
Collaborator Author

Superseded: the fixture generator, the measurement instrument and the write-up are now one branch, bench/dogfooding (7 commits). No content is lost — this PR described one third of a deliverable.

@lloydsk lloydsk closed this Sep 26, 2026
@lloydsk
lloydsk deleted the measure-sessions branch September 26, 2026 02:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant