Skip to content

A stand-in gh in the test image: the pull request steps played, recorded and measured - #91

Merged
PierreMardon merged 1 commit into
mainfrom
feat/stand-in-gh
Oct 5, 2026
Merged

PierreMardon merged 1 commit into
mainfrom
feat/stand-in-gh

Conversation

@PierreMardon

@PierreMardon PierreMardon commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

What changes

The image of the end to end tests and of the evaluations holds a stand-in for gh, first on its PATH. A session there plays the pull request steps of the chain, the stand-in records every call, and the harness asserts and measures what was asked. This is what #75 asks for the evaluations and #8 for the end to end tests: one stand-in serves both.

  • The stand-in, tests/e2e/toy/gh_stand_in.py, standard library only, copied in the image as /opt/gh-stand-in/gh. It answers, as gh 2.78.0 does when its output is not a terminal: pr create, pr edit (--body, --body-file <file>, --body-file -), pr list --json, pr view, pr ready, repo view, auth status. A second opening for a branch is refused with the address of the first, an edit on a branch with no pull request fails with no pull requests found for branch "<branch>", an opening is refused while the remote does not hold the last commit of the branch. Every other call gets an error that says the stand-in does not play it, and is recorded with played false.
  • The record: the pull requests and one line per call, with its arguments, its standard input, the description given, its exit code and what it printed, in the git directory of the project, .git/gh-stand-in/. A call that asks to mark a pull request ready is recorded as such (marks_ready).
  • The prepared states of the toy with a drafted plan hold the draft pull request /surface-plan opens at the first plan-drafted, with the description of that moment, and no call recorded for it.
  • The nominal scenario of the end to end tests holds that the stop refreshed the description to what surface-status pr-body prints, by an edit and no second opening, that the pull request is still a draft, and that no call marked it ready.
  • The evaluations: each stop of a run keeps the pull request of the branch then, and five measures follow. pr_draft_at_hand_over: at every hand over, from the first, the branch has a pull request and it is a draft. pr_described: at the stops of planning, its description is the one surface-status pr-body prints then. pr_refreshed: the same at the stops of the loop. pr_marked_ready, which must stay false. gh_not_played: the calls the stand-in could not answer as gh would.

Why

No session of the tests ever opened a draft pull request, refreshed its description or read one. Every hand over said that gh was missing, at up to four stops of a run, and the judge counted it as noise: the scores of the messages held what the chain says badly, what it says because of the container, and what the container hides, all at once (#75). On the path most developers live, with gh, none of those sentences is written, and that path had never run (#8).

Decisions

  • Where the record lives: the git directory of the project, and not a folder next to its bare remote. Every repository has one, with or without a remote, so a call refused for want of a remote is recorded too. It is one project's alone, where --jobs plays several at the same time in one container. And it is out of the work tree: nothing shows in git status, nothing can be committed, and a session that explores the code does not meet it.
  • A prepared state holds an open draft, as a real /surface-plan would have left, and not none. With none, /surface-execute would find no pull request at any stop and take the path of a branch that has none, which is the only path the runs ever took and the one A hand back at conformity that fits a branch with no pull request #78 fixed. A state past the approval keeps the description of the draft, since the loop it prepares has not stopped yet.
  • The pull request is read at each stop, by the harness, and not from timestamps afterwards: only then can the description be compared with what surface-status pr-body prints at that moment. The description counts at the stops that leave the plan to the developer, awaiting approval, blocked, with a plan change proposed or conformant, where the prompts push and refresh. A session that ends on a question while it drafts a revision has pushed nothing, and one killed at the timeout reached no stop.
  • A run that stops at the hand over (corpus --stop-at-hand-over) keeps what planning gives: pr_draft_at_hand_over, pr_described, and gh_not_played, which says how far to trust the run as contaminated does. pr_refreshed and pr_marked_ready are among the measures of the execution, empty for such a run. Never marked ready is a promise to the end of the chain, and the hand back at conformity, where a chain would break it, is not reached: a false would claim what was not played, and would sit with the runs played whole. A pull request that planning marked ready still shows, since it is no draft at the hand over.
  • What the stand-in does not know, it refuses and records, and never guesses. gh_not_played then says how often a run met its limits, so the list of calls can grow from what sessions really ask.

Files

  • tests/e2e/toy/gh_stand_in.py (new), tests/e2e/Dockerfile, tests/e2e/toy/toy.py, tests/e2e/test_toy.py, tests/e2e/README.md
  • evals/surface_evals/runner.py, measure.py, report.py, evals/README.md
  • tests/gh/test_gh_stand_in.py (new): the stand-in run as a script in temporary repositories, in the gates, with no container and no session
  • tests/evals/support.py, test_runner.py, test_measure.py, test_report.py
  • ARCHITECTURE.md: the codemap, the tests, the evaluations

Nothing under skills/, agents/ or install.py changes: the stand-in belongs to the test image only. ADR 0025 holds as it is.

What the stand-in does not imitate

GitHub itself: no network, no CI, no check, no review, no comment, nothing merges or closes. The remote: it takes origin, wherever it points, for the GitHub repository. A terminal: it never prompts. Every other call and flag: gh api, gh pr merge, gh pr checks, gh pr status, --fill, --web, --repo, --template, and a --jq that is more than a path. The two refusals that are GitHub's own answers to an opening, a branch with nothing to merge and a head the remote does not hold, are worded as remembered and were not checked against GitHub. The whole list is in tests/e2e/README.md.

Not measured

No session was played with the stand-in: nothing here says yet that a session opens its draft through it, nor that the hand overs stop speaking of gh. What was run: every gate; the image built under another tag; the stand-in played by hand in that image on a toy project, with the two gh pr list calls of the branch step, an opening, the refresh of a stop on its standard input, and a gh pr ready; and the two end to end tests that start no session.

To validate, the opening first, on planning alone, then the refresh by the loop:

EVALS_OUT=.evals/stand-in scripts/gate.sh evals corpus --case 01-overdue-list --stop-at-hand-over
python3 evals/run.py --out .evals/stand-in report
scripts/gate.sh e2e -k nominal

The first must leave, in .evals/stand-in/runs/01-overdue-list/run-01/, a run.json whose stops on awaiting-approval hold "pull": {"draft": true, "described": true}, and a work/lending/.git/gh-stand-in/calls.jsonl with a pr create --draft that exits 0 and no pr ready. Its final messages must not speak of a missing gh. The report must show pr_draft_at_hand_over 1, pr_described 1 and gh_not_played 0, with pr_refreshed and pr_marked_ready empty. When gh_not_played is not 0, the lines of calls.jsonl with "played": false name the calls a session made and the stand-in does not answer yet. The end to end scenario must pass with its new assertions: the stop of the loop refreshed the description. A run played whole, the same corpus without --stop-at-hand-over, must then show pr_refreshed 1 and pr_marked_ready 0.

A campaign kept before this one was played with no gh: against it, a move of the scores of the messages may come from the stand-in, which the hash of the chain does not show.

Refs #75, #8

…ded and measured

The image of the end to end tests and of the evaluations has no `gh`, and a project there
pushes to a bare repository on disk. So no session ever opened a draft pull request,
refreshed its description or read one, every hand over said that `gh` was missing, and the
judge of the evaluations scored that as noise, over three campaigns (#75). What the README
promises of the pull request had never run (#8).

`tests/e2e/toy/gh_stand_in.py` is a stand-in for `gh`, standard library only, copied in the
image as `/opt/gh-stand-in/gh`, first on the `PATH`. It answers the calls the prompts make
as `gh` 2.78.0 does when its output is not a terminal: `pr create`, `pr edit` with `--body`,
`--body-file <file>` and `--body-file -`, `pr list --json`, `pr view`, `pr ready`,
`repo view` and `auth status`. A second opening for a branch is refused with the address
of the first, an edit on a branch that has no pull request fails with "no pull requests
found for branch", and an opening is refused while the remote does not hold the last
commit of the branch. Every other call gets an error that says the stand-in does not play
it.

It keeps the pull requests and one line per call, with its arguments, its standard input,
the description given, its exit code and what it printed, in the git directory of the
project, under `gh-stand-in/`. Every repository has a git directory, with or without a
remote; it is one project's alone, where several run at the same time in one container;
and it is out of the work tree, so nothing shows in `git status` and no session meets it.
A call that asks to mark a pull request ready is recorded as such, since the chain must
never make one.

- `tests/e2e/Dockerfile`: the stand-in copied, its folder first on the `PATH`.
- `tests/e2e/toy/toy.py`: a prepared state with a drafted plan holds the draft pull request
  `/surface-plan` opens at the first `plan-drafted`, with the description of that moment
  and no call recorded, so that `/surface-execute` meets what a developer with `gh` has.
  With none, every stop would take the path of a branch without a pull request, the only
  one the runs ever took.
- `tests/e2e/test_toy.py`: the nominal path holds that the stop refreshed the description
  to what `surface-status pr-body` prints, by an edit and no second opening, that the pull
  request is still a draft, and that no call marked it ready.
- `evals/surface_evals/runner.py`: each stop of a run keeps the pull request of the branch
  then, a draft or not, its description current or not.
- `evals/surface_evals/measure.py` and `report.py`: five measures. `pr_draft_at_hand_over`:
  at every hand over, from the first, the branch has a pull request and it is a draft.
  `pr_described`: at the stops of planning, its description is the one
  `surface-status pr-body` prints then. `pr_refreshed`: the same at the stops of the loop.
  `pr_marked_ready`, which must stay false. `gh_not_played`: the calls the stand-in could
  not answer as `gh` would. The description is read where a stop leaves the plan to the
  developer, since that is where the prompts push and refresh it. A run kept before has
  none of the five.
- A run that stops at the hand over keeps what planning gives: the draft, its description,
  and the calls not played, which say how far to trust the run as `contaminated` does. It
  holds neither `pr_refreshed` nor `pr_marked_ready`: never marked ready is a promise to
  the end of the chain, and a false would claim what the loop never played. A pull request
  that planning marked ready still shows, since it is no draft at the hand over.
- `tests/gh/`: the unit tests of the stand-in, run as a script in temporary repositories,
  with no container and no session. `tests/evals/`: the stops and the measures.
- `tests/e2e/README.md`, `evals/README.md` and `ARCHITECTURE.md`: what the stand-in
  answers, where it records and why, and what it does not imitate: GitHub itself, the
  remote, a terminal, and every call outside those above.

Nothing under `skills/`, `agents/` or `install.py` changes: the stand-in belongs to the
test image.

Not measured: no session was played with it. The image builds, the stand-in answers in
it, and the tests that start no session pass there.

Refs #75, #8
@PierreMardon
PierreMardon merged commit 7fcf947 into main Oct 5, 2026
3 checks passed
@PierreMardon
PierreMardon deleted the feat/stand-in-gh branch October 5, 2026 10:24
PierreMardon added a commit that referenced this pull request Oct 5, 2026
Kept for the next campaign to be compared with: `evals/reports/2026-10-05-d139a1ca0dbe/`.
It measures the chain on `main` after the two fixes of the day (#87, #88), with the harness
of #89, #90 and #91, on the four cases that no campaign had played again, or once:
`04-suspension`, `05-fine-cap` and `06-borrow-limit`, one run played whole and one stopped
at the hand over each, and `02-reservations`, one run stopped at the hand over. 108
sessions, 14.23 USD at list price.

`report.md` and `summary.json` are as `report --against <summary> --keep` wrote them, and
nothing else changes. The summary it is compared with is the one of the first campaign,
computed again from its raw data by the harness as it stands, with the count of its six
blueprints: no measure the kept summary of that campaign holds differs in it.

The first run of `06-borrow-limit` was played in another folder, before the stand-in `gh`
was in the image, to check the option that stops a run at the hand over: its sessions and
its cost were added to the ledger of the campaign.

What it shows:

- The three runs played whole reached `conformant` and passed all of their hidden
  acceptance tests.
- The pull request steps are played for the first time: a draft at every hand over, its
  description refreshed at conformity, never marked ready. In 3 runs of 7 the draft is
  opened with a line of attribution under the description the state script prints, which
  the next refresh removes.
- The count of the judge follows the fix of the blueprint (#83) on these cases: the facts
  said twice in prose fall from 19 to 9, from 13 to 4 and from 8 to 1.5 on `02`, `04` and
  `05`, and the details of implementation from 5 to 11 a page to 1 or 2, while
  `blueprint.no-padding` is 2 on every run of both campaigns.
- On `06-borrow-limit` the interview asks 2 and 3 questions where it asked 1, and its
  developer hands three of them back: `interview.right-number` falls from 5 to 3.5.
- Planning costs 0.94 to 2.20 USD a run here, and a run played whole 1.53 to 1.88.

The conclusions, issue by issue, are in the issues #77 lists.
PierreMardon added a commit that referenced this pull request Oct 5, 2026
… and nothing of the session's

In 3 of the 6 runs of the fourth campaign of the evaluations that had a `gh` to call, the
draft pull request was opened with the text `surface-status pr-body` prints followed by a
line of attribution, where the prompt says "and nothing else" (#93). The stand-in `gh` of
#91 let it be seen for the first time.

The two commands did not ask for the same thing. `/surface-plan` named `gh pr create --draft`
"with [...] that description" and `gh pr edit --body`, and left the session to carry the
text: every session wrote it to a file of its own first, and three added a footer to it.
`/surface-execute` gives the output of the script to `gh pr edit --body-file -` on its
standard input, and no refresh of the three runs played whole carries one.

`/surface-plan` now does the same: the output of the script is given to `gh` on its standard
input, with `gh pr create --draft --body-file -` when the branch has no pull request and
`gh pr edit --body-file -` when it has one.

Measured on `05-fine-cap`, where the footer was written in two runs of two: three runs
played to the hand over, three drafts opened from the standard input with the description
alone, and at every hand over the description is the one the script prints. 19 sessions,
4.94 USD.
PierreMardon added a commit that referenced this pull request Oct 5, 2026
… and nothing of the session's (#95)

In 3 of the 6 runs of the fourth campaign of the evaluations that had a `gh` to call, the
draft pull request was opened with the text `surface-status pr-body` prints followed by a
line of attribution, where the prompt says "and nothing else" (#93). The stand-in `gh` of
#91 let it be seen for the first time.

The two commands did not ask for the same thing. `/surface-plan` named `gh pr create --draft`
"with [...] that description" and `gh pr edit --body`, and left the session to carry the
text: every session wrote it to a file of its own first, and three added a footer to it.
`/surface-execute` gives the output of the script to `gh pr edit --body-file -` on its
standard input, and no refresh of the three runs played whole carries one.

`/surface-plan` now does the same: the output of the script is given to `gh` on its standard
input, with `gh pr create --draft --body-file -` when the branch has no pull request and
`gh pr edit --body-file -` when it has one.

Measured on `05-fine-cap`, where the footer was written in two runs of two: three runs
played to the hand over, three drafts opened from the standard input with the description
alone, and at every hand over the description is the one the script prints. 19 sessions,
4.94 USD.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant