A stand-in gh in the test image: the pull request steps played, recorded and measured - #91
Merged
Merged
Conversation
PierreMardon
force-pushed
the
feat/stand-in-gh
branch
from
October 5, 2026 10:01
ef2040b to
aa161fe
Compare
…ded and measured The image of the end to end tests and of the evaluations has no `gh`, and a project there pushes to a bare repository on disk. So no session ever opened a draft pull request, refreshed its description or read one, every hand over said that `gh` was missing, and the judge of the evaluations scored that as noise, over three campaigns (#75). What the README promises of the pull request had never run (#8). `tests/e2e/toy/gh_stand_in.py` is a stand-in for `gh`, standard library only, copied in the image as `/opt/gh-stand-in/gh`, first on the `PATH`. It answers the calls the prompts make as `gh` 2.78.0 does when its output is not a terminal: `pr create`, `pr edit` with `--body`, `--body-file <file>` and `--body-file -`, `pr list --json`, `pr view`, `pr ready`, `repo view` and `auth status`. A second opening for a branch is refused with the address of the first, an edit on a branch that has no pull request fails with "no pull requests found for branch", and an opening is refused while the remote does not hold the last commit of the branch. Every other call gets an error that says the stand-in does not play it. It keeps the pull requests and one line per call, with its arguments, its standard input, the description given, its exit code and what it printed, in the git directory of the project, under `gh-stand-in/`. Every repository has a git directory, with or without a remote; it is one project's alone, where several run at the same time in one container; and it is out of the work tree, so nothing shows in `git status` and no session meets it. A call that asks to mark a pull request ready is recorded as such, since the chain must never make one. - `tests/e2e/Dockerfile`: the stand-in copied, its folder first on the `PATH`. - `tests/e2e/toy/toy.py`: a prepared state with a drafted plan holds the draft pull request `/surface-plan` opens at the first `plan-drafted`, with the description of that moment and no call recorded, so that `/surface-execute` meets what a developer with `gh` has. With none, every stop would take the path of a branch without a pull request, the only one the runs ever took. - `tests/e2e/test_toy.py`: the nominal path holds that the stop refreshed the description to what `surface-status pr-body` prints, by an edit and no second opening, that the pull request is still a draft, and that no call marked it ready. - `evals/surface_evals/runner.py`: each stop of a run keeps the pull request of the branch then, a draft or not, its description current or not. - `evals/surface_evals/measure.py` and `report.py`: five measures. `pr_draft_at_hand_over`: at every hand over, from the first, the branch has a pull request and it is a draft. `pr_described`: at the stops of planning, its description is the one `surface-status pr-body` prints then. `pr_refreshed`: the same at the stops of the loop. `pr_marked_ready`, which must stay false. `gh_not_played`: the calls the stand-in could not answer as `gh` would. The description is read where a stop leaves the plan to the developer, since that is where the prompts push and refresh it. A run kept before has none of the five. - A run that stops at the hand over keeps what planning gives: the draft, its description, and the calls not played, which say how far to trust the run as `contaminated` does. It holds neither `pr_refreshed` nor `pr_marked_ready`: never marked ready is a promise to the end of the chain, and a false would claim what the loop never played. A pull request that planning marked ready still shows, since it is no draft at the hand over. - `tests/gh/`: the unit tests of the stand-in, run as a script in temporary repositories, with no container and no session. `tests/evals/`: the stops and the measures. - `tests/e2e/README.md`, `evals/README.md` and `ARCHITECTURE.md`: what the stand-in answers, where it records and why, and what it does not imitate: GitHub itself, the remote, a terminal, and every call outside those above. Nothing under `skills/`, `agents/` or `install.py` changes: the stand-in belongs to the test image. Not measured: no session was played with it. The image builds, the stand-in answers in it, and the tests that start no session pass there. Refs #75, #8
PierreMardon
force-pushed
the
feat/stand-in-gh
branch
from
October 5, 2026 10:15
aa161fe to
7c070e7
Compare
This was referenced Oct 5, 2026
PierreMardon
added a commit
that referenced
this pull request
Oct 5, 2026
Kept for the next campaign to be compared with: `evals/reports/2026-10-05-d139a1ca0dbe/`. It measures the chain on `main` after the two fixes of the day (#87, #88), with the harness of #89, #90 and #91, on the four cases that no campaign had played again, or once: `04-suspension`, `05-fine-cap` and `06-borrow-limit`, one run played whole and one stopped at the hand over each, and `02-reservations`, one run stopped at the hand over. 108 sessions, 14.23 USD at list price. `report.md` and `summary.json` are as `report --against <summary> --keep` wrote them, and nothing else changes. The summary it is compared with is the one of the first campaign, computed again from its raw data by the harness as it stands, with the count of its six blueprints: no measure the kept summary of that campaign holds differs in it. The first run of `06-borrow-limit` was played in another folder, before the stand-in `gh` was in the image, to check the option that stops a run at the hand over: its sessions and its cost were added to the ledger of the campaign. What it shows: - The three runs played whole reached `conformant` and passed all of their hidden acceptance tests. - The pull request steps are played for the first time: a draft at every hand over, its description refreshed at conformity, never marked ready. In 3 runs of 7 the draft is opened with a line of attribution under the description the state script prints, which the next refresh removes. - The count of the judge follows the fix of the blueprint (#83) on these cases: the facts said twice in prose fall from 19 to 9, from 13 to 4 and from 8 to 1.5 on `02`, `04` and `05`, and the details of implementation from 5 to 11 a page to 1 or 2, while `blueprint.no-padding` is 2 on every run of both campaigns. - On `06-borrow-limit` the interview asks 2 and 3 questions where it asked 1, and its developer hands three of them back: `interview.right-number` falls from 5 to 3.5. - Planning costs 0.94 to 2.20 USD a run here, and a run played whole 1.53 to 1.88. The conclusions, issue by issue, are in the issues #77 lists.
This was referenced Oct 5, 2026
Open
PierreMardon
added a commit
that referenced
this pull request
Oct 5, 2026
… and nothing of the session's In 3 of the 6 runs of the fourth campaign of the evaluations that had a `gh` to call, the draft pull request was opened with the text `surface-status pr-body` prints followed by a line of attribution, where the prompt says "and nothing else" (#93). The stand-in `gh` of #91 let it be seen for the first time. The two commands did not ask for the same thing. `/surface-plan` named `gh pr create --draft` "with [...] that description" and `gh pr edit --body`, and left the session to carry the text: every session wrote it to a file of its own first, and three added a footer to it. `/surface-execute` gives the output of the script to `gh pr edit --body-file -` on its standard input, and no refresh of the three runs played whole carries one. `/surface-plan` now does the same: the output of the script is given to `gh` on its standard input, with `gh pr create --draft --body-file -` when the branch has no pull request and `gh pr edit --body-file -` when it has one. Measured on `05-fine-cap`, where the footer was written in two runs of two: three runs played to the hand over, three drafts opened from the standard input with the description alone, and at every hand over the description is the one the script prints. 19 sessions, 4.94 USD.
PierreMardon
added a commit
that referenced
this pull request
Oct 5, 2026
… and nothing of the session's (#95) In 3 of the 6 runs of the fourth campaign of the evaluations that had a `gh` to call, the draft pull request was opened with the text `surface-status pr-body` prints followed by a line of attribution, where the prompt says "and nothing else" (#93). The stand-in `gh` of #91 let it be seen for the first time. The two commands did not ask for the same thing. `/surface-plan` named `gh pr create --draft` "with [...] that description" and `gh pr edit --body`, and left the session to carry the text: every session wrote it to a file of its own first, and three added a footer to it. `/surface-execute` gives the output of the script to `gh pr edit --body-file -` on its standard input, and no refresh of the three runs played whole carries one. `/surface-plan` now does the same: the output of the script is given to `gh` on its standard input, with `gh pr create --draft --body-file -` when the branch has no pull request and `gh pr edit --body-file -` when it has one. Measured on `05-fine-cap`, where the footer was written in two runs of two: three runs played to the hand over, three drafts opened from the standard input with the description alone, and at every hand over the description is the one the script prints. 19 sessions, 4.94 USD.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes
The image of the end to end tests and of the evaluations holds a stand-in for
gh, first on itsPATH. A session there plays the pull request steps of the chain, the stand-in records every call, and the harness asserts and measures what was asked. This is what #75 asks for the evaluations and #8 for the end to end tests: one stand-in serves both.tests/e2e/toy/gh_stand_in.py, standard library only, copied in the image as/opt/gh-stand-in/gh. It answers, asgh2.78.0 does when its output is not a terminal:pr create,pr edit(--body,--body-file <file>,--body-file -),pr list --json,pr view,pr ready,repo view,auth status. A second opening for a branch is refused with the address of the first, an edit on a branch with no pull request fails withno pull requests found for branch "<branch>", an opening is refused while the remote does not hold the last commit of the branch. Every other call gets an error that says the stand-in does not play it, and is recorded withplayedfalse..git/gh-stand-in/. A call that asks to mark a pull request ready is recorded as such (marks_ready)./surface-planopens at the firstplan-drafted, with the description of that moment, and no call recorded for it.surface-status pr-bodyprints, by an edit and no second opening, that the pull request is still a draft, and that no call marked it ready.pr_draft_at_hand_over: at every hand over, from the first, the branch has a pull request and it is a draft.pr_described: at the stops of planning, its description is the onesurface-status pr-bodyprints then.pr_refreshed: the same at the stops of the loop.pr_marked_ready, which must stay false.gh_not_played: the calls the stand-in could not answer asghwould.Why
No session of the tests ever opened a draft pull request, refreshed its description or read one. Every hand over said that
ghwas missing, at up to four stops of a run, and the judge counted it as noise: the scores of the messages held what the chain says badly, what it says because of the container, and what the container hides, all at once (#75). On the path most developers live, withgh, none of those sentences is written, and that path had never run (#8).Decisions
--jobsplays several at the same time in one container. And it is out of the work tree: nothing shows ingit status, nothing can be committed, and a session that explores the code does not meet it./surface-planwould have left, and not none. With none,/surface-executewould find no pull request at any stop and take the path of a branch that has none, which is the only path the runs ever took and the one A hand back at conformity that fits a branch with no pull request #78 fixed. A state past the approval keeps the description of the draft, since the loop it prepares has not stopped yet.surface-status pr-bodyprints at that moment. The description counts at the stops that leave the plan to the developer, awaiting approval, blocked, with a plan change proposed or conformant, where the prompts push and refresh. A session that ends on a question while it drafts a revision has pushed nothing, and one killed at the timeout reached no stop.corpus --stop-at-hand-over) keeps what planning gives:pr_draft_at_hand_over,pr_described, andgh_not_played, which says how far to trust the run ascontaminateddoes.pr_refreshedandpr_marked_readyare among the measures of the execution, empty for such a run. Never marked ready is a promise to the end of the chain, and the hand back at conformity, where a chain would break it, is not reached: a false would claim what was not played, and would sit with the runs played whole. A pull request that planning marked ready still shows, since it is no draft at the hand over.gh_not_playedthen says how often a run met its limits, so the list of calls can grow from what sessions really ask.Files
tests/e2e/toy/gh_stand_in.py(new),tests/e2e/Dockerfile,tests/e2e/toy/toy.py,tests/e2e/test_toy.py,tests/e2e/README.mdevals/surface_evals/runner.py,measure.py,report.py,evals/README.mdtests/gh/test_gh_stand_in.py(new): the stand-in run as a script in temporary repositories, in the gates, with no container and no sessiontests/evals/support.py,test_runner.py,test_measure.py,test_report.pyARCHITECTURE.md: the codemap, the tests, the evaluationsNothing under
skills/,agents/orinstall.pychanges: the stand-in belongs to the test image only. ADR 0025 holds as it is.What the stand-in does not imitate
GitHub itself: no network, no CI, no check, no review, no comment, nothing merges or closes. The remote: it takes
origin, wherever it points, for the GitHub repository. A terminal: it never prompts. Every other call and flag:gh api,gh pr merge,gh pr checks,gh pr status,--fill,--web,--repo,--template, and a--jqthat is more than a path. The two refusals that are GitHub's own answers to an opening, a branch with nothing to merge and a head the remote does not hold, are worded as remembered and were not checked against GitHub. The whole list is intests/e2e/README.md.Not measured
No session was played with the stand-in: nothing here says yet that a session opens its draft through it, nor that the hand overs stop speaking of
gh. What was run: every gate; the image built under another tag; the stand-in played by hand in that image on a toy project, with the twogh pr listcalls of the branch step, an opening, the refresh of a stop on its standard input, and agh pr ready; and the two end to end tests that start no session.To validate, the opening first, on planning alone, then the refresh by the loop:
The first must leave, in
.evals/stand-in/runs/01-overdue-list/run-01/, arun.jsonwhose stops onawaiting-approvalhold"pull": {"draft": true, "described": true}, and awork/lending/.git/gh-stand-in/calls.jsonlwith apr create --draftthat exits 0 and nopr ready. Its final messages must not speak of a missinggh. The report must showpr_draft_at_hand_over1,pr_described1 andgh_not_played0, withpr_refreshedandpr_marked_readyempty. Whengh_not_playedis not 0, the lines ofcalls.jsonlwith"played": falsename the calls a session made and the stand-in does not answer yet. The end to end scenario must pass with its new assertions: the stop of the loop refreshed the description. A run played whole, the samecorpuswithout--stop-at-hand-over, must then showpr_refreshed1 andpr_marked_ready0.A campaign kept before this one was played with no
gh: against it, a move of the scores of the messages may come from the stand-in, which the hash of the chain does not show.Refs #75, #8