Skip to content

Evaluations of the chain: whether it meets its goals, and what it is like to work with - #58

Merged
PierreMardon merged 2 commits into
mainfrom
feat/evaluations-of-the-chain
Oct 3, 2026
Merged

PierreMardon merged 2 commits into
mainfrom
feat/evaluations-of-the-chain

Conversation

@PierreMardon

@PierreMardon PierreMardon commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

The tests hold the state script, the wording of the prompts and, end to end, that a real session reaches the expected state on a toy with one behavior. None says whether the chain delivers what the README promises, nor whether its documents and messages serve the developer who works with it. evals/ answers both after a change of the prompts, on real sessions and with no human in them, and its first campaign is kept in this pull request.

What was done

  • A synthetic host and a corpus. evals/host/ is lending, a command line tool for a library: several components, state in files, two critical zones declared in its AGENTS.md. evals/cases/ holds six needs chosen for their shapes: one behavior, several flows, one flow across components, a trade-off, a critical zone touched, an ambiguous need. Each case keeps from the agents its acceptance tests, through the command line, and a reference implementation that proves them fair: the unit tests check that they fail on the host and pass on the reference.
  • A run plays the developer's side. It types the need, has a model answer each question from a written brief of what the developer knows, then read the blueprint handed over and send back what contradicts the brief, as an amendment. It gives the amendment of the case when it has one, replies with a sentence that agrees, which must approve nothing, and launches /surface-execute. It keeps the project, every stream, each blueprint handed over, and what was said at each stop.
  • Part 1, by a script, from the journal, the plan folder, the streams and the delivered code: the outcome, the passes, the acceptance, the critical zones named at approval and their files listed at conformity, the form of the blueprint, the plan drafted by the built-in agent, and whether a session read what it must not.
  • Faults put there on purpose, on the toy, whose documents and code are written out and so can be spoiled exactly. A fresh checker cross-checks a blueprint that hides a column, a critical zone, an irreversible effect, and a faithful one. /surface-execute is launched on a branch that holds a defect the tests do not see, a deviation that keeps the blueprint true, a commit against the contract, and work as planned, and is stopped at its first review. The toy gains three reviewable states for it.
  • Part 2, by a judge against evals/judge/rubric.md: the blueprint for a developer who must decide, the interview, what the agents say, following along. Each judgement has its score, its reason and the passage it rests on, and a passage found in no document is marked. The judge is checked without a human: a padded blueprint, one stripped of the rules the interview settled, an interview with questions already answered must each score below its original.
  • A report that compares. A measure is kept as its mean and its spread over the runs. Against an earlier campaign, a measure moved when its values lie wholly outside those of the campaign before. Nothing fails and there is no threshold. --keep copies the report under evals/reports/, named after the date and a hash of agents/ and skills/.
  • A ledger stops a command before a ceiling, in USD at list price or in points of the weekly gauge of a subscription, which the streams report.
  • scripts/gate.sh evals runs it in the container of the end to end tests. ADR 0036 records the decision; ARCHITECTURE.md, the contributing guide and evals/README.md, the operating guide, follow.

Decisions

On the questions #55 left open. The developer chose the host and who plays the developer; the others were proposed to them and not objected to.

  • The host: synthetic, written for the purpose. It has what the cases need and nothing else, costs nothing to build, never drifts, and raises no question of licence or of privacy.
  • The developer: a model that holds a written brief. Written answers alone cannot meet questions that change from run to run. A trial run added the reading of the blueprint: without it, a rule the plan guessed and the blueprint showed as an assumption was never corrected, and the acceptance tests blamed the chain for what a developer would have caught.
  • The rubric and the spoiled documents: written here, the rubric after the four questions of the issue. The padded blueprint is padded from a lean one of its own: the judge scored the toy's blueprint as low as its padded version, since that page already repeats itself.
  • The reports: streams and projects stay in the campaign folder, out of git; summary.json and report.md are committed when a campaign is kept. A regression is a measure that lies outside the spread of the campaign before, in the wrong direction: it is reported, never a failure.

The first campaign

Six cases, one run each, then the probes, the judge and its check: 101 sessions, 18.71 USD at list price, kept in evals/reports/2026-10-03-6bf7c06609a6/.

  • Goals. Every case reached conformant and passed all of its acceptance tests. No sentence approved a revision. Every plan was drafted by the built-in agent on opus and held the minimum. The critical zone of the case that touches one was named at approval and its file listed at conformity; the others said none. No session read the clone.
  • The interview. It asked from none to four questions. Twice it asked none where a rule was the developer's, the order of ties and how often a member is reminded: the plan guessed, the blueprint showed the guess, and the developer sent it back, three corrections in all.
  • The form of the blueprint. The frame held everywhere, no heading carried a number, and the three revised blueprints kept their cut. The body went from one section to seven, outside what the case expected three times out of six, and two blueprints drew a diagram where the case expected none.
  • The probes. The cross-check found each of the three hidden things. On the faithful blueprint it first found two omissions, and it was right: the toy's blueprint did not show the line ending its plan assumes. The toy is fixed here, and the probe then found nothing. The review found the defect, the deviation and the break, each in its class, and nothing on work as planned.
  • The judge. High on a blueprint that lets the developer decide, 4.7 of 5, and on stops that say why they stopped, 5. Low on padding, 2 on every blueprint: rules said in the criteria, again in the body, again in the sensitive zones, and details that are not the developer's. Low on the messages, 2 to 3: the vocabulary of the chain, and a hand back that says to mark ready a pull request that was never opened. Every passage it quoted stands in a document it was given, and the three spoiled documents scored below their originals.

Tests

scripts/gate.sh is green. tests/evals/ holds the harness with no session: the streams and the ledger, the form of a blueprint against the template the extractor is given, the measures on folders as a run leaves them, the rubric and the judge's answers, the report and what moved, the probes built on the toy, the command line, and a run played against the real state script with the sessions replaced by a script. For each case, the acceptance tests fail on the host and pass on its reference.

Risks

  • One run per case: the spread between runs is not measured yet, so the next campaign has little to be compared with. --runs is there for it.
  • The corpus is small and synthetic. It says nothing of a large code base.
  • The judge is a model. Its scores are a trend between two versions of the prompts, on the criteria its check clears, never a grade. Its check rests on three pairs.
  • The container has no gh: every hand over says so, which the judge counts as noise. That part of the low scores on the messages is the environment's.
  • A session can read the clone the container mounts, where the cases are. A run that does is marked, not prevented.
  • What the body of a blueprint should look like, per case, is this corpus's opinion: body_in_range and diagram_as_expected are to be read with the judge's reasons, not alone.

Closes #55

…like to work with

The tests hold the state script, the wording of the prompts and, end to end, that a real
session reaches the expected state on a toy with one behavior. None says whether the
chain delivers what the README promises, nor whether its documents and messages serve
the developer. `evals/` answers both after a change of the prompts, on real sessions and
with no human in them.

- A synthetic host, `lending`, with several components, state and two declared critical
  zones, and a corpus of six needs chosen for their shapes: one behavior, several flows,
  one flow across components, a trade-off, a critical zone touched, an ambiguous need.
  Each case keeps from the agents its acceptance tests, and a reference implementation
  that proves them fair.
- A run plays the developer's side: the need, an answer to each question by a model
  that holds a written brief, a reading of the blueprint whose corrections go back as
  amendments, a sentence that agrees, which must approve nothing, then the launch of
  `/surface-execute`.
- A script measures what a script can: the outcome, the passes, the acceptance, the
  critical zones named at approval and their files listed at conformity, the form of the
  blueprint. Faults put there on purpose, on the toy, show whether the cross-check finds
  what a blueprint hides and whether the review classifies a defect, a deviation and a
  break. The toy gains three reviewable states for it.
- A model judges the rest against a written rubric, each judgement with its reason and
  the passage it rests on, and is checked: a document spoiled in a known way must score
  below its original.
- The report keeps each measure as its spread over the runs and says what lies outside
  the spread of the campaign before. No threshold, and not in CI. A ledger stops a
  campaign before a ceiling, in USD or in points of the weekly gauge of a subscription.
- `scripts/gate.sh evals` runs it in the container of the end to end tests. ADR 0036
  records the decision, and `ARCHITECTURE.md` and the contributing guide follow.
- The report of the first campaign is kept under `evals/reports/`: six cases, the probes,
  the judge and its check.
- The toy's blueprint shows the line ending its plan assumes, and says what data the
  feature adds: a fresh checker found both missing on the blueprint held to be faithful.
- The padded blueprint of the judge's check is padded from a lean one of its own: the
  judge scored the toy's blueprint as low as its padded version.
- A blueprint that says the plan touches neither critical zone says none.
- The report gives the lowest judgement of each criterion, with its reason and passage.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Evaluations of the chain: whether it meets its goals, and what it is like to work with

1 participant