Evaluations of the chain: whether it meets its goals, and what it is like to work with - #58
Merged
Merged
Conversation
…like to work with The tests hold the state script, the wording of the prompts and, end to end, that a real session reaches the expected state on a toy with one behavior. None says whether the chain delivers what the README promises, nor whether its documents and messages serve the developer. `evals/` answers both after a change of the prompts, on real sessions and with no human in them. - A synthetic host, `lending`, with several components, state and two declared critical zones, and a corpus of six needs chosen for their shapes: one behavior, several flows, one flow across components, a trade-off, a critical zone touched, an ambiguous need. Each case keeps from the agents its acceptance tests, and a reference implementation that proves them fair. - A run plays the developer's side: the need, an answer to each question by a model that holds a written brief, a reading of the blueprint whose corrections go back as amendments, a sentence that agrees, which must approve nothing, then the launch of `/surface-execute`. - A script measures what a script can: the outcome, the passes, the acceptance, the critical zones named at approval and their files listed at conformity, the form of the blueprint. Faults put there on purpose, on the toy, show whether the cross-check finds what a blueprint hides and whether the review classifies a defect, a deviation and a break. The toy gains three reviewable states for it. - A model judges the rest against a written rubric, each judgement with its reason and the passage it rests on, and is checked: a document spoiled in a known way must score below its original. - The report keeps each measure as its spread over the runs and says what lies outside the spread of the campaign before. No threshold, and not in CI. A ledger stops a campaign before a ceiling, in USD or in points of the weekly gauge of a subscription. - `scripts/gate.sh evals` runs it in the container of the end to end tests. ADR 0036 records the decision, and `ARCHITECTURE.md` and the contributing guide follow.
- The report of the first campaign is kept under `evals/reports/`: six cases, the probes, the judge and its check. - The toy's blueprint shows the line ending its plan assumes, and says what data the feature adds: a fresh checker found both missing on the blueprint held to be faithful. - The padded blueprint of the judge's check is padded from a lean one of its own: the judge scored the toy's blueprint as low as its padded version. - A blueprint that says the plan touches neither critical zone says none. - The report gives the lowest judgement of each criterion, with its reason and passage.
PierreMardon
marked this pull request as ready for review
October 3, 2026 17:41
This was referenced Oct 3, 2026
Closed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The tests hold the state script, the wording of the prompts and, end to end, that a real session reaches the expected state on a toy with one behavior. None says whether the chain delivers what the README promises, nor whether its documents and messages serve the developer who works with it.
evals/answers both after a change of the prompts, on real sessions and with no human in them, and its first campaign is kept in this pull request.What was done
evals/host/islending, a command line tool for a library: several components, state in files, two critical zones declared in itsAGENTS.md.evals/cases/holds six needs chosen for their shapes: one behavior, several flows, one flow across components, a trade-off, a critical zone touched, an ambiguous need. Each case keeps from the agents its acceptance tests, through the command line, and a reference implementation that proves them fair: the unit tests check that they fail on the host and pass on the reference./surface-execute. It keeps the project, every stream, each blueprint handed over, and what was said at each stop./surface-executeis launched on a branch that holds a defect the tests do not see, a deviation that keeps the blueprint true, a commit against the contract, and work as planned, and is stopped at its first review. The toy gains three reviewable states for it.evals/judge/rubric.md: the blueprint for a developer who must decide, the interview, what the agents say, following along. Each judgement has its score, its reason and the passage it rests on, and a passage found in no document is marked. The judge is checked without a human: a padded blueprint, one stripped of the rules the interview settled, an interview with questions already answered must each score below its original.--keepcopies the report underevals/reports/, named after the date and a hash ofagents/andskills/.scripts/gate.sh evalsruns it in the container of the end to end tests. ADR 0036 records the decision;ARCHITECTURE.md, the contributing guide andevals/README.md, the operating guide, follow.Decisions
On the questions #55 left open. The developer chose the host and who plays the developer; the others were proposed to them and not objected to.
summary.jsonandreport.mdare committed when a campaign is kept. A regression is a measure that lies outside the spread of the campaign before, in the wrong direction: it is reported, never a failure.The first campaign
Six cases, one run each, then the probes, the judge and its check: 101 sessions, 18.71 USD at list price, kept in
evals/reports/2026-10-03-6bf7c06609a6/.conformantand passed all of its acceptance tests. No sentence approved a revision. Every plan was drafted by the built-in agent onopusand held the minimum. The critical zone of the case that touches one was named at approval and its file listed at conformity; the others said none. No session read the clone.Tests
scripts/gate.shis green.tests/evals/holds the harness with no session: the streams and the ledger, the form of a blueprint against the template the extractor is given, the measures on folders as a run leaves them, the rubric and the judge's answers, the report and what moved, the probes built on the toy, the command line, and a run played against the real state script with the sessions replaced by a script. For each case, the acceptance tests fail on the host and pass on its reference.Risks
--runsis there for it.gh: every hand over says so, which the judge counts as noise. That part of the low scores on the messages is the environment's.body_in_rangeanddiagram_as_expectedare to be read with the judge's reasons, not alone.Closes #55