diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md index 4bbcf38..0d6155a 100644 --- a/ARCHITECTURE.md +++ b/ARCHITECTURE.md @@ -23,7 +23,7 @@ host CI --runs----> surface-status check --require conformant ``` - **Commands** are three skills the developer types. `/surface-plan` explores, interviews, has the plan drafted and writes it with its gates, has the blueprint drawn and cross-checked, then commits, pushes and opens a draft pull request. It then stays in the conversation and takes amendments there, as it does after a block during planning, since a reply that approves nothing needs no new launch. `/surface-execute` approves the drafted revision by being launched, then dispatches slices, gate runs, reviews and fixes. It puts a plan change proposal and the ceiling to the developer in the conversation: a refusal or a resumption goes on in the session, while an accepted change or an amendment goes to `/surface-plan`, since this command cannot load it and never writes the plan. `/surface-status` reports, runs the conformity check, and abandons a plan on confirmation. -- **Agents** are four roles of the chain, and one agent built into Claude Code. The `Plan` agent drafts the plan, in the form it chooses: a good plan is what it already does well, so the chain ships no definition for it and holds the plan only to the minimum it reads, the numbered acceptance criteria, the slices behind their markers, each enough for an agent that starts fresh, and the `gates` block ([ADR 0035](docs/adr/0035-plan-drafted-by-the-built-in-plan-agent.md)). It has no write tool: it returns the plan, and `/surface-plan` writes `plan.md` as returned, but for two repairs it makes without a word to the developer, who never reads the plan: a return that arrived escaped, whose markers the script would not read, and a path of the machine. The extractor draws `blueprint.md` from `plan.md`, as long as the feature needs, since the developer reads all of it. A blueprint has a fixed frame and a free body: it opens with the idea, the acceptance criteria and the scope, and closes with the sensitive zones, while the extractor cuts the body by what the developer decides separately, and keeps the cut of the previous revision, so that a revision is read against the one before. No heading carries a number: a section is cited by its title, which an amendment does not move, where a number would shift and send a report to another section. Five aspects are gone through whatever the cut, the data schema, the boundaries, the sequences, the state machines and the algorithms, and one closing line names those the plan leaves alone, since silence must not read as "unchanged". A diagram is drawn where what it shows has a shape that prose flattens, with the existing elements the change attaches to. The checker looks for what the plan does and the blueprint does not show, never for a cut, a short section or a diagram that is not there. The executor carries out a slice or a fix. The reviewer judges the branch against the approved blueprint, writes `conformity.md` when it finds nothing, with the changed files of the critical zones, judges a break an executor suspects, and corrects that list of files when the script refuses it. +- **Agents** are four roles of the chain, and one agent built into Claude Code. The `Plan` agent drafts the plan, in the form it chooses: a good plan is what it already does well, so the chain ships no definition for it and holds the plan only to the minimum it reads, the numbered acceptance criteria, the slices behind their markers, each enough for an agent that starts fresh, and the `gates` block ([ADR 0035](docs/adr/0035-plan-drafted-by-the-built-in-plan-agent.md)). It has no write tool: it returns the plan, and `/surface-plan` writes `plan.md` as returned, but for two repairs it makes without a word to the developer, who never reads the plan: a return that arrived escaped, whose markers the script would not read, and a path of the machine. The extractor draws `blueprint.md` from `plan.md`, as long as the feature needs, since the developer reads all of it. A fact is said once in prose on the whole page: the criteria state the rules, and the body shows what no criterion states, as behavior and not as code, since a page that tells a rule three times is a page the developer skims. So an aspect the plan changes is shown by a criterion or in the body, and the checker reads both. A blueprint has a fixed frame and a free body: it opens with the idea, the acceptance criteria and the scope, and closes with the sensitive zones, while the extractor cuts the body by what the developer decides separately, and keeps the cut of the previous revision, so that a revision is read against the one before. No heading carries a number: a section is cited by its title, which an amendment does not move, where a number would shift and send a report to another section. Five aspects are gone through whatever the cut, the data schema, the boundaries, the sequences, the state machines and the algorithms, and one closing line names those the plan leaves alone, since silence must not read as "unchanged". A diagram is drawn where what it shows has a shape that prose flattens, with the existing elements the change attaches to. The checker looks for what the plan does and the blueprint does not show, never for a cut, a short section or a diagram that is not there. The executor carries out a slice or a fix. The reviewer judges the branch against the approved blueprint, writes `conformity.md` when it finds nothing, with the changed files of the critical zones, judges a break an executor suspects, and corrects that list of files when the script refuses it. - **The state script** derives the state, refuses illegal steps, runs the gates, and answers the skills and the host's CI. - **A plan folder**, `docs/plans/-/` by default, holds the specs, the exploration, the interview, the plan, the blueprint, the reports (`checks/`, `reviews/`, `plan-changes/`, `gates/`), `conformity.md` and `journal.jsonl`. It is the whole state of a plan. - **The installer** copies the chain into the host, updates it, and checks it for drift. diff --git a/agents/surface-checker.md b/agents/surface-checker.md index 6deef1f..784e2ae 100644 --- a/agents/surface-checker.md +++ b/agents/surface-checker.md @@ -22,7 +22,7 @@ Count only what would change the decision of the person who validates the bluepr The closing section of the blueprint, the sensitive zones, names each critical zone the repository's agent instructions declare that the plan touches, or says that the plan touches none: the developer reads the code of those zones themselves, and learns there which ones. A critical zone the plan touches and the sensitive zones do not name is an omission, even when another section shows the change. So is a closing section that says nothing of the critical zones, or says none while the plan touches one. -The blueprint is as long as the feature needs, and its body is cut for the feature. The cut, the titles and the order of its sections are never an omission, and neither is a short section or the absence of a diagram: only what the blueprint does not show counts. Whatever the cut, five aspects must not be left in the dark: the data schema, the architecture and its boundaries, the sequences, the state machines, the algorithms. The closing line of the blueprint names those the plan leaves alone. An aspect the closing line names while the plan changes it is an omission. So is an aspect neither shown in the body nor named by the closing line: the developer cannot tell that the feature leaves it alone. +The blueprint is as long as the feature needs, and its body is cut for the feature. The cut, the titles and the order of its sections are never an omission, and neither is a short section or the absence of a diagram: only what the blueprint does not show counts. Whatever the cut, five aspects must not be left in the dark: the data schema, the architecture and its boundaries, the sequences, the state machines, the algorithms. The closing line of the blueprint names those the plan leaves alone. An aspect the closing line names while the plan changes it is an omission. So is an aspect neither shown on the page, by a criterion or in the body, nor named by the closing line: the developer cannot tell that the feature leaves it alone. If you find nothing, say so: zero omissions is an answer. diff --git a/agents/surface-extractor.md b/agents/surface-extractor.md index 28a4941..eddb5bd 100644 --- a/agents/surface-extractor.md +++ b/agents/surface-extractor.md @@ -32,6 +32,8 @@ And the repository's agent instructions, `AGENTS.md` or `CLAUDE.md` at its root, Write in the language `exploration.md` names in its repository rules. The blueprint is as long as the feature needs, and no longer: the developer reads all of it, so a small change gets a short page. No section and no diagram is written for its own sake. +A fact is said once in prose on the whole page, the frame included: the idea sums the feature up, and a diagram may draw what the prose says. The acceptance criteria state the rules, carried whole, and nothing tells them again: the body shows what no criterion states, the shape of the data, the boundaries between components, an order between actors, the states, an algorithm, an irreversible effect, an assumption the plan takes, the case that explains a rule, and names a criterion instead of saying it again. The scope says what is left out: what is done stands in the criteria and in the body, never there. The sensitive zones name each zone, point to the criterion or the section that holds its rule, and say in full only what stands nowhere else. The statement of the critical zones and the closing line are always written, even when a criterion says the same. + ### The frame Every blueprint opens and closes the same way, so the developer always finds what the work is judged against. No heading carries a number: a section is cited by its title, which an amendment does not move. @@ -50,7 +52,7 @@ Sensitive zones opens with the critical zones because their code is what the dev ### The body -Between them, the body shows what will be built. You choose how to cut it: by what the developer has to decide separately, never by the slices of the plan nor by the layout of the code. Take the first cut that fits: +Between them, the body shows what will be built. It shows behavior, which the developer decides, not code. What a user, a file or another program sees stays on the page: a name they type or read, a message, the value of a limit, a dependency the feature adds, which component calls which, each component under the name the repository gives it, with what it answers for. What only the code sees stays in the plan: the names of functions, constants and helpers, the calls to a library, the layout of the code. You choose how to cut it: by what the developer has to decide separately, never by the slices of the plan nor by the layout of the code. Take the first cut that fits: 1. One behavior, a small change: no cut, a single section. 2. Several flows or visible behaviors, largely independent: one section per flow, each with its own data, order and states. @@ -64,7 +66,7 @@ When `blueprint.md` already exists, read it before you write: keep its cut and i ### The aspects -Whatever the cut, five aspects must not be left in the dark: the data schema, the architecture and its boundaries, the sequences, the state machines, the algorithms. Go through each: what the plan changes of it is shown in the body, in the section it belongs to. The blueprint ends with one closing line that names every aspect the plan leaves alone, as the template shows, so the developer sees at a glance what the feature does not touch. There is no closing line when the plan changes all five. +Whatever the cut, five aspects must not be left in the dark: the data schema, the architecture and its boundaries, the sequences, the state machines, the algorithms. Go through each: what the plan changes of it is shown on the page, by a criterion or in the body, in the section it belongs to. The blueprint ends with one closing line that names every aspect the plan leaves alone, as the template shows, so the developer sees at a glance what the feature does not touch. There is no closing line when the plan changes all five. ### The diagrams diff --git a/evals/reports/2026-10-03-d765f898c9ad/report.md b/evals/reports/2026-10-03-d765f898c9ad/report.md new file mode 100644 index 0000000..ab1119f --- /dev/null +++ b/evals/reports/2026-10-03-d765f898c9ad/report.md @@ -0,0 +1,81 @@ +# Evaluation of the chain + +Chain `d765f898c9ad`, measured on 2026-10-03T21:39:49Z: 5 runs over 3 cases. A value is the mean over the runs of a case, with its lowest and highest when they differ; a truth counts 1. An empty cell does not apply. + +## Does it meet its goals + +| | 01-overdue-list | 02-reservations | 03-overdue-reminders | all | +|---|---|---|---|---| +| `handed_over` | 1 | 1 | 1 | 1 | +| `conformant` | 0.5 (0 to 1) | 1 | 1 | 0.8 (0 to 1) | +| `acceptance` | 1 | 1 | 1 | 1 | +| `approved_by_sentence` | 0 | 0 | 0 | 0 | +| `questions` | 2 | 4 | 0 | 1.6 (0 to 4) | +| `corrections` | 1 | 0 | 1 | 0.8 (0 to 1) | +| `planning_passes` | 1 (0 to 2) | 0 | 0 | 0.4 (0 to 2) | +| `reviews` | 1 | 1 | 1 | 1 | +| `fixes` | 0 | 0 | 0 | 0 | +| `zones_none_said` | 1 | 1 | 1 | 1 | +| `zones_files_listed` | 1 | 1 | 1 | 1 | +| `zones_consistent` | 1 | 1 | 1 | 1 | +| `frame` | 1 | 1 | 1 | 1 | +| `body_sections` | 1.5 (1 to 2) | 5 | 1 | 2 (1 to 5) | +| `body_in_range` | 0.5 (0 to 1) | 1 | 0 | 0.4 (0 to 1) | +| `diagrams` | 1 | 3 | 1 | 1.4 (1 to 3) | +| `diagram_as_expected` | 0 | 1 | 1 | 0.6 (0 to 1) | +| `cut_kept` | 0.5 (0 to 1) | 1 | 1 | 0.8 (0 to 1) | +| `plan_minimum` | 1 | 1 | 1 | 1 | +| `contaminated` | 0 | 0 | 0 | 0 | +| `usd` | 3.4204 (2.2259 to 4.6149) | 4.2718 | 2.9079 (2.8248 to 2.991) | 3.3857 (2.2259 to 4.6149) | + +Outcomes: `01-overdue-list` handed-back, conformant; `02-reservations` conformant; `03-overdue-reminders` conformant, conformant. + +## Is it good to work with + +Scores from 1 to 5, by the judge. + +| | 01-overdue-list | 02-reservations | 03-overdue-reminders | all | +|---|---|---|---|---| +| `blueprint.cut` | 3 (2 to 4) | 3 | 3 | 3 (2 to 4) | +| `blueprint.decidable` | 4.5 (4 to 5) | 4 | 5 | 4.6 (4 to 5) | +| `blueprint.diagrams` | 4 (3 to 5) | 3 | 3.5 (3 to 4) | 3.6 (3 to 5) | +| `blueprint.no-padding` | 2 | 2 | 2 | 2 | +| `following.no-noise` | 2 | 2 | 2 | 2 | +| `following.whose-turn` | 4.5 (4 to 5) | 4 | 5 | 4.6 (4 to 5) | +| `following.why-stopped` | 5 | 5 | 5 | 5 | +| `grounded` | 1 | 1 | 1 | 1 | +| `interview.no-invented-rule` | 4 | 4 | 3 (2 to 4) | 3.6 (2 to 4) | +| `interview.not-already-answered` | 4.5 (4 to 5) | 5 | 5 | 4.8 (4 to 5) | +| `interview.questions-matter` | 4 (3 to 5) | 4 | 2.5 (2 to 3) | 3.4 (2 to 5) | +| `interview.right-number` | 4 (3 to 5) | 4 | 1.5 (1 to 2) | 3 (1 to 5) | +| `messages.clear` | 2.5 (2 to 3) | 2 | 2 | 2.2 (2 to 3) | +| `messages.next-step` | 4.5 (4 to 5) | 3 | 3.5 (3 to 4) | 3.8 (3 to 5) | +| `messages.right-length` | 3.5 (3 to 4) | 3 | 3 | 3.2 (3 to 4) | + +The lowest score of each criterion, with the judge's reason and its passage: + +- `blueprint.cut`, 2 in `01-overdue-list`: One section titled for the example output holds unrelated matters: module design, the missing file, extra arguments, tie ordering, and help and README text. None of these gets its own section in the order the developer would meet it. The behaviors the developer decides separately, such as ties, usage errors and documentation, are not cut apart. Passage: "## The overdue list at the desk" +- `blueprint.decidable`, 4 in `01-overdue-list`: The blueprint carries every rule and interview answer: text-ordered ties, the empty or missing file, and exit 2 on an extra argument. One lapse of detail: criterion 8 says listing 'always ends with exit code 0' while exit 2 for wrong usage only appears later in prose, so the contract reads as self-contradictory on its own. Passage: "The command takes no argument of its own and never refuses: listing the loans always ends with exit code 0, never 1." +- `blueprint.diagrams`, 3 in `01-overdue-list`: The single flowchart does show several modules in relation, including the existing import between them. But what it shows is internal module wiring that serves an implementation argument, not a behavior the developer decides, so it stands where no decision needs it. Passage: "fines -.->|already imports| loans" +- `blueprint.no-padding`, 2 in `01-overdue-list`: Several facts appear three times: a `--today` after `overdue` being ignored, a misspelled `--today` before it exiting 2, `B10` before `B2`, and the parsing reaching every command each recur across the criteria, the "In practice" bullets and "Sensitive zones". The page also carries implementation detail that is not the developer's to decide, such as the Python 3.11 library behaviour, the rejected parsing alternatives and the reason why `loans.py` cannot sort. Passage: "A `--today` written after `overdue` is ignored without a word, and the listing then uses the real date. Only what follows the command is ignored: a misspelled `--today` before it still exits 2 (criterion 10). See \"Extra arguments and exit c [...]" +- `following.no-noise`, 2 in `01-overdue-list`: The missing `gh` and the absent pull request are reported at three stops, and the chain's internals are narrated for their own sake: the drafting agent's stray sentence, its later removal, and "the third check found no omissions". Reassurances such as "no code has been written yet" add nothing the developer needs. Passage: "Revision 2 is drafted and cross-checked, and the third check found no omissions. It's committed and pushed to `feat/overdue-command`. `gh` isn't installed, so there is still no pull request, and no code has been written yet." +- `following.whose-turn`, 4 in `01-overdue-list`: Every stop makes the developer's turn clear, and most say exactly what to do. Stop 3 is the exception: it lays out options A and B but never asks for a pick, ending only on a side question about scope. Passage: "Errors raised before the command is chosen, such as a malformed `--today` or a missing command, belong to the whole tool and keep exiting 2, as the README says. Tell me if you meant something wider." +- `following.why-stopped`, 5 in `01-overdue-list`: Each stop names its reason: a question at stops 1 and 3, a plan ready for approval at stops 2 and 4, a non-approving reply at stop 5, and a plan change proposal at stop 6. There is no lapse. Passage: "The loop has stopped on a plan change proposal, and it's your turn." +- `interview.no-invented-rule`, 2 in `03-overdue-reminders`: The blueprint states rules as settled that come from no source. A member unknown to the library is still reminded, although the README says an unknown member is refused. A reminder dated after today counts as recent. The member order and the wording of the notice appear as criteria, not as assumptions to confirm. Passage: "The member ids come from `loans.jsonl` alone: `members.jsonl` is not read, so a member id unknown there is still reminded, and the command never refuses." +- `interview.not-already-answered`, 4 in `01-overdue-list`: The code already settles Q2: argparse rejects an unknown argument with exit 2, and the README gives 2 as the wrong-usage code. Asking was still justified only because the amendment cast doubt on that, so this is a small lapse rather than a redundant question. Passage: "The command line is built on `argparse`, which rejects an unknown argument with exit code 2, the \"wrong usage\" code of the README." +- `interview.questions-matter`, 2 in `03-overdue-reminders`: The interview asked no questions, so it never brought out the one rule only the developer knew: how often a member may be reminded. The developer had to correct that rule in an amendment after reading the blueprint, and the contract with the mailer was also left unasked. Passage: "None. The specs and the code settle every rule a member or a program reading the outbox would see" +- `interview.right-number`, 1 in `03-overdue-reminders`: Zero questions was too few. The reminder cadence was guessed as once a day and only fixed by a later amendment, and whether the mailer accepts the new `reminder` kind was never asked. Passage: "The mailer, which reads the outbox, will find lines of a kind it has never received, `reminder`, and every line it sends reaches a real member. The plan does not look at the mailer." +- `messages.clear`, 2 in `01-overdue-list`: Several messages use the chain's internal vocabulary instead of plain words. Stop 2 says "Cut", "Gates", "Critical zones", "the drafting agent" and "cross-checked with no omissions"; Stop 4 says "the third check"; Stop 6 says "slice 1", "the loop", "the gate", "contract break" and "A fixer". Passage: "- **Cut:** the blueprint has one section, \"Listing the late loans\", because this is a small change with a single behavior." +- `messages.next-step`, 3 in `02-reservations`: Stops 1 to 7 each end on the right step: answer the question, amend, or launch `/surface-execute`. The final message ends on a remark that the critical zones are untouched, and the real next step, opening the pull request and marking it ready, is buried in the middle. Passage: "Here there are none: `lending/fines.py` and the `loans.jsonl` format are untouched." +- `messages.right-length`, 3 in `01-overdue-list`: The developer gave a custom answer to Q1 (ties by book id as text, no member-id rule), but Stop 2 never confirms how the plan recorded it. Stop 4, by contrast, restates each change it made. Passage: "The plan is drafted, committed and pushed to `feat/overdue-command`. It is waiting for your approval." + +## Cost + +65 sessions, 18.31 USD at list price. The weekly gauge of the subscription rose by 5 points while they ran, every other use of the account included. + +## What moved since the campaign before + +- `01-overdue-list` `questions`: 1 then 2, moved. +- `03-overdue-reminders` `words`: 1309.67 (1235 to 1425) then 945.5 (905 to 986), moved. +- `03-overdue-reminders` `blueprint.cut`: 4 then 3, worse. +- `03-overdue-reminders` `messages.clear`: 3 then 2, worse. diff --git a/evals/reports/2026-10-03-d765f898c9ad/summary.json b/evals/reports/2026-10-03-d765f898c9ad/summary.json new file mode 100644 index 0000000..6f91e61 --- /dev/null +++ b/evals/reports/2026-10-03-d765f898c9ad/summary.json @@ -0,0 +1,1336 @@ +{ + "v": 1, + "at": "2026-10-03T21:39:49Z", + "chain": "d765f898c9ad", + "cases": { + "01-overdue-list": { + "shape": "one-behavior", + "runs": 2, + "outcomes": [ + "handed-back", + "conformant" + ], + "measures": { + "acceptance": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "acceptance_all": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "approvals": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "approved_by_sentence": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "body_in_range": { + "mean": 0.5, + "low": 0.0, + "high": 1.0, + "n": 2 + }, + "body_sections": { + "mean": 1.5, + "low": 1.0, + "high": 2.0, + "n": 2 + }, + "checks": { + "mean": 3.0, + "low": 2.0, + "high": 4.0, + "n": 2 + }, + "conformant": { + "mean": 0.5, + "low": 0.0, + "high": 1.0, + "n": 2 + }, + "contaminated": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "corrections": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "criteria": { + "mean": 10.0, + "low": 10.0, + "high": 10.0, + "n": 2 + }, + "cut_kept": { + "mean": 0.5, + "low": 0.0, + "high": 1.0, + "n": 2 + }, + "diagram_as_expected": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "diagrams": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "findings": { + "mean": 1.0, + "low": 0.0, + "high": 2.0, + "n": 2 + }, + "fixes": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "frame": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "gate": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "gate_runs_failed": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "handed_over": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "killed": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "numbered_headings": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "omissions": { + "mean": 1.0, + "low": 0.0, + "high": 2.0, + "n": 2 + }, + "plan_drafts": { + "mean": 2.5, + "low": 2.0, + "high": 3.0, + "n": 2 + }, + "plan_minimum": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "planning_passes": { + "mean": 1.0, + "low": 0.0, + "high": 2.0, + "n": 2 + }, + "questions": { + "mean": 2.0, + "low": 2.0, + "high": 2.0, + "n": 2 + }, + "relaunches": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "reviews": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "usd": { + "mean": 3.4204, + "low": 2.2259, + "high": 4.6149, + "n": 2 + }, + "words": { + "mean": 1096.5, + "low": 730.0, + "high": 1463.0, + "n": 2 + }, + "zones_consistent": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 1 + }, + "zones_files_listed": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 1 + }, + "zones_none_said": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + } + }, + "judge": { + "blueprint.cut": { + "mean": 3.0, + "low": 2.0, + "high": 4.0, + "n": 2 + }, + "blueprint.decidable": { + "mean": 4.5, + "low": 4.0, + "high": 5.0, + "n": 2 + }, + "blueprint.diagrams": { + "mean": 4.0, + "low": 3.0, + "high": 5.0, + "n": 2 + }, + "blueprint.no-padding": { + "mean": 2.0, + "low": 2.0, + "high": 2.0, + "n": 2 + }, + "following.no-noise": { + "mean": 2.0, + "low": 2.0, + "high": 2.0, + "n": 2 + }, + "following.whose-turn": { + "mean": 4.5, + "low": 4.0, + "high": 5.0, + "n": 2 + }, + "following.why-stopped": { + "mean": 5.0, + "low": 5.0, + "high": 5.0, + "n": 2 + }, + "grounded": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "interview.no-invented-rule": { + "mean": 4.0, + "low": 4.0, + "high": 4.0, + "n": 2 + }, + "interview.not-already-answered": { + "mean": 4.5, + "low": 4.0, + "high": 5.0, + "n": 2 + }, + "interview.questions-matter": { + "mean": 4.0, + "low": 3.0, + "high": 5.0, + "n": 2 + }, + "interview.right-number": { + "mean": 4.0, + "low": 3.0, + "high": 5.0, + "n": 2 + }, + "messages.clear": { + "mean": 2.5, + "low": 2.0, + "high": 3.0, + "n": 2 + }, + "messages.next-step": { + "mean": 4.5, + "low": 4.0, + "high": 5.0, + "n": 2 + }, + "messages.right-length": { + "mean": 3.5, + "low": 3.0, + "high": 4.0, + "n": 2 + } + } + }, + "02-reservations": { + "shape": "several-flows", + "runs": 1, + "outcomes": [ + "conformant" + ], + "measures": { + "acceptance": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 1 + }, + "acceptance_all": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 1 + }, + "approvals": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 1 + }, + "approved_by_sentence": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 1 + }, + "body_in_range": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 1 + }, + "body_sections": { + "mean": 5.0, + "low": 5.0, + "high": 5.0, + "n": 1 + }, + "checks": { + "mean": 2.0, + "low": 2.0, + "high": 2.0, + "n": 1 + }, + "conformant": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 1 + }, + "contaminated": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 1 + }, + "corrections": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 1 + }, + "criteria": { + "mean": 11.0, + "low": 11.0, + "high": 11.0, + "n": 1 + }, + "cut_kept": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 1 + }, + "diagram_as_expected": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 1 + }, + "diagrams": { + "mean": 3.0, + "low": 3.0, + "high": 3.0, + "n": 1 + }, + "findings": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 1 + }, + "fixes": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 1 + }, + "frame": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 1 + }, + "gate": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 1 + }, + "gate_runs_failed": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 1 + }, + "handed_over": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 1 + }, + "killed": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 1 + }, + "numbered_headings": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 1 + }, + "omissions": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 1 + }, + "plan_drafts": { + "mean": 2.0, + "low": 2.0, + "high": 2.0, + "n": 1 + }, + "plan_minimum": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 1 + }, + "planning_passes": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 1 + }, + "questions": { + "mean": 4.0, + "low": 4.0, + "high": 4.0, + "n": 1 + }, + "relaunches": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 1 + }, + "reviews": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 1 + }, + "usd": { + "mean": 4.2718, + "low": 4.2718, + "high": 4.2718, + "n": 1 + }, + "words": { + "mean": 1572.0, + "low": 1572.0, + "high": 1572.0, + "n": 1 + }, + "zones_consistent": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 1 + }, + "zones_files_listed": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 1 + }, + "zones_none_said": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 1 + } + }, + "judge": { + "blueprint.cut": { + "mean": 3.0, + "low": 3.0, + "high": 3.0, + "n": 1 + }, + "blueprint.decidable": { + "mean": 4.0, + "low": 4.0, + "high": 4.0, + "n": 1 + }, + "blueprint.diagrams": { + "mean": 3.0, + "low": 3.0, + "high": 3.0, + "n": 1 + }, + "blueprint.no-padding": { + "mean": 2.0, + "low": 2.0, + "high": 2.0, + "n": 1 + }, + "following.no-noise": { + "mean": 2.0, + "low": 2.0, + "high": 2.0, + "n": 1 + }, + "following.whose-turn": { + "mean": 4.0, + "low": 4.0, + "high": 4.0, + "n": 1 + }, + "following.why-stopped": { + "mean": 5.0, + "low": 5.0, + "high": 5.0, + "n": 1 + }, + "grounded": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 1 + }, + "interview.no-invented-rule": { + "mean": 4.0, + "low": 4.0, + "high": 4.0, + "n": 1 + }, + "interview.not-already-answered": { + "mean": 5.0, + "low": 5.0, + "high": 5.0, + "n": 1 + }, + "interview.questions-matter": { + "mean": 4.0, + "low": 4.0, + "high": 4.0, + "n": 1 + }, + "interview.right-number": { + "mean": 4.0, + "low": 4.0, + "high": 4.0, + "n": 1 + }, + "messages.clear": { + "mean": 2.0, + "low": 2.0, + "high": 2.0, + "n": 1 + }, + "messages.next-step": { + "mean": 3.0, + "low": 3.0, + "high": 3.0, + "n": 1 + }, + "messages.right-length": { + "mean": 3.0, + "low": 3.0, + "high": 3.0, + "n": 1 + } + } + }, + "03-overdue-reminders": { + "shape": "across-components", + "runs": 2, + "outcomes": [ + "conformant", + "conformant" + ], + "measures": { + "acceptance": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "acceptance_all": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "approvals": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "approved_by_sentence": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "body_in_range": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "body_sections": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "checks": { + "mean": 2.0, + "low": 2.0, + "high": 2.0, + "n": 2 + }, + "conformant": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "contaminated": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "corrections": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "criteria": { + "mean": 9.0, + "low": 9.0, + "high": 9.0, + "n": 2 + }, + "cut_kept": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "diagram_as_expected": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "diagrams": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "findings": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "fixes": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "frame": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "gate": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "gate_runs_failed": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "handed_over": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "killed": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "numbered_headings": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "omissions": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "plan_drafts": { + "mean": 2.0, + "low": 2.0, + "high": 2.0, + "n": 2 + }, + "plan_minimum": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "planning_passes": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "questions": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "relaunches": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 2 + }, + "reviews": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "usd": { + "mean": 2.9079, + "low": 2.8248, + "high": 2.991, + "n": 2 + }, + "words": { + "mean": 945.5, + "low": 905.0, + "high": 986.0, + "n": 2 + }, + "zones_consistent": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "zones_files_listed": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "zones_none_said": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + } + }, + "judge": { + "blueprint.cut": { + "mean": 3.0, + "low": 3.0, + "high": 3.0, + "n": 2 + }, + "blueprint.decidable": { + "mean": 5.0, + "low": 5.0, + "high": 5.0, + "n": 2 + }, + "blueprint.diagrams": { + "mean": 3.5, + "low": 3.0, + "high": 4.0, + "n": 2 + }, + "blueprint.no-padding": { + "mean": 2.0, + "low": 2.0, + "high": 2.0, + "n": 2 + }, + "following.no-noise": { + "mean": 2.0, + "low": 2.0, + "high": 2.0, + "n": 2 + }, + "following.whose-turn": { + "mean": 5.0, + "low": 5.0, + "high": 5.0, + "n": 2 + }, + "following.why-stopped": { + "mean": 5.0, + "low": 5.0, + "high": 5.0, + "n": 2 + }, + "grounded": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 2 + }, + "interview.no-invented-rule": { + "mean": 3.0, + "low": 2.0, + "high": 4.0, + "n": 2 + }, + "interview.not-already-answered": { + "mean": 5.0, + "low": 5.0, + "high": 5.0, + "n": 2 + }, + "interview.questions-matter": { + "mean": 2.5, + "low": 2.0, + "high": 3.0, + "n": 2 + }, + "interview.right-number": { + "mean": 1.5, + "low": 1.0, + "high": 2.0, + "n": 2 + }, + "messages.clear": { + "mean": 2.0, + "low": 2.0, + "high": 2.0, + "n": 2 + }, + "messages.next-step": { + "mean": 3.5, + "low": 3.0, + "high": 4.0, + "n": 2 + }, + "messages.right-length": { + "mean": 3.0, + "low": 3.0, + "high": 3.0, + "n": 2 + } + } + } + }, + "all": { + "runs": 5, + "measures": { + "acceptance": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 5 + }, + "acceptance_all": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 5 + }, + "approvals": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 5 + }, + "approved_by_sentence": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 5 + }, + "body_in_range": { + "mean": 0.4, + "low": 0.0, + "high": 1.0, + "n": 5 + }, + "body_sections": { + "mean": 2.0, + "low": 1.0, + "high": 5.0, + "n": 5 + }, + "checks": { + "mean": 2.4, + "low": 2.0, + "high": 4.0, + "n": 5 + }, + "conformant": { + "mean": 0.8, + "low": 0.0, + "high": 1.0, + "n": 5 + }, + "contaminated": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 5 + }, + "corrections": { + "mean": 0.8, + "low": 0.0, + "high": 1.0, + "n": 5 + }, + "criteria": { + "mean": 9.8, + "low": 9.0, + "high": 11.0, + "n": 5 + }, + "cut_kept": { + "mean": 0.8, + "low": 0.0, + "high": 1.0, + "n": 5 + }, + "diagram_as_expected": { + "mean": 0.6, + "low": 0.0, + "high": 1.0, + "n": 5 + }, + "diagrams": { + "mean": 1.4, + "low": 1.0, + "high": 3.0, + "n": 5 + }, + "findings": { + "mean": 0.4, + "low": 0.0, + "high": 2.0, + "n": 5 + }, + "fixes": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 5 + }, + "frame": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 5 + }, + "gate": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 5 + }, + "gate_runs_failed": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 5 + }, + "handed_over": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 5 + }, + "killed": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 5 + }, + "numbered_headings": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 5 + }, + "omissions": { + "mean": 0.4, + "low": 0.0, + "high": 2.0, + "n": 5 + }, + "plan_drafts": { + "mean": 2.2, + "low": 2.0, + "high": 3.0, + "n": 5 + }, + "plan_minimum": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 5 + }, + "planning_passes": { + "mean": 0.4, + "low": 0.0, + "high": 2.0, + "n": 5 + }, + "questions": { + "mean": 1.6, + "low": 0.0, + "high": 4.0, + "n": 5 + }, + "relaunches": { + "mean": 0.0, + "low": 0.0, + "high": 0.0, + "n": 5 + }, + "reviews": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 5 + }, + "usd": { + "mean": 3.3857, + "low": 2.2259, + "high": 4.6149, + "n": 5 + }, + "words": { + "mean": 1131.2, + "low": 730.0, + "high": 1572.0, + "n": 5 + }, + "zones_consistent": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 4 + }, + "zones_files_listed": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 4 + }, + "zones_none_said": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 5 + } + }, + "judge": { + "blueprint.cut": { + "mean": 3.0, + "low": 2.0, + "high": 4.0, + "n": 5 + }, + "blueprint.decidable": { + "mean": 4.6, + "low": 4.0, + "high": 5.0, + "n": 5 + }, + "blueprint.diagrams": { + "mean": 3.6, + "low": 3.0, + "high": 5.0, + "n": 5 + }, + "blueprint.no-padding": { + "mean": 2.0, + "low": 2.0, + "high": 2.0, + "n": 5 + }, + "following.no-noise": { + "mean": 2.0, + "low": 2.0, + "high": 2.0, + "n": 5 + }, + "following.whose-turn": { + "mean": 4.6, + "low": 4.0, + "high": 5.0, + "n": 5 + }, + "following.why-stopped": { + "mean": 5.0, + "low": 5.0, + "high": 5.0, + "n": 5 + }, + "grounded": { + "mean": 1.0, + "low": 1.0, + "high": 1.0, + "n": 5 + }, + "interview.no-invented-rule": { + "mean": 3.6, + "low": 2.0, + "high": 4.0, + "n": 5 + }, + "interview.not-already-answered": { + "mean": 4.8, + "low": 4.0, + "high": 5.0, + "n": 5 + }, + "interview.questions-matter": { + "mean": 3.4, + "low": 2.0, + "high": 5.0, + "n": 5 + }, + "interview.right-number": { + "mean": 3.0, + "low": 1.0, + "high": 5.0, + "n": 5 + }, + "messages.clear": { + "mean": 2.2, + "low": 2.0, + "high": 3.0, + "n": 5 + }, + "messages.next-step": { + "mean": 3.8, + "low": 3.0, + "high": 5.0, + "n": 5 + }, + "messages.right-length": { + "mean": 3.2, + "low": 3.0, + "high": 4.0, + "n": 5 + } + } + }, + "remarks": { + "blueprint.cut": { + "case": "01-overdue-list", + "score": 2, + "reason": "One section titled for the example output holds unrelated matters: module design, the missing file, extra arguments, tie ordering, and help and README text. None of these gets its own section in the order the developer would meet it. The behaviors the developer decides separately, such as ties, usage errors and documentation, are not cut apart.", + "passage": "## The overdue list at the desk", + "grounded": true + }, + "blueprint.decidable": { + "case": "01-overdue-list", + "score": 4, + "reason": "The blueprint carries every rule and interview answer: text-ordered ties, the empty or missing file, and exit 2 on an extra argument. One lapse of detail: criterion 8 says listing 'always ends with exit code 0' while exit 2 for wrong usage only appears later in prose, so the contract reads as self-contradictory on its own.", + "passage": "The command takes no argument of its own and never refuses: listing the loans always ends with exit code 0, never 1.", + "grounded": true + }, + "blueprint.diagrams": { + "case": "01-overdue-list", + "score": 3, + "reason": "The single flowchart does show several modules in relation, including the existing import between them. But what it shows is internal module wiring that serves an implementation argument, not a behavior the developer decides, so it stands where no decision needs it.", + "passage": "fines -.->|already imports| loans", + "grounded": true + }, + "blueprint.no-padding": { + "case": "01-overdue-list", + "score": 2, + "reason": "Several facts appear three times: a `--today` after `overdue` being ignored, a misspelled `--today` before it exiting 2, `B10` before `B2`, and the parsing reaching every command each recur across the criteria, the \"In practice\" bullets and \"Sensitive zones\". The page also carries implementation detail that is not the developer's to decide, such as the Python 3.11 library behaviour, the rejected parsing alternatives and the reason why `loans.py` cannot sort.", + "passage": "A `--today` written after `overdue` is ignored without a word, and the listing then uses the real date. Only what follows the command is ignored: a misspelled `--today` before it still exits 2 (criterion 10). See \"Extra arguments and exit c [...]", + "grounded": true + }, + "following.no-noise": { + "case": "01-overdue-list", + "score": 2, + "reason": "The missing `gh` and the absent pull request are reported at three stops, and the chain's internals are narrated for their own sake: the drafting agent's stray sentence, its later removal, and \"the third check found no omissions\". Reassurances such as \"no code has been written yet\" add nothing the developer needs.", + "passage": "Revision 2 is drafted and cross-checked, and the third check found no omissions. It's committed and pushed to `feat/overdue-command`. `gh` isn't installed, so there is still no pull request, and no code has been written yet.", + "grounded": true + }, + "following.whose-turn": { + "case": "01-overdue-list", + "score": 4, + "reason": "Every stop makes the developer's turn clear, and most say exactly what to do. Stop 3 is the exception: it lays out options A and B but never asks for a pick, ending only on a side question about scope.", + "passage": "Errors raised before the command is chosen, such as a malformed `--today` or a missing command, belong to the whole tool and keep exiting 2, as the README says. Tell me if you meant something wider.", + "grounded": true + }, + "following.why-stopped": { + "case": "01-overdue-list", + "score": 5, + "reason": "Each stop names its reason: a question at stops 1 and 3, a plan ready for approval at stops 2 and 4, a non-approving reply at stop 5, and a plan change proposal at stop 6. There is no lapse.", + "passage": "The loop has stopped on a plan change proposal, and it's your turn.", + "grounded": true + }, + "interview.no-invented-rule": { + "case": "03-overdue-reminders", + "score": 2, + "reason": "The blueprint states rules as settled that come from no source. A member unknown to the library is still reminded, although the README says an unknown member is refused. A reminder dated after today counts as recent. The member order and the wording of the notice appear as criteria, not as assumptions to confirm.", + "passage": "The member ids come from `loans.jsonl` alone: `members.jsonl` is not read, so a member id unknown there is still reminded, and the command never refuses.", + "grounded": true + }, + "interview.not-already-answered": { + "case": "01-overdue-list", + "score": 4, + "reason": "The code already settles Q2: argparse rejects an unknown argument with exit 2, and the README gives 2 as the wrong-usage code. Asking was still justified only because the amendment cast doubt on that, so this is a small lapse rather than a redundant question.", + "passage": "The command line is built on `argparse`, which rejects an unknown argument with exit code 2, the \"wrong usage\" code of the README.", + "grounded": true + }, + "interview.questions-matter": { + "case": "03-overdue-reminders", + "score": 2, + "reason": "The interview asked no questions, so it never brought out the one rule only the developer knew: how often a member may be reminded. The developer had to correct that rule in an amendment after reading the blueprint, and the contract with the mailer was also left unasked.", + "passage": "None. The specs and the code settle every rule a member or a program reading the outbox would see", + "grounded": true + }, + "interview.right-number": { + "case": "03-overdue-reminders", + "score": 1, + "reason": "Zero questions was too few. The reminder cadence was guessed as once a day and only fixed by a later amendment, and whether the mailer accepts the new `reminder` kind was never asked.", + "passage": "The mailer, which reads the outbox, will find lines of a kind it has never received, `reminder`, and every line it sends reaches a real member. The plan does not look at the mailer.", + "grounded": true + }, + "messages.clear": { + "case": "01-overdue-list", + "score": 2, + "reason": "Several messages use the chain's internal vocabulary instead of plain words. Stop 2 says \"Cut\", \"Gates\", \"Critical zones\", \"the drafting agent\" and \"cross-checked with no omissions\"; Stop 4 says \"the third check\"; Stop 6 says \"slice 1\", \"the loop\", \"the gate\", \"contract break\" and \"A fixer\".", + "passage": "- **Cut:** the blueprint has one section, \"Listing the late loans\", because this is a small change with a single behavior.", + "grounded": true + }, + "messages.next-step": { + "case": "02-reservations", + "score": 3, + "reason": "Stops 1 to 7 each end on the right step: answer the question, amend, or launch `/surface-execute`. The final message ends on a remark that the critical zones are untouched, and the real next step, opening the pull request and marking it ready, is buried in the middle.", + "passage": "Here there are none: `lending/fines.py` and the `loans.jsonl` format are untouched.", + "grounded": true + }, + "messages.right-length": { + "case": "01-overdue-list", + "score": 3, + "reason": "The developer gave a custom answer to Q1 (ties by book id as text, no member-id rule), but Stop 2 never confirms how the plan recorded it. Stop 4, by contrast, restates each change it made.", + "passage": "The plan is drafted, committed and pushed to `feat/overdue-command`. It is waiting for your approval.", + "grounded": true + } + }, + "probes": [], + "judge_check": null, + "spent": { + "usd": 18.3091, + "points": 5.0, + "sessions": 65 + }, + "against": { + "chain": "e830bc236e9b", + "at": "2026-10-03T21:09:36Z" + } +} diff --git a/skills/surface-plan/templates/blueprint.md b/skills/surface-plan/templates/blueprint.md index 7ae989a..12ca7ef 100644 --- a/skills/surface-plan/templates/blueprint.md +++ b/skills/surface-plan/templates/blueprint.md @@ -1,6 +1,6 @@ # : blueprint -Revision , drawn from `plan.md`. The page the developer reads to approve the work, then the contract the work is judged against: frozen at approval. Written in the language `exploration.md` names in its repository rules: translate the headings. As long as the feature needs, and no longer. No heading carries a number. The three opening sections and the closing one are always written. Between them, the body is cut for the feature: as many sections as it has things to decide separately, titled in its own words. The closing line names the aspects the plan leaves alone, among the data schema, the architecture and its boundaries, the sequences, the state machines and the algorithms, and goes when the plan changes all five. A mermaid diagram where what it shows has a shape that prose flattens, in the section it serves. Neither the order of construction nor the distribution of tests: they belong to the plan. +Revision , drawn from `plan.md`. The page the developer reads to approve the work, then the contract the work is judged against: frozen at approval. Written in the language `exploration.md` names in its repository rules: translate the headings. As long as the feature needs, and no longer. A fact is said once in prose on the whole page: what a criterion states is not told again in the scope, the body or the sensitive zones. No heading carries a number. The three opening sections and the closing one are always written. Between them, the body is cut for the feature: as many sections as it has things to decide separately, titled in its own words. The closing line names the aspects the plan leaves alone, among the data schema, the architecture and its boundaries, the sequences, the state machines and the algorithms, and goes when the plan changes all five. A mermaid diagram where what it shows has a shape that prose flattens, in the section it serves. Neither the order of construction nor the distribution of tests: they belong to the plan. ## The idea in one sentence @@ -12,11 +12,11 @@ Numbered, taken from the specs and the interview. ## Scope and out of scope -What the feature does, and what it explicitly does not. +What the feature explicitly does not do. What it does stands in the criteria and in the body. ## -What will be built there, and what the developer decides by approving it. +What will be built there that no criterion states, as behavior and not as code, and what the developer decides by approving it. ## Sensitive zones diff --git a/tests/prompts/test_agent_mandates.py b/tests/prompts/test_agent_mandates.py index 0b9491e..c897bc8 100644 --- a/tests/prompts/test_agent_mandates.py +++ b/tests/prompts/test_agent_mandates.py @@ -163,6 +163,39 @@ def test_extractor_cuts_the_body_by_what_the_developer_decides_separately() -> N assert "Say a fact once" in writing +def test_extractor_says_a_fact_once_on_the_whole_page_and_shows_behavior_not_code() -> None: + _, body = read_agent("surface-extractor") + writing = section(body, "What you write") + # The criteria state the rules: the body, the scope and the sensitive zones do not tell them + # again, so the developer skips nothing of the page they approve. + said = [ + "A fact is said once in prose on the whole page, the frame included", + "the body shows what no criterion states", + "The scope says what is left out: what is done stands in the criteria and in the body", + "point to the criterion or the section that holds its rule", + ] + positions = [writing.index(sentence) for sentence in said] + assert positions == sorted(positions) + assert "names a criterion instead of saying it again" in writing + # What must not be lost to brevity: the frame, a diagram, what stands nowhere else. + assert "and a diagram may draw what the prose says" in writing + assert "The acceptance criteria state the rules, carried whole" in writing + assert "and say in full only what stands nowhere else" in writing + always = "The statement of the critical zones and the closing line are always written" + assert always in writing + assert "shown on the page, by a criterion or in the body" in writing + # Behavior, not code: what is seen from outside stays, what only the code sees goes. + assert "It shows behavior, which the developer decides, not code" in writing + assert "What a user, a file or another program sees stays on the page" in writing + assert "which component calls which" in writing + assert "What only the code sees stays in the plan" in writing + template = Path(__file__).resolve().parents[2] / "skills/surface-plan/templates/blueprint.md" + form = template.read_text(encoding="utf-8") + assert "A fact is said once in prose on the whole page" in form + assert "What it does stands in the criteria and in the body" in form + assert "that no criterion states, as behavior and not as code" in form + + def test_extractor_keeps_the_cut_of_the_previous_revision() -> None: _, body = read_agent("surface-extractor") writing = section(body, "What you write") @@ -208,7 +241,11 @@ def test_checker_never_counts_a_cut_a_short_section_or_the_absence_of_a_diagram( assert "only what the blueprint does not show counts" in counting assert f"five aspects must not be left in the dark: {', '.join(BLUEPRINT_ASPECTS)}" in counting assert "An aspect the closing line names while the plan changes it is an omission" in counting - assert "neither shown in the body nor named by the closing line" in counting + # A rule a criterion states is not told again in the body: the page shows it all the same. + shown = ( + "neither shown on the page, by a criterion or in the body, nor named by the closing line" + ) + assert shown in counting # The rule shared with the reviewer's amendment check stays as it was. rule = marked_block(counting, "checker-rule").lower() assert "closing line" not in rule