Skip to content

A blueprint that says a fact once, and shows behavior, not code - #83

Merged
PierreMardon merged 2 commits into
mainfrom
fix/blueprint-says-a-fact-once
Oct 3, 2026
Merged

PierreMardon merged 2 commits into
mainfrom
fix/blueprint-says-a-fact-once

Conversation

@PierreMardon

@PierreMardon PierreMardon commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

The blueprints of the first campaign of the evaluations ran from 700 to 2400 words for small features. A blind reader counted 10 to 29 facts said more than once on each, mostly a body that tells the acceptance criteria again, a scope that names again what the criteria fix, and sensitive zones that restate both, with 4 to 8 details of implementation per page. Two judges score every one of them 2 of 5 on padding. The frame does not force it: a lean page that keeps the whole frame says each fact once (#62).

What was done

  • agents/surface-extractor.md: a fact is said once in prose on the whole page, the frame included. The criteria state the rules, carried whole; the body shows what no criterion states, and names a criterion instead of saying it again; the scope says what is left out, and nothing of what is done; the sensitive zones name each zone and point to the criterion or the section that holds its rule. What stays whatever it repeats: the criteria, the statement of the critical zones, the closing line, and a diagram, which may draw what the prose says.
  • The body shows behavior, not code: what a user, a file or another program sees stays on the page, a name they type or read, a message, the value of a limit, an added dependency, which component calls which; what only the code sees stays in the plan, the names of functions, constants and helpers, the calls to a library, the layout of the code.
  • An aspect the plan changes is shown on the page, by a criterion or in the body. The extractor and agents/surface-checker.md both said "in the body": a rule that a criterion states and the body does not tell again would have been counted as an omission, and the rework would have put the repetition back.
  • templates/blueprint.md and ARCHITECTURE.md say the same.

Measured before merging

From this branch, with no merge: the probes of the cross-check on the toy, then a third campaign of the evaluations, five runs, compared with the second campaign, which holds the four first fixes and an unchanged extractor. Its report is kept here, evals/reports/2026-10-03-d765f898c9ad/.

  • The cross-check still sees. The faithful blueprint: 0 omission. The hidden column, critical zone and effect: each found.
  • The pages are shorter where they were long. 03-overdue-reminders: 1310 words (1235 to 1425) then 946 (905 to 986), outside the spread of the campaign before. 02-reservations, one run each: 2296 then 1572. 01-overdue-list did not move: 984 (669 to 1545) then 1097 (730 to 1463), one of its two runs carrying an amendment that changed how every command parses its arguments.
  • They repeat far less. A blind reader, given the blueprints of both campaigns under neutral names with the lean and padded pages of the judge's check among them, ranked every page of this branch as less redundant than every page of the campaign before. Facts said more than once in prose: 13 to 15 then 2 and 4 on 03-overdue-reminders, 7 to 18 then 3 and 6 on 01-overdue-list. Details of implementation: 5 to 9 then 1 on 03-overdue-reminders. Share of the page it would skip: 35 to 40% then 10 to 15% on 03-overdue-reminders.
  • Nothing more is missing to decide. blueprint.decidable: 4.33 then 4.5, 4.67 then 5, and 4 then 4. The blind reader finds 2 things missing over the 5 pages of this branch, 4 over the 6 pages before.
  • Every run passes its hidden acceptance tests. Four are conformant. The fifth handed back on a contract break that is not about the blueprint's length: the name of a model in a commit trailer (A commit attribution written in the plan reaches the blueprint, and the review raises a contract break on it #86).

What did not get better, or got worse:

  • blueprint.no-padding, by the judge of the harness, is 2 on every run, as before. That score has one point of range (Evaluations: the judge's blueprint.no-padding has one point of range, and no judgement was checked against a human #73), and part of what the judge faults is what the frame asks for: the closing line, and which component calls which.
  • blueprint.cut is 3 on both runs of 03-overdue-reminders, where it was 4 on the three runs before. With less to say, the body is a single section that mixes what is left. Twice the judge found a visible behavior, the help text, filed under the scope.
  • One run of 01-overdue-list took two planning passes on omissions. Both came from the amendment named above, a closing line left as it was and a behavior the plan itself did not state, and not from a rule left out for brevity.

Changed after the measure

One sentence, for the lapse the judge found twice: the first wording let the scope hold "of what is done only what no criterion says", and the help text landed there. The scope now says what is left out, and nothing of what is done. This sentence was not measured.

Tests

A test holds the rule and its guards: the criteria carried whole, the frame always written, what stands nowhere else said in full, what is seen from outside kept. The test of the checker holds the new sentence on the aspects. The lint, the prompt tests and the tests of the evaluations pass locally; the CI runs every gate.

Risks

Refs #62, #77

@PierreMardon

Copy link
Copy Markdown
Contributor Author

Measured on this branch, before any merge: the four probes of the cross-check on the toy, with the checker's new sentence on the aspects. The faithful blueprint: 0 omission. The hidden column, the hidden critical zone and the hidden effect: each found, 1 omission each. 4 sessions, 0.34 USD, kept in .evals/probes-blueprint-fix/ on the developer's machine. So the change does not blind the cross-check on the toy. Next: runs of the corpus from this branch, compared with campaign 2, once that campaign is over.

The blueprints of the first campaign of the evaluations ran from 700 to 2400 words for
small features. A blind reader counted 10 to 29 facts said more than once on each, mostly
a body that tells the acceptance criteria again, a scope that names again what the
criteria fix, and sensitive zones that restate both, with 4 to 8 details of implementation
per page. Two judges score every one of them 2 of 5 on padding. The frame does not force
it: a lean page that keeps the whole frame says each fact once (#62).

Refs #62
The third campaign of the evaluations, played from this branch, is kept in
evals/reports/2026-10-03-d765f898c9ad/. Its judge found twice a visible behavior, the help
text, filed under the scope: the scope now says what is left out, and what is done stands
in the criteria and in the body.

Refs #62
@PierreMardon
PierreMardon force-pushed the fix/blueprint-says-a-fact-once branch from 6054d80 to 313b603 Compare October 3, 2026 21:41
@PierreMardon
PierreMardon merged commit 32e0525 into main Oct 3, 2026
3 checks passed
@PierreMardon
PierreMardon deleted the fix/blueprint-says-a-fact-once branch October 3, 2026 21:58
PierreMardon added a commit that referenced this pull request Oct 5, 2026
Kept for the next campaign to be compared with: `evals/reports/2026-10-05-d139a1ca0dbe/`.
It measures the chain on `main` after the two fixes of the day (#87, #88), with the harness
of #89, #90 and #91, on the four cases that no campaign had played again, or once:
`04-suspension`, `05-fine-cap` and `06-borrow-limit`, one run played whole and one stopped
at the hand over each, and `02-reservations`, one run stopped at the hand over. 108
sessions, 14.23 USD at list price.

`report.md` and `summary.json` are as `report --against <summary> --keep` wrote them, and
nothing else changes. The summary it is compared with is the one of the first campaign,
computed again from its raw data by the harness as it stands, with the count of its six
blueprints: no measure the kept summary of that campaign holds differs in it.

The first run of `06-borrow-limit` was played in another folder, before the stand-in `gh`
was in the image, to check the option that stops a run at the hand over: its sessions and
its cost were added to the ledger of the campaign.

What it shows:

- The three runs played whole reached `conformant` and passed all of their hidden
  acceptance tests.
- The pull request steps are played for the first time: a draft at every hand over, its
  description refreshed at conformity, never marked ready. In 3 runs of 7 the draft is
  opened with a line of attribution under the description the state script prints, which
  the next refresh removes.
- The count of the judge follows the fix of the blueprint (#83) on these cases: the facts
  said twice in prose fall from 19 to 9, from 13 to 4 and from 8 to 1.5 on `02`, `04` and
  `05`, and the details of implementation from 5 to 11 a page to 1 or 2, while
  `blueprint.no-padding` is 2 on every run of both campaigns.
- On `06-borrow-limit` the interview asks 2 and 3 questions where it asked 1, and its
  developer hands three of them back: `interview.right-number` falls from 5 to 3.5.
- Planning costs 0.94 to 2.20 USD a run here, and a run played whole 1.53 to 1.88.

The conclusions, issue by issue, are in the issues #77 lists.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant