A blueprint that says a fact once, and shows behavior, not code - #83
Merged
Merged
Conversation
Contributor
Author
|
Measured on this branch, before any merge: the four probes of the cross-check on the toy, with the checker's new sentence on the aspects. The faithful blueprint: 0 omission. The hidden column, the hidden critical zone and the hidden effect: each found, 1 omission each. 4 sessions, 0.34 USD, kept in |
This was referenced Oct 3, 2026
The blueprints of the first campaign of the evaluations ran from 700 to 2400 words for small features. A blind reader counted 10 to 29 facts said more than once on each, mostly a body that tells the acceptance criteria again, a scope that names again what the criteria fix, and sensitive zones that restate both, with 4 to 8 details of implementation per page. Two judges score every one of them 2 of 5 on padding. The frame does not force it: a lean page that keeps the whole frame says each fact once (#62). Refs #62
The third campaign of the evaluations, played from this branch, is kept in evals/reports/2026-10-03-d765f898c9ad/. Its judge found twice a visible behavior, the help text, filed under the scope: the scope now says what is left out, and what is done stands in the criteria and in the body. Refs #62
PierreMardon
force-pushed
the
fix/blueprint-says-a-fact-once
branch
from
October 3, 2026 21:41
6054d80 to
313b603
Compare
This was referenced Oct 3, 2026
Open
PierreMardon
added a commit
that referenced
this pull request
Oct 5, 2026
Kept for the next campaign to be compared with: `evals/reports/2026-10-05-d139a1ca0dbe/`. It measures the chain on `main` after the two fixes of the day (#87, #88), with the harness of #89, #90 and #91, on the four cases that no campaign had played again, or once: `04-suspension`, `05-fine-cap` and `06-borrow-limit`, one run played whole and one stopped at the hand over each, and `02-reservations`, one run stopped at the hand over. 108 sessions, 14.23 USD at list price. `report.md` and `summary.json` are as `report --against <summary> --keep` wrote them, and nothing else changes. The summary it is compared with is the one of the first campaign, computed again from its raw data by the harness as it stands, with the count of its six blueprints: no measure the kept summary of that campaign holds differs in it. The first run of `06-borrow-limit` was played in another folder, before the stand-in `gh` was in the image, to check the option that stops a run at the hand over: its sessions and its cost were added to the ledger of the campaign. What it shows: - The three runs played whole reached `conformant` and passed all of their hidden acceptance tests. - The pull request steps are played for the first time: a draft at every hand over, its description refreshed at conformity, never marked ready. In 3 runs of 7 the draft is opened with a line of attribution under the description the state script prints, which the next refresh removes. - The count of the judge follows the fix of the blueprint (#83) on these cases: the facts said twice in prose fall from 19 to 9, from 13 to 4 and from 8 to 1.5 on `02`, `04` and `05`, and the details of implementation from 5 to 11 a page to 1 or 2, while `blueprint.no-padding` is 2 on every run of both campaigns. - On `06-borrow-limit` the interview asks 2 and 3 questions where it asked 1, and its developer hands three of them back: `interview.right-number` falls from 5 to 3.5. - Planning costs 0.94 to 2.20 USD a run here, and a run played whole 1.53 to 1.88. The conclusions, issue by issue, are in the issues #77 lists.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The blueprints of the first campaign of the evaluations ran from 700 to 2400 words for small features. A blind reader counted 10 to 29 facts said more than once on each, mostly a body that tells the acceptance criteria again, a scope that names again what the criteria fix, and sensitive zones that restate both, with 4 to 8 details of implementation per page. Two judges score every one of them 2 of 5 on padding. The frame does not force it: a lean page that keeps the whole frame says each fact once (#62).
What was done
agents/surface-extractor.md: a fact is said once in prose on the whole page, the frame included. The criteria state the rules, carried whole; the body shows what no criterion states, and names a criterion instead of saying it again; the scope says what is left out, and nothing of what is done; the sensitive zones name each zone and point to the criterion or the section that holds its rule. What stays whatever it repeats: the criteria, the statement of the critical zones, the closing line, and a diagram, which may draw what the prose says.agents/surface-checker.mdboth said "in the body": a rule that a criterion states and the body does not tell again would have been counted as an omission, and the rework would have put the repetition back.templates/blueprint.mdandARCHITECTURE.mdsay the same.Measured before merging
From this branch, with no merge: the probes of the cross-check on the toy, then a third campaign of the evaluations, five runs, compared with the second campaign, which holds the four first fixes and an unchanged extractor. Its report is kept here,
evals/reports/2026-10-03-d765f898c9ad/.03-overdue-reminders: 1310 words (1235 to 1425) then 946 (905 to 986), outside the spread of the campaign before.02-reservations, one run each: 2296 then 1572.01-overdue-listdid not move: 984 (669 to 1545) then 1097 (730 to 1463), one of its two runs carrying an amendment that changed how every command parses its arguments.03-overdue-reminders, 7 to 18 then 3 and 6 on01-overdue-list. Details of implementation: 5 to 9 then 1 on03-overdue-reminders. Share of the page it would skip: 35 to 40% then 10 to 15% on03-overdue-reminders.blueprint.decidable: 4.33 then 4.5, 4.67 then 5, and 4 then 4. The blind reader finds 2 things missing over the 5 pages of this branch, 4 over the 6 pages before.What did not get better, or got worse:
blueprint.no-padding, by the judge of the harness, is 2 on every run, as before. That score has one point of range (Evaluations: the judge'sblueprint.no-paddinghas one point of range, and no judgement was checked against a human #73), and part of what the judge faults is what the frame asks for: the closing line, and which component calls which.blueprint.cutis 3 on both runs of03-overdue-reminders, where it was 4 on the three runs before. With less to say, the body is a single section that mixes what is left. Twice the judge found a visible behavior, the help text, filed under the scope.01-overdue-listtook two planning passes on omissions. Both came from the amendment named above, a closing line left as it was and a behavior the plan itself did not state, and not from a rule left out for brevity.Changed after the measure
One sentence, for the lapse the judge found twice: the first wording let the scope hold "of what is done only what no criterion says", and the help text landed there. The scope now says what is left out, and nothing of what is done. This sentence was not measured.
Tests
A test holds the rule and its guards: the criteria carried whole, the frame always written, what stands nowhere else said in full, what is seen from outside kept. The test of the checker holds the new sentence on the aspects. The lint, the prompt tests and the tests of the evaluations pass locally; the CI runs every gate.
Risks
blueprint.no-paddinghas one point of range, and no judgement was checked against a human #73).Refs #62, #77