Repository navigation
The body of a blueprint is cut by the layout of the code, or left uncut where several things are decided separately #63
Description
Activity
Read again in campaign 2
Campaign 2:
01-overdue-listand03-overdue-reminders, three runs each, onmainwith the four first fixes (#78, #79, #80, #81). Report kept:evals/reports/2026-10-03-e830bc236e9b/report.md. Raw data, out of git:.evals/campaign-2/. In the tables, a value is the mean over the three runs, with its lowest and highest when they differ.Campaign 1 Campaign 2 01-overdue-listblueprint.cut3 3.67 (3 to 4) 01-overdue-listsections of the body, 1 expected1 1.33 (1 to 2) 01-overdue-listcut_kept1 0.5 (0 to 1) 03-overdue-remindersblueprint.cut3 4 03-overdue-reminderssections of the body, 2 to 4 expected1 1.67 (1 to 3) - The score is one point better on
03-overdue-remindersin every run, which the report counts as better. No prompt of the blueprint changed between the two campaigns: with a judge that moves by one point on this criterion in half of its cells (Evaluations: one run per case, so the spread between runs is unknown and the next report cannot say what moved #74), this is not to be read as a gain. - New: on
01-overdue-list, one of the two runs that drew a second revision did not keep the cut of the first, which the extractor is told to keep. - The lowest judgement is the same kind as before: "its one feature section, 'The overdue list', restates the acceptance criteria. That section also mixes the visible ordering rule with an architecture choice".
Status
Observed again, over three runs per case. Not fixed. The count of sections moves from run to run on the same need (1 to 3 on
03-overdue-reminders): the cut is a path the model takes that day more than a rule it follows.- The score is one point better on
Campaign 3: one point lower with the shorter pages
With the fix of the extractor (#83), which keeps in the body only what no criterion states:
03-overdue-reminders:blueprint.cut3 and 3, against 4, 4 and 4 in campaign 2. The body is one section on both runs, where the case expects 2 to 4. The judge: "'A remind run' mixes the notice format, the module architecture and rule edge cases in one section, so they cannot be decided separately."01-overdue-list: 4 and 2, against 3, 4 and 4. The 2: "One section titled for the example output holds unrelated matters: module design, the missing file, extra arguments, tie ordering, and help and README text."02-reservations: 3, as in campaign 1, with 5 sections where there were 7, now inside what the case expects.- Twice the help text of the command was filed under the scope. The sentence of the fix that allowed it was reworded before the merge; the new wording was not measured.
One point on this criterion is within the noise of the judge (#74), and two runs are few. But the direction is plausible: a body that no longer tells the criteria again has less to cut, and what is left lands in one section.
Status
Observed over three campaigns, possibly one point worse since #83. Not fixed. The next thing to try is in the extractor's first cut, "One behavior, a small change: no cut, a single section": a page of one section is where rules, boundaries and edge cases get mixed.
Campaign 4
Campaign 4: seven runs on
mainwith every fix so far, #87 and #88 included, and the harness of #89, #90 and #91, on the cases no campaign had played again, or once.04-suspension,05-fine-capand06-borrow-limit: one run played whole and one stopped at the hand over each.02-reservations: one run stopped at the hand over. Report kept:evals/reports/2026-10-05-d139a1ca0dbe/report.md, compared with campaign 1, the only earlier runs of three of these cases. Raw data, out of git:.evals/campaign-4/. A value is the mean over the runs of a case, with its lowest and highest when they differ.Campaign 1 Campaign 4 02-reservationsblueprint.cut3 4 02-reservationssections of the body, 2 to 5 expected7 6 04-suspensionblueprint.cut3 3.5 (3 to 4) 04-suspensionsections of the body, 2 to 4 expected4 3 05-fine-capblueprint.cut4 3 05-fine-capsections of the body, 1 expected3 1.5 (1 to 2) 06-borrow-limitblueprint.cut3 3.5 (3 to 4) 06-borrow-limitsections of the body, 1 expected1 1 The same two kinds as before:
- A section titled by the code, before the behaviors.
04-suspension: "a section about where the code lives comes before both, against the order in which the developer would discover the feature".05-fine-cap: "one section is cut by implementation ("Where the price comes from")". - One section or one list for several things decided separately.
05-fine-cap: "Grace, cap and the unknown-price case are decided separately, yet they are packed into one numbered list rather than given their own sections."06-borrow-limit: "The behaviors the developer decided separately (limits per category, refusal wording, unknown category) sit in one flat list of criteria instead of having their own sections."
The bodies are shorter than in campaign 1 on three cases of four, which is #83. The score does not follow: within one point everywhere, which is the noise of the judge on this criterion (#74).
Status
Observed over four campaigns, on every case. Not fixed. Neither better nor worse than before #83 on these four cases.
- A section titled by the code, before the behaviors.
What the developer's reading of two pages shows of the cut
Recorded on 2026-10-07, from the reading and the comparison told in #62.
- The body of the page drawn after A blueprint that says a fact once, and shows behavior, not code #83 is one section, "A reminder run": a worked example, then three paragraphs with no title. A rule is harder to find there than under the titles in bold of the page before it, "Who is overdue", "Who is skipped: the 7-day interval", "What is written", "What is printed".
- Three rules of that page stand in its scope and not among its criteria: the command takes no argument, a member unknown to
members.jsonlis still reminded, the command never refuses.
The developer chose to add nothing to the changes already decided, two of which bear on every blueprint to come (#60, #64): a page with the text of one and the titles of the other waits until those are measured.
Status
Unchanged: open, not fixed, and left out of the pass under way (#77).
Campaign 5, then the titles the developer asked for: #111, measured before its merge
Campaign 5, twelve runs on
mainwith the nine pull requests of 2026-10-07, the six cases twice each (report kept:evals/reports/2026-10-07-9d5de55f57a9/report.md). It is the measure this issue waited for, with #60 and #64 in it:Campaign 4 Campaign 5 blueprint.cut, all the pages3.4 (3 to 4) on seven 3.8 (3 to 5) on twelve sections of the body inside what the case expects 0.7 1 titles inside the sections of a body none on any page none on any page pages with a section of more than two paragraphs and no title 7 of 7 9 of 12 The two last lines are new measures,
inner_titlesanduntitled_long_sections(#109). They tell apart the two pages the developer read: 5 titles and no long section without one on the page drawn before #83, no title and one such section on the page drawn after.The change, #111, merged on 2026-10-07: inside a section of the body, each thing the developer would come back for on its own opens with a short title in bold, which names the thing and states no rule. A section taken in at a glance has none, a title never stands in for a section the cut asks for, and inside a section what the feature does comes before which component answers for it.
Measured with the plan held still. The twelve plans of campaign 5 were drawn again by an extractor alone, with the mandate
/surface-plangives it: once by the extractor ofmain, the control, once by the one of #111. Only the prompt differs between the two, which a campaign cannot give, since the plan changes from run to run. The judge then scored and counted each page. Means over twelve pages:Extractor of mainExtractor of #111 titles inside the sections 0 3.25 (0 to 10) sections of more than two paragraphs with no title 0.75 0.17 blueprint.cut3.7 (3 to 4) 4.1 (3 to 5) pages where the judge says that one section mixes several things 4 0 facts said twice in prose 4.5 (2 to 8) 5.0 (2 to 8) share of the page to skip 8.8% 8.3% details of implementation 1.3 1.7 blueprint.decidable4.9 4.6 words 789 899 - The titles are there where a section is long and not where it is short, and they name: "The command.", "Who is reminded.", "Where the work is done.", "A run cut short." on a page for the need the developer read. Three of 39 state something instead.
- Nothing more is said twice or to skip. The pages are 14% longer than the control, and as long as those of the campaign.
- The judge no longer says of any page that one section mixes several things, which was the second half of this issue.
- Which component answers for what came before the behavior on five pages, against two: the sentence on the order inside a section was added for that, and checked on four pages, where the component comes last on three.
Two pages to read side by side, for the need of the two the developer read, out of git:
.evals/redraw-main/runs/03-overdue-reminders/run-02/blueprint-rev-02.mdand.evals/redraw-63/runs/03-overdue-reminders/run-02/blueprint-rev-02.md.What is left as it is. The first half of this issue, a body that opens on a section about where the code lives: the judge names it on two pages of twelve in campaign 5 and on one of the control. The extractor is asked to give what several sections share a section before them, which is what puts it there, and the page the developer preferred for its titles has a "Where it lives" of its own, last. A rule on the order of the sections was drafted and taken out on an independent reading: it went against three sentences of the prompt.
What a campaign will say.
inner_titles,untitled_long_sectionsandblueprint.cut, against campaign 5. The drawing again was a one-off of the session that did this, 84 sessions and 11.28 USD with the judge, and is not a command of the harness.Status
Closed: done by #111 and measured before its merge. A rule is found by its title in a long section. The body that opens on where the code lives is left as it is.
- added a commit that references this issue
on Oct 7, 2026
Status
Closed: done by #111, measured before its merge. Observed over five campaigns. The developer's reading of two pages gave the direction: the text of the page drawn since #83, under titles like those of the page before it. #111 asks the extractor for a title in bold on each thing the developer comes back for in a long section. On the twelve plans of campaign 5 drawn again, the pages hold 3.25 such titles where they held none, the judge no longer says that one section mixes several things, and nothing more is said twice. Left as it is: a body that opens on a section about where the code lives. The measures and their sources are in the comments below, campaign by campaign.
What was seen
02-reservations), "## Where the debt is computed" (04-suspension), "Where the logic lives." (03-overdue-reminders).01-overdue-listruns together "selection, order and ties, the source of today, argument refusal, and empty or missing-file output";03-overdue-reminders"mixes the notice shape, the 7-day rule and code placement under bold run-in headings".06-borrow-limit: "the developer's real decisions sit under a generic 'Sensitive zones' title, away from the behaviour they qualify".body_in_rangeblueprint.cut01-overdue-list02-reservations03-overdue-reminders04-suspension05-fine-cap06-borrow-limitThe judge and the corpus disagree on
01-overdue-list: the case expects one section and gets it, and the judge faults that single section. On05-fine-capit is the reverse: three sections where the case expects one, and the best score of the six.Sources
body_sectionsandbody_in_range, scores, lowest judgement.evals/cases/*/case.json,body_sections..evals/campaign-1/runs/*/run-01/judge.json,measures.json.Suspected cause
None in the wording:
agents/surface-extractor.mdlines 53 to 61 already say "never by the slices of the plan nor by the layout of the code" and "Title each section in the words of the feature". The sections "where it lives" may come from the aspect "the architecture and its boundaries", which the extractor must not leave in the dark (line 67).Next
Read again once several runs of one case exist: with one run, a cut cannot be told from the path a model took that day.