Status
Closed. Confirmed, fixed by #83, measured. Three readings confirmed it, a blind count located its cause, and with the fix the facts said more than once fall from 13 to 15 per page to 2 to 4 on the case measured over several runs. The measures and their sources are in the comments below, campaign by campaign.
What was seen
The judge's reasons, one per case:
01-overdue-list: "The text tie rule is stated in criterion 3, again in the body, again in the 'Compared as text' paragraph, and once more under sensitive zones; the missing-file case is likewise repeated in the criteria, the body and two sensitive-zone bullets. Implementation details (argparse, counting 'once', main falling back to the clock) and draft history also appear."
02-reservations: "Many facts are said three times: the acceptance criteria restate the body sections, and the untouched fines.py/loans.jsonl and the 'holds do not run out on their own' rule recur in scope, the body and sensitive zones. The page also spends a section on implementation choices that are not the developer's to decide: the wrapper, the import cycle and a helper in __main__.py."
03-overdue-reminders: "The 7-day rule and the silent skip are restated in the intro, criteria 3 and 7, scope, the rule section, the diagram note and the sensitive zones. The page also spends space on function names, constants and Decimal handling."
04-suspension: "Several facts are stated three times: returned loans counting for nothing, the owed output and refusal, and the check order and nothing written appear in the criteria, the body and the sensitive zones."
05-fine-cap: "the uncapped book with no price appears in criteria 4-5, the table note, the cap section, the diagram and the sensitive zones [...] Code details such as the constant name and the required-parameter design are not the developer's to decide."
06-borrow-limit: "fines.py untouched and the five fields of loans.jsonl appear in criterion 7, out of scope and sensitive zones [...] The page also carries implementation detail (where the constant sits, when the file is saved) and a boilerplate 'No change' line."
Two things are mixed in that score: a fact said several times, and a detail of implementation that is not the developer's to decide.
Sources
- Report of campaign 1, committed: scores, lowest judgement, the judge, checked.
- Raw data, out of git:
.evals/campaign-1/runs/*/run-01/judge.json, and the blueprints themselves, .evals/campaign-1/runs/*/run-01/blueprint-rev-*.md.
Suspected cause (a hypothesis)
- The frame asks for it: the acceptance criteria state the rules, the body shows what will be built, the sensitive zones name what touches the developer's control (
agents/surface-extractor.md lines 39 to 47). "Say a fact once" (line 61) speaks of the sections of the body only.
- The cross-check counts as an omission what the plan does and the blueprint does not show: the extractor is pushed to show everything, details of implementation included.
Why not fix it now
A blueprint that shows less may raise the omissions of the cross-check, so the planning passes and the cost, and may lower blueprint.decidable, the criterion the chain scores best on (4.7). With one point of range, the judge cannot say whether a change helped.
Next
- Settle what the judge is worth on this criterion (see the issue on the judge's range).
- Then decide: the rubric, the frame of the blueprint, or the extractor.
Status
Closed. Confirmed, fixed by #83, measured. Three readings confirmed it, a blind count located its cause, and with the fix the facts said more than once fall from 13 to 15 per page to 2 to 4 on the case measured over several runs. The measures and their sources are in the comments below, campaign by campaign.
What was seen
The judge's reasons, one per case:
01-overdue-list: "The text tie rule is stated in criterion 3, again in the body, again in the 'Compared as text' paragraph, and once more under sensitive zones; the missing-file case is likewise repeated in the criteria, the body and two sensitive-zone bullets. Implementation details (argparse, counting 'once',mainfalling back to the clock) and draft history also appear."02-reservations: "Many facts are said three times: the acceptance criteria restate the body sections, and the untouchedfines.py/loans.jsonland the 'holds do not run out on their own' rule recur in scope, the body and sensitive zones. The page also spends a section on implementation choices that are not the developer's to decide: the wrapper, the import cycle and a helper in__main__.py."03-overdue-reminders: "The 7-day rule and the silent skip are restated in the intro, criteria 3 and 7, scope, the rule section, the diagram note and the sensitive zones. The page also spends space on function names, constants and Decimal handling."04-suspension: "Several facts are stated three times: returned loans counting for nothing, theowedoutput and refusal, and the check order and nothing written appear in the criteria, the body and the sensitive zones."05-fine-cap: "the uncapped book with no price appears in criteria 4-5, the table note, the cap section, the diagram and the sensitive zones [...] Code details such as the constant name and the required-parameter design are not the developer's to decide."06-borrow-limit: "fines.py untouched and the five fields of loans.jsonl appear in criterion 7, out of scope and sensitive zones [...] The page also carries implementation detail (where the constant sits, when the file is saved) and a boilerplate 'No change' line."Two things are mixed in that score: a fact said several times, and a detail of implementation that is not the developer's to decide.
Sources
.evals/campaign-1/runs/*/run-01/judge.json, and the blueprints themselves,.evals/campaign-1/runs/*/run-01/blueprint-rev-*.md.Suspected cause (a hypothesis)
agents/surface-extractor.mdlines 39 to 47). "Say a fact once" (line 61) speaks of the sections of the body only.Why not fix it now
A blueprint that shows less may raise the omissions of the cross-check, so the planning passes and the cost, and may lower
blueprint.decidable, the criterion the chain scores best on (4.7). With one point of range, the judge cannot say whether a change helped.Next