Repository navigation
Evaluations: one run per case, so the spread between runs is unknown and the next report cannot say what moved #74
Description
Activity
Measured: what the judge itself adds to the spread
The six runs of campaign 1 were judged again with no session of the chain: twice more by the same judge (
opus), once by another model (sonnet). Same documents, same rubric: whatever moves here is the judge, not the chain.Each cell: the three passes of the same judge (the campaign's, then the two new ones), then the other model after the slash.
noneis an answer that was not the JSON asked for, which the harness refused.Criterion 010203040506Cells that moved blueprint.cut3 2 3 / 3 3 4 3 / 4 3 4 4 / 3 3 3 3 / 4 4 4 4 / 4 3 3 3 / 3 3 of 6 blueprint.decidable5 5 5 / 4 4 4 4 / 4 5 5 5 / 4 5 5 5 / 4 4 4 4 / 4 5 5 5 / 4 0 of 6 blueprint.diagrams5 5 5 / 4 4 4 4 / 3 3 4 3 / 3 4 4 4 / 2 3 3 4 / 3 5 5 5 / 3 2 of 6 blueprint.no-padding2 2 2 / 2 2 2 2 / 2 2 2 2 / 2 2 2 2 / 2 2 2 2 / 2 2 2 2 / 3 0 of 6 following.no-noise2 2 2 / none 2 2 2 / 3 2 2 2 / 3 2 2 2 / 3 2 2 2 / 2 3 2 2 / 3 1 of 6 following.whose-turn4 4 3 / none 3 3 3 / 4 4 4 5 / 4 4 4 4 / 4 5 5 5 / 4 4 5 3 / 4 3 of 6 following.why-stopped5 5 5 / none 5 5 5 / 5 5 5 5 / 5 5 5 5 / 5 5 5 5 / 5 5 5 5 / 5 0 of 6 interview.no-invented-rule4 4 4 / 3 4 4 3 / none 4 4 4 / 4 5 5 5 / 4 4 4 4 / none 5 5 5 / 5 1 of 6 interview.not-already-answered5 5 5 / 5 5 4 4 / none 5 5 5 / 5 5 5 5 / 5 5 5 5 / none 4 5 4 / 5 2 of 6 interview.questions-matter4 4 4 / 4 5 5 5 / none 3 3 3 / 3 5 5 5 / 5 5 5 5 / none 5 4 5 / 5 1 of 6 interview.right-number3 3 3 / 2 4 4 3 / none 2 2 2 / 2 5 5 5 / 4 4 4 5 / none 5 5 5 / 5 2 of 6 messages.clear2 2 2 / 3 2 2 2 / 3 2 2 2 / 3 3 2 2 / 3 2 2 2 / 3 3 2 3 / 3 2 of 6 messages.next-step3 3 3 / 4 3 3 3 / 3 3 3 3 / 3 2 3 3 / 4 3 3 4 / 4 3 3 3 / 4 2 of 6 messages.right-length3 3 3 / 2 3 3 3 / 3 3 2 3 / 2 3 3 3 / 3 3 3 3 / 3 3 3 3 / 4 1 of 6 What it says:
- Over the 84 cells, the three passes of the same judge give the same score in 64, differ by one point in 19, and by two points in one (
following.whose-turnon06-borrow-limit). - Never moved:
blueprint.no-padding,blueprint.decidable,following.why-stopped. Moved most:blueprint.cutandfollowing.whose-turn, in 3 cells of 6. - So one point of difference on one run is inside the judge's own noise for most criteria. A change between two campaigns is worth reading when it holds over several runs, or exceeds one point.
- The other model agrees within one point nearly everywhere. It is a little kinder on the messages (
messages.clear3 on every case, where the first judge gives 2 on four) and harder on the diagrams.
Cost: 48 sessions and 3.14 USD for the two passes of
opus, 24 sessions and 0.59 USD for the pass ofsonnet. Raw data, out of git:.evals/rejudge-opus/,.evals/rejudge-sonnet/.Still unknown: the spread between runs of the chain itself, which the next campaign measures on a few cases.
- Over the 84 cells, the three passes of the same judge give the same score in 64, differ by one point in 19, and by two points in one (
Measured in campaign 2: a first spread between runs, on two cases
Campaign 2:
01-overdue-listand03-overdue-reminders, three runs each, onmainwith the four first fixes (#78, #79, #80, #81). Report kept:evals/reports/2026-10-03-e830bc236e9b/report.md. Raw data, out of git:.evals/campaign-2/. In the tables, a value is the mean over the three runs, with its lowest and highest when they differ.Three runs of the same need, on the same version of the chain:
01-overdue-list03-overdue-remindersconformant,acceptance1, 1 1, 1 questions1 0.33 (0 to 1) corrections1 (0 to 2) 1 drafts of the plan 2 (1 to 3) 2 sections of the body 1.33 (1 to 2) 1.67 (1 to 3) diagrams 0.33 (0 to 1) 1.67 (1 to 3) words of the blueprint 984 (669 to 1545) 1310 (1235 to 1425) acceptance criteria 8 (6 to 10) 10 (9 to 12) usd2.34 (1.41 to 3.60) 2.84 (2.65 to 3.11) scores of the judge within one point, on every criterion but interview.right-numberof01(2 to 5) andinterview.no-invented-rule(3 to 5, and 2 to 4)What it says:
- The outcome is steady: six runs, six conformant, every hidden test passed.
- The path is not: the same need gives a blueprint of 669 or of 1545 words, one section or three, none or two corrections, and costs from one to two and a half times as much.
- So the form of the blueprint and the count of corrections cannot be read on one run. The first campaign's "1 section, as expected" and "2 corrections" were one draw each.
- With the noise of the judge measured above, the report's rule now has something to stand on for these two cases: it listed 7 moves between the two campaigns. Three are what the fixes were written for: the question asked on
01-overdue-list, the score of that interview, and the next step of the messages on03-overdue-reminders. The four others, the cut and the clarity on03-overdue-reminders, its count of criteria and of words, have no fix to explain them: each is within one point of the judge or one draw of the model, and three runs are still few.
Status
Measured on two cases of six, three runs each. The four other cases are still one run.
Campaign 3 adds two runs of each case, on another version
Not the same version of the chain as campaign 2, so not the same spread, but the same lesson: on
01-overdue-listone run cost 2.23 USD and the other 4.61, one page had 730 words and the other 1463, one run was conformant and the other handed back. The difference between the two runs is an answer of the simulated developer (#82), then the name of a model in a commit (#86).So part of the spread is not the chain's: it comes from the developer played by a model, whose answers change from run to run on what its brief leaves open. A spread between runs measures the chain and its developer together.
Campaign 4: two runs of planning on the cases never played again
Campaign 4: seven runs on
mainwith every fix so far, #87 and #88 included, and the harness of #89, #90 and #91, on the cases no campaign had played again, or once.04-suspension,05-fine-capand06-borrow-limit: one run played whole and one stopped at the hand over each.02-reservations: one run stopped at the hand over. Report kept:evals/reports/2026-10-05-d139a1ca0dbe/report.md, compared with campaign 1, the only earlier runs of three of these cases. Raw data, out of git:.evals/campaign-4/. A value is the mean over the runs of a case, with its lowest and highest when they differ.The three runs played whole are conformant and pass every hidden test: 12 of 12, 11 of 11 and 9 of 9.
02-reservations, one run04-suspension05-fine-cap06-borrow-limitquestions3 2.5 (2 to 3) 1 2.5 (2 to 3) corrections0 0 0 0.5 (0 to 1) drafts of the plan 2 1 1 1.5 (1 to 2) sections of the body 6 3 1.5 (1 to 2) 1 diagrams 3 1.5 (1 to 2) 1 1 acceptance criteria 9 10 (9 to 11) 8.5 (8 to 9) 9 (8 to 10) words of the blueprint 1448 928.5 (858 to 999) 780 (684 to 876) 713 (621 to 805) facts repeated in prose 9 4 (3 to 5) 1.5 (1 to 2) 6 (5 to 7) planning_usd2.20 1.05 (1.02 to 1.09) 1.02 (1.00 to 1.04) 1.28 (0.94 to 1.61) usd, the one run played whole1.88 1.85 1.53 What it says:
- The outcome is steady here too.
- These three cases move less from one run to the next than
01-overdue-listand03-overdue-remindersdid in campaign 2: within 200 words, one section and one diagram, and a few cents of planning on two of them. The run that stands apart is the one that took a correction, on06-borrow-limit. - Planning alone costs 0.94 to 2.20 USD a run here, and the execution adds 0.6 to 0.9: a run stopped at the hand over (A run of the evaluations that stops at the hand over, for a campaign on planning alone #90) is two thirds of a run, as Evaluations: a run cannot stop at the hand over, so measuring a change of the interview pays for the whole loop #76 found.
- For the execution, the spread of these four cases is still one run a version:
usd, the reviews, the fixes.
Status
Measured for planning on the six cases: three runs of two cases on one version in campaign 2, two runs of three others and one of
02-reservationsin campaign 4. The execution of these four is still one run a case.Campaign 5: every case twice on one version
Campaign 5: twelve runs on
mainwith the nine pull requests of 2026-10-07 (#97 to #105).02-reservations,04-suspension,05-fine-capand06-borrow-limitare played whole, twice each;01-overdue-listand03-overdue-remindersare stopped at the hand over, twice each. Report kept:evals/reports/2026-10-07-9d5de55f57a9/report.md, compared with campaign 4. Raw data, out of git:.evals/campaign-5/. A value is the mean over the two runs of a case, with its lowest and highest when they differ.The eight runs played whole are conformant and pass every hidden test.
01-overdue-list02-reservations03-overdue-reminders04-suspension05-fine-cap06-borrow-limitquestions1 6 2 4 1 3 corrections1 0 0 0 0 0.5 (0 to 1) drafts of the plan 2 2 2 1 1 1.5 (1 to 2) sections of the body 1 4 (3 to 5) 3 2.5 (2 to 3) 1 1 diagrams 0 2.5 (2 to 3) 1 1 0 1 acceptance criteria 8.5 (8 to 9) 11.5 (11 to 12) 8 (7 to 9) 10 11.5 (10 to 13) 9.5 (7 to 12) words of the blueprint 590.5 (589 to 592) 1440 (1414 to 1466) 1186 (1120 to 1252) 861.5 (784 to 939) 722 (709 to 735) 649 (542 to 756) facts repeated in prose 3.5 (3 to 4) 11 (9 to 13) 5 (2 to 8) 3 (1 to 5) 5.5 (4 to 7) 4 reviews 1 1 1.5 (1 to 2) 1 fixes 0 0 0.5 (0 to 1) 0 planning_usd1.42 (1.37 to 1.47) 2.51 (2.42 to 2.60) 2.12 (2.12 to 2.13) 1.18 (1.10 to 1.26) 0.99 (0.98 to 1.01) 1.36 (0.98 to 1.74) usd, played whole3.62 (3.50 to 3.75) 1.91 (1.85 to 1.98) 1.86 (1.67 to 2.05) 1.97 (1.59 to 2.35) What it says:
- The outcome is steady. Over the five campaigns, 27 runs of the 28 played whole are conformant with every hidden test passed; the 28th handed back on the name of a model in a commit (A commit attribution written in the plan reaches the blueprint, and the review raises a contract break on it #86).
- On one version, the two runs of a case ask the same number of questions on the six cases, and draw pages within 4% of each other's length on three cases, 12 to 39% apart on the three others. The run that stands apart is again one that took a correction, on
06-borrow-limit. - The execution moves little: one review and no fix in seven runs of eight, a second review and one fix in a run of
05-fine-cap. A run played whole costs 1.59 to 2.35 USD on three cases, and 3.50 to 3.75 on02-reservations. - The judge, between the two runs of a case: 51 cells of 84 hold the same score, 32 are one point apart, one is two points apart (
messages.right-lengthon05-fine-cap). That is the noise of the judge alone, measured above on the same documents: 64 of 84 the same, 19 one point apart. Two runs of a case add little to it. - What moves from run to run is the form of the page, its sections, its criteria and the facts it says twice, more than what the developer is asked or what gets built.
So the report has what its rule needs: a spread on the six cases for planning, and for the execution on these four on one version, the two others by the three runs of campaign 2.
To say with it: the ledger of the campaigns over-counted the points of the weekly gauge when runs were played together, which #106 fixed. It bears on no measure of the runs.
Status
Closed: measured. Every case has two runs or three on one version, for planning and for the execution. One run a case no longer stands for a case: campaign 5 is the one the next is compared with.
Status
Closed: measured. The noise of the judge is known on every criterion. Since campaign 5 every case has two runs or three on one version, for planning and for the execution:
02-reservations,04-suspension,05-fine-capand06-borrow-limittwice in campaign 5,01-overdue-listand03-overdue-remindersthree times in campaign 2. The measures and their sources are in the comments below, campaign by campaign.What was seen
Campaign 1 is 6 runs over 6 cases. The report compares two campaigns by their spreads: "a measure moved when its values lie wholly outside those of the campaign before". With one run per case the spread is a point, so any difference in the next campaign reads as a move.
What one case already shows of the spread:
01-overdue-listwas run twice on the same version, in the trial and in the campaign. It drew one diagram, then none; went through two reviews and one fix, then one review and none.Sources
evals/README.md, "The report"..evals/trial-1/report.md,.evals/campaign-1/report.md.Why it matters
Every other issue of this campaign that says "not confirmed" says so partly for this reason.
Next
Three runs each of
01-overdue-listand03-overdue-reminderson the chain once the first fixes are merged: the two cases where the interview asked nothing, so the same runs measure the fix of the interview. The spread of the four other cases stays unknown until a campaign pays for it.