Skip to content

Evaluations: one run per case, so the spread between runs is unknown and the next report cannot say what moved #74

Description

@PierreMardon

Status

Closed: measured. The noise of the judge is known on every criterion. Since campaign 5 every case has two runs or three on one version, for planning and for the execution: 02-reservations, 04-suspension, 05-fine-cap and 06-borrow-limit twice in campaign 5, 01-overdue-list and 03-overdue-reminders three times in campaign 2. The measures and their sources are in the comments below, campaign by campaign.

What was seen

Campaign 1 is 6 runs over 6 cases. The report compares two campaigns by their spreads: "a measure moved when its values lie wholly outside those of the campaign before". With one run per case the spread is a point, so any difference in the next campaign reads as a move.

What one case already shows of the spread: 01-overdue-list was run twice on the same version, in the trial and in the campaign. It drew one diagram, then none; went through two reviews and one fix, then one review and none.

Sources

Why it matters

Every other issue of this campaign that says "not confirmed" says so partly for this reason.

Next

Three runs each of 01-overdue-list and 03-overdue-reminders on the chain once the first fixes are merged: the two cases where the interview asked nothing, so the same runs measure the fix of the interview. The spread of the four other cases stays unknown until a campaign pays for it.

Activity

  1. PierreMardon commented on Oct 3, 2026

    @PierreMardon
    ContributorAuthor

    Measured: what the judge itself adds to the spread

    The six runs of campaign 1 were judged again with no session of the chain: twice more by the same judge (opus), once by another model (sonnet). Same documents, same rubric: whatever moves here is the judge, not the chain.

    Each cell: the three passes of the same judge (the campaign's, then the two new ones), then the other model after the slash. none is an answer that was not the JSON asked for, which the harness refused.

    Criterion 01 02 03 04 05 06 Cells that moved
    blueprint.cut 3 2 3 / 3 3 4 3 / 4 3 4 4 / 3 3 3 3 / 4 4 4 4 / 4 3 3 3 / 3 3 of 6
    blueprint.decidable 5 5 5 / 4 4 4 4 / 4 5 5 5 / 4 5 5 5 / 4 4 4 4 / 4 5 5 5 / 4 0 of 6
    blueprint.diagrams 5 5 5 / 4 4 4 4 / 3 3 4 3 / 3 4 4 4 / 2 3 3 4 / 3 5 5 5 / 3 2 of 6
    blueprint.no-padding 2 2 2 / 2 2 2 2 / 2 2 2 2 / 2 2 2 2 / 2 2 2 2 / 2 2 2 2 / 3 0 of 6
    following.no-noise 2 2 2 / none 2 2 2 / 3 2 2 2 / 3 2 2 2 / 3 2 2 2 / 2 3 2 2 / 3 1 of 6
    following.whose-turn 4 4 3 / none 3 3 3 / 4 4 4 5 / 4 4 4 4 / 4 5 5 5 / 4 4 5 3 / 4 3 of 6
    following.why-stopped 5 5 5 / none 5 5 5 / 5 5 5 5 / 5 5 5 5 / 5 5 5 5 / 5 5 5 5 / 5 0 of 6
    interview.no-invented-rule 4 4 4 / 3 4 4 3 / none 4 4 4 / 4 5 5 5 / 4 4 4 4 / none 5 5 5 / 5 1 of 6
    interview.not-already-answered 5 5 5 / 5 5 4 4 / none 5 5 5 / 5 5 5 5 / 5 5 5 5 / none 4 5 4 / 5 2 of 6
    interview.questions-matter 4 4 4 / 4 5 5 5 / none 3 3 3 / 3 5 5 5 / 5 5 5 5 / none 5 4 5 / 5 1 of 6
    interview.right-number 3 3 3 / 2 4 4 3 / none 2 2 2 / 2 5 5 5 / 4 4 4 5 / none 5 5 5 / 5 2 of 6
    messages.clear 2 2 2 / 3 2 2 2 / 3 2 2 2 / 3 3 2 2 / 3 2 2 2 / 3 3 2 3 / 3 2 of 6
    messages.next-step 3 3 3 / 4 3 3 3 / 3 3 3 3 / 3 2 3 3 / 4 3 3 4 / 4 3 3 3 / 4 2 of 6
    messages.right-length 3 3 3 / 2 3 3 3 / 3 3 2 3 / 2 3 3 3 / 3 3 3 3 / 3 3 3 3 / 4 1 of 6

    What it says:

    • Over the 84 cells, the three passes of the same judge give the same score in 64, differ by one point in 19, and by two points in one (following.whose-turn on 06-borrow-limit).
    • Never moved: blueprint.no-padding, blueprint.decidable, following.why-stopped. Moved most: blueprint.cut and following.whose-turn, in 3 cells of 6.
    • So one point of difference on one run is inside the judge's own noise for most criteria. A change between two campaigns is worth reading when it holds over several runs, or exceeds one point.
    • The other model agrees within one point nearly everywhere. It is a little kinder on the messages (messages.clear 3 on every case, where the first judge gives 2 on four) and harder on the diagrams.

    Cost: 48 sessions and 3.14 USD for the two passes of opus, 24 sessions and 0.59 USD for the pass of sonnet. Raw data, out of git: .evals/rejudge-opus/, .evals/rejudge-sonnet/.

    Still unknown: the spread between runs of the chain itself, which the next campaign measures on a few cases.

  2. PierreMardon commented on Oct 3, 2026

    @PierreMardon
    ContributorAuthor

    Measured in campaign 2: a first spread between runs, on two cases

    Campaign 2: 01-overdue-list and 03-overdue-reminders, three runs each, on main with the four first fixes (#78, #79, #80, #81). Report kept: evals/reports/2026-10-03-e830bc236e9b/report.md. Raw data, out of git: .evals/campaign-2/. In the tables, a value is the mean over the three runs, with its lowest and highest when they differ.

    Three runs of the same need, on the same version of the chain:

    01-overdue-list 03-overdue-reminders
    conformant, acceptance 1, 1 1, 1
    questions 1 0.33 (0 to 1)
    corrections 1 (0 to 2) 1
    drafts of the plan 2 (1 to 3) 2
    sections of the body 1.33 (1 to 2) 1.67 (1 to 3)
    diagrams 0.33 (0 to 1) 1.67 (1 to 3)
    words of the blueprint 984 (669 to 1545) 1310 (1235 to 1425)
    acceptance criteria 8 (6 to 10) 10 (9 to 12)
    usd 2.34 (1.41 to 3.60) 2.84 (2.65 to 3.11)
    scores of the judge within one point, on every criterion but interview.right-number of 01 (2 to 5) and interview.no-invented-rule (3 to 5, and 2 to 4)

    What it says:

    • The outcome is steady: six runs, six conformant, every hidden test passed.
    • The path is not: the same need gives a blueprint of 669 or of 1545 words, one section or three, none or two corrections, and costs from one to two and a half times as much.
    • So the form of the blueprint and the count of corrections cannot be read on one run. The first campaign's "1 section, as expected" and "2 corrections" were one draw each.
    • With the noise of the judge measured above, the report's rule now has something to stand on for these two cases: it listed 7 moves between the two campaigns. Three are what the fixes were written for: the question asked on 01-overdue-list, the score of that interview, and the next step of the messages on 03-overdue-reminders. The four others, the cut and the clarity on 03-overdue-reminders, its count of criteria and of words, have no fix to explain them: each is within one point of the judge or one draw of the model, and three runs are still few.

    Status

    Measured on two cases of six, three runs each. The four other cases are still one run.

  3. PierreMardon commented on Oct 3, 2026

    @PierreMardon
    ContributorAuthor

    Campaign 3 adds two runs of each case, on another version

    Not the same version of the chain as campaign 2, so not the same spread, but the same lesson: on 01-overdue-list one run cost 2.23 USD and the other 4.61, one page had 730 words and the other 1463, one run was conformant and the other handed back. The difference between the two runs is an answer of the simulated developer (#82), then the name of a model in a commit (#86).

    So part of the spread is not the chain's: it comes from the developer played by a model, whose answers change from run to run on what its brief leaves open. A spread between runs measures the chain and its developer together.

  4. PierreMardon commented on Oct 5, 2026

    @PierreMardon
    ContributorAuthor

    Campaign 4: two runs of planning on the cases never played again

    Campaign 4: seven runs on main with every fix so far, #87 and #88 included, and the harness of #89, #90 and #91, on the cases no campaign had played again, or once. 04-suspension, 05-fine-cap and 06-borrow-limit: one run played whole and one stopped at the hand over each. 02-reservations: one run stopped at the hand over. Report kept: evals/reports/2026-10-05-d139a1ca0dbe/report.md, compared with campaign 1, the only earlier runs of three of these cases. Raw data, out of git: .evals/campaign-4/. A value is the mean over the runs of a case, with its lowest and highest when they differ.

    The three runs played whole are conformant and pass every hidden test: 12 of 12, 11 of 11 and 9 of 9.

    02-reservations, one run 04-suspension 05-fine-cap 06-borrow-limit
    questions 3 2.5 (2 to 3) 1 2.5 (2 to 3)
    corrections 0 0 0 0.5 (0 to 1)
    drafts of the plan 2 1 1 1.5 (1 to 2)
    sections of the body 6 3 1.5 (1 to 2) 1
    diagrams 3 1.5 (1 to 2) 1 1
    acceptance criteria 9 10 (9 to 11) 8.5 (8 to 9) 9 (8 to 10)
    words of the blueprint 1448 928.5 (858 to 999) 780 (684 to 876) 713 (621 to 805)
    facts repeated in prose 9 4 (3 to 5) 1.5 (1 to 2) 6 (5 to 7)
    planning_usd 2.20 1.05 (1.02 to 1.09) 1.02 (1.00 to 1.04) 1.28 (0.94 to 1.61)
    usd, the one run played whole 1.88 1.85 1.53

    What it says:

    Status

    Measured for planning on the six cases: three runs of two cases on one version in campaign 2, two runs of three others and one of 02-reservations in campaign 4. The execution of these four is still one run a case.

  5. PierreMardon commented on Oct 7, 2026

    @PierreMardon
    ContributorAuthor

    Campaign 5: every case twice on one version

    Campaign 5: twelve runs on main with the nine pull requests of 2026-10-07 (#97 to #105). 02-reservations, 04-suspension, 05-fine-cap and 06-borrow-limit are played whole, twice each; 01-overdue-list and 03-overdue-reminders are stopped at the hand over, twice each. Report kept: evals/reports/2026-10-07-9d5de55f57a9/report.md, compared with campaign 4. Raw data, out of git: .evals/campaign-5/. A value is the mean over the two runs of a case, with its lowest and highest when they differ.

    The eight runs played whole are conformant and pass every hidden test.

    01-overdue-list 02-reservations 03-overdue-reminders 04-suspension 05-fine-cap 06-borrow-limit
    questions 1 6 2 4 1 3
    corrections 1 0 0 0 0 0.5 (0 to 1)
    drafts of the plan 2 2 2 1 1 1.5 (1 to 2)
    sections of the body 1 4 (3 to 5) 3 2.5 (2 to 3) 1 1
    diagrams 0 2.5 (2 to 3) 1 1 0 1
    acceptance criteria 8.5 (8 to 9) 11.5 (11 to 12) 8 (7 to 9) 10 11.5 (10 to 13) 9.5 (7 to 12)
    words of the blueprint 590.5 (589 to 592) 1440 (1414 to 1466) 1186 (1120 to 1252) 861.5 (784 to 939) 722 (709 to 735) 649 (542 to 756)
    facts repeated in prose 3.5 (3 to 4) 11 (9 to 13) 5 (2 to 8) 3 (1 to 5) 5.5 (4 to 7) 4
    reviews 1 1 1.5 (1 to 2) 1
    fixes 0 0 0.5 (0 to 1) 0
    planning_usd 1.42 (1.37 to 1.47) 2.51 (2.42 to 2.60) 2.12 (2.12 to 2.13) 1.18 (1.10 to 1.26) 0.99 (0.98 to 1.01) 1.36 (0.98 to 1.74)
    usd, played whole 3.62 (3.50 to 3.75) 1.91 (1.85 to 1.98) 1.86 (1.67 to 2.05) 1.97 (1.59 to 2.35)

    What it says:

    • The outcome is steady. Over the five campaigns, 27 runs of the 28 played whole are conformant with every hidden test passed; the 28th handed back on the name of a model in a commit (A commit attribution written in the plan reaches the blueprint, and the review raises a contract break on it #86).
    • On one version, the two runs of a case ask the same number of questions on the six cases, and draw pages within 4% of each other's length on three cases, 12 to 39% apart on the three others. The run that stands apart is again one that took a correction, on 06-borrow-limit.
    • The execution moves little: one review and no fix in seven runs of eight, a second review and one fix in a run of 05-fine-cap. A run played whole costs 1.59 to 2.35 USD on three cases, and 3.50 to 3.75 on 02-reservations.
    • The judge, between the two runs of a case: 51 cells of 84 hold the same score, 32 are one point apart, one is two points apart (messages.right-length on 05-fine-cap). That is the noise of the judge alone, measured above on the same documents: 64 of 84 the same, 19 one point apart. Two runs of a case add little to it.
    • What moves from run to run is the form of the page, its sections, its criteria and the facts it says twice, more than what the developer is asked or what gets built.

    So the report has what its rule needs: a spread on the six cases for planning, and for the execution on these four on one version, the two others by the three runs of campaign 2.

    To say with it: the ledger of the campaigns over-counted the points of the weekly gauge when runs were played together, which #106 fixed. It bears on no measure of the runs.

    Status

    Closed: measured. Every case has two runs or three on one version, for planning and for the execution. One run a case no longer stands for a case: campaign 5 is the one the next is compared with.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions