Skip to content

The report of the fourth campaign of the evaluations, kept - #92

Merged
PierreMardon merged 1 commit into
mainfrom
docs/campaign-4-report
Oct 5, 2026
Merged

PierreMardon merged 1 commit into
mainfrom
docs/campaign-4-report

Conversation

@PierreMardon

@PierreMardon PierreMardon commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

The report of the fourth campaign of the evaluations, kept for the next one to be compared with: evals/reports/2026-10-05-d139a1ca0dbe/. It measures the chain on main after the two fixes of the day (#87, #88), with the harness of #89, #90 and #91, on the four cases that no campaign had played again, or once: 04-suspension, 05-fine-cap and 06-borrow-limit, one run played whole and one stopped at the hand over each, and 02-reservations, one run stopped at the hand over. 108 sessions, 14.23 USD at list price.

What was done

  • report.md and summary.json of the campaign, as report --against <summary> --keep wrote them. Nothing else changes.
  • The summary it is compared with is the one of the first campaign, computed again from its raw data by the harness as it stands, with the count of its six blueprints (6 sessions, 0.82 USD, outside the cost of this report). No measure the kept summary of that campaign holds differs in it.
  • The first run of 06-borrow-limit was played in another folder, before the stand-in gh was in the image, to check the option that stops a run at the hand over. Its sessions and its cost were added to the ledger of the campaign, and its messages still speak of a missing gh.

What it shows

  • The three runs played whole reached conformant and passed all of their hidden acceptance tests.
  • The pull request steps are played for the first time: a draft at every hand over, its description refreshed at conformity, never marked ready. In 3 of the 6 runs that had a gh to call, the draft is opened with a line of attribution under the description the state script prints, which the next refresh removes.
  • The count of the judge follows the fix of the blueprint (A blueprint that says a fact once, and shows behavior, not code #83) on these cases: the facts said twice in prose fall from 19 to 9, from 13 to 4 and from 8 to 1.5 on 02, 04 and 05, and the details of implementation from 5 to 11 a page to 1 or 2, while blueprint.no-padding is 2 on every run of both campaigns.
  • On 06-borrow-limit the interview asks 2 and 3 questions where it asked 1, and its developer hands three of them back: interview.right-number falls from 5 to 3.5.
  • Planning costs 0.94 to 2.20 USD a run here, and a run played whole 1.53 to 1.88.

The conclusions, issue by issue, are in the issues #77 lists.

Tests

No code changes. The text rules of the repository pass on the two files; the CI runs every gate.

Refs #77

Kept for the next campaign to be compared with: `evals/reports/2026-10-05-d139a1ca0dbe/`.
It measures the chain on `main` after the two fixes of the day (#87, #88), with the harness
of #89, #90 and #91, on the four cases that no campaign had played again, or once:
`04-suspension`, `05-fine-cap` and `06-borrow-limit`, one run played whole and one stopped
at the hand over each, and `02-reservations`, one run stopped at the hand over. 108
sessions, 14.23 USD at list price.

`report.md` and `summary.json` are as `report --against <summary> --keep` wrote them, and
nothing else changes. The summary it is compared with is the one of the first campaign,
computed again from its raw data by the harness as it stands, with the count of its six
blueprints: no measure the kept summary of that campaign holds differs in it.

The first run of `06-borrow-limit` was played in another folder, before the stand-in `gh`
was in the image, to check the option that stops a run at the hand over: its sessions and
its cost were added to the ledger of the campaign.

What it shows:

- The three runs played whole reached `conformant` and passed all of their hidden
  acceptance tests.
- The pull request steps are played for the first time: a draft at every hand over, its
  description refreshed at conformity, never marked ready. In 3 runs of 7 the draft is
  opened with a line of attribution under the description the state script prints, which
  the next refresh removes.
- The count of the judge follows the fix of the blueprint (#83) on these cases: the facts
  said twice in prose fall from 19 to 9, from 13 to 4 and from 8 to 1.5 on `02`, `04` and
  `05`, and the details of implementation from 5 to 11 a page to 1 or 2, while
  `blueprint.no-padding` is 2 on every run of both campaigns.
- On `06-borrow-limit` the interview asks 2 and 3 questions where it asked 1, and its
  developer hands three of them back: `interview.right-number` falls from 5 to 3.5.
- Planning costs 0.94 to 2.20 USD a run here, and a run played whole 1.53 to 1.88.

The conclusions, issue by issue, are in the issues #77 lists.
@PierreMardon
PierreMardon merged commit b285ab2 into main Oct 5, 2026
3 checks passed
@PierreMardon
PierreMardon deleted the docs/campaign-4-report branch October 5, 2026 11:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant