Skip to content

Messages speak in the vocabulary of the chain #68

Description

@PierreMardon

Status

Closed: measured in campaign 5. Which words of the chain a developer is taught, and which stay inside, was decided by the developer on 2026-10-07. #104 wrote it in the rubric of the judge and #105 in the two commands that speak to the developer. In the twelve runs of campaign 5, one final message names an agent and none gives the path of a report, and messages.clear goes from 2 to 4.4. The measures and their sources are in the comments below, campaign by campaign.

What was seen

The rubric asks for messages "in words the developer can act on, without the vocabulary of the chain's insides". The judge quotes, across the six cases: "the checker", "slice", "gate", "the loop", "critical zones", "Amendment 2", "the cross-check", "the drafting agent", "autonomous passes", "executor", gates/run-02.txt, conformity.md, "0 defects, 0 deviations and 0 breaks", "revision 1".

05-fine-cap: "Review: the reviewer found 0 defects, 0 deviations and 0 breaks (reviews/pass-01.md). conformity.md proves all 9 acceptance criteria."

Why it is not confirmed

Some of these words are the developer's: the README and the guide teach the blueprint, the gates, the critical zones and the revision, and the developer approves the gates with the plan. Others belong to the chain's insides: the checker, the drafting agent, the count of passes, the paths of the reports. The rubric draws no line between the two, and the judge faults both.

Sources

Next

A decision of the developer, which nothing blocks on: which words of the chain a developer is taught, and which stay inside. Then the rubric says it, and the prompts follow if needed.

Activity

  1. PierreMardon commented on Oct 3, 2026

    @PierreMardon
    ContributorAuthor

    Read again in campaign 2

    Campaign 2: 01-overdue-list and 03-overdue-reminders, three runs each, on main with the four first fixes (#78, #79, #80, #81). Report kept: evals/reports/2026-10-03-e830bc236e9b/report.md. Raw data, out of git: .evals/campaign-2/. In the tables, a value is the mean over the three runs, with its lowest and highest when they differ.

    messages.clear: 2 on the three runs of 01-overdue-list, as before; 3 on the three runs of 03-overdue-reminders, where it was 2. The words the judge quotes are the same: "no cut", "Gate the loop will run", "cross-check found no omissions", "Slice 1", "critical-zone files", "defects, deviations and breaks".

    Status

    Unchanged, and still a choice of the developer: which words of the chain a developer is taught, and which stay inside.

  2. PierreMardon commented on Oct 5, 2026

    @PierreMardon
    ContributorAuthor

    Read again in campaign 3

    Campaign 3: five runs on three cases, played from the branch of #83. Report kept: evals/reports/2026-10-03-d765f898c9ad/report.md. Raw data, out of git: .evals/campaign-3/. Written on 2026-10-05 from the data of that campaign: this issue had not been read again after it.

    messages.clear: 2.5 (2 to 3) on 01-overdue-list, 2 on 02-reservations and on 03-overdue-reminders. The words the judge quotes are the same: "Stop 2 says "Cut", "Gates", "Critical zones", "the drafting agent" and "cross-checked with no omissions"; Stop 4 says "the third check"; Stop 6 says "slice 1", "the loop", "the gate", "contract break" and "A fixer"."

    Status

    Unchanged over three campaigns, and still a choice of the developer.

  3. PierreMardon commented on Oct 5, 2026

    @PierreMardon
    ContributorAuthor

    Campaign 4

    Campaign 4: seven runs on main with every fix so far, #87 and #88 included, and the harness of #89, #90 and #91, on the cases no campaign had played again, or once. 04-suspension, 05-fine-cap and 06-borrow-limit: one run played whole and one stopped at the hand over each. 02-reservations: one run stopped at the hand over. Report kept: evals/reports/2026-10-05-d139a1ca0dbe/report.md, compared with campaign 1, the only earlier runs of three of these cases. Raw data, out of git: .evals/campaign-4/. A value is the mean over the runs of a case, with its lowest and highest when they differ.

    messages.clear: 2 on the three runs played whole, where it was 3, 2 and 3 in campaign 1. Of planning alone, on the four runs stopped at the hand over: 3, 2, 2 and 2. The words the judge quotes are the same: "an 'extractor', a 'cross-check', a 'gate' the 'loop' will run", "slices", "conformant", "critical zones", "deviations or breaks".

    Status

    Unchanged over four campaigns, and still a choice of the developer.

  4. PierreMardon commented on Oct 7, 2026

    @PierreMardon
    ContributorAuthor

    Decided by the developer

    Recorded on 2026-10-07. The words of the chain a message may use are those the README and the guide teach the developer: the blueprint, the gates, the critical zones, a revision, a slice, conformant, the cross-check, the loop, the passes. The chain's insides stay out of the messages: the names of the agents, such as the checker, the extractor or the executor, and the paths of the reports.

    Status

    Open, decided, not done yet. What is left to write: the line in the criterion messages.clear of evals/judge/rubric.md, then the prompts where a message still names an agent or the path of a report, and a campaign to read the score again.

  5. PierreMardon commented on Oct 7, 2026

    @PierreMardon
    ContributorAuthor

    Done by #104 and #105, not measured

    Merged on 2026-10-07.

    The rubric, #104. messages.clear lists the words of the chain that its README and its guide teach, which a message may use, and names the insides a final message keeps out: the name of any agent but the reviewer, such as the checker, the extractor or the executor, and the paths of the reports. The judge has neither document, so the criterion holds the list.

    The prompts, #105. /surface-plan and /surface-execute hold the same rule and the same list, where each says in which language it speaks: "Say what a review found or which gate failed, not which agent said so nor in which file." The steps that asked for a path no longer do, at a stop of the loop and at the ceiling, and two steps of /surface-plan no longer hand the session an agent to name. A test holds the list equal in the two commands and in the rubric. /surface-status is left as it is: it relays what the state script answers.

    Beyond the words of the decision, to say when read

    • The list. The decision gives the rule and names nine words. The list holds seven more, each read in the guide: the plan, the cut, an amendment, a review and its reviewer, a defect, the ceiling, a contract break and its plan change proposal.
    • The reviewer. The decision gives the checker, the extractor and the executor as the insides. The guide speaks of "the reviewer" twice, and the description of the pull request, which the state script writes, says "dismissed by a reviewer": so the reviewer is among the words taught. That description is not faulted for the files it links.
    • A report the developer has to open. A report is named by its path when the developer has to open it to go on, or asks for it: without it, a session that stops on a report that fails a gate could not say which one. A plan change proposal is no report: it is the developer's to decide on, and the guide names its folder.
    • A word not taught. "Deviation" is in neither document. The prompt of /surface-execute has it said as what it is: code that departs from the plan while the blueprint stays true.

    What to read on the next campaign

    messages.clear, by the rubric as it stands, and the messages themselves: whether an agent or the path of a report is still named, and whether a stop leaves out something the developer needed.

    Status

    Open: done by #104 and #105, not measured.

  6. PierreMardon commented on Oct 7, 2026

    @PierreMardon
    ContributorAuthor

    Measured in campaign 5

    Campaign 5: twelve runs on main with the nine pull requests of 2026-10-07 (#97 to #105), the six cases twice each: 02-reservations, 04-suspension, 05-fine-cap and 06-borrow-limit played whole, 01-overdue-list and 03-overdue-reminders stopped at the hand over. The eight runs played whole are conformant and pass every hidden test. Report kept: evals/reports/2026-10-07-9d5de55f57a9/report.md, compared with campaign 4. Raw data, out of git: .evals/campaign-5/. A value is the mean over the runs, with its lowest and highest when they differ.

    Campaign 4 Campaign 5
    messages.clear, runs played whole 2 on the three runs 4.4 (3 to 5) on eight
    messages-in-planning.clear, runs stopped at the hand over 2.25 (2 to 3) on four 3.75 (3 to 4) on four

    Read in the final messages of the twelve runs, 73 stops:

    • One names an agent. 01-overdue-list, a hand over: "The plan also warns the executor not to add a test or README sentence about it."
    • None gives the path of a report. In the three runs of campaign 4 played whole, each hand back at conformity gave one or named an agent: "The review is in docs/plans/2026-10-05-fine-limit/reviews/pass-01.md", "The slice 2 executor ran one cd", "conformity.md keeps that proof".
    • The judge no longer faults a word the guide teaches. Its lowest judgement on the criterion is another matter: the hand back says that marking the pull request ready "triggers your CI", in a host whose hand over said that it has none. It is in the eight runs, and it is told in Hand overs repeat what the developer is sent to read, and tell the chain's own steps #67.

    The score moved by the rubric and by the prompts together, since #104 changed what the judge counts: what tells the prompts apart is the count above, one name in twelve runs.

    Nothing a developer needed was left out for want of a path, as far as the messages show: no stop of these runs was one where a report had to be opened to go on.

    Status

    Closed: measured. The messages speak in the words the README and the guide teach, and keep the chain's insides out.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions