Skip to content

Since the interview asks for the rules a user would see, it asks some its developer hands back #94

Description

@PierreMardon

Status

Closed: no fault of the chain. Decided by the developer on 2026-10-07: the stop rule of the interview stays as it is, and so does the brief of 06-borrow-limit, which plays a developer who delegates. questions_handed_back stays in the report as a sign to follow, with no better way since #98. Where such a question is not asked, the blueprint shows the assumption of the plan, since #100 (#60). The measures and their sources are in the comments below.

What was seen

06-borrow-limit, two runs:

  • Run 1, Q2, "When a borrow breaks several rules at once, which refusal is reported?": "That's the assistant's call, since the order of the checks isn't mine to decide. I'll take your recommendation, A."
  • Run 2, Q2, the same question with the wording of the refusal: "That's the assistant's call, and I'm fine with its recommendation, A. The wording of the refusal is free as long as it tells the desk the member holds too many books."
  • Run 2, Q3, "What happens for a member whose category is neither student nor staff?": "There are only the two categories, student and staff, so that case shouldn't come up. What happens with a bad value is your call, and I'll take your recommendation: A."
06-borrow-limit Campaign 1, before #81 Campaign 4
questions 1 2.5 (2 to 3)
interview.questions-matter 5 3
interview.right-number 5 3.5 (3 to 4)
planning_usd 0.85 1.28 (0.94 to 1.61)

The judge, interview.questions-matter 3: "Q2 asks about the internal order of the checks, which the developer said was not theirs to decide, so it should have been an implementation decision rather than a question."

Counted in the answers kept in the interview.md of every run since #81: a question handed back whole in 0 answers of 4 in campaign 2, 1 of 8 in campaign 3 (01-overdue-list: "That's your call, and I'm going with your recommendation, A"), 3 of 15 in campaign 4, all three on this case. An answer also hands back a part of its question in each campaign, such as "What the three commands print is your call, since I only rely on their exit codes" on 02-reservations, twice.

On the three other cases of campaign 4 nothing of the kind: interview.questions-matter is 5 on 04-suspension and on 05-fine-cap, 4 on 02-reservations.

Why it is not confirmed

  • The questions follow the rule An interview that asks for the rules a user of the feature would see #81 wrote. Which refusal the desk reads when a borrow breaks several rules is "a rule that a user of the feature [...] would see applied", and where the new check goes among the existing ones is settled by nothing in the code.
  • The judge pulls both ways on that very rule. On 04-suspension, where the order of the new check was not asked, interview.no-invented-rule is 4 on both runs: "One visible rule was not asked and is not offered as an assumption to confirm: the debt check runs after the existing checks". On 06-borrow-limit, where it was asked, the question is faulted.
  • What a simulated developer hands back is what its brief says. The brief of this case lists, under "Not mine to decide", "The order of this check among the other checks of borrow", and says of the categories "There are only these two". So the prompt and the case do not hold the same idea of whose that order is: by the rule of An interview that asks for the rules a user of the feature would see #81 it is the developer's, by the case it is the implementer's. A developer at a real desk may care which refusal is read. A run measures the chain and its developer together (Evaluations: one run per case, so the spread between runs is unknown and the next report cannot say what moved #74).
  • One case, two runs.

So before anything is fixed there is a choice, and it is the developer's: whether the order in which refusals are reported is theirs to decide. Then either the stop rule of the interview or the brief of the case says so.

Why it matters

A question is a turn of the developer, and here a third more of the cost of planning. An interview that asks what its developer does not care about also teaches them to answer "your call", which is the answer a guessed rule hides behind.

Sources

Next

Read questions_handed_back per case on the next campaign: the harness counts the answers that hand a question back since #96, from what run.json keeps. If the sign holds on other cases, what the stop rule of the interview lacks is a word on what the precedent of the code nearly settles, such as the place of a new check among the existing ones. Tied to #60, which asks the opposite question: how a rule nobody asked reaches the developer.

Activity

  1. PierreMardon commented on Oct 5, 2026

    @PierreMardon
    ContributorAuthor

    Counted by the harness since #96, on every run kept

    questions_handed_back (#96, merged): the answers of the developer of a case that say of a question, or of a part of it, that it is "the assistant's call" or "your call", read by the script from run.json. The rules of the developer did not change, and no session was played: the measure is read on the runs kept under .evals/, out of git.

    Answers that hand a question back
    Campaign 1, before #81 0 of 8
    Campaign 2 1 of 4
    Campaign 3 2 of 8
    Campaign 4 4 of 15
    The runs that measured the fix of #93 1 of 4

    By case, since #81:

    Case Runs Questions Handed back Runs with one
    01-overdue-list 5 7 2 2
    02-reservations 2 7 2 2
    03-overdue-reminders 5 1 0 0
    04-suspension 2 5 0 0
    05-fine-cap 5 6 1 1
    06-borrow-limit 2 5 3 2

    The eight answers are of three kinds, which the measure counts alike:

    • A whole question of the interview: three, all on 06-borrow-limit, in its two runs. They are what this issue opened on.
    • A part of a question: three. 02-reservations, in both of its runs: "What the three commands print is your call, since I only rely on their exit codes", to a question that also asked what return prints (campaign 3) or when cancel is refused (campaign 4), which the developer did answer. 01-overdue-list, campaign 2, on the tie beyond the book id: "It doesn't say anything about member id, so the rest is your call."
    • A question the developer's own correction raised: two. 01-overdue-list in campaign 3 and 05-fine-cap in the runs of the fix of A draft pull request is opened with a line of attribution under the description the state script prints #93. The developer sent the blueprint back, the session asked what the correction left open, what becomes of an extra argument, and the fine of a book whose price is 0, where "Point 1 contradicts itself", and the developer handed that back. They are asked while a revision is drafted, not in the interview, and are counted like the others, as questions counts them.

    One answer leaves the choice without the words, and is not counted: the only question of 03-overdue-reminders since #81, campaign 2, run 3, "I don't rely on the order of the printed lines, so take your recommendation."

    What it says:

    • Nothing of the kind before An interview that asks for the rules a user of the feature would see #81, in 8 answers. Since, 8 answers of 31 hand something back.
    • A whole question of the interview handed back is still seen on one case only. What shows elsewhere is weaker: a question worth asking that carries a part its developer does not care about, what a command prints, on 02-reservations in 2 runs of 2.
    • Two of the eight come from the developer the harness plays, whose correction opened the question. They say nothing of the interview.

    Sources: .evals/campaign-1/ to campaign-4/ and .evals/fix-93/, runs/*/run-*/run.json, the stops of kind answer. The measure: evals/surface_evals/developer.py, hands_back.

    Status

    Open, not confirmed: a whole question of the interview handed back on one case, a part of one on two others. The harness counts it from now on, in every campaign.

  2. PierreMardon commented on Oct 7, 2026

    @PierreMardon
    ContributorAuthor

    Decided by the developer: the chain does not change

    Recorded on 2026-10-07. The stop rule of the interview stays as it is. Two wordings of an exception were put to the developer, who set both aside:

    • "what the precedent of the code nearly settles is not asked": nearly asks the session for a judgement, which is what makes one run ask and another not;
    • "the place of a new check among the existing ones is not asked": precise, and specific to the host of the evaluations. It comes from two cases of the same synthetic project, both on the refusals of borrow. Made general, "where the new rule and an existing one apply to the same case, the existing one goes first", it is harmless for a refusal and a business rule for two discounts or two access rights.

    So the sign does not call for a change of the prompts:

    The brief of 06-borrow-limit stays: it plays a developer who delegates. questions_handed_back stays in the report as a sign to follow, and is no longer called better or worse (#96 gave it a better way, fewer).

    Status

    Decided: no fault of the chain. To close once the measure has no better way in evals/surface_evals/report.py.

  3. PierreMardon commented on Oct 7, 2026

    @PierreMardon
    ContributorAuthor

    Done by #98

    Merged on 2026-10-07. questions_handed_back leaves the measures that have a better way in evals/surface_evals/report.py: a move of it is told as moved, like a move of questions, and never called better or worse. evals/README.md says why: a developer may leave a choice to the session, and the brief of a case may play one who does.

    No session was played: nothing of the chain changes, and the measure is read from the runs as before.

    The other half of the decision is done too: where the interview does not ask such a rule, the plan marks it and the blueprint shows it as an assumption of the plan, which approving confirms (#100, for #60).

    Status

    Closed: no fault of the chain.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions