Skip to content

The questions a developer hands back, counted in every run of the evaluations - #96

Merged
PierreMardon merged 1 commit into
mainfrom
feat/questions-handed-back
Oct 5, 2026
Merged

PierreMardon merged 1 commit into
mainfrom
feat/questions-handed-back

Conversation

@PierreMardon

Copy link
Copy Markdown
Contributor

What changes

A run of the evaluations now measures questions_handed_back: among the questions of the interview, how many the developer of the case handed back as not theirs to decide. The report lists it after questions.

Why

Since the interview asks for the rules a user of the feature would see (#81), it asks some that its developer hands back (#94). That issue counted the answers by hand, from the interview.md of each run, and named the count as what the harness lacked: it measured how many questions an interview asked, and nothing of what became of them.

How it is counted

By the script, from the answers run.json keeps, with no session and no judge.

  • The rules of the developer of a case already tell them to say of such a question "that it is the assistant's call". The answers kept say it in those words, or as "your call". The measure reads the two, and nothing else.
  • So the rules of the developer do not change: a campaign stays comparable with the ones before, and every run kept has the measure.
  • It counts an answer that hands back a part of its question like one that hands it back whole: a script cannot tell them apart.
  • It does not count an answer that leaves the choice in other words. One of the 31 answers kept since An interview that asks for the rules a user of the feature would see #81 does.
  • Fewer is better in the report, like corrections: the two faults of an interview, asking too much and asking too little. Whether a question handed back on 06-borrow-limit is a fault of the interview or of the case is still the choice Since the interview asks for the rules a user would see, it asks some its developer hands back #94 leaves to the developer, and one line of DIRECTION to change.

Files

  • evals/surface_evals/developer.py: the words, and hands_back.
  • evals/surface_evals/measure.py, report.py: the measure, its place in the goals table, its better way.
  • evals/README.md: the measure, and what it does not count.
  • tests/evals/test_developer.py, new, test_measure.py, test_report.py.

Measured

Read by the harness of this branch on the runs kept under .evals/, out of git, with no session:

Answers that hand a question back
Campaign 1, before #81 0 of 8
Campaign 2 1 of 4
Campaign 3 2 of 8
Campaign 4 4 of 15
The runs that measured the fix of #93 1 of 4

By case, since #81:

Case Runs Questions Handed back Runs with one
01-overdue-list 5 7 2 2
02-reservations 2 7 2 2
03-overdue-reminders 5 1 0 0
04-suspension 2 5 0 0
05-fine-cap 5 6 1 1
06-borrow-limit 2 5 3 2

Against the count made by hand in #94:

Not measured: a campaign played with this harness. The reports kept in evals/reports/ are not written again, so they do not hold the measure.

Refs #94

…luations

Since the interview asks for the rules a user of the feature would see (#81), it asks some
that its developer hands back as not theirs to decide (#94). The answers were counted by
hand, from the `interview.md` of each run: the harness measured how many questions an
interview asked, and nothing of what became of them.

`questions_handed_back` is that count, made by the script from the answers `run.json`
keeps. The rules of the developer of a case tell them to say of such a question that it is
the assistant's call, and the answers kept say it in those words or as "your call": the
measure reads the two, and nothing else. So the rules of the developer do not change, a
campaign stays comparable with the ones before, and every run kept has the measure. It
counts an answer that hands back a part of its question like one that hands it back whole,
since a script cannot tell them apart, and it does not count an answer that leaves the
choice in other words.

The report lists it after `questions`, as a measure where fewer is better, like
`corrections`: the two faults of an interview, asking too much and asking too little.

Read by the harness on the runs kept, with no session:

- Before #81, campaign 1: 0 answers of 8.
- Since #81: 1 of 4 in campaign 2, 2 of 8 in campaign 3, 4 of 15 in campaign 4, 1 of 4
  in the runs that measured the fix of #93. 8 of 31.
- By case since #81: `06-borrow-limit` 3 questions of 5, in 2 runs of 2;
  `02-reservations` 2 of 7, in 2 runs of 2; `01-overdue-list` 2 of 7, in 2 runs of 5;
  `05-fine-cap` 1 of 6, in 1 run of 5; `04-suspension` 0 of 5; `03-overdue-reminders`
  0 of 1.
- Five of the eight hand the whole question back: the four #94 counted by hand, and
  one in the runs of the fix of #93, played since. Three hand back a part, which #94
  told apart and the measure does not: what the new commands print, on
  `02-reservations` in both runs, and the tie beyond the book id, on `01-overdue-list`.
- One answer of the 31 leaves the choice without the words, and is not counted:
  "I don't rely on the order of the printed lines, so take your recommendation", on
  `03-overdue-reminders` in campaign 2.

Refs #94
@PierreMardon
PierreMardon merged commit c9d926d into main Oct 5, 2026
3 checks passed
@PierreMardon
PierreMardon deleted the feat/questions-handed-back branch October 5, 2026 13:53
PierreMardon added a commit that referenced this pull request Oct 7, 2026
… called better or worse (#98)

Since the interview asks for the rules a user of the feature would see (#81), its developer
hands some of its questions back as not theirs to decide (#94), and the harness counts them
as `questions_handed_back` (#96). The report gave that measure a better way, fewer, like
`corrections`.

The developer decided on 2026-10-07 that this is no fault of the chain. A rule of the
interview precise enough to decide without judgement was specific to the host of the
evaluations, and a general one was vague or wrong elsewhere: the stop rule of the interview
stays as it is, and so does the brief of `06-borrow-limit`, which plays a developer who
delegates. Where such a question is not asked, the blueprint is to show the assumption to
confirm (#60).

So the measure stays in the report as a sign to follow, and loses its better way: a move of
it is told as moved, like a move of `questions`.

- `evals/surface_evals/report.py`: `questions_handed_back` leaves `DIRECTION`.
- `evals/README.md`: the report tells it and never calls it better or worse, and why.
- `tests/evals/test_report.py`: holds both.

No session was played: nothing of the chain changes, and the measure is read as before.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant