Repository navigation
The questions a developer hands back, counted in every run of the evaluations - #96
Merged
Merged
Conversation
…luations Since the interview asks for the rules a user of the feature would see (#81), it asks some that its developer hands back as not theirs to decide (#94). The answers were counted by hand, from the `interview.md` of each run: the harness measured how many questions an interview asked, and nothing of what became of them. `questions_handed_back` is that count, made by the script from the answers `run.json` keeps. The rules of the developer of a case tell them to say of such a question that it is the assistant's call, and the answers kept say it in those words or as "your call": the measure reads the two, and nothing else. So the rules of the developer do not change, a campaign stays comparable with the ones before, and every run kept has the measure. It counts an answer that hands back a part of its question like one that hands it back whole, since a script cannot tell them apart, and it does not count an answer that leaves the choice in other words. The report lists it after `questions`, as a measure where fewer is better, like `corrections`: the two faults of an interview, asking too much and asking too little. Read by the harness on the runs kept, with no session: - Before #81, campaign 1: 0 answers of 8. - Since #81: 1 of 4 in campaign 2, 2 of 8 in campaign 3, 4 of 15 in campaign 4, 1 of 4 in the runs that measured the fix of #93. 8 of 31. - By case since #81: `06-borrow-limit` 3 questions of 5, in 2 runs of 2; `02-reservations` 2 of 7, in 2 runs of 2; `01-overdue-list` 2 of 7, in 2 runs of 5; `05-fine-cap` 1 of 6, in 1 run of 5; `04-suspension` 0 of 5; `03-overdue-reminders` 0 of 1. - Five of the eight hand the whole question back: the four #94 counted by hand, and one in the runs of the fix of #93, played since. Three hand back a part, which #94 told apart and the measure does not: what the new commands print, on `02-reservations` in both runs, and the tie beyond the book id, on `01-overdue-list`. - One answer of the 31 leaves the choice without the words, and is not counted: "I don't rely on the order of the printed lines, so take your recommendation", on `03-overdue-reminders` in campaign 2. Refs #94
This was referenced Oct 5, 2026
Closed
PierreMardon
added a commit
that referenced
this pull request
Oct 7, 2026
… called better or worse (#98) Since the interview asks for the rules a user of the feature would see (#81), its developer hands some of its questions back as not theirs to decide (#94), and the harness counts them as `questions_handed_back` (#96). The report gave that measure a better way, fewer, like `corrections`. The developer decided on 2026-10-07 that this is no fault of the chain. A rule of the interview precise enough to decide without judgement was specific to the host of the evaluations, and a general one was vague or wrong elsewhere: the stop rule of the interview stays as it is, and so does the brief of `06-borrow-limit`, which plays a developer who delegates. Where such a question is not asked, the blueprint is to show the assumption to confirm (#60). So the measure stays in the report as a sign to follow, and loses its better way: a move of it is told as moved, like a move of `questions`. - `evals/surface_evals/report.py`: `questions_handed_back` leaves `DIRECTION`. - `evals/README.md`: the report tells it and never calls it better or worse, and why. - `tests/evals/test_report.py`: holds both. No session was played: nothing of the chain changes, and the measure is read as before.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes
A run of the evaluations now measures
questions_handed_back: among the questions of the interview, how many the developer of the case handed back as not theirs to decide. The report lists it afterquestions.Why
Since the interview asks for the rules a user of the feature would see (#81), it asks some that its developer hands back (#94). That issue counted the answers by hand, from the
interview.mdof each run, and named the count as what the harness lacked: it measured how many questions an interview asked, and nothing of what became of them.How it is counted
By the script, from the answers
run.jsonkeeps, with no session and no judge.corrections: the two faults of an interview, asking too much and asking too little. Whether a question handed back on06-borrow-limitis a fault of the interview or of the case is still the choice Since the interview asks for the rules a user would see, it asks some its developer hands back #94 leaves to the developer, and one line ofDIRECTIONto change.Files
evals/surface_evals/developer.py: the words, andhands_back.evals/surface_evals/measure.py,report.py: the measure, its place in the goals table, its better way.evals/README.md: the measure, and what it does not count.tests/evals/test_developer.py, new,test_measure.py,test_report.py.Measured
Read by the harness of this branch on the runs kept under
.evals/, out of git, with no session:By case, since #81:
01-overdue-list02-reservations03-overdue-reminders04-suspension05-fine-cap06-borrow-limitAgainst the count made by hand in #94:
01-overdue-listand three on06-borrow-limit, and one on05-fine-capin the runs of the fix of A draft pull request is opened with a line of attribution under the description the state script prints #93, played since.02-reservations, in both runs: "What the three commands print is your call, since I only rely on their exit codes." On01-overdue-list: "my rule only says book id. It doesn't say anything about member id, so the rest is your call."03-overdue-remindersin campaign 2.Not measured: a campaign played with this harness. The reports kept in
evals/reports/are not written again, so they do not hold the measure.Refs #94