Repository navigation
An interview that asks for the rules a user of the feature would see - #81
Merged
Merged
Conversation
In the first campaign of the evaluations the interview asked nothing in two cases where rules were the developer's to decide. On `01-overdue-list` the order of loans equally late was filed as an assumption and guessed. On `03-overdue-reminders` who is reminded, which notices count and the order of members were filed as "implementation matters" (#59). Refs #59
PierreMardon
force-pushed
the
fix/interview-asks-what-a-user-observes
branch
from
October 3, 2026 20:27
fa2bb66 to
f337dd8
Compare
PierreMardon
added a commit
that referenced
this pull request
Oct 3, 2026
The report of the second campaign of the evaluations, kept for the next one to be compared with: `evals/reports/2026-10-03-e830bc236e9b/`. It measures the chain after the four first fixes drawn from the first campaign (#78, #79, #80, #81), on `01-overdue-list` and `03-overdue-reminders`, three runs each: 67 sessions, 17.07 USD at list price. - `report.md` and `summary.json` of the campaign, as `report --against evals/reports/2026-10-03-6bf7c06609a6/summary.json --keep` wrote them. Nothing else changes. Refs #77
This was referenced Oct 3, 2026
Closed
Closed
PierreMardon
added a commit
that referenced
this pull request
Oct 5, 2026
…luations (#96) Since the interview asks for the rules a user of the feature would see (#81), it asks some that its developer hands back as not theirs to decide (#94). The answers were counted by hand, from the `interview.md` of each run: the harness measured how many questions an interview asked, and nothing of what became of them. `questions_handed_back` is that count, made by the script from the answers `run.json` keeps. The rules of the developer of a case tell them to say of such a question that it is the assistant's call, and the answers kept say it in those words or as "your call": the measure reads the two, and nothing else. So the rules of the developer do not change, a campaign stays comparable with the ones before, and every run kept has the measure. It counts an answer that hands back a part of its question like one that hands it back whole, since a script cannot tell them apart, and it does not count an answer that leaves the choice in other words. The report lists it after `questions`, as a measure where fewer is better, like `corrections`: the two faults of an interview, asking too much and asking too little. Read by the harness on the runs kept, with no session: - Before #81, campaign 1: 0 answers of 8. - Since #81: 1 of 4 in campaign 2, 2 of 8 in campaign 3, 4 of 15 in campaign 4, 1 of 4 in the runs that measured the fix of #93. 8 of 31. - By case since #81: `06-borrow-limit` 3 questions of 5, in 2 runs of 2; `02-reservations` 2 of 7, in 2 runs of 2; `01-overdue-list` 2 of 7, in 2 runs of 5; `05-fine-cap` 1 of 6, in 1 run of 5; `04-suspension` 0 of 5; `03-overdue-reminders` 0 of 1. - Five of the eight hand the whole question back: the four #94 counted by hand, and one in the runs of the fix of #93, played since. Three hand back a part, which #94 told apart and the measure does not: what the new commands print, on `02-reservations` in both runs, and the tie beyond the book id, on `01-overdue-list`. - One answer of the 31 leaves the choice without the words, and is not counted: "I don't rely on the order of the printed lines, so take your recommendation", on `03-overdue-reminders` in campaign 2. Refs #94
PierreMardon
added a commit
that referenced
this pull request
Oct 7, 2026
… called better or worse (#98) Since the interview asks for the rules a user of the feature would see (#81), its developer hands some of its questions back as not theirs to decide (#94), and the harness counts them as `questions_handed_back` (#96). The report gave that measure a better way, fewer, like `corrections`. The developer decided on 2026-10-07 that this is no fault of the chain. A rule of the interview precise enough to decide without judgement was specific to the host of the evaluations, and a general one was vague or wrong elsewhere: the stop rule of the interview stays as it is, and so does the brief of `06-borrow-limit`, which plays a developer who delegates. Where such a question is not asked, the blueprint is to show the assumption to confirm (#60). So the measure stays in the report as a sign to follow, and loses its better way: a move of it is told as moved, like a move of `questions`. - `evals/surface_evals/report.py`: `questions_handed_back` leaves `DIRECTION`. - `evals/README.md`: the report tells it and never calls it better or worse, and why. - `tests/evals/test_report.py`: holds both. No session was played: nothing of the chain changes, and the measure is read as before.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
In the first campaign of the evaluations the interview asked nothing in two cases where rules were the developer's to decide. On
01-overdue-listthe order of loans equally late was filed as an assumption and guessed. On03-overdue-reminderswho is reminded, which notices count and the order of members were filed as "implementation matters" (#59).What was done
Step 5 of
skills/surface-plan/SKILL.mdstopped the interview "when what remains is a matter of implementation you can decide", and nothing said what such a matter is. It now says, in this order:docs/guide.mdsays it to the developer.Tests
A test holds these sentences in their order, the guard against asking what is settled among them. The lint and the prompt tests pass locally; the CI runs every gate.
Risks
questions,interview.not-already-answeredandinterview.right-number.Planagent's to guess: its template lets it take "the assumptions taken instead of a question". Closing that path needs a way for the agent to send a question back, which this change does not add.03-overdue-remindersthe correction of the campaign came from the need itself, which states a rule its brief contradicts: no interview would have asked. That case will still count one correction.Refs #59, #77