Skip to content

An interview that asks for the rules a user of the feature would see - #81

Merged
PierreMardon merged 1 commit into
mainfrom
fix/interview-asks-what-a-user-observes
Oct 3, 2026
Merged

PierreMardon merged 1 commit into
mainfrom
fix/interview-asks-what-a-user-observes

Conversation

@PierreMardon

Copy link
Copy Markdown
Contributor

In the first campaign of the evaluations the interview asked nothing in two cases where rules were the developer's to decide. On 01-overdue-list the order of loans equally late was filed as an assumption and guessed. On 03-overdue-reminders who is reminded, which notices count and the order of members were filed as "implementation matters" (#59).

What was done

Step 5 of skills/surface-plan/SKILL.md stopped the interview "when what remains is a matter of implementation you can decide", and nothing said what such a matter is. It now says, in this order:

  • what makes a point the developer's: a rule that a user of the feature, or a program that reads what it writes, would see applied: an order, ties included, a number, a frequency or a limit, who or what is counted or left out, what is refused;
  • what is not asked: what the specs or the code settle, its closest precedent included;
  • what neither settles is a question, however obvious its answer looks: the pick of the session is its recommendation, never an assumption of the plan;
  • a matter of implementation is what neither a user nor such a program would see: where the code lives, how the code is named, how a settled result is computed.

docs/guide.md says it to the developer.

Tests

A test holds these sentences in their order, the guard against asking what is settled among them. The lint and the prompt tests pass locally; the CI runs every gate.

Risks

  • The opposite fault, an interview that asks to be thorough: the next campaign of the evaluations watches questions, interview.not-already-answered and interview.right-number.
  • A rule first met when the plan is drafted is still the Plan agent's to guess: its template lets it take "the assumptions taken instead of a question". Closing that path needs a way for the agent to send a question back, which this change does not add.
  • On 03-overdue-reminders the correction of the campaign came from the need itself, which states a rule its brief contradicts: no interview would have asked. That case will still count one correction.

Refs #59, #77

In the first campaign of the evaluations the interview asked nothing in two cases where
rules were the developer's to decide. On `01-overdue-list` the order of loans equally late
was filed as an assumption and guessed. On `03-overdue-reminders` who is reminded, which
notices count and the order of members were filed as "implementation matters" (#59).

Refs #59
@PierreMardon
PierreMardon force-pushed the fix/interview-asks-what-a-user-observes branch from fa2bb66 to f337dd8 Compare October 3, 2026 20:27
@PierreMardon
PierreMardon merged commit 065e1d3 into main Oct 3, 2026
3 checks passed
@PierreMardon
PierreMardon deleted the fix/interview-asks-what-a-user-observes branch October 3, 2026 20:35
PierreMardon added a commit that referenced this pull request Oct 3, 2026
The report of the second campaign of the evaluations, kept for the next one to be compared
with: `evals/reports/2026-10-03-e830bc236e9b/`. It measures the chain after the four first
fixes drawn from the first campaign (#78, #79, #80, #81), on `01-overdue-list` and
`03-overdue-reminders`, three runs each: 67 sessions, 17.07 USD at list price.

- `report.md` and `summary.json` of the campaign, as `report --against
  evals/reports/2026-10-03-6bf7c06609a6/summary.json --keep` wrote them. Nothing else
  changes.

Refs #77
This was referenced Oct 3, 2026
PierreMardon added a commit that referenced this pull request Oct 5, 2026
…luations (#96)

Since the interview asks for the rules a user of the feature would see (#81), it asks some
that its developer hands back as not theirs to decide (#94). The answers were counted by
hand, from the `interview.md` of each run: the harness measured how many questions an
interview asked, and nothing of what became of them.

`questions_handed_back` is that count, made by the script from the answers `run.json`
keeps. The rules of the developer of a case tell them to say of such a question that it is
the assistant's call, and the answers kept say it in those words or as "your call": the
measure reads the two, and nothing else. So the rules of the developer do not change, a
campaign stays comparable with the ones before, and every run kept has the measure. It
counts an answer that hands back a part of its question like one that hands it back whole,
since a script cannot tell them apart, and it does not count an answer that leaves the
choice in other words.

The report lists it after `questions`, as a measure where fewer is better, like
`corrections`: the two faults of an interview, asking too much and asking too little.

Read by the harness on the runs kept, with no session:

- Before #81, campaign 1: 0 answers of 8.
- Since #81: 1 of 4 in campaign 2, 2 of 8 in campaign 3, 4 of 15 in campaign 4, 1 of 4
  in the runs that measured the fix of #93. 8 of 31.
- By case since #81: `06-borrow-limit` 3 questions of 5, in 2 runs of 2;
  `02-reservations` 2 of 7, in 2 runs of 2; `01-overdue-list` 2 of 7, in 2 runs of 5;
  `05-fine-cap` 1 of 6, in 1 run of 5; `04-suspension` 0 of 5; `03-overdue-reminders`
  0 of 1.
- Five of the eight hand the whole question back: the four #94 counted by hand, and
  one in the runs of the fix of #93, played since. Three hand back a part, which #94
  told apart and the measure does not: what the new commands print, on
  `02-reservations` in both runs, and the tie beyond the book id, on `01-overdue-list`.
- One answer of the 31 leaves the choice without the words, and is not counted:
  "I don't rely on the order of the printed lines, so take your recommendation", on
  `03-overdue-reminders` in campaign 2.

Refs #94
PierreMardon added a commit that referenced this pull request Oct 7, 2026
… called better or worse (#98)

Since the interview asks for the rules a user of the feature would see (#81), its developer
hands some of its questions back as not theirs to decide (#94), and the harness counts them
as `questions_handed_back` (#96). The report gave that measure a better way, fewer, like
`corrections`.

The developer decided on 2026-10-07 that this is no fault of the chain. A rule of the
interview precise enough to decide without judgement was specific to the host of the
evaluations, and a general one was vague or wrong elsewhere: the stop rule of the interview
stays as it is, and so does the brief of `06-borrow-limit`, which plays a developer who
delegates. Where such a question is not asked, the blueprint is to show the assumption to
confirm (#60).

So the measure stays in the report as a sign to follow, and loses its better way: a move of
it is told as moved, like a move of `questions`.

- `evals/surface_evals/report.py`: `questions_handed_back` leaves `DIRECTION`.
- `evals/README.md`: the report tells it and never calls it better or worse, and why.
- `tests/evals/test_report.py`: holds both.

No session was played: nothing of the chain changes, and the measure is read as before.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant