From d9b53b6796ea00b262424ae8bb636e87c5406c1b Mon Sep 17 00:00:00 2001 From: Agent Norton Date: Sun, 20 Sep 2026 06:15:07 +0000 Subject: [PATCH 1/7] Record the register eval review: noise or bad evals Five CI runs of the prose-register suite at the tip of the eval fixes branch, recomputed from the latest.json artifacts, with every failing subject output and judge reason read. The memo attributes the failure mass to eval-side defects (contested keys, non-minimal pairs, grader substring bugs, notes the judge reads as requirements), separates the judge noise from subject variance, and records what the suite does not measure. The fixture pass that follows cites it case by case. Co-Authored-By: Claude Fable 5.1 --- .../specs/2026-09-20-register-eval-review.md | 186 ++++++++++++++++++ 1 file changed, 186 insertions(+) create mode 100644 docs/superpowers/specs/2026-09-20-register-eval-review.md diff --git a/docs/superpowers/specs/2026-09-20-register-eval-review.md b/docs/superpowers/specs/2026-09-20-register-eval-review.md new file mode 100644 index 00000000..2b894434 --- /dev/null +++ b/docs/superpowers/specs/2026-09-20-register-eval-review.md @@ -0,0 +1,186 @@ +# Register eval review: noise or bad evals? + +**Date:** 2026-09-20\ +**Scope:** `prose-register` eval suite at the tip of PR #68 (`73a9271`), five CI result sets (runs 35468566695, 35475298553, 35475698801, 35476167110, 35476580577), `code-comment-register` as a control.\ +**Method:** recomputed every number below from the `latest.json` artifacts with a local script; read every failing subject output and judge reason; no model calls were made. + +## Conclusion + +Mostly bad evals, with a smaller layer of judge noise on top, and almost no evidence that the model under the skill is unpredictable on the cases that fail. + +On the tip run (0577, 10 of 22), 11 of the 12 failures are eval-side: contested or wrong answer keys (4), non-minimal or context-stripped case construction (2), grader defects (3, including both detection cases that fail by design), and a rubric or grading note that the judge applies as a requirement (2). One failure (trans-05) is the subject producing a different rewrite that the judge fairly rejected. Seven cases fail in every run with the same reasoning each time; that is not sampling noise, that is a key the model disagrees with consistently. + +The noise that does exist lives almost entirely in the LLM-judged components. Between the two pairs of runs with identical prompts and fixtures, 6 and then 2 rows flipped; of those 8 flips, 1 was a discrimination choice, 7 were judge or grader verdicts. The deterministic choice assert was stable within identical conditions (0 flips in the first pair, 1 in the second). + +Rough proportions of the failure mass per run: about 60% deterministic eval defects (keys, construction, graders), about 15% by-design detection failures, about 15 to 20% judge inconsistency, about 5 to 10% subject variance. Confidence: high on the deterministic defects (five identical outcomes with identical reasoning; P(5 of 5 same by chance at p = 0.5) is 3%), medium on the judge share (each row has one judge sample), low on anything about the CI-versus-local gap (no local result files survive). + +The suite also does not measure what the operator wants. Fifteen of 22 rows test the model as a critic of the author's edit history. Seven test generation on one-to-three-sentence snippets with the rule named in the task. Zero rows test unprompted drafting under the skill, and zero compare skill-on with skill-off. "Predictable generated prose" is unmeasured today, so fixing the suite's noise would not answer the operator's question either. + +## Evidence + +### Per-run totals (prose-register) + +| run | fixtures / prompt | success | choice | rule (judged) | rubric (judged) | detection | wrong choices | +|---|---|---|---|---|---|---|---| +| 6695 (#67) | disc-01 has "just"; no flags; two-line prompt | 12/22 | 8/12 | 7/12 | 5/7 | 1/3 | disc-01, 09, 11, 12 | +| 8553 | disc-01 fixed; 4 flags; two-line prompt | 10/22 | 9/12 | 7/12 | 5/7 | 0/3 | disc-06, 09, 12 | +| 8801 | same as 8553 | 12/22 | 9/12 | 8/12 | 5/7 | 1/3 | disc-06, 09, 12 | +| 7110 | compare-first prompt | 10/22 | 9/12 | 6/12 | 5/7 | 0/3 | disc-04, 09, 12 | +| 0577 (tip) | same as 7110 plus disc-12 flag | 10/22 | 8/12 | 9/12 | 4/7 | 1/3 | disc-04, 09, 11, 12 | + +Discrimination output length tripled after the compare-first change (median 218 to 258 visible characters before, 737 to 762 after). Model was `claude-sonnet-5` on every row that recorded it. Every run shows the skill loaded: 7 of 12 discrimination outputs per run quote SKILL.md wording the prompt never contains ("hedges of cowardice", "joints, not filler", "drumbeat"), and the subject refers to SKILL.md's section heading "The baseline" by name. + +### Per-case outcomes across the five runs + +| case | pass | choice | rule | rubric/det | what is happening | +|---|---|---|---|---|---| +| disc-01 | 4/5 | 4/5 | 4/5 | | The one miss is the pre-fix fixture (the "just" in `after`). Fixed. | +| disc-02 | 5/5 | 5/5 | 5/5 | | Stable pass. | +| disc-03 | 5/5 | 5/5 | 5/5 | | Stable pass. | +| disc-04 | 3/5 | 3/5 | 5/5 | | Wrong on both compare-first runs, for the reason the case's own note predicts: the connective has nothing to connect to. Construction defect. | +| disc-05 | 5/5 | 5/5 | 5/5 | | Stable pass. | +| disc-06 | 3/5 | 3/5 | 5/5 | | Both misses are self-contradictions under the two-line prompt: `ANSWER: A` with a RULE line saying "B (en-dashes) is on-register". 2 of 2 correct after compare-first. Harness defect, likely fixed. | +| disc-07 | 2/5 | 5/5 | 2/5 | | Judge passed RULE lines that cite only "just" (6695, 8801) and failed ones that cite both "just" and hedges (8553, 7110), the reverse of what the grading note asks. Judge inconsistency, fed by a human-facing note passed to the judge as a rubric. | +| disc-08 | 0/5 | 5/5 | 0/5 | | Choice right every time. RULE fails every time: subject cites the breath rule; key says "Open with substance", which the note itself says both variants satisfy. Key defect. | +| disc-09 | 0/5 | 0/5 | 1/5 | | Ranks `restored` last every run (reported 0 of 10 with local), same reasoning each time: `restored` is five consecutive same-pressure declaratives, which SKILL.md's Lint fails. The key conflicts with the skill. Key defect, deterministic. | +| disc-10 | 0/5 | 5/5 | 0/5 | | Choice right every time. RULE fails every time (breath cited, key says concrete nouns). The master rule absorbs the case. Key or skill-text defect. | +| disc-11 | 0/5 | 3/5 | 1/5 | | The pair differs in two places ("as you assume" and "underleveled"); the subject decides on the first in 3 of 5 runs and the RULE judge fails it. When it picks `before` it argues "underleveled with great reviews" is the concrete detail, which is a defensible reading. Non-minimal pair plus contested key. | +| disc-12 | 0/5 | 0/5 | 4/5 | | Picks `before` every CI run with a coherent argument ("These are hard won lessons" is a summarizing label). Reported 6 of 6 the other way locally. Contested key plus an unexplained environment effect. | +| trans-01 | 5/5 | | | 5/5 | Stable pass. | +| trans-02 | 1/5 | | | 1/5 | Three of four failures are `voice_match` demanding the reference's specific color word ("unsettling"); one is a fair "throat-clearing swapped, not removed". Over-specified rubric. | +| trans-03 | 4/5 | | | 4/5 | The most generative case; one judge rejection on a partial reframe. Judge on a subjective rubric. | +| trans-04 | 5/5 | | | 5/5 | Stable pass. | +| trans-05 | 3/5 | | | 3/5 | Failures cite real differences in the rewrite (drops "character"; reintroduces "caring" as a noun). Subject variance, fairly judged. | +| trans-06 | 1/5 | | | 1/5 | Task asks for "an actual mechanism" only the author knows; the subject invents a Java billing service with a rounding bug and the judge rejects the invention. Case demands facts the subject cannot have. | +| trans-07 | 5/5 | | | 5/5 | Stable pass. | +| det-01 | 0/5 | | | 0/5 | Subject finds 5 to 6 of 12 listed violations each run; pass requires all 12 by substring. By design. | +| det-02 | 0/5 | | | 0/5 | Subject finds 0 to 3 of 7 by substring; the note itself says the list is not exhaustive and the matcher under-scores. By design. | +| det-03 | 3/5 | | | 3/5 | Subject flagged the em-dash for the em-dash rule in all 5 runs. The two "misses" are the substring matcher: the key quote begins "same artifact — one rewrite" and the subject quoted from "artifact — one rewrite". Grader defect. | + +### Identical-condition pairs (pure sampling noise) + +Runs 8553 and 8801 share prompt and fixtures; so do 7110 and 0577. + +- 8553 vs 8801: 6 row flips. Choice 0, rule 1 (disc-07), rubric 4 (trans-02, 03, 05, 06), detection 1 (det-03, grader). +- 7110 vs 0577: 2 row flips. Choice 1 (disc-11), rule 3 (disc-09, 11, 12), rubric 1 (trans-05), detection 1 (det-03, grader). The disc-09 rule flip is the judge passing (0.85) essentially the same breath-rule sentence it failed four times. + +Observed SD of run totals: 1.1 rows. Predicted from the per-case rates under independence: 1.36 rows. The runs behave like a fixed set of per-case probabilities, not like a suite whose cases drift. + +### Control: code-comment-register + +Same five runs: 15, 15, 13, 13, 16 of 16. Choice 10 of 10 on every run; every row flip came from the rule judge (7 to 10 of 10 across runs), one rubric verdict, and one detection trap hit. That is the judge-noise floor on well-constructed, single-rule, anonymized cases: about plus or minus 2 rows in 16. The same judge that is stable enough there is noisy on prose-register because the prose keys give it contestable references and human-facing notes. + +## Question 1: noise or bad evals, by source + +- **Sampling noise (same case, same conditions, different outcome).** Small on the deterministic assert (1 choice flip in 44 identical-condition comparisons). Present on judged rows: 4 of 7 rubric rows flipped in one identical pair, 1 of 7 in the other. Estimated share of a run's failures: 15 to 25%, all in judged components. Confidence: medium. +- **Key defects.** disc-08 (rule), disc-09 (ranking contradicts the skill's own Lint), disc-10 (rule), disc-11 (contested and non-minimal), disc-12 (contested), disc-04 (context stripped). Six cases, four of them 0 of 5 with identical reasoning each time. Share: about half of each run's failures. Confidence: high. +- **Grader defects.** det-01 and det-02 fail by construction; det-03 fails on quote-boundary substring matching while the subject finds the violation every time. Share: 2 to 3 rows per run. Confidence: high. +- **Judge strictness and inconsistency.** disc-07 (passes the weaker answer, fails the stronger), trans-02 (voice_match reads the reference's color word as required), disc-09's one rule pass. The RULE rubric says "grade only the RULE line" and the judge reads the body anyway (7110 disc-07 reason quotes it). Share: 2 to 3 rows per run. Confidence: medium. +- **Harness and environment.** The two-line prompt produced self-contradictory answers (disc-06, 2 of 3 runs under it); compare-first removed that on 2 of 2 runs. The CI-versus-local gap is real for disc-12 (0 of 5 CI, reported 6 of 6 local; not a coincidence under any single rate) but no local result files survive, so its aggregate size cannot be decomposed here. Share: one row per run today, more before the prompt change. Confidence: low. +- **Inherent unpredictability of the model under the skill.** Visible in trans-05 (different rewrites, fairly judged), disc-11's choice (3 of 5, but on a two-difference pair), and trans-03 (one rejection in five on the most open task). One to two rows per run. Confidence: medium. There is no data at all on generation-side predictability, which is the operator's actual question. + +## Question 2: what the suite measures versus what the operator wants + +The suite answers "does Sonnet, told to use the skill, agree with the author's revision history when shown the before and after?" It answers that well on sentence-level mechanical cases (disc-02, 03, 05, 06 after the prompt fix, disc-07 choice) and badly on taste cases (disc-08, 09, 11, 12), where the author's final call is one defensible reading among two. + +The operator wants "when I draft under this skill, does the prose come out on-register, and does it come out the same way each time?" Nothing in the 22 rows samples a draft. The seven transformation rows name the rule and the fix in the task, so they measure instruction-following on a snippet, not the register's effect on writing. `promptfooconfig.compare.yaml` (skill on versus off) exists and has never run in CI. + +A generation-side measurement that answers the operator's question: + +- **Inputs:** 6 to 8 drafting prompts of the kind the operator actually issues (a 300 to 500 word post from a bullet outline; an edit pass over a clipped draft; a rewrite of an op-ed paragraph). det-01's and det-02's documents are ready-made edit-pass inputs. +- **Conditions:** skill on and skill off, the existing two providers. +- **Samples:** k = 5 per prompt per condition. +- **Tier 1 scoring, deterministic, zero judge sessions:** em-dash count; "just" as a modifier; the banned-phrase list from SKILL.md; five consecutive sentences within a length band; first-person pronoun present; adverb density; sentence-length distribution (median, share under six words). +- **Tier 2 scoring, judged, four binary items:** flatters the writer; claims the reader's reaction; summary close; floral. One judge call per sample as an informational score, or majority of three if it ever gates. +- **Metrics:** effect = tier-1 violation rate on versus off; predictability = share of prompts where all k samples pass tier 1, plus per-prompt variance of the tier-1 count and tier-2 agreement across the k samples. +- **Cost:** 8 prompts × 5 samples × 2 conditions = 80 subject sessions, plus 80 judge sessions at one call each. About 3.4 full prose runs. + +## Question 3: power + +With per-case rates taken from the five runs, a run's total has expected value 10.8 and SD 1.36 rows (about 6 points). The 0.4 floor (9 of 22) trips on unchanged code about 4% of the time from the floor alone. The hard choice gate trips far more often: under the two-line prompt, disc-06 alone failed 2 of 3 runs; at the tip flags, the unflagged wrong-choice rate is near zero only because every contested case is now flagged. + +Regression detection at one sample per case: if 1 of the 6 reliably-passing cases becomes an always-fail, the floor trips 17% of the time; 2 cases, 41%; 3 cases, 70%; 4 cases, 90%. If 4 reliable cases degrade to coin flips, 43%. The floor detects a collapse of half the reliable cases and nothing smaller. + +Per case, 5 samples cannot separate a 90% case from a 60% case (one-sided power 0.32 at alpha 0.05); 10 samples gives 0.62; 20 gives 0.95. Every per-case rate in this memo carries a 95% interval about 0.5 wide (0 of 5 is [0, 0.52]; 3 of 5 is [0.15, 0.95]). + +## Question 4: which SKILL.md rules are checkable + +Objectively checkable by regex or structure, on generated text: + +- No em-dashes (U+2014). +- "Just" as a modifier (with a short adjective exclusion list). +- The named phrases: "In today's world", "It's important to note", "in this essay", "as a writer". +- Five consecutive sentences of similar length (the "or similar pressure" half is judgment). +- Keep the "I" (first-person pronoun count over a text of N words). +- Short sentences never the baseline (median length and share of very short sentences as a proxy). +- Adverb density as a proxy for "strong verbs over adverbs". +- Partial: opening and closing (first sentence contains no listed throat-clearing phrase; last paragraph contains no "In sum", "So,", "In short"). Whether a close "summarizes" is judgment. +- Partial: a hedge phrase list ("I think maybe", "because I believe that", "a decent amount", "really", "sort of"); precision versus cowardice on unlisted hedges is judgment. + +Judgment only: the master rule (a hard line needs a breath; a paragraph of jabs), connectives as joints versus filler, the triplet's third item doing work, hedges of cowardice beyond the list, reversal for argument versus cleverness, no flex, never claim the reader's reaction, no grandiosity, nothing floral, no jargon without definition, therapy voice, concrete nouns over abstractions, questions that warm. + +The master rule is the problem for the eval. It is broad enough that the subject cites it as the deciding rule on disc-08, 09, 10, 11 and 12, five cases whose keys name five other rules. A judge cannot apply "the same principle" consistently when one principle subsumes the rest. Two fixes, either or both: the eval stops asking for a free-text rule and asks for a pick from SKILL.md's numbered rule list against an `accepted_rules` set (deterministic, removes the judge from 12 rows); and SKILL.md's Lint section splits into "checkable" and "judged" items so a linter enforces the first half. The skill text otherwise reads as a voice guide and should stay one; the eval should stop pretending its judgment rules have single right answers. + +## Question 5: the menu + +- **Tiering (deterministic hard, judged advisory).** Keep; cheap. Insufficient alone, because the current deterministic tier asks the model to detect em-dashes and the model self-contradicted on that 2 of 3 times under the old prompt. The deterministic tier should be lint over generated prose, not model detection. +- **Variance study first (5 repeats × on/off on the existing suite).** Drop in that form. It costs about 470 sessions to learn what five CI runs already show: seven cases fail identically every time and the flips sit in judged rows. Replace with the generation pilot above, which is the variance study that matters. +- **Minimal synthetic pairs.** Keep for the coverage gaps and to replace disc-04 and disc-11. Zero sessions to author. They pin critic behaviour, not generation, so they belong in the advisory tier. +- **Structured `accepted_rules`.** Keep, and go further to numbered-rule selection so the check is deterministic. Also stop passing `grading_note` to the judge; disc-07 shows the judge treating a human-facing note as a requirement. +- **Judge calibration (20 rows against a human).** Defer. After the changes above the judged surface is the four tier-2 items and trans rubrics; majority-of-three on those costs about 22 extra judge sessions per run and acts on every run, which a one-off calibration does not. A cheaper diagnostic exists first: re-judge the stored transformation outputs from the artifacts three times each (21 judge sessions, zero subject sessions) to split judge noise from subject noise on the 7 trans rows. + +### Recommended sequence + +1. **Fixture pass, zero model sessions.** Re-key or drop disc-09 (key contradicts the Lint); accept the breath rule or drop the rule check on disc-08 and disc-10; restore a one-difference pair for disc-11; prepend the prior paragraph to disc-04 as its note suggests; put the facts into trans-06's task; rewrite trans-02's `voice_match` to not name the reference's word; state in disc-07's note that either rule passes; fix the detection matcher to search for a shorter core of each quote, and score det-01 and det-02 as recall percentages that never pass or fail. Verify with one CI run (47 sessions). Expected: 16 to 19 of 22 with no model change. If it stays at 10 to 12, the noise is inherent and this memo's conclusion is wrong. +2. **Numbered-rule deterministic RULE check with `accepted_rules`.** Zero sessions to build; verify in the same CI run as step 1 or the next. Removes about 12 judge sessions per run. +3. **Generation pilot** (about 160 sessions, or 80 for tier 1 only). Decides whether the skill changes output and how stable the output is. This is the first measurement of what the operator asked for. +4. **Gate redesign.** Tier-1 lint on generated prose gates hard; the critic suite and tier-2 judgments report a trend and never gate. Keep the `evals` job advisory until five CI runs of the new shape exist. +5. **CI-versus-local gap.** Either the operator runs one local suite with a `claude setup-token` token exported as `CLAUDE_CODE_OAUTH_TOKEN` (47 sessions), or local runs stop being a reference at all. The second is cheaper and the spec already calibrates on CI. + +## Ranked recommendations + +1. Do the zero-session fixture pass and one CI run before anything else. It settles noise-versus-evals for 47 sessions. +2. Build the generation-side set with skill on and off, k = 5, deterministic lint first. Nothing in the current suite measures predictability. +3. Replace free-text RULE judging with numbered-rule selection and `accepted_rules`; stop feeding `grading_note` to the judge. +4. Move the author's-taste cases (disc-08, disc-09, disc-11, disc-12) out of any gate for good. A key that encodes one defensible reading cannot gate a model that finds the other one. +5. Split SKILL.md's Lint into checkable and judged items, and let a linter own the first half. +6. Drop the 5×2 repeat study of the existing suite and the judge calibration study; spend those sessions on the pilot. + +## Decisions for the operator + +- Whether the purpose of `evals/` is a regression pin for the critic (keep, tiered, advisory) or a measurement of generated prose (build the pilot). Both can coexist; only the second answers the question in the handoff. +- Re-key or drop disc-09, disc-11 and disc-12; accept the breath rule on disc-08 and disc-10 or drop their rule checks. These are the author's own editorial calls. +- Which model is the subject. The suite runs Sonnet; if the operator drafts with a different model, the generation pilot should run on that one. +- Whether to spend a token on the CI-versus-local test, or drop local as a reference. +- Whether SKILL.md's Lint section may be restructured (checkable versus judged), since it is the author's voice document. + +## What I could not determine + +- The size and cause of the CI-versus-local gap beyond disc-12. No local result files survive and I ran nothing. The five CI runs are internally consistent, so calibration on CI is sound; the open question is only whether disc-12's local record means anything. +- The split between judge noise and subject noise on the transformation rows. One subject sample and one judge sample per row cannot separate them; re-judging stored outputs would. +- Any number about generation-side predictability. No row measures it. +- Whether the compare-first prompt fixed disc-06 for good: 2 of 2 is all the evidence there is. + +## What would change my mind + +If, after the fixture pass alone (no prompt or model change), five CI runs still land at 10 to 12 of 22, with the re-keyed cases flipping between runs rather than passing together, the variance is in the model and not in the keys. Equally, if a five-repeat of the tip suite showed disc-08's rule, disc-09, disc-10's rule or disc-12 passing in 2 or more of 5 runs, those are noisy rather than mis-keyed and the "deterministic defect" share above is overstated. On the operator's real question, the pilot could show skill-on and skill-off tier-1 rates indistinguishable at k = 5; then the skill has no measurable effect on drafting and the eval question is secondary to the skill itself. + +## Where the handoff's sketch was wrong + +- disc-08 is not environment-sensitive; its choice is 5 of 5 in CI and the 0 of 5 is the RULE key. +- disc-06 is not environment-sensitive; the misses are self-contradictions under the two-line prompt, corrected 2 of 2 after compare-first. +- trans-04 is a stable pass (5 of 5), not a coin flip; trans-02 and trans-06 (1 of 5 each) are near-stable fails; trans-05 (3 of 5) is the only transformation row behaving like a coin flip. +- det-03's flips are a grader substring defect; the subject found the em-dash all five times. +- disc-11 is not a "coin-flip judge" case; its choice is 3 of 5 because the pair has two differences. +- The stable-pass class also holds trans-01, trans-04 and trans-07, not only disc-02, 03 and 05. +- disc-07 (rule 2 of 5, judge-inconsistent) was missing from the sketch entirely. + +## Errata (2026-09-20, found while acting on this memo) + +Rechecked against the same five artifacts during review of the fixture pass. The sections above stand as written; these are the corrections. + +- disc-10: the RULE line cited the breath rule on four of five runs, not every run. Run 8553 cited "No flex". +- disc-11: the subject decided on the first difference ("as you assume") in 2 of 5 runs, not 3, and both of those chose correctly. The two wrong choices (6695, 0577) engaged the jargon clause and read "underleveled with great reviews" as the concrete detail, so a one-difference pair leaves that failure mode in place. The same correction applies to the line above about disc-11. +- trans-02: two of four failures are voice_match on the missing color word (6695, 7110), not three. 8553 failed criterion 1 for dropping the developer-backlash claim; 0577 failed it for swapping the throat-clearing rather than removing it. +- trans-06: one run was rejected for an invented mechanism (0577). Two were rejected for naming no mechanism (8801, 7110), one failed voice_match with criteria 1 and 2 passing (6695), and one passed (8553, with an invented ledger service the judge accepted). The task still asked for facts the subject cannot have; "invents a mechanism" describes one run, not the pattern. +- The eight row flips across the two identical-condition pairs are all rule, rubric or detection verdicts. The one choice flip (disc-11, second pair) did not flip its row, because that row failed on the rule in the other run. From 07447f454c1672ca17a4e0fd58a8810a0f2f0db4 Mon Sep 17 00:00:00 2001 From: Agent Norton Date: Sun, 20 Sep 2026 06:17:39 +0000 Subject: [PATCH 2/7] Match detection quotes by overlap and let a case floor its recall The detection grader required the first 60 characters of each key quote to appear verbatim in the subject's output. det-03's subject found the em-dash on all five CI runs and was scored a miss on two of them because it dropped the key's first word. Both strings are verbatim runs of the same document, so the grader now counts a subject quote that contains the key, sits inside it, or overlaps one of its ends by at least 20 characters. One subject quote is one finding: it is credited to every key it holds whole, or else to the one key it overlaps most, so a quote that runs from one violation into the opening of the next is not a find of both (seen on one CI run of det-02). Traps use the same rule, and a fixture test keeps any violation quote from overlapping a trap. Overlap matching needs every key quote to appear in the document verbatim; the validator now enforces that, and det-01's one trap that ended in an ellipsis is quoted in full. A detection case may set min_recall in [0, 1]. Below 1 the case passes at that recall and records recall as its score, for lists the author says are not exhaustive (prose-register's det-01 and det-02 failed every run by construction). A flagged trap still fails at any floor. The default of 1 keeps every existing case's behaviour. Co-Authored-By: Claude Fable 5.1 --- evals/asserts/detection.mjs | 87 +++++++++++--- evals/bin/check-gate.mjs | 6 +- evals/lib/load-evals.mjs | 21 ++++ evals/test/asserts.test.mjs | 113 +++++++++++++++++- evals/test/load-evals.test.mjs | 28 +++++ evals/tests.mjs | 1 + home/.agents/skills/prose-register/evals.json | 2 +- 7 files changed, 234 insertions(+), 24 deletions(-) diff --git a/evals/asserts/detection.mjs b/evals/asserts/detection.mjs index 62696f85..524778f5 100644 --- a/evals/asserts/detection.mjs +++ b/evals/asserts/detection.mjs @@ -1,28 +1,77 @@ import { normalize } from './heuristics.mjs'; +// A subject quote shorter than this would sit inside several key quotes at +// once ("just", "the code"), so it never counts on its own. +const MIN_OVERLAP = 20; + +// Both strings are verbatim runs of the same document, so they refer to the +// same passage when one contains the other or the end of one is the start of +// the other. Returns the length of the shared run (0 when none counts) and +// whether the subject quote holds the whole key. A prefix-only match called +// det-03 a miss when the subject began its quote one word into the key and +// ran past the key's end. +export function overlap(keyQuote, subjectQuote) { + const key = normalize(keyQuote); + const quote = normalize(subjectQuote); + const min = Math.min(MIN_OVERLAP, key.length); + if (quote.length < min) return { length: 0, whole: false }; + if (quote.includes(key)) return { length: key.length, whole: true }; + if (key.includes(quote)) return { length: quote.length, whole: false }; + for (let len = Math.min(key.length, quote.length); len >= min; len--) { + if (key.endsWith(quote.slice(0, len)) || quote.endsWith(key.slice(0, len))) { + return { length: len, whole: false }; + } + } + return { length: 0, whole: false }; +} + +export const quotesOverlap = (keyQuote, subjectQuote) => overlap(keyQuote, subjectQuote).length > 0; + +// One subject quote is one finding. It is credited to every target it holds +// whole and, when it holds none, to the single target it overlaps most; a +// quote that runs from one violation into the opening of the next is not a +// find of both. +function credited(targets, quote) { + const scored = targets.map((target) => ({ target, ...overlap(target.quote, quote) })); + const whole = scored.filter((s) => s.whole).map((s) => s.target); + if (whole.length) return whole; + const best = scored.filter((s) => s.length > 0).sort((a, b) => b.length - a.length)[0]; + return best ? [best.target] : []; +} + +function subjectQuotes(output) { + const quotes = []; + for (const line of output.split('\n')) { + const match = line.match(/^-\s*QUOTE:\s*(.*)$/i); + if (!match) continue; + const [text] = match[1].split(/\|\s*RULE:/i); + const quote = normalize(text).replace(/^"|"$/g, '').trim(); + if (quote) quotes.push(quote); + } + return quotes; +} + export default function assertDetection(output, context) { - const { violations, traps } = context.vars; - const violationLines = output - .split('\n') - .filter((line) => /^-\s*QUOTE:/i.test(line)) - .join('\n'); - const normalized = normalize(violationLines); - - const missed = violations.filter( - (violation) => !normalized.includes(normalize(violation.quote).slice(0, 60)), - ); - const trapHits = traps.filter( - (trap) => normalized.includes(normalize(trap.quote).slice(0, 40)), - ); + const { violations, traps, min_recall: minRecall = 1 } = context.vars; + const hit = new Set(); + for (const quote of subjectQuotes(output)) { + for (const target of credited([...violations, ...traps], quote)) hit.add(target); + } + + const missed = violations.filter((violation) => !hit.has(violation)); + const trapHits = traps.filter((trap) => hit.has(trap)); const found = violations.length - missed.length; - const pass = missed.length === 0 && trapHits.length === 0; + const recall = found / violations.length; + // A case whose violation list is not exhaustive sets min_recall below 1 and + // records recall as its score; a flagged trap fails at any floor. + const pass = recall >= minRecall && trapHits.length === 0; const score = Math.max(0, (found - trapHits.length) / violations.length); - const reason = pass - ? `${found}/${violations.length} violations, 0 traps` - : `${found}/${violations.length} violations, ${trapHits.length} trap(s) flagged` + - (missed.length ? `; missed: ${missed.map((m) => m.quote.slice(0, 40)).join(' | ')}` : '') + - (trapHits.length ? `; traps: ${trapHits.map((t) => t.quote.slice(0, 40)).join(' | ')}` : ''); + const floor = minRecall < 1 ? `, floor ${minRecall}` : ''; + const reason = + `${found}/${violations.length} violations, ${trapHits.length} trap(s) flagged${floor}` + + (missed.length ? `; missed: ${missed.map((m) => m.quote.slice(0, 40)).join(' | ')}` : '') + + (trapHits.length ? `; traps: ${trapHits.map((t) => t.quote.slice(0, 40)).join(' | ')}` : ''); return { pass, score, reason }; } diff --git a/evals/bin/check-gate.mjs b/evals/bin/check-gate.mjs index 1d100d68..c59a5303 100644 --- a/evals/bin/check-gate.mjs +++ b/evals/bin/check-gate.mjs @@ -51,9 +51,9 @@ const isHardFailure = (row) => components(row).some((c) => !c.pass && c.assertion?.metric === 'choice'))); // Precedence: argv, then the skill's own `min_pass_rate` in evals.json, then -// the default. A skill whose cases are soft by design (prose-register's -// detection cases fail by construction) can carry a lower floor than one -// whose cases all have a single right answer. +// the default. A skill whose keys are contested or whose detection cases +// record recall instead of failing (prose-register) can carry a lower floor +// than one whose cases all have a single right answer. // // Only the named skill's file is parsed (the generator locates it the same // way), so a sibling's broken evals.json cannot fail this gate; an unreadable diff --git a/evals/lib/load-evals.mjs b/evals/lib/load-evals.mjs index ecb0cacc..beb71a6f 100644 --- a/evals/lib/load-evals.mjs +++ b/evals/lib/load-evals.mjs @@ -1,5 +1,6 @@ import { readFileSync, readdirSync, existsSync, lstatSync } from 'node:fs'; import path from 'node:path'; +import { normalize } from '../asserts/heuristics.mjs'; export const SUPPORTED_TYPES = new Set([ 'discrimination', @@ -107,6 +108,17 @@ export function validateData(data) { } } + // Below 1 the case records recall as its score instead of failing on a + // non-exhaustive violation list; only detection has recall to floor. + if (item.min_recall !== undefined) { + if (item.type !== 'detection') { + throw new Error(`${item.id}.min_recall applies only to detection cases`); + } + if (typeof item.min_recall !== 'number' || !(item.min_recall >= 0 && item.min_recall <= 1)) { + throw new Error(`${item.id}.min_recall must be a number in [0, 1]`); + } + } + if (item.type === 'detection') { requireField(item.prompt, `${item.id}.prompt`); requireField(item.input_document, `${item.id}.input_document`); @@ -123,6 +135,15 @@ export function validateData(data) { for (const trap of item.traps) { requireField(trap.quote, `${item.id}.traps[].quote`); } + // The grader matches by text overlap, so a quote the document does not + // contain verbatim (an ellipsis, a paraphrase) can never be found or + // tripped. + const document = normalize(item.input_document); + for (const { quote } of [...item.violations, ...item.traps]) { + if (!document.includes(normalize(quote))) { + throw new Error(`${item.id} quote is not in input_document verbatim: "${quote.slice(0, 40)}"`); + } + } } } diff --git a/evals/test/asserts.test.mjs b/evals/test/asserts.test.mjs index d64a67ea..ec02eb01 100644 --- a/evals/test/asserts.test.mjs +++ b/evals/test/asserts.test.mjs @@ -1,7 +1,12 @@ import test from 'node:test'; import assert from 'node:assert/strict'; +import path from 'node:path'; +import { fileURLToPath } from 'node:url'; import assertDiscrimination from '../asserts/discrimination.mjs'; -import assertDetection from '../asserts/detection.mjs'; +import assertDetection, { quotesOverlap } from '../asserts/detection.mjs'; +import { findEvalFiles, loadEvals } from '../lib/load-evals.mjs'; + +const REPO_ROOT = path.resolve(path.dirname(fileURLToPath(import.meta.url)), '../..'); const discVars = { letter_to_key: { A: 'generic_comment', B: 'no_comment' }, @@ -102,3 +107,109 @@ test('detection ignores lines outside the QUOTE format', () => { const result = assertDetection(output, { vars: detVars }); assert.equal(result.pass, true, 'prose mention of a trap outside violation lines must not count'); }); + +// det-03's miss: the key begins "same artifact — one rewrite", the subject +// quoted from "artifact — one rewrite" onward and past the key's end. +test('detection counts a quote that starts inside the key and runs past it', () => { + const vars = { + violations: [{ quote: 'same artifact — one rewrite by strangers', rule: 'No em-dashes.' }], + traps: [], + }; + const output = '- QUOTE: "artifact — one rewrite by strangers, one by a machine" | RULE: em-dash'; + const result = assertDetection(output, { vars }); + assert.equal(result.pass, true, result.reason); + assert.equal(result.score, 1); +}); + +test('detection counts a quote that wraps the key in context on both sides', () => { + const output = '- QUOTE: "Then: Larger chunks use more memory. Smaller chunks use more CPU. And so on." | RULE: generic'; + const result = assertDetection(output, { vars: { ...detVars, violations: detVars.violations.slice(0, 1) } }); + assert.equal(result.pass, true, result.reason); +}); + +test('detection does not count a fragment shorter than the overlap floor', () => { + const output = '- QUOTE: "more memory" | RULE: generic tradeoff'; + const result = assertDetection(output, { vars: { ...detVars, violations: detVars.violations.slice(0, 1) } }); + assert.equal(result.pass, false); + assert.match(result.reason, /0\/1 violations/); +}); + +test('detection flags a trap quoted from its middle, not only from its start', () => { + const output = [ + '- QUOTE: "Larger chunks use more memory. Smaller chunks use more CPU." | RULE: generic tradeoff', + '- QUOTE: "Parse all records and drop the header before writing." | RULE: narration', + '- QUOTE: "returns a view of the buffer" | RULE: jargon', + ].join('\n'); + const result = assertDetection(output, { vars: { ...detVars, traps: [{ quote: '`binary_part/3` returns a view of the buffer' }] } }); + assert.equal(result.pass, false); + assert.match(result.reason, /1 trap/); +}); + +test('detection with min_recall 0 passes on partial recall and records recall as the score', () => { + const output = '- QUOTE: "Larger chunks use more memory. Smaller chunks use more CPU." | RULE: generic tradeoff'; + const result = assertDetection(output, { vars: { ...detVars, min_recall: 0 } }); + assert.equal(result.pass, true, result.reason); + assert.equal(result.score, 0.5); + assert.match(result.reason, /1\/2 violations, 0 trap\(s\) flagged, floor 0/); +}); + +test('detection with min_recall 0 still fails on a trap hit', () => { + const output = '- QUOTE: "`binary_part/3` returns a view" | RULE: needless jargon'; + const result = assertDetection(output, { vars: { ...detVars, min_recall: 0 } }); + assert.equal(result.pass, false); + assert.match(result.reason, /1 trap/); +}); + +test('detection with a fractional min_recall gates at that floor', () => { + const output = '- QUOTE: "Larger chunks use more memory. Smaller chunks use more CPU." | RULE: generic tradeoff'; + assert.equal(assertDetection(output, { vars: { ...detVars, min_recall: 0.5 } }).pass, true); + assert.equal(assertDetection(output, { vars: { ...detVars, min_recall: 0.6 } }).pass, false); +}); + +const adjacentVars = { + violations: [ + { quote: 'The creator of Bun spent the tokens. It took eleven days. A model wrote the commits.', rule: 'drumbeat' }, + { quote: 'Developers did not take it well. Soon it will run on that rewrite. We are users now.', rule: 'drumbeat' }, + ], + traps: [], +}; + +// Seen on a CI run: one subject line quoting the first paragraph and the +// opening sentence of the next was credited to both violations. +test('detection credits a quote that runs from one key into the next to the key it holds whole', () => { + const output = '- QUOTE: "The creator of Bun spent the tokens. It took eleven days. A model wrote the commits. Developers did not take it well." | RULE: drumbeat'; + const result = assertDetection(output, { vars: adjacentVars }); + assert.equal(result.pass, false); + assert.match(result.reason, /1\/2 violations/); +}); + +test('detection credits a fragment straddling two keys to the one it overlaps most', () => { + const output = '- QUOTE: "It took eleven days. A model wrote the commits. Developers did not take" | RULE: drumbeat'; + const result = assertDetection(output, { vars: adjacentVars }); + assert.match(result.reason, /1\/2 violations.*missed: Developers/); +}); + +test('detection credits a quote holding two whole keys to both', () => { + const output = `- QUOTE: "${adjacentVars.violations[0].quote} ${adjacentVars.violations[1].quote}" | RULE: drumbeat`; + const result = assertDetection(output, { vars: adjacentVars }); + assert.equal(result.pass, true, result.reason); +}); + +// Overlap matching is symmetric enough that a violation quote sharing a run of +// text with a trap would flag the trap on a correct answer; keep the fixtures +// free of that. +test('no detection case has a violation quote that would itself trip one of its traps', () => { + for (const file of findEvalFiles(REPO_ROOT)) { + const data = loadEvals(file); + for (const item of data.cases.filter((c) => c.type === 'detection')) { + for (const violation of item.violations) { + for (const trap of item.traps) { + assert.ok( + !quotesOverlap(trap.quote, violation.quote), + `${data.skill}/${item.id}: violation "${violation.quote.slice(0, 40)}" overlaps trap "${trap.quote.slice(0, 40)}"`, + ); + } + } + } + } +}); diff --git a/evals/test/load-evals.test.mjs b/evals/test/load-evals.test.mjs index 94b14d4a..6cfed6fb 100644 --- a/evals/test/load-evals.test.mjs +++ b/evals/test/load-evals.test.mjs @@ -96,6 +96,34 @@ test('validateData accepts a min_pass_rate in (0, 1]', () => { assert.doesNotThrow(() => validateData(ok)); }); +const detectionCase = (extra) => ({ + skill: 'x', + cases: [{ id: 'x1', type: 'detection', prompt: 'p', input_document: 'the quoted line', + violations: [{ quote: 'quoted line', rule: 'r' }], traps: [], ...extra }], +}); + +test('validateData accepts a min_recall in [0, 1] on a detection case', () => { + for (const ok of [0, 0.5, 1]) { + assert.doesNotThrow(() => validateData(detectionCase({ min_recall: ok }))); + } +}); + +test('validateData rejects a min_recall outside [0, 1]', () => { + for (const bad of [-0.1, 1.5, '0.5']) { + assert.throws(() => validateData(detectionCase({ min_recall: bad })), /x1\.min_recall must be a number in \[0, 1\]/); + } +}); + +test('validateData rejects a detection quote the document does not contain verbatim', () => { + const bad = detectionCase({ input_document: 'Therefore, trust the pilot.', traps: [{ quote: 'Therefore, trust...' }] }); + bad.cases[0].violations = [{ quote: 'trust the pilot', rule: 'r' }]; + assert.throws(() => validateData(bad), /x1 quote is not in input_document verbatim: "Therefore, trust\.\.\."/); +}); + +test('validateData rejects min_recall on a case type with no recall to floor', () => { + assert.throws(() => validateData(arguableCase({ min_recall: 0.5 })), /d1\.min_recall applies only to detection cases/); +}); + test('validateData rejects a min_pass_rate outside (0, 1]', () => { for (const bad of [0, 1.5, '0.5', -1]) { const data = { ...arguableCase({}), min_pass_rate: bad }; diff --git a/evals/tests.mjs b/evals/tests.mjs index ad6fc3b3..a1c99378 100644 --- a/evals/tests.mjs +++ b/evals/tests.mjs @@ -73,6 +73,7 @@ export default async function generateTests() { } else if (item.type === 'detection') { base.vars.violations = item.violations; base.vars.traps = item.traps; + if (item.min_recall !== undefined) base.vars.min_recall = item.min_recall; base.assert = [{ type: 'javascript', value: 'file://asserts/detection.mjs' }]; } else { base.assert = [judged(buildTransformationRubric(item), 'rubric')]; diff --git a/home/.agents/skills/prose-register/evals.json b/home/.agents/skills/prose-register/evals.json index 110133ad..869e1fde 100644 --- a/home/.agents/skills/prose-register/evals.json +++ b/home/.agents/skills/prose-register/evals.json @@ -240,7 +240,7 @@ ], "traps": [ {"quote": "the AI's Competence is usually sufficient", "why_not_a_violation": "SKILL.md lists 'usually sufficient' verbatim as a hedge of precision to KEEP."}, - {"quote": "Therefore, if you are trusting an AI with a high-risk task...", "why_not_a_violation": "SKILL.md lists 'Therefore,' verbatim as a good connective/joint."}, + {"quote": "Therefore, if you are trusting an AI with a high-risk task, the human pilot must entirely absorb the responsibility for those two traits.", "why_not_a_violation": "SKILL.md lists 'Therefore,' verbatim as a good connective/joint."}, {"quote": "Recently, the creator of Bun burned through roughly $200,000 worth of tokens to rewrite their core runtime from Zig into Rust.", "why_not_a_violation": "The opening correctly leads with substance (no throat-clearing). The dollar figure itself is factually wrong ($200,000 vs. the real $165,000), but that is a fact-check issue, not a register violation -- don't let a model conflate the two when grading detection."} ], "grading_note": "12 violations, 3 traps. Grade recall (how many of the 12 were found) and precision (were any traps wrongly flagged, or was a genuinely fine sentence flagged for no stated reason)." From 40f7d7a3b6b358d70d5ef302349725f7db7d19fc Mon Sep 17 00:00:00 2001 From: Agent Norton Date: Sun, 20 Sep 2026 06:18:25 +0000 Subject: [PATCH 3/7] Judge the stated rule against accepted_rules, not the grading note The rule rubric passed each case's grading_note to the judge as the source of accepted alternatives. disc-07's note asks for a precise answer that cites both the "just" prohibition and the hedge rule; across five CI runs the judge passed the RULE lines that cited one rule and failed the ones that cited both, reading a human-facing note as a requirement. A discrimination case now lists its alternatives in accepted_rules, a non-empty array of rule strings, and the rubric shows the judge only the reference rule, the skill's rule text and that list. grading_note is human-facing again and never reaches the rule judge. The transformation rubric keeps its note: those notes are written as grading guidance and none of the five runs showed the judge misreading one. code-comment-register shares the rubric. Its ten discrimination notes explain the key rather than list alternatives, so none migrates, and its rule judge loses that context; the CI run on this branch measures the effect against its 7 to 10 of 10 across the five prior runs. Co-Authored-By: Claude Fable 5.1 --- evals/lib/load-evals.mjs | 12 ++++++++++++ evals/lib/prompts.mjs | 13 ++++++++++--- evals/test/generator.test.mjs | 1 + evals/test/load-evals.test.mjs | 19 +++++++++++++++++++ evals/test/prompts.test.mjs | 13 ++++++++----- 5 files changed, 50 insertions(+), 8 deletions(-) diff --git a/evals/lib/load-evals.mjs b/evals/lib/load-evals.mjs index beb71a6f..5c63f1e7 100644 --- a/evals/lib/load-evals.mjs +++ b/evals/lib/load-evals.mjs @@ -108,6 +108,18 @@ export function validateData(data) { } } + // Alternatives the rule judge accepts besides expected_rule; only + // discrimination cases have a stated rule to judge. + if (item.accepted_rules !== undefined) { + if (!item.type.startsWith('discrimination')) { + throw new Error(`${item.id}.accepted_rules applies only to discrimination cases`); + } + const rules = item.accepted_rules; + if (!Array.isArray(rules) || rules.length === 0 || rules.some((rule) => typeof rule !== 'string' || rule === '')) { + throw new Error(`${item.id}.accepted_rules must be a non-empty array of rule strings`); + } + } + // Below 1 the case records recall as its score instead of failing on a // non-exhaustive violation list; only detection has recall to floor. if (item.min_recall !== undefined) { diff --git a/evals/lib/prompts.mjs b/evals/lib/prompts.mjs index e656ae3e..1c02347e 100644 --- a/evals/lib/prompts.mjs +++ b/evals/lib/prompts.mjs @@ -112,10 +112,17 @@ Notes: ${item.grading_note || 'None.'}`; } +// The judge sees only the reference rule and the alternatives the case lists; +// grading_note stays human-facing. Passed as free text, a note that asked for +// a precise two-rule answer was read as a requirement and failed the answers +// that met it. export function buildRuleRubric(item) { + const accepted = item.accepted_rules?.length + ? item.accepted_rules.map((rule) => `- ${rule}`).join('\n') + : 'None.'; return `Grade only the line in the output that begins with "RULE:". Ignore which option was chosen and anything else in the output. -The RULE line passes if it names the same principle as the reference rule in any wording, names the part of a compound reference rule that decides this case, or names an alternative the author's notes accept. It fails if it names a different principle, is missing, or is too vague to tell this rule apart from the skill's other rules. +The RULE line passes if it names the same principle as the reference rule in any wording, names the part of a compound reference rule that decides this case, or names one of the accepted alternative rules. It fails if it names a different principle, is missing, or is too vague to tell this rule apart from the skill's other rules. Reference rule: ${item.expected_rule ?? item.expected_rule_for_worst} @@ -123,6 +130,6 @@ ${item.expected_rule ?? item.expected_rule_for_worst} Rule text from the skill: ${item.rule_quote || 'None.'} -Author's notes (may list accepted alternative rules): -${item.grading_note || 'None.'}`; +Accepted alternative rules (any one of these also passes): +${accepted}`; } diff --git a/evals/test/generator.test.mjs b/evals/test/generator.test.mjs index f03d57f1..2a14ca37 100644 --- a/evals/test/generator.test.mjs +++ b/evals/test/generator.test.mjs @@ -29,6 +29,7 @@ function answerKeyValues(item) { } if (item.rubric && typeof item.rubric === 'object') values.push(...Object.values(item.rubric)); if (Array.isArray(item.correct_ranking)) values.push(item.correct_ranking.join(',')); + if (Array.isArray(item.accepted_rules)) values.push(...item.accepted_rules); return values; } diff --git a/evals/test/load-evals.test.mjs b/evals/test/load-evals.test.mjs index 6cfed6fb..f3da22d4 100644 --- a/evals/test/load-evals.test.mjs +++ b/evals/test/load-evals.test.mjs @@ -96,6 +96,25 @@ test('validateData accepts a min_pass_rate in (0, 1]', () => { assert.doesNotThrow(() => validateData(ok)); }); +test('validateData accepts accepted_rules as a non-empty list of rule strings on a discrimination case', () => { + assert.doesNotThrow(() => validateData(arguableCase({ accepted_rules: ['Another rule.'] }))); +}); + +test('validateData rejects an empty or non-string accepted_rules list', () => { + for (const bad of [[], 'Another rule.', ['ok', ''], [1]]) { + assert.throws(() => validateData(arguableCase({ accepted_rules: bad })), /d1\.accepted_rules must be a non-empty array of rule strings/); + } +}); + +test('validateData rejects accepted_rules on a case type with no stated rule to judge', () => { + const bad = { + skill: 'x', + cases: [{ id: 't1', type: 'transformation', input: 'i', task: 't', reference_after: 'r', + rubric: { violation_fixed: 'v' }, accepted_rules: ['x'] }], + }; + assert.throws(() => validateData(bad), /t1\.accepted_rules applies only to discrimination cases/); +}); + const detectionCase = (extra) => ({ skill: 'x', cases: [{ id: 'x1', type: 'detection', prompt: 'p', input_document: 'the quoted line', diff --git a/evals/test/prompts.test.mjs b/evals/test/prompts.test.mjs index c01ab182..876e00c7 100644 --- a/evals/test/prompts.test.mjs +++ b/evals/test/prompts.test.mjs @@ -72,10 +72,11 @@ test('transformation rubric embeds task, reference, and all three criteria', () } }); -test('rule rubric grades only the RULE line against the reference rule', () => { +test('rule rubric grades only the RULE line against the reference rule and the accepted alternatives', () => { const item = { id: 'd', type: 'discrimination', expected_rule: 'Explain why, not what.', rule_quote: 'Quote.', - grading_note: 'Also accept X.', + accepted_rules: ['Placement.', 'Brevity.'], + grading_note: 'A precise answer cites both rules.', }; const rubric = buildRuleRubric(item); assert.match(rubric, /RULE:/); @@ -83,13 +84,15 @@ test('rule rubric grades only the RULE line against the reference rule', () => { assert.ok(rubric.includes('Quote.')); assert.match(rubric, /Ignore which option was chosen/); assert.match(rubric, /part of a compound reference rule/); - assert.ok(rubric.includes('Also accept X.')); + assert.ok(rubric.includes('Accepted alternative rules (any one of these also passes):\n- Placement.\n- Brevity.')); + // The note is for humans; passed to the judge it was read as a requirement. + assert.ok(!rubric.includes('A precise answer cites both rules.')); }); -test('rule rubric defaults the author notes to None when a case sets no grading_note', () => { +test('rule rubric lists None when a case accepts no alternative rule', () => { const item = { id: 'd', type: 'discrimination', expected_rule: 'Explain why, not what.', rule_quote: 'Quote.' }; const rubric = buildRuleRubric(item); - assert.ok(rubric.includes('Author\'s notes (may list accepted alternative rules):\nNone.')); + assert.ok(rubric.includes('Accepted alternative rules (any one of these also passes):\nNone.')); }); test('rank prompt letters stages and asks for a best-first ordering', () => { From a154459406974c7b33983ebc2a7dffa0353697be Mon Sep 17 00:00:00 2001 From: Agent Norton Date: Sun, 20 Sep 2026 06:23:54 +0000 Subject: [PATCH 4/7] Re-cut the prose-register keys the subject disagrees with every run Five CI runs at the tip of the fixes branch failed the same seven prose-register cases with the same reasoning each time. A key the model disagrees with consistently is a fixture defect, not sampling noise, and each one here is fixed at the fixture with zero model sessions. - disc-09 becomes a two-way pick, original over over_tight. The subject ranked the hand-edited 'restored' stage last on all five CI runs because it is five declaratives of similar pressure, which SKILL.md's own Lint fails, and ranked original above over_tight on every one of them. 'restored' survives as trans-06's reference. The pair has not run yet, so it stays arguable until CI shows the choice holds. - disc-11 keeps 'as you assume' in both variants, and disc-12's 'after' drops 'These are hard won lessons, and' while both variants share the shipped closing sentence, so each pair differs in one place. Both are hand edits from the shipped text, recorded in the grading notes and the provenance block, as disc-01's was. - disc-04's variants open with the paragraph that precedes them in c8fa24c, so the connective has something to join. - disc-02, disc-07, disc-08, disc-10 and disc-11 list their accepted alternative rules in accepted_rules. disc-08 and disc-10 accept the breath rule, which the subject cited on every disc-08 run and four of five disc-10 runs; disc-07's note no longer asks the judge for a two-rule answer. - trans-06's task states the mechanism and duration the reference uses, in other words than the reference's, instead of asking for facts only the author knows; trans-02's voice rubric stops naming the reference's words as required. - det-01 and det-02 set min_recall 0: their violation lists are not exhaustive, so recall is recorded as the score and only a trap hit fails them. The floor stays at 0.4 until five CI runs of the new shape exist. Co-Authored-By: Claude Fable 5.1 --- evals/test/generator.test.mjs | 9 ++- home/.agents/skills/prose-register/evals.json | 71 ++++++++++--------- 2 files changed, 46 insertions(+), 34 deletions(-) diff --git a/evals/test/generator.test.mjs b/evals/test/generator.test.mjs index 2a14ca37..d424f6ce 100644 --- a/evals/test/generator.test.mjs +++ b/evals/test/generator.test.mjs @@ -24,7 +24,7 @@ function subjectVisibleValues(item) { function answerKeyValues(item) { const values = []; - for (const key of ['expected_rule', 'expected_rule_for_worst', 'rule_quote', 'grading_note']) { + for (const key of ['expected_rule', 'expected_rule_for_worst', 'rule_quote', 'grading_note', 'reference_after']) { if (typeof item[key] === 'string' && item[key].length > 0) values.push(item[key]); } if (item.rubric && typeof item.rubric === 'object') values.push(...Object.values(item.rubric)); @@ -140,7 +140,7 @@ test('generates all 22 prose-register tests, rank and structural included', asyn ); assert.equal(tests.length, 22); const byType = Object.groupBy(tests, (t) => t.metadata.case_type); - assert.equal(byType['discrimination-rank'].length, 2); + assert.equal(byType['discrimination-rank'].length, 1); assert.equal(byType['discrimination-structural'].length, 2); for (const t of [...byType.discrimination, ...byType['discrimination-structural'], ...byType['discrimination-rank']]) { @@ -179,6 +179,11 @@ test('generates all 22 prose-register tests, rank and structural included', asyn } const det02 = tests.find((t) => t.metadata.case_id === 'det-02'); assert.ok(det02.vars.subject_prompt.includes('{{<')); + + // det-01 and det-02 record recall on a non-exhaustive list; det-03's single em-dash still gates. + for (const t of byType.detection) { + assert.equal(t.vars.min_recall, t.metadata.case_id === 'det-03' ? undefined : 0, t.metadata.case_id); + } }); test('subject prompts never leak an answer-key value the subject cannot already see', async () => { diff --git a/home/.agents/skills/prose-register/evals.json b/home/.agents/skills/prose-register/evals.json index 869e1fde..e3270cda 100644 --- a/home/.agents/skills/prose-register/evals.json +++ b/home/.agents/skills/prose-register/evals.json @@ -11,7 +11,7 @@ "source_file": "content/posts/trust-in-the-age-of-ai.md", "commit_range_mined": "5edc968..5b4cbde", "cases": "disc-01..disc-08, trans-01..trans-05, det-01..det-03", - "verification": "All quoted text was re-pulled directly from `git show :content/posts/trust-in-the-age-of-ai.md` in the source repo on 2026-07-18, not copied from the handoff doc's paraphrases alone. One exception: disc-01's 'after' variant had the modifier 'just' removed by hand on 2026-09-19 (see that case's grading_note). Detection-case violation lists (det-01, det-02) were hand-built against SKILL.md by reading the full draft text end to end, per the handoff doc's own caution not to infer them only from the diff.", + "verification": "All quoted text was re-pulled directly from `git show :content/posts/trust-in-the-age-of-ai.md` in the source repo on 2026-07-18, not copied from the handoff doc's paraphrases alone. Two exceptions, both hand edits (see each case's grading_note): disc-01's 'after' variant had the modifier 'just' removed on 2026-09-19, and disc-04's two variants were each prefixed on 2026-09-20 with the paragraph that precedes the sentence in c8fa24c, so the connective has something to join. Detection-case violation lists (det-01, det-02) were hand-built against SKILL.md by reading the full draft text end to end, per the handoff doc's own caution not to infer them only from the diff.", "excluded_material": "None of the text comes from the 'Workshop and the Job' section (added in f08e2fe, cut in 069e841) -- that section was drafted but never shipped, so it is excluded from these fixtures. See the handoff doc's Caveats section if you want to add it back deliberately." }, { @@ -20,7 +20,7 @@ "source_file": "content/posts/things-ive-changed-my-mind-about.md", "commit_range_mined": "d847318 (polished draft, still carrying raw dictation-transcript blockquotes) -> ef020b0 (\"shorten everything\", automated simplification) -> working tree (author's manual hand-edit on top of ef020b0, uncommitted as of 2026-07-19)", "cases": "disc-09..disc-12, trans-06..trans-07", - "verification": "All quoted text was re-pulled directly from `git show :content/posts/things-ive-changed-my-mind-about.md` (or the working tree, for the uncommitted hand-edit stage) in the source repo on 2026-07-19.", + "verification": "All quoted text was re-pulled directly from `git show :content/posts/things-ive-changed-my-mind-about.md` (or the working tree, for the uncommitted hand-edit stage) in the source repo on 2026-07-19. Three departures since, all on 2026-09-20 and each recorded in its case's grading_note: disc-11's 'after' keeps 'as you assume' from 'before' and disc-12's 'after' drops 'These are hard won lessons, and', so each pair differs in one place; disc-09 was cut from a three-stage rank to the original-versus-over_tight pair, and its 'restored' stage survives as trans-06's reference. trans-06's task now states the facts the reference uses, in other words than the reference's, because the subject cannot know them: on the five CI runs to 2026-09-20 the judge rejected two rewrites for naming no mechanism and one for inventing one. Transformation notes reach the judge, so this history lives here rather than in the case.", "note": "Three stages, not two. ef020b0's automated pass conflates legitimate cleanup (stripping leftover dictation-transcript scaffolding) with genuine over-tightening of the already-polished paragraphs -- only the latter is register-relevant and mined here. The working-tree hand-edit restores some, not all, of what the automated pass flattened; disc-10 documents a concrete detail neither pass restored. One hand-edit was deliberately excluded from these fixtures rather than mined as a positive example -- see the top-level `notes` field below." } ] @@ -30,7 +30,7 @@ "discrimination-rank": "Same as discrimination, but 3 variants ranked instead of a 2-way pick. Can be decomposed into 3 pairwise comparisons if a runner prefers binary grading.", "discrimination-structural": "Same as discrimination, but the distinction is structural (an opening or closing move) rather than sentence-level. Grade more loosely -- see each case's grading_note.", "transformation": "Given a before-passage, produce an on-register rewrite. Grade with a 3-item rubric against the reference after-text: (1) is the specific violation fixed, (2) does the rewrite avoid introducing a new Lint violation, (3) does it read like the author's voice (subjective -- human or LLM-judge spot check, not exact-match).", - "detection": "Given a full passage or document, list register violations with the specific SKILL.md rule each breaks. Grade on recall/precision against the case's hand-built violation list. Each detection case also lists deliberate 'traps' -- lines that look like violations but are explicitly sanctioned by SKILL.md -- to catch over-flagging, not just under-flagging." + "detection": "Given a full passage or document, list register violations with the specific SKILL.md rule each breaks. Grade on recall/precision against the case's hand-built violation list; a case whose list is not exhaustive sets min_recall below 1 and records recall as its score instead of failing on it. Each detection case also lists deliberate 'traps' -- lines that look like violations but are explicitly sanctioned by SKILL.md -- to catch over-flagging, not just under-flagging." }, "cases": [ { @@ -59,7 +59,8 @@ }, "correct": "after", "expected_rule": "Connectives are joints, not filler.", - "grading_note": "'This begs a massive question' inflates the stakes of a simple transition (also touches the 'No grandiosity' prohibition); 'Which raises the question' is the plain joint SKILL.md names verbatim as a good example. Accept either rule as the cited one, but 'Connectives are joints' is the primary match." + "accepted_rules": ["No grandiosity. Don't grant a machine a stake it cannot have."], + "grading_note": "'This begs a massive question' inflates the stakes of a simple transition (also touches the 'No grandiosity' prohibition); 'Which raises the question' is the plain joint SKILL.md names verbatim as a good example. Either rule passes the RULE check; 'Connectives are joints' is the primary match." }, { "id": "disc-03", @@ -79,17 +80,17 @@ { "id": "disc-04", "type": "discrimination", - "prompt": "Which variant is on-register, and which SKILL.md rule decides it?", + "prompt": "Which variant is on-register, and which SKILL.md rule decides it? The opening paragraph is the same in both; the difference is the sentence that follows it.", "rule_quote": "Connectives are joints, not filler.", "source_commits": ["f98e637 (over-cut)", "3c2f396 (restored)"], "variants": { - "before": "How do you trust code a machine wrote?", - "after": "Which raises the question: how do you establish *trust* in code written by a machine?" + "before": "Developer communities did not take it well, and an AI rewriting a low-level runtime out from under everyone is an unsettling thing to sit with. But very soon, if you run tools like Claude locally, you will be executing on that very Rust rewrite. Like it or not, we are becoming users of AI-authored infrastructure.\n\nHow do you trust code a machine wrote?", + "after": "Developer communities did not take it well, and an AI rewriting a low-level runtime out from under everyone is an unsettling thing to sit with. But very soon, if you run tools like Claude locally, you will be executing on that very Rust rewrite. Like it or not, we are becoming users of AI-authored infrastructure.\n\nWhich raises the question: how do you establish *trust* in code written by a machine?" }, "correct": "after", "expected_rule": "Connectives are joints, not filler.", "arguable": true, - "grading_note": "Same underlying sentence as disc-02, one revision cycle later: c8fa24c had already fixed the melodrama, then f98e637's over-tightening pass cut the connective back down to a bare question, then 3c2f396 restored it. Good for showing the rule isn't 'shorter is always better' -- the cut version isn't wrong on any other axis, it's wrong specifically for dropping the joint. KNOWN LIMITATION (found running sonnet against this case on 2026-07-18): this pair is shown with nothing preceding it, so 'Which raises the question:' has nothing textual to connect back to inside the case itself. A model judging it in isolation can reasonably call it 'decorative preamble' rather than recognize it as the connective-joint example -- sonnet did exactly this and picked 'before'. In the real essay the sentence follows two paragraphs of setup, so the joint has something to join. If reusing this case, consider prepending the prior paragraph (see e55f75c in the source repo) so the joint has something to connect to before grading a model's answer as wrong." + "grading_note": "Same underlying sentence as disc-02, one revision cycle later: c8fa24c had already fixed the melodrama, then f98e637's over-tightening pass cut the connective back down to a bare question, then 3c2f396 restored it. Good for showing the rule isn't 'shorter is always better' -- the cut version isn't wrong on any other axis, it's wrong specifically for dropping the joint. Fixture edit (2026-09-20): both variants now open with the paragraph that precedes this sentence in c8fa24c (the same text as trans-02's reference), so 'Which raises the question:' has something to join. Shown alone, the pair lost the subject on 3 of 12 keyed runs (the 2026-07-18 local probe and 2 of the 5 CI runs, both on the compare-first prompt; the six local suite runs all picked the key), each time by reading the connective as decorative preamble with nothing to connect to. In f98e637 the preceding paragraph was itself tighter (see det-02); the c8fa24c text is used for both variants so the pair differs in one place. Kept arguable until CI runs with the context show the choice holds." }, { "id": "disc-05", @@ -131,7 +132,8 @@ }, "correct": "after", "expected_rule": "\"Just\" is banned as a modifier.", - "grading_note": "Watch for a grader that credits all three clauses to the same rule: the first two cuts are 'just' as a modifier (the named prohibition); the third cut ('really') is a hedge-of-cowardice removal under a different rule ('Hedges of precision stay. Hedges of cowardice go.'), not the 'just' prohibition. A precise answer cites both rules, not one rule for all three." + "accepted_rules": ["Hedges of precision stay. Hedges of cowardice go."], + "grading_note": "The first two cuts are 'just' as a modifier (the named prohibition); the third cut ('really') is a hedge-of-cowardice removal under a different rule ('Hedges of precision stay. Hedges of cowardice go.'). Either rule passes the RULE check, and an answer that cites both is the most precise. Until 2026-09-20 this note asked for both and reached the judge as free text; across five CI runs the judge then passed the answers citing one rule and failed the ones citing both." }, { "id": "disc-08", @@ -145,8 +147,9 @@ }, "correct": "headline_device", "expected_rule": "Open with substance. Close without summary.", + "accepted_rules": ["A hard line needs a breath next to it. Vary the pressure, not only the length."], "arguable": true, - "grading_note": "Both openings arguably 'open with substance' -- neither throat-clears. This is the doc's flagged coarse-grained case: every commit before e55f75c used the direct opening; e55f75c replaced it with the headline device and it shipped that way. Treat 'correct' as the author's actual final call, not a strict rule violation in the 'direct' version -- a model arguing 'direct' is also acceptable on baseline grounds is making a defensible case. Grade on whether the answer identifies the real tradeoff (immediate fact vs. rhetorical framing that still front-loads substance) rather than marking 'direct' wrong outright." + "grading_note": "Both openings arguably 'open with substance' -- neither throat-clears. This is the doc's flagged coarse-grained case: every commit before e55f75c used the direct opening; e55f75c replaced it with the headline device and it shipped that way. Treat 'correct' as the author's actual final call, not a strict rule violation in the 'direct' version -- a model arguing 'direct' is also acceptable on baseline grounds is making a defensible case. Grade on whether the answer identifies the real tradeoff (immediate fact vs. rhetorical framing that still front-loads substance) rather than marking 'direct' wrong outright. CI evidence (five runs to 2026-09-20): the subject chose 'headline_device' on every run and cited the breath rule on every run, reading the two block quotes and the run of argument after them as pressure varied on purpose; that rule is accepted, since the headline device is a breath move as much as an opening one." }, { "id": "trans-01", @@ -172,7 +175,7 @@ "rubric": { "violation_fixed": "Is 'here is the hard truth' (throat-clearing) gone? Is the dramatized reaction ('immediate uproar... feels inherently unsettling') replaced with a plainer claim? Is 'forced to become' softened to 'becoming'?", "no_new_violation": "No em-dash, no 'just', no new grandiosity introduced while flattening.", - "voice_match": "Subjective -- the reference keeps 'Like it or not' and 'unsettling'; a rewrite that scrubs all color, not just the op-ed tells, has over-corrected (see det-02 for what over-correction looks like)." + "voice_match": "Subjective -- the paragraph should keep some of its color: the reference holds on to one plain word for the discomfort and one wry aside, in its own words. A rewrite that scrubs every trace of feeling along with the op-ed tells has over-corrected (see det-02 for what over-correction looks like). No particular word from the reference is required." }, "grading_note": "SKILL.md doesn't name 'op-ed tells' as a rule verbatim; the closest textual anchors are baseline 'No throat-clearing' and prohibition 'No grandiosity. Don't grant a machine a stake it cannot have.' Accept either as the cited rule." }, @@ -223,6 +226,7 @@ "prompt": "List every register violation in this draft, quoting the offending text and naming the specific SKILL.md rule it breaks. Note anything that looks like a violation but is not.", "source_commit": "5edc968", "source_label": "Initial draft -- op-ed register throughout, before any register editing pass.", + "min_recall": 0, "input_document": "Recently, the creator of Bun burned through roughly $200,000 worth of tokens to rewrite their core runtime from Zig into Rust. The migration was completed in mere weeks, with 99% of the commits authored by a pre-release AI model.\n\nIt caused an immediate uproar in developer communities. The idea of an AI unilaterally rewriting a low-level systems programming language runtime feels inherently unsettling. But here is the hard truth: very soon, if you run tools like Claude locally on your machine, it will be executing on that very Rust rewrite. Like it or not, we are being forced to become users of AI-authored infrastructure.\n\nThis begs a massive question: How do we establish *trust* with code written by a machine?\n\n### The Formula for Trust\n\nYears ago, a consultant named Chris from Middle Path Consulting gave me a formula that has stuck in my head ever since. As a math nerd, I love it.\n\n**Trust = (Competence + Character + Caring) / Risk**\n\nTrust is not absolute; it is entirely dependent on context and risk. If I ask you to toss an empty coffee cup into a nearby trash can, the risk is practically zero. If you miss, the cup is just sitting on the floor. I can trust literally anyone on the street with that task.\n\nBut if I ask you to manage a critical production database, or watch my child, the risk is immense. To maintain trust in a high-risk scenario, the three variables in the numerator must be overwhelmingly high:\n\n- **Competence:** Do you have the actual capability to handle the task?\n- **Character:** Are you forthright? Do your values align with mine?\n- **Caring:** Do you actually care if the job is done well, or are you just phoning it in?\n\nIf any one of those three variables goes to zero, trust evaporates.\n\n### AI as the New Teammate\n\nThis formula is the perfect lens for viewing our current adoption of AI tooling.\n\nFor the last year, my AI usage has skewed heavily toward low-risk tasks—throwaway games, little toys, boilerplate scripts. For these tasks, the AI's Competence is usually sufficient, and because the Risk is so low, I don't really need to worry about Character or Caring. The coffee cup makes it into the trash, or it doesn't.\n\nBut rewriting a runtime that will execute on millions of machines? That is incredibly high-risk. And while an AI might have the **Competence** to write idiomatic Rust, how do we evaluate the other two Cs?\n\n- **Caring:** For an AI, caring translates to diligence and attention to detail. Does it thoughtfully review its own code for edge cases, or does it just confidently blast out the first statistically probable solution?\n- **Character:** For an AI, character is alignment and safety. Does it follow strict architectural values, or does it quietly introduce subtle vulnerabilities?\n\nAn AI cannot intrinsically \"care,\" and its \"character\" is limited to the guardrails placed upon it by its creators. Therefore, if you are trusting an AI with a high-risk task, the human pilot must entirely absorb the responsibility for those two traits.\n\n### The End of Solo Authorship\n\nIf I subcontract the Competence of writing code to an AI, I am still 100% on the hook for the Character and Caring of the final product.\n\nThis is exactly why I've recently started committing under an alter-ego GitHub profile (\"non-reagent\") when working heavily with AI. It allows me to explicitly separate my human, solo-authored commits from my AI-assisted workflows.\n\nThe era of solo authorship is over. Moving forward, the \"integrity of the author\" no longer means proving that you typed every single character by hand. It means being completely transparent about *how* you collaborate with AI. It means proving to your users that while the machine provided the Competence, a human being was there to provide the Character and the Caring.\n\n*(I'm currently formalizing these thoughts into an AI manifesto, which I am drafting in a separate agent session and will publish here soon).*", "violations": [ {"quote": "It caused an immediate uproar in developer communities. The idea of an AI unilaterally rewriting a low-level systems programming language runtime feels inherently unsettling.", "rule": "No grandiosity. Don't grant a machine a stake it cannot have.", "fixed_in": "c8fa24c"}, @@ -243,7 +247,7 @@ {"quote": "Therefore, if you are trusting an AI with a high-risk task, the human pilot must entirely absorb the responsibility for those two traits.", "why_not_a_violation": "SKILL.md lists 'Therefore,' verbatim as a good connective/joint."}, {"quote": "Recently, the creator of Bun burned through roughly $200,000 worth of tokens to rewrite their core runtime from Zig into Rust.", "why_not_a_violation": "The opening correctly leads with substance (no throat-clearing). The dollar figure itself is factually wrong ($200,000 vs. the real $165,000), but that is a fact-check issue, not a register violation -- don't let a model conflate the two when grading detection."} ], - "grading_note": "12 violations, 3 traps. Grade recall (how many of the 12 were found) and precision (were any traps wrongly flagged, or was a genuinely fine sentence flagged for no stated reason)." + "grading_note": "12 violations, 3 traps. The list is one reading of the draft, not an exhaustive one: the five CI runs to 2026-09-20 score 6 to 8 of the 12 under this matcher while flagging other real instances the list omits. Recall is recorded as the score (min_recall 0) rather than gated, and a flagged trap still fails the case. Grade precision by reading the output: were any traps wrongly flagged, or was a genuinely fine sentence flagged for no stated reason?" }, { "id": "det-02", @@ -251,6 +255,7 @@ "prompt": "List every register violation in this draft, quoting the offending text and naming the specific SKILL.md rule it breaks. Note anything that looks like a violation but is not.", "source_commit": "f98e637", "source_label": "Over-tightening pass -- the op-ed tells from the initial draft are already gone (fixed in c8fa24c); this commit's failure mode is the opposite one: cutting connectives and breath until the prose is all jab.", + "min_recall": 0, "input_document": "The creator of Bun spent $165,000 in tokens to rewrite his runtime from Zig to Rust. It took eleven days. A pre-release Claude Fable 5 wrote nearly all of the 6,778 commits.\n\nDevelopers did not take it well. Soon, though, if you run Claude on your own machine, it will run on that rewrite. We are becoming users of infrastructure no human wrote.\n\nHow do you trust code a machine wrote?\n\n### The Formula for Trust\n\nYears ago a consultant named Chris, from Middle Path Consulting, gave me a formula that stuck. I am a math nerd. I love it.\n\n{{< math >}}T = \\frac{3C}{R}{{< /math >}}\n\nThree C's over Risk: Competence, Character, and Caring on top, Risk underneath. Trust is not absolute. It bends to context and to risk. Ask me to toss an empty cup into a trash can and the risk is nothing; miss, and the cup sits on the floor. I would trust anyone on the street with that.\n\nAsk me to let you run a production database, or watch my child, and the risk is enormous. Now the three terms on top have to be overwhelming.\n\n- **Competence.** Can you actually do it?\n- **Character.** Are you honest? Do your values match mine?\n- **Caring.** Do you care whether it is done well, or are you phoning it in?\n\nZero out any one of them and trust is gone.\n\n### AI as the New Teammate\n\nThe formula fits a machine as well as it fits a person.\n\nFor a year I have pointed AI at low-risk work: throwaway games, toys, boilerplate. The Competence is enough, and the Risk runs so low that Character and Caring never come up. The cup lands in the trash or it doesn't.\n\nA runtime on millions of machines is another matter. The machine may have the Competence to write clean Rust. The other two C's are the hard ones.\n\n- **Caring.** For a machine, caring is diligence. Does it hunt its own edge cases, or fire off the first answer that scores well?\n- **Character.** For a machine, character is alignment. Does it hold the architecture, or slip in a quiet hole?\n\nA machine cannot care. Its character stops at the guardrails its makers gave it. On a high-risk job, the human supplies both terms. All of it.\n\n### A Hundred People for Eleven Days\n\nRun the counterfactual. Same rewrite, same eleven days, but a hundred engineers do it on nights and weekends. We would argue about the crunch. We would not feel what we feel about the machine. Why is the machine worse?\n\nThe speed is the same. The risk is the same. What the hundred bring that the machine cannot is themselves. A hundred engineers are a hundred minds, and each one understood a piece, chose it, and can answer for it later. The knowledge lives in people who were there. When the machine does the work, the code exists and the understanding does not. A working runtime, and no one who holds it. That is the new thing, and the cold one: knowledge with no knower.\n\nThere is a second thing. Nights and weekends are Caring you can see. They cost someone something, and we read the cost as belief. The machine's version cost $165,000. You buy that. No one bled for it. We know how to trust care that hurt to give. We have not learned to trust care that came off an invoice.\n\nPut the formula on it and the unease has a name. A human team hands you the whole numerator in one body: Competence, bound to the Character and Caring of the people who did the work. The machine hands you Competence alone. No one stands behind it. For the first time you can buy that term by itself. The work no longer implies a worker.\n\n### The End of Solo Authorship\n\nHand the Competence to a machine and the Character and the Caring are still yours. Every bit.\n\nSo I have started committing under a second name, \"non-reagent,\" when I lean hard on AI. It keeps my solo work and my machine-assisted work on separate ledgers.\n\nSolo authorship is over. The integrity of the author no longer means you typed every character by hand. It means you are honest about how you worked with the machine. It means showing your readers that a person supplied the Character and the Caring the machine could not.\n\n*(I'm currently formalizing these thoughts into an AI manifesto, which I am drafting in a separate agent session and will publish here soon).*", "violations": [ {"quote": "The creator of Bun spent $165,000 in tokens to rewrite his runtime from Zig to Rust. It took eleven days. A pre-release Claude Fable 5 wrote nearly all of the 6,778 commits.", "rule": "A hard line needs a breath next to it... Three declaratives in a row are a drumbeat. / Lint: \"Five consecutive sentences of similar length or similar pressure.\"", "fixed_in": "24d944d (paragraph rewritten with breath restored)"}, @@ -265,7 +270,7 @@ {"quote": "Zero out any one of them and trust is gone.", "why_not_a_violation": "A single hard line landing after a bulleted list is fine -- it is not preceded by other short declaratives in a row. The rule targets runs of three or more, not any short sentence."}, {"quote": "The speed is the same. The risk is the same.", "why_not_a_violation": "Only two short declaratives before a longer explanatory sentence follows ('What the hundred bring...'). SKILL.md's stated threshold is three declaratives in a row as a drumbeat; two stays under that line."} ], - "grading_note": "7 violations (5 originally documented + 2 added after a live run exposed the gap -- see their 'note' fields; likely still not exhaustive, the jab pattern recurs elsewhere in this draft too), 2 traps. This case is a harder detection exercise than det-01: most violations here are about pacing and breath, not a bright-line Lint item like an em-dash or a banned word, so grading precision (not over-flagging every short sentence) matters as much as recall. METHODOLOGY NOTE (2026-07-18): run-evals.mjs grades detection cases by exact substring match against this violations list. A sonnet run on the original 5-item list scored 0/5 despite correctly identifying 3 real violations (including the 2 added above), because it reasonably chose different valid instances of the same rule categories than the ones originally listed -- substring matching against one fixed quote list underscores recall whenever a document has more true violations than the list captures. Read the raw subjectResponse for this case type rather than trusting the found/total count alone." + "grading_note": "7 violations (5 originally documented + 2 added after a live run exposed the gap -- see their 'note' fields; likely still not exhaustive, the jab pattern recurs elsewhere in this draft too), 2 traps. This case is a harder detection exercise than det-01: most violations here are about pacing and breath, not a bright-line Lint item like an em-dash or a banned word, so grading precision (not over-flagging every short sentence) matters as much as recall. METHODOLOGY NOTE (2026-07-18, updated 2026-09-20): the grader matches each listed quote against the subject's quotes by text overlap. A sonnet run on the original 5-item list scored 0/5 despite correctly identifying 3 real violations (including the 2 added above), because it reasonably chose different valid instances of the same rule categories than the ones originally listed -- a fixed quote list under-scores recall whenever a document has more true violations than the list captures, and the five CI runs to 2026-09-20 score 0 to 4 of 7 under this matcher. Recall is therefore recorded as the score (min_recall 0) rather than gated; a flagged trap still fails the case. Read the raw output for this case type rather than trusting the found/total count alone." }, { "id": "det-03", @@ -282,21 +287,21 @@ }, { "id": "disc-09", - "type": "discrimination-rank", - "prompt": "Rank these three versions of the same paragraph from most to least on-register, and name the rule the worst version breaks.", + "type": "discrimination", + "prompt": "Which variant is on-register, and which SKILL.md rule decides it?", "rule_quote": "Concrete nouns over abstractions. Strong verbs over adverbs. / Hedges of precision stay. Hedges of cowardice go.", - "source_commits": ["d847318 (original)", "ef020b0 (over-tight, \"shorten everything\")", "working tree (restored, hand-edited, uncommitted as of 2026-07-19)"], + "source_commits": ["d847318 (before)", "ef020b0 (after, \"shorten everything\")"], "source_repo": "~/src/alannorton.com", "source_file": "content/posts/things-ive-changed-my-mind-about.md", - "stages": { - "original": "I doubted that in 2013, when Betterment hired a Ruby enthusiast, let him rewrite the core of the investor experience in Rails, and watched him become CTO a few years later. Nobody talked me out of my skepticism; the business ran the numbers and decided the return was worth the risk regardless. Custody and trading stayed in Java and Scala, and the Ruby services ran reliably for years anyway. Ifs are ifs, logic is logic, and correct is correct no matter what language spells it out. What isn't fixed is whether writing it feels like a chore or a delight, and that joy turns out to matter exactly as much as the function it serves.", - "over_tight": "At Betterment, plenty of us were sure Ruby had no business near money, yet big pieces of the investor experience ran on Ruby for years without trouble. Safety comes from how you prove the code correct, not from the language you wrote it in.", - "restored": "Betterment began as a Java shop. When we considered replatforming in Ruby, plenty of us were wary. Yes, we decomposed our Java monolith, bit by bit, into distributed domain services on Rails. The biggest pieces of the investor experience have been running on Ruby for nearly a decade. Safety comes from how you prove the code correct, not from the language you write it in." + "variants": { + "before": "I doubted that in 2013, when Betterment hired a Ruby enthusiast, let him rewrite the core of the investor experience in Rails, and watched him become CTO a few years later. Nobody talked me out of my skepticism; the business ran the numbers and decided the return was worth the risk regardless. Custody and trading stayed in Java and Scala, and the Ruby services ran reliably for years anyway. Ifs are ifs, logic is logic, and correct is correct no matter what language spells it out. What isn't fixed is whether writing it feels like a chore or a delight, and that joy turns out to matter exactly as much as the function it serves.", + "after": "At Betterment, plenty of us were sure Ruby had no business near money, yet big pieces of the investor experience ran on Ruby for years without trouble. Safety comes from how you prove the code correct, not from the language you wrote it in." }, - "correct_ranking": ["restored", "original", "over_tight"], - "expected_rule_for_worst": "Concrete nouns over abstractions, strong verbs over adverbs -- 'plenty of us were sure Ruby had no business near money' and 'ran on Ruby for years without trouble' assert the outcome with no mechanism and a loose duration where the sibling versions both have one.", + "correct": "before", + "expected_rule": "Concrete nouns over abstractions. Strong verbs over adverbs.", + "accepted_rules": ["Hedges of precision stay. Hedges of cowardice go."], "arguable": true, - "grading_note": "Unlike disc-03, 'restored' is not a reconstruction of 'original' -- it's an independent rewrite with different concrete details (a decomposed monolith, distributed domain services, 'nearly a decade' instead of 'a few years later... for years anyway'). Both clearly beat 'over_tight', which strips out every proper noun, mechanism, and duration in favor of a generic claim. Ranking 'restored' first is defensible on precision grounds ('nearly a decade' is a tighter hedge than original's looser timeline) and because this essay's own editorial history (commit 7957f74, 'cut changed-my-mind-about essay by 25%') suggests 'original's five flowing sentences, however good the closing aphorism ('Ifs are ifs, logic is logic, and correct is correct'), may be too indulgent for this compressed format. A model that ranks 'original' first is not wrong -- it's defending the aphorism and the CTO narrative beat, both genuinely strong -- but should still place 'over_tight' last and name the same rule for why." + "grading_note": "Like disc-10, the later version is the violation: the 'shorten everything' pass strips every proper noun, mechanism, and duration in favor of a generic claim, where 'before' backs the same claim with the hire, the Rails rewrite, the CTO beat, and the Java and Scala services that stayed. The hedge rule is accepted as an alternative: 'for years without trouble' against 'a few years later... ran reliably for years' is a precision difference as well as a concreteness one. Until 2026-09-20 this was a three-stage rank with the author's hand-edited paragraph ('restored', now trans-06's reference) keyed first. The subject ranked it last on all five CI runs, with the same reasoning each time: five consecutive declaratives of similar pressure, which SKILL.md's own Lint fails (the six local runs' outputs do not survive; their 0 of 10 with CI records only that the key ordering never appeared). The key contradicted the skill, so the case was cut to the pair the subject and the key agree on: the subject ranked 'original' above 'over_tight' on all five CI runs. The pair itself has not run; it stays arguable until CI runs on it show the choice holds." }, { "id": "disc-10", @@ -312,7 +317,8 @@ }, "correct": "before", "expected_rule": "Concrete nouns over abstractions. Strong verbs over adverbs.", - "grading_note": "The rare case in this set where the automated pass makes things worse, not better -- worth keeping precisely because a model shouldn't assume 'after' is always the answer. 'before' backs its claim with an actual, slightly funny, real method name; 'after' asserts the identical claim ('the code read like a sentence') with no evidence at all. This cut was never restored in the later hand-edit pass either (checked against the source repo's working tree as of 2026-07-19) -- flag that to a human if this case is used to test whether a model notices an opportunity to restore, not just judges two fixed variants." + "accepted_rules": ["A hard line needs a breath next to it. Vary the pressure, not only the length."], + "grading_note": "One of two cases in this set (with disc-09) where the automated pass makes things worse, not better -- worth keeping precisely because a model shouldn't assume 'after' is always the answer. 'before' backs its claim with an actual, slightly funny, real method name; 'after' asserts the identical claim ('the code read like a sentence') with no evidence at all. This cut was never restored in the later hand-edit pass either (checked against the source repo's working tree as of 2026-07-19) -- flag that to a human if this case is used to test whether a model notices an opportunity to restore, not just judges two fixed variants. CI evidence (five runs to 2026-09-20): the choice was right on every run, and the RULE line cited the breath rule on four runs (the method name being the breath in a paragraph of claims) and 'No flex' on one; the breath rule is accepted." }, { "id": "disc-11", @@ -324,12 +330,13 @@ "source_file": "content/posts/things-ive-changed-my-mind-about.md", "variants": { "before": "No one is watching your work as closely as you assume. Waiting to be noticed is a strategy for being underleveled with great reviews.", - "after": "No one is watching your work as closely as you. Waiting to be noticed is a strategy for staying exactly where you are." + "after": "No one is watching your work as closely as you assume. Waiting to be noticed is a strategy for staying exactly where you are." }, "correct": "after", "expected_rule": "No jargon without definition.", + "accepted_rules": ["Concrete nouns over abstractions. Strong verbs over adverbs."], "arguable": true, - "grading_note": "'Underleveled' is HR/performance-review jargon the essay never defines anywhere, for a general readership that shouldn't need to already know it. Fills a documented coverage gap: the trust essay's mined history never caught a live jargon violation being fixed (see the coverage_gaps entry this closes). Note the fix here is CUT, not DEFINE -- 'underleveled' disappears rather than getting a gloss. Accept an answer that names 'No jargon without definition' even though the resolution technique differs from an in-text definition; also accept 'Concrete nouns over abstractions' for 'staying exactly where you are' being more vivid than the jargon term, since both readings are defensible. This term survived unchanged through the automated 'shorten everything' pass (ef020b0) -- only the hand-edit caught it, worth noting if this case is used to argue automated tightening alone isn't sufficient register QA." + "grading_note": "Fixture edit (2026-09-20): the hand-edited source also changed 'as closely as you assume' to 'as closely as you'; 'after' here keeps 'as you assume' so the pair differs only in the jargon clause. Shown with both differences, the subject decided the case on the first one on 2 of 5 CI runs (both chose 'after'). The two wrong choices came from runs that engaged the jargon clause and read 'underleveled with great reviews' as the concrete detail; the one-difference pair does not change that reading, which is why the key stays arguable. 'Underleveled' is HR/performance-review jargon the essay never defines anywhere, for a general readership that shouldn't need to already know it. Fills a documented coverage gap: the trust essay's mined history never caught a live jargon violation being fixed (see the coverage_gaps entry this closes). Note the fix here is CUT, not DEFINE -- 'underleveled' disappears rather than getting a gloss. Accept an answer that names 'No jargon without definition' even though the resolution technique differs from an in-text definition; also accept 'Concrete nouns over abstractions' for 'staying exactly where you are' being more vivid than the jargon term, since both readings are defensible. This term survived unchanged through the automated 'shorten everything' pass (ef020b0) -- only the hand-edit caught it, worth noting if this case is used to argue automated tightening alone isn't sufficient register QA." }, { "id": "disc-12", @@ -340,19 +347,19 @@ "source_repo": "~/src/alannorton.com", "source_file": "content/posts/things-ive-changed-my-mind-about.md", "variants": { - "before": "Changing your mind is the work. Read these back to back and they rhyme: I overvalued the artifact and undervalued the people in the room. Correctness over understanding, the code over the company, sounding smart over making anyone else smarter.\n\nThere will be more. The next version of me is already reading this and shaking his head, adding entries to the README.", - "after": "Changing your mind is the work. These are hard won lessons, and there will be more. The next version of me is already re-reading this, shaking his head, and adding entries to the README." + "before": "Changing your mind is the work. Read these back to back and they rhyme: I overvalued the artifact and undervalued the people in the room. Correctness over understanding, the code over the company, sounding smart over making anyone else smarter.\n\nThere will be more. The next version of me is already re-reading this, shaking his head, and adding entries to the README.", + "after": "Changing your mind is the work. There will be more. The next version of me is already re-reading this, shaking his head, and adding entries to the README." }, "correct": "after", "expected_rule": "Open with substance. Close without summary.", "arguable": true, - "grading_note": "'before' spells out the essay's own throughline in plain restatement ('I overvalued the artifact and undervalued the people in the room... Correctness over understanding, the code over the company...') -- a textbook summary-close, the exact move the baseline rule bans. 'after' cuts that sentence entirely and lands on the same closing image (the next version of me shaking his head) without re-explaining what the reader just read. This is a cleaner, essay-sourced example of 'close without summary' than anything in the trust-essay mining, which only had an open-with-substance example (disc-08). CI evidence (2026-09-19): the subject chose 'before' on all four runs on a bare CI runner, reading A's recap triplet as a hard line with a breath, against 6 of 6 correct in local bare-$HOME suite runs. The cause of that gap is open, so the key does not gate hard until it is understood." + "grading_note": "'before' spells out the essay's own throughline in plain restatement ('I overvalued the artifact and undervalued the people in the room... Correctness over understanding, the code over the company...') -- a textbook summary-close, the exact move the baseline rule bans. 'after' cuts that sentence entirely and lands on the same closing image (the next version of me shaking his head) without re-explaining what the reader just read. This is a cleaner, essay-sourced example of 'close without summary' than anything in the trust-essay mining, which only had an open-with-substance example (disc-08). Fixture edit (2026-09-20): the hand-edited source reads 'These are hard won lessons, and there will be more.' On all five CI runs to that date the subject chose 'before', reading 'These are hard won lessons' as a summarizing label of its own and 'before's recap triplet as a hard line with a breath; 'after' here drops that clause, and the closing sentence uses the shipped wording in both variants (ef020b0 read 'already reading this and shaking his head, adding entries'), so the pair differs only by the recap sentence the rule bans and the paragraph break that goes with it. Six local bare-$HOME suite runs on the unedited pair had all chosen 'after'; the cause of that CI-versus-local gap is open (see the eval spec), so the key stays arguable until CI runs on the edited pair show the choice holds." }, { "id": "trans-06", "type": "transformation", "input": "At Betterment, plenty of us were sure Ruby had no business near money, yet big pieces of the investor experience ran on Ruby for years without trouble. Safety comes from how you prove the code correct, not from the language you wrote it in.", - "task": "This paragraph makes its claim in the abstract -- no mechanism, no duration more precise than 'for years'. Rewrite it with a concrete mechanism (what actually happened to the Java code) and a duration that reads as exact, per 'concrete nouns over abstractions' and 'hedges of precision stay, hedges of cowardice go'.", + "task": "This paragraph makes its claim in the abstract -- no mechanism, no duration more precise than 'for years'. The facts: the company started on Java; its Java monolith was later broken up piece by piece into separate domain services written in Rails; the largest parts of the investor experience have run on Ruby for close to ten years. Rewrite the paragraph so it names that mechanism and a duration that reads as exact, per 'concrete nouns over abstractions' and 'hedges of precision stay, hedges of cowardice go'.", "reference_after": "Betterment began as a Java shop. When we considered replatforming in Ruby, plenty of us were wary. Yes, we decomposed our Java monolith, bit by bit, into distributed domain services on Rails. The biggest pieces of the investor experience have been running on Ruby for nearly a decade. Safety comes from how you prove the code correct, not from the language you write it in.", "source_commits": ["ef020b0 (before)", "working tree (after, hand-edited)"], "source_repo": "~/src/alannorton.com", @@ -362,7 +369,7 @@ "no_new_violation": "No em-dash; watch for the rewrite drifting into a decorative triplet or an unearned 'not X, but Y' -- the reversal here (Java shop -> Ruby) is factual, not rhetorical, so it doesn't need the signature move.", "voice_match": "Subjective -- the reference uses 'Yes,' as a connective/joint (an affirming beat, not a transition-into-a-question, the only flavor SKILL.md's own examples show); a rewrite that finds a different joint is fine, one with no joint at all is missing the texture." }, - "grading_note": "Same underlying paragraph as disc-09's three-stage case, isolated as a single before/after transformation. The reference answer is the author's own hand-edit, not a synthetic ideal -- treat divergent-but-equally-concrete rewrites as passing, per trans-04's precedent." + "grading_note": "Same underlying paragraph as disc-09 (a three-stage rank until 2026-09-20, now a pair), isolated as a single before/after transformation. The reference answer is the author's own hand-edit, not a synthetic ideal -- treat divergent-but-equally-concrete rewrites as passing, per trans-04's precedent." }, { "id": "trans-07", @@ -391,5 +398,5 @@ {"rule": "Lint: \"A triplet whose third item adds nothing the first two didn't\"", "gap": "The essay only shows the good triplet (disc-05); no essay-sourced bad triplet exists (SKILL.md's own contrast example, 'remix, retool, reimagine', is not drawn from this essay -- it would need to be authored synthetically for a negative example)."}, {"rule": "\"A hedge that protects the writer rather than the fact\" / \"A reversal used for cleverness rather than argument\"", "gap": "No bad examples survive in this history -- only good ones (hedges of precision, argument-driven reversals) made it to the shipped text, which makes sense for a piece mined after it shipped. Negative examples for these two would need to be authored synthetically or mined from a rougher draft of a different essay."} ], - "notes": "No eval convention existed in this repo before this file. Format choice (structured JSON over a narrative markdown rubric) was made 2026-07-18 so cases are machine-gradable by a future runner script or fed directly into an LLM-judge prompt; grading logic itself is not implemented here, only the cases and rubrics. A second source was mined 2026-07-19 from content/posts/things-ive-changed-my-mind-about.md (see provenance.sources) -- a genuine three-stage revision arc (polished draft -> automated 'shorten everything' pass -> manual hand-edit) rather than the trust essay's long single-commit history. One hand-edit from that pass was initially flagged rather than mined: the 'appearing smart' section's closing paragraph replaces 'They used small words and showed up curious' with 'They focused on expressing ideas accessibly.' On review with the author (2026-07-19) this was a misapplication of the register, not a real violation -- the rule was never 'no big words', it's 'no jargon without definition' / don't wall the reader out; SKILL.md's Prohibitions section now says so explicitly ('Big words are fine when they're accessible'). 'Focused on expressing ideas accessibly' names the actual mechanism (accessibility) rather than a proxy for it (word length), which is arguably the more precise choice, not a regression. Not mined as a discrimination case even so -- there's no clean single right answer to rank, since both phrasings are defensible once the 'word-size is the metric' framing is dropped. 'without being committed co-owners in the solution' (replacing 'not like they built it with you') is a separate, weaker clause in the same sentence and wasn't part of the author's clarification -- still worth a second look, just not for the jargon reason. Case count: 12 discrimination (14 gradable comparisons, since disc-03 and disc-09 are 3-way ranks), 7 transformation, 3 detection (20 hand-built violations + 5 traps total across the three documents)." + "notes": "No eval convention existed in this repo before this file. Format choice (structured JSON over a narrative markdown rubric) was made 2026-07-18 so cases are machine-gradable by a future runner script or fed directly into an LLM-judge prompt; grading logic itself is not implemented here, only the cases and rubrics. A second source was mined 2026-07-19 from content/posts/things-ive-changed-my-mind-about.md (see provenance.sources) -- a genuine three-stage revision arc (polished draft -> automated 'shorten everything' pass -> manual hand-edit) rather than the trust essay's long single-commit history. One hand-edit from that pass was initially flagged rather than mined: the 'appearing smart' section's closing paragraph replaces 'They used small words and showed up curious' with 'They focused on expressing ideas accessibly.' On review with the author (2026-07-19) this was a misapplication of the register, not a real violation -- the rule was never 'no big words', it's 'no jargon without definition' / don't wall the reader out; SKILL.md's Prohibitions section now says so explicitly ('Big words are fine when they're accessible'). 'Focused on expressing ideas accessibly' names the actual mechanism (accessibility) rather than a proxy for it (word length), which is arguably the more precise choice, not a regression. Not mined as a discrimination case even so -- there's no clean single right answer to rank, since both phrasings are defensible once the 'word-size is the metric' framing is dropped. 'without being committed co-owners in the solution' (replacing 'not like they built it with you') is a separate, weaker clause in the same sentence and wasn't part of the author's clarification -- still worth a second look, just not for the jargon reason. Case count: 12 discrimination (13 gradable comparisons, since disc-03 is a 3-way rank), 7 transformation, 3 detection (20 hand-built violations + 5 traps total across the three documents)." } From ce57e630124acbb55cf0d0a625af23da3ae90868 Mon Sep 17 00:00:00 2001 From: Agent Norton Date: Sun, 20 Sep 2026 06:24:43 +0000 Subject: [PATCH 5/7] Record decision 12, the fixture pass, in the eval spec accepted_rules, min_recall and the overlap matcher join the Phase 2 revisions as decision 12, with the review memo as their evidence. The grading section describes what the rule judge and the detection assert see now, and the After Phase 2 paragraph drops the two keys the pass settled and names the measurement the suite still does not make: no row samples a draft written under the skill. Co-Authored-By: Claude Fable 5.1 --- .../specs/2026-08-29-skill-eval-suite-design.md | 15 ++++++++------- 1 file changed, 8 insertions(+), 7 deletions(-) diff --git a/docs/superpowers/specs/2026-08-29-skill-eval-suite-design.md b/docs/superpowers/specs/2026-08-29-skill-eval-suite-design.md index 4b84c5cd..d0481ec5 100644 --- a/docs/superpowers/specs/2026-08-29-skill-eval-suite-design.md +++ b/docs/superpowers/specs/2026-08-29-skill-eval-suite-design.md @@ -1,7 +1,7 @@ # Skill Eval Suite: Promptfoo Runner for Per-Skill Evals **Date:** 2026-08-29\ -**Status:** Approved design; Phase 1 shipped (#58); Phase 2 revision approved 2026-09-19 +**Status:** Approved design; Phase 1 shipped (#58); Phase 2 revision approved 2026-09-19; fixture pass (decision 12) 2026-09-20 ## TL;DR @@ -35,8 +35,9 @@ Phase 1's keyed runs surfaced one flaky grader, and CI planning surfaced an auth 7. **The CI full-run job starts advisory.** Its thresholds are still Phase 1 guesses; it reports on skill PRs without blocking merge until real runs tune them. 8. **Only the discrimination choice gates hard.** Detection and every judged assert count toward the pass-rate floor; see Pass criteria. 9. **Transformation rubrics list whatever criteria the case defines.** `prose-register` uses `voice_match` where `code-comment-register` uses `placement`; the rubric builder stops hardcoding three keys. -10. **A case can mark its answer key arguable.** `arguable: true` in evals.json, with a grading note saying why, makes a wrong choice count against the pass-rate floor instead of hard-failing the job; a subject that returns no answer still gates. It applies only to discrimination cases. The flag is for a key that cannot gate hard: the author's note concedes another answer is defensible, or the key does not reproduce across environments. Five prose-register cases carry it. Four are flagged on their notes (disc-04, disc-08, disc-09, disc-11). Across ten keyed runs (six local repeats under a bare `$HOME`, four CI runs) the subject picked the key in 6 of 10 for disc-11, 7 of 10 for disc-08, 9 of 10 for disc-04 and 0 of 10 for disc-09; disc-04 is flagged on its note and on the parent PR's report that it flipped once in four earlier runs. disc-12 is flagged on evidence: the CI subject picked the wrong answer on all four bare-runner runs (reading the recap triplet in `before` as a hard line with a breath), while all six local suite runs got it right. See "CI versus local" under After Phase 2. -11. **Each skill can set its own pass-rate floor.** `min_pass_rate` at the top of evals.json overrides the 90% default; argv overrides both, and a floor outside (0, 1] is rejected. The values come from CI runs only, at one row below the worst run so far. prose-register: 12, 10, 12 and 10 of 22 (54.5, 45.5, 54.5, 45.5%), so 0.4 (9 of 22). code-comment-register: 15, 15, 13 and 13 of 16 (93.8, 93.8, 81.3, 81.3%), so 0.75 (12 of 16); its 90% default failed two of four runs on judge-rule and detection misses alone. prose-register's two detection cases fail on every run by design, which alone leaves 20 of 22 rows. These floors catch a collapse, not a drift, and rest on four runs each. Local runs are not a valid source: they scored 59 to 73% on prose-register. +10. **A case can mark its answer key arguable.** `arguable: true` in evals.json, with a grading note saying why, makes a wrong choice count against the pass-rate floor instead of hard-failing the job; a subject that returns no answer still gates. It applies only to discrimination cases. The flag is for a key that cannot gate hard: the author's note concedes another answer is defensible, or the key does not reproduce across environments. Five prose-register cases carry it. Four are flagged on their notes (disc-04, disc-08, disc-09, disc-11). Across ten keyed runs (six local repeats under a bare `$HOME`, four CI runs) the subject picked the key in 6 of 10 for disc-11, 7 of 10 for disc-08, 9 of 10 for disc-04 and 0 of 10 for disc-09; disc-04 is flagged on its note and on the parent PR's report that it flipped once in four earlier runs. disc-12 is flagged on evidence: the CI subject picked the wrong answer on all four bare-runner runs (reading the recap triplet in `before` as a hard line with a breath), while all six local suite runs got it right. See "CI versus local" under After Phase 2. Decision 12 re-cut disc-09 to a two-way pair and edited disc-04, disc-11 and disc-12; all five flags stay until CI runs on the edited cases show the choices hold. +11. **Each skill can set its own pass-rate floor.** `min_pass_rate` at the top of evals.json overrides the 90% default; argv overrides both, and a floor outside (0, 1] is rejected. The values come from CI runs only, at one row below the worst run so far. prose-register: 12, 10, 12 and 10 of 22 (54.5, 45.5, 54.5, 45.5%), so 0.4 (9 of 22). code-comment-register: 15, 15, 13 and 13 of 16 (93.8, 93.8, 81.3, 81.3%), so 0.75 (12 of 16); its 90% default failed two of four runs on judge-rule and detection misses alone. prose-register's two detection cases failed on every run by design, which alone left 20 of 22 rows (decision 12 has them record recall and pass instead). These floors catch a collapse, not a drift, and rest on four runs each. Local runs are not a valid source: they scored 59 to 73% on prose-register. +12. **A key the subject disagrees with on every run is re-cut, not flagged.** A review of five CI runs at the tip of the fixes branch ([2026-09-20-register-eval-review.md](2026-09-20-register-eval-review.md)) attributed 11 of the tip run's 12 failures to the eval side: seven cases failed on every run with the same reasoning, which is a key the model disagrees with, not sampling noise. Three mechanisms follow, all built with zero model sessions. (a) `accepted_rules`, a non-empty string array on discrimination cases, lists the alternatives the rule judge accepts; the rubric shows the judge the reference rule, the skill's rule text and that list, and `grading_note` is human-facing again (disc-07's note asked for a two-rule answer, and the judge read it as a requirement and failed the answers that met it). code-comment-register's discrimination notes explain the key rather than list alternatives, so none migrates, and its rule judge loses that context; its run on the fixture-pass PR measures the effect. (b) `min_recall` on detection cases, a number in [0, 1] defaulting to 1, floors recall; below 1 the case records recall as its score, and a flagged trap still fails it. det-01 and det-02 set 0, since their lists are not exhaustive. The detection grader also matches each key quote against the subject's quotes by overlap rather than by the key's first 60 characters, which scored det-03 a miss on two runs where the subject had dropped the key's first word; one subject quote is one finding, credited to every key it holds whole or else to the one it overlaps most, so a quote that runs from one violation into the next is not a find of both. Detection quotes must appear in the document verbatim, which the validator now enforces. (c) Fixture edits, each recorded in the case's grading note and the provenance block: disc-09 is cut from a three-stage rank to the original-versus-over_tight pair, because the subject ranked the hand-edited stage last on all five CI runs on Lint grounds the skill supports and ranked original above over_tight on every one of them (the six local runs' outputs do not survive; decision 10's 0 of 10 records only that the key ordering never appeared); disc-11 and disc-12 become one-difference pairs by hand edit; disc-04's variants open with their preceding paragraph; trans-06's task states the facts its reference uses, in its own words; trans-02's voice rubric names no required word; disc-08 and disc-10 accept the breath rule, which the subject cited on every run for disc-08 and on four of five for disc-10. The four edited cases (disc-04, disc-09, disc-11, disc-12) stay `arguable` until CI runs on the edited pairs show the choices hold, and disc-08 keeps its flag on its note. The floors from decision 11 stand until five CI runs of the new shape exist; det-01 and det-02 now pass on every run, so the 0.4 floor is two rows looser than when it was calibrated, and it is recalibrated from those runs rather than guessed now. The review predicted 16 to 19 of 22 on the first run with no model or prompt change, and a run at 10 to 12 would have meant the variance is in the model, not the keys. The first run (35495133404, 2026-09-20) scored 15 of 22: choice 12 of 12 with all four edited pairs picking the key, detection 3 of 3, rule 8 of 12, transformation rubric 4 of 7. Every remaining failure is a judged row. code-comment-register scored 14 of 16 with its rule judge at 8 of 10, inside its prior range. ## Architecture @@ -74,10 +75,10 @@ The baseline condition (skill hidden) is a second provider entry passing `--disa The existing heuristics port into per-case asserts the generator emits: -- **discrimination** — two asserts. Deterministic: regex-extract the `ANSWER:` line and check the chosen variant against `correct`. Judged: an `llm-rubric` asking whether the `RULE:` line names the same principle as `expected_rule`. (Phase 1 used keyword overlap for the rule; see decision 5.) +- **discrimination** — two asserts. Deterministic: regex-extract the `ANSWER:` line and check the chosen variant against `correct`. Judged: an `llm-rubric` asking whether the `RULE:` line names the same principle as `expected_rule` or one of the case's `accepted_rules`; the case's `grading_note` never reaches the judge (decision 12). (Phase 1 used keyword overlap for the rule; see decision 5.) - **discrimination-structural** — identical to discrimination; the variants are whole openings or closings rather than sentences. - **discrimination-rank** — the prompt lists `stages` and asks for an ordering. Deterministic: the full ordering must equal `correct_ranking`. Judged: the rule is compared against `expected_rule_for_worst`. -- **detection** — deterministic: the existing quote-matching-with-traps logic as a javascript assert (recall against violations, precision against traps). +- **detection** — deterministic: quote matching with traps as a javascript assert (recall against violations, precision against traps). A key quote counts as found when a subject quote contains it, sits inside it, or overlaps one of its ends by at least 20 characters. A case may set `min_recall` below 1 to record recall as its score instead of failing on it; a flagged trap fails at any floor (decision 12). - **transformation** — `llm-rubric` built from whatever rubric fields the case defines (`code-comment-register`: `violation_fixed`, `placement`, `no_new_violation`; `prose-register`: `voice_match` in place of `placement`). **Judge provider.** Every `llm-rubric` names `providers/judge.mjs`, a custom provider spawning `claude -p` with skills and tools disabled and cwd in a scratch directory, so the repo's skills and its `CLAUDE.md` do not reach the judge; the developer's user config still does locally (see After Phase 2). Its model is pinned in the provider. An offline probe (2026-09-19) confirmed promptfoo 0.122.2 accepts a `file://` custom provider as an `llm-rubric` grader and echoes each assertion, `metric` included, into the result's `componentResults`. The judge cannot use `claude --bare`: bare mode reads only `ANTHROPIC_API_KEY`, never an OAuth token. @@ -90,7 +91,7 @@ Every discrimination choice must be correct, except on a case that marks its key Because a discrimination case now carries both kinds of assert, `check-gate` classifies a failure by which assert failed (promptfoo's per-assert component results), not by case type. The generator tags the one hard assert, the discrimination choice, with `metric: choice`. A case whose `choice` assert failed gates, as does a case whose provider errored before any assert ran. Every other failure counts against the floor. -**Detection gates soft (decision 8).** Detection grades by exact quote match against a fixed violation list. That holds for `code-comment-register`'s short documents, but `prose-register`'s essays carry more true violations than their lists capture, so a correct reviewer quoting a different valid instance fails (its `det-02` grading note records a 0/5 run that found 3 real violations). The old prose runner never failed these cases. Detection failures therefore count against the floor in every skill rather than gating. +**Detection gates soft (decision 8).** Detection grades by quote match against a fixed violation list. That holds for `code-comment-register`'s short documents, but `prose-register`'s essays carry more true violations than their lists capture, so a correct reviewer quoting a different valid instance fails (its `det-02` grading note records a 0/5 run that found 3 real violations). The old prose runner never failed these cases. Detection failures therefore count against the floor in every skill rather than gating, and a case whose list is not exhaustive sets `min_recall` so that recall is recorded rather than failed (decision 12). ### Make and CI @@ -106,7 +107,7 @@ Each skill's `run-evals.mjs` is deleted once that skill passes through promptfoo - **Phase 1** — `code-comment-register` scoring end to end through promptfoo: layout, generator, subject provider, deterministic asserts, `llm-rubric` for the 4 transformation cases, parity check against the old runner, delete its `run-evals.mjs`. No CI. - **Phase 2** — in order: verify `claude -p` auth with `CLAUDE_CODE_OAUTH_TOKEN` in CI; the `claude -p` judge provider; split discrimination grading and the assert-aware gate; generator support for `discrimination-structural` and `discrimination-rank`; a keyed `prose-register` run and deletion of its `run-evals.mjs`; CI wiring (`eval-validate` and unit tests into `preflight`, the advisory path-filtered full-run job). -- **After Phase 2** — tune the 90% floor and 0.75 rubric threshold from repeated CI runs, then make the `evals` job a required check. Each is a one-line change once the data exists. Separately, author an `evals.json` for `code-review-register`: its review comments fit the same discrimination, transformation, and detection shapes, and the generator and CI pick it up with no suite changes. Two prose-register keys still want the author's eye. disc-09: every run ranked `restored` last, against the key's worst stage, and the case is `arguable` for now rather than re-keyed because a re-key needs text the source essay does not supply. disc-10: every run names the breath rule where the key names concrete nouns, and the judge rejects it, so it costs one row of the floor on each run. A single wrong sample of any unflagged choice can still fail the job (disc-06 missed twice in four CI runs); repeating each choice and gating on the majority would cost 2 to 3 times the sessions. +- **After Phase 2** — tune the 90% floor and 0.75 rubric threshold from repeated CI runs, then make the `evals` job a required check. Each is a one-line change once the data exists. Separately, author an `evals.json` for `code-review-register`: its review comments fit the same discrimination, transformation, and detection shapes, and the generator and CI pick it up with no suite changes. disc-09 and disc-10, the two keys that wanted the author's eye, were settled by decision 12: disc-09 is a two-way pair now, and disc-10 accepts the breath rule. A single wrong sample of any unflagged choice can still fail the job (disc-06 missed twice in four CI runs); repeating each choice and gating on the majority would cost 2 to 3 times the sessions. The next measurement is the one the suite does not make today: the review found that no row samples a draft written under the skill, and that `promptfooconfig.compare.yaml` (skill on versus off) has never run in CI. A generation-side pilot is a separate design. **CI versus local.** The same suite scores about 17 points lower on the CI runner than under a local bare `$HOME`: prose-register 45.5 to 54.5% in CI against 59 to 73% locally, and CI's discrimination sessions answered with the two required lines alone (about 210 visible characters, against about 640 locally) until the subject prompt asked for a comparison first (about 700 after). The recorded model is `claude-sonnet-5` in both. Ruled out locally, on the same disc-12 prompt: the developer's inherited shell environment, `CI` and `GITHUB_ACTIONS` variables, and every `--effort` level (correct in 31 of 34 probes, against 0 of 4 in CI). Not yet tested, and needing the CI token or an equivalent from `claude setup-token`: authentication through `CLAUDE_CODE_OAUTH_TOKEN` rather than a credentials file (a token without profile scope may change server-side defaults), the npm-installed CLI against a native one, and a first-run `$HOME` without a cached `.claude.json`. A bare `$HOME` with a symlinked credentials file therefore does not reproduce CI; use CI runs for calibration. Isolate eval sessions from the developer's user config locally: in CI the runner's bare $HOME keeps user rules, plugins, and hooks away from subject and judge, but locally every session loads them (user rules, plugin SessionStart preambles, Stop hooks such as a session-cost logger that gains a row per session). Candidate: claude's --setting-sources or --settings flags; verify what each suppresses before relying on it. A temporary `$HOME` with symlinks to the skill under test and to `~/.claude/.credentials.json` removes the user-config confound but does not reproduce CI's behavior (see above). From a4d989632cfc9b29cc12de68fd07c3e605d3c037 Mon Sep 17 00:00:00 2001 From: Agent Norton Date: Sun, 20 Sep 2026 18:04:58 +0000 Subject: [PATCH 6/7] Count a rule-only failure as informational in the gate On the first run after the fixture pass every deterministic assert passed (choice 12 of 12, detection 3 of 3) and four of the seven remaining failures were the rule judge rejecting a rule the subject named that is true of the passage: the "Not X, but Y" move on disc-08, the breath rule on disc-09, "No flex" on disc-10, the reversal rule on disc-11. A passage usually breaks more than one rule, the skill's master rule absorbs cases keyed to others, and widening accepted_rules run by run ends with any true rule passing. The judge still runs and its verdict is recorded per case, and the gate prints the rule count on its summary line, but a row whose only failed asserts are the rule judgment now counts as passed for the floor. The choice still gates hard, and an arguable case whose choice is wrong still counts against the floor whatever its rule verdict. Decision 13 in the spec records this and names rule identifiers as the follow-on that makes the check deterministic. Co-Authored-By: Claude Fable 5.1 --- .../2026-08-29-skill-eval-suite-design.md | 9 +++++---- evals/bin/check-gate.mjs | 17 ++++++++++++++-- evals/test/check-gate.test.mjs | 20 ++++++++++++++----- 3 files changed, 35 insertions(+), 11 deletions(-) diff --git a/docs/superpowers/specs/2026-08-29-skill-eval-suite-design.md b/docs/superpowers/specs/2026-08-29-skill-eval-suite-design.md index d0481ec5..06e78323 100644 --- a/docs/superpowers/specs/2026-08-29-skill-eval-suite-design.md +++ b/docs/superpowers/specs/2026-08-29-skill-eval-suite-design.md @@ -33,11 +33,12 @@ Phase 1's keyed runs surfaced one flaky grader, and CI planning surfaced an auth 5. **Discrimination grades the choice hard and the rule soft.** The model picks the correct letter every run, but the ported keyword-overlap check on its stated rule (threshold 0.2) failed a different case each run: a correct paraphrase shares few words with the answer key. The choice (letter, or full ordering for rank cases) stays a deterministic assert that always gates. The stated rule moves to an `llm-rubric` judgment that counts only toward the pass-rate floor. This departs from the byte-for-byte port on purpose. 6. **Every model call runs through `claude -p` on a subscription token.** CI authenticates with a single `CLAUDE_CODE_OAUTH_TOKEN` secret; no `ANTHROPIC_API_KEY` anywhere. The judge becomes a second `claude -p` provider rather than a direct API provider, trading some latency and Claude Code's system prompt in the judge's context for one credential and no API billing. 7. **The CI full-run job starts advisory.** Its thresholds are still Phase 1 guesses; it reports on skill PRs without blocking merge until real runs tune them. -8. **Only the discrimination choice gates hard.** Detection and every judged assert count toward the pass-rate floor; see Pass criteria. +8. **Only the discrimination choice gates hard.** Detection and the transformation rubric count toward the pass-rate floor, and the rule judgment did until decision 13; see Pass criteria. 9. **Transformation rubrics list whatever criteria the case defines.** `prose-register` uses `voice_match` where `code-comment-register` uses `placement`; the rubric builder stops hardcoding three keys. 10. **A case can mark its answer key arguable.** `arguable: true` in evals.json, with a grading note saying why, makes a wrong choice count against the pass-rate floor instead of hard-failing the job; a subject that returns no answer still gates. It applies only to discrimination cases. The flag is for a key that cannot gate hard: the author's note concedes another answer is defensible, or the key does not reproduce across environments. Five prose-register cases carry it. Four are flagged on their notes (disc-04, disc-08, disc-09, disc-11). Across ten keyed runs (six local repeats under a bare `$HOME`, four CI runs) the subject picked the key in 6 of 10 for disc-11, 7 of 10 for disc-08, 9 of 10 for disc-04 and 0 of 10 for disc-09; disc-04 is flagged on its note and on the parent PR's report that it flipped once in four earlier runs. disc-12 is flagged on evidence: the CI subject picked the wrong answer on all four bare-runner runs (reading the recap triplet in `before` as a hard line with a breath), while all six local suite runs got it right. See "CI versus local" under After Phase 2. Decision 12 re-cut disc-09 to a two-way pair and edited disc-04, disc-11 and disc-12; all five flags stay until CI runs on the edited cases show the choices hold. 11. **Each skill can set its own pass-rate floor.** `min_pass_rate` at the top of evals.json overrides the 90% default; argv overrides both, and a floor outside (0, 1] is rejected. The values come from CI runs only, at one row below the worst run so far. prose-register: 12, 10, 12 and 10 of 22 (54.5, 45.5, 54.5, 45.5%), so 0.4 (9 of 22). code-comment-register: 15, 15, 13 and 13 of 16 (93.8, 93.8, 81.3, 81.3%), so 0.75 (12 of 16); its 90% default failed two of four runs on judge-rule and detection misses alone. prose-register's two detection cases failed on every run by design, which alone left 20 of 22 rows (decision 12 has them record recall and pass instead). These floors catch a collapse, not a drift, and rest on four runs each. Local runs are not a valid source: they scored 59 to 73% on prose-register. 12. **A key the subject disagrees with on every run is re-cut, not flagged.** A review of five CI runs at the tip of the fixes branch ([2026-09-20-register-eval-review.md](2026-09-20-register-eval-review.md)) attributed 11 of the tip run's 12 failures to the eval side: seven cases failed on every run with the same reasoning, which is a key the model disagrees with, not sampling noise. Three mechanisms follow, all built with zero model sessions. (a) `accepted_rules`, a non-empty string array on discrimination cases, lists the alternatives the rule judge accepts; the rubric shows the judge the reference rule, the skill's rule text and that list, and `grading_note` is human-facing again (disc-07's note asked for a two-rule answer, and the judge read it as a requirement and failed the answers that met it). code-comment-register's discrimination notes explain the key rather than list alternatives, so none migrates, and its rule judge loses that context; its run on the fixture-pass PR measures the effect. (b) `min_recall` on detection cases, a number in [0, 1] defaulting to 1, floors recall; below 1 the case records recall as its score, and a flagged trap still fails it. det-01 and det-02 set 0, since their lists are not exhaustive. The detection grader also matches each key quote against the subject's quotes by overlap rather than by the key's first 60 characters, which scored det-03 a miss on two runs where the subject had dropped the key's first word; one subject quote is one finding, credited to every key it holds whole or else to the one it overlaps most, so a quote that runs from one violation into the next is not a find of both. Detection quotes must appear in the document verbatim, which the validator now enforces. (c) Fixture edits, each recorded in the case's grading note and the provenance block: disc-09 is cut from a three-stage rank to the original-versus-over_tight pair, because the subject ranked the hand-edited stage last on all five CI runs on Lint grounds the skill supports and ranked original above over_tight on every one of them (the six local runs' outputs do not survive; decision 10's 0 of 10 records only that the key ordering never appeared); disc-11 and disc-12 become one-difference pairs by hand edit; disc-04's variants open with their preceding paragraph; trans-06's task states the facts its reference uses, in its own words; trans-02's voice rubric names no required word; disc-08 and disc-10 accept the breath rule, which the subject cited on every run for disc-08 and on four of five for disc-10. The four edited cases (disc-04, disc-09, disc-11, disc-12) stay `arguable` until CI runs on the edited pairs show the choices hold, and disc-08 keeps its flag on its note. The floors from decision 11 stand until five CI runs of the new shape exist; det-01 and det-02 now pass on every run, so the 0.4 floor is two rows looser than when it was calibrated, and it is recalibrated from those runs rather than guessed now. The review predicted 16 to 19 of 22 on the first run with no model or prompt change, and a run at 10 to 12 would have meant the variance is in the model, not the keys. The first run (35495133404, 2026-09-20) scored 15 of 22: choice 12 of 12 with all four edited pairs picking the key, detection 3 of 3, rule 8 of 12, transformation rubric 4 of 7. Every remaining failure is a judged row. code-comment-register scored 14 of 16 with its rule judge at 8 of 10, inside its prior range. +13. **The rule judgment is informational.** On that first run every deterministic assert passed, and four of the seven remaining failures were the rule judge rejecting a rule the subject named that is true of the passage: the "Not X, but Y" move on disc-08, the breath rule on disc-09, "No flex" on disc-10, the reversal rule on disc-11. A passage usually breaks more than one rule, the skill's master rule absorbs cases keyed to others, and widening `accepted_rules` run by run ends with any true rule passing, which the case description itself warns against. The judge still runs and its verdict is recorded per case, and `check-gate` prints the rule count on the summary line, but a row whose only failed asserts are the rule judgment counts as passed for the floor. The choice still gates hard. A judge outage is indistinguishable from a judge disagreement in the results, so it shows as a collapsed rule count on the summary line rather than as a failed gate; read that line. The follow-on is rule identifiers: a stable slug per SKILL.md rule (light tags in the file, or a sidecar the validator checks against the file's text), the RULE line answered with a slug and a sentence, and a deterministic membership check against the key's slug and its accepted set. That removes the judge session, and it makes each case's link to its rule checkable when the rule's wording changes, where today's `rule_quote` goes stale silently; the same slugs name the checks in a generation-side lint. The floor is recalibrated from the first five runs after this decision, since rule-only failures no longer count against it. ## Architecture @@ -75,7 +76,7 @@ The baseline condition (skill hidden) is a second provider entry passing `--disa The existing heuristics port into per-case asserts the generator emits: -- **discrimination** — two asserts. Deterministic: regex-extract the `ANSWER:` line and check the chosen variant against `correct`. Judged: an `llm-rubric` asking whether the `RULE:` line names the same principle as `expected_rule` or one of the case's `accepted_rules`; the case's `grading_note` never reaches the judge (decision 12). (Phase 1 used keyword overlap for the rule; see decision 5.) +- **discrimination** — two asserts. Deterministic: regex-extract the `ANSWER:` line and check the chosen variant against `correct`. Judged: an `llm-rubric` asking whether the `RULE:` line names the same principle as `expected_rule` or one of the case's `accepted_rules`; the case's `grading_note` never reaches the judge (decision 12). Its verdict is recorded and never counts toward the floor (decision 13). (Phase 1 used keyword overlap for the rule; see decision 5.) - **discrimination-structural** — identical to discrimination; the variants are whole openings or closings rather than sentences. - **discrimination-rank** — the prompt lists `stages` and asks for an ordering. Deterministic: the full ordering must equal `correct_ranking`. Judged: the rule is compared against `expected_rule_for_worst`. - **detection** — deterministic: quote matching with traps as a javascript assert (recall against violations, precision against traps). A key quote counts as found when a subject quote contains it, sits inside it, or overlaps one of its ends by at least 20 characters. A case may set `min_recall` below 1 to record recall as its score instead of failing on it; a flagged trap fails at any floor (decision 12). @@ -87,9 +88,9 @@ The answer-key separation holds: the subject prompt never contains `rule_quote`, ### Pass criteria -Every discrimination choice must be correct, except on a case that marks its key `arguable` (decision 10). Judged asserts carry a per-test threshold, with a suite gate at 90% overall unless the skill sets its own `min_pass_rate` (decision 11), so one flaky rubric call does not block a PR. Both numbers are Phase-1 guesses to be tuned after the first real runs. +Every discrimination choice must be correct, except on a case that marks its key `arguable` (decision 10). Judged asserts carry a per-test threshold, with a suite gate at 90% overall unless the skill sets its own `min_pass_rate` (decision 11), so one flaky rubric call does not block a PR. A row whose only failed assert is the rule judgment counts as passed for the floor (decision 13). Both numbers are Phase-1 guesses to be tuned after the first real runs. -Because a discrimination case now carries both kinds of assert, `check-gate` classifies a failure by which assert failed (promptfoo's per-assert component results), not by case type. The generator tags the one hard assert, the discrimination choice, with `metric: choice`. A case whose `choice` assert failed gates, as does a case whose provider errored before any assert ran. Every other failure counts against the floor. +Because a discrimination case now carries both kinds of assert, `check-gate` classifies a failure by which assert failed (promptfoo's per-assert component results), not by case type. The generator tags the one hard assert, the discrimination choice, with `metric: choice`. A case whose `choice` assert failed gates, as does a case whose provider errored before any assert ran. Every other failure counts against the floor, except a failure of the rule judgment alone (decision 13). **Detection gates soft (decision 8).** Detection grades by quote match against a fixed violation list. That holds for `code-comment-register`'s short documents, but `prose-register`'s essays carry more true violations than their lists capture, so a correct reviewer quoting a different valid instance fails (its `det-02` grading note records a 0/5 run that found 3 real violations). The old prose runner never failed these cases. Detection failures therefore count against the floor in every skill rather than gating, and a case whose list is not exhaustive sets `min_recall` so that recall is recorded rather than failed (decision 12). diff --git a/evals/bin/check-gate.mjs b/evals/bin/check-gate.mjs index c59a5303..6d35fb03 100644 --- a/evals/bin/check-gate.mjs +++ b/evals/bin/check-gate.mjs @@ -50,6 +50,15 @@ const isHardFailure = (row) => (meta(row).arguable !== true && components(row).some((c) => !c.pass && c.assertion?.metric === 'choice'))); +// The rule judgment is recorded but never counts: a passage breaks more than +// one rule, and the judge rejects a true rule the key did not name. A row +// whose only failed asserts are informational passes for the floor. +const INFORMATIONAL_METRICS = new Set(['rule']); +const isInformational = (component) => INFORMATIONAL_METRICS.has(component.assertion?.metric); +const countsAsPassed = (row) => + row.success || + (components(row).length > 0 && components(row).every((c) => c.pass || isInformational(c))); + // Precedence: argv, then the skill's own `min_pass_rate` in evals.json, then // the default. A skill whose keys are contested or whose detection cases // record recall instead of failing (prose-register) can carry a lower floor @@ -80,10 +89,14 @@ if (!(minRate > 0 && minRate <= 1)) { } const hardFailures = rows.filter(isHardFailure); -const passed = rows.filter((row) => row.success).length; +const passed = rows.filter(countsAsPassed).length; const passRate = passed / rows.length; -console.log(`${passed}/${rows.length} passed (rate ${(passRate * 100).toFixed(1)}%, floor ${(minRate * 100).toFixed(0)}%)`); +const informational = rows.flatMap((row) => components(row).filter(isInformational)); +const informationalNote = informational.length + ? `; rule ${informational.filter((c) => c.pass).length}/${informational.length} informational` + : ''; +console.log(`${passed}/${rows.length} passed (rate ${(passRate * 100).toFixed(1)}%, floor ${(minRate * 100).toFixed(0)}%)${informationalNote}`); let failed = false; if (hardFailures.length > 0) { diff --git a/evals/test/check-gate.test.mjs b/evals/test/check-gate.test.mjs index d6cbc346..b8132e1b 100644 --- a/evals/test/check-gate.test.mjs +++ b/evals/test/check-gate.test.mjs @@ -28,10 +28,20 @@ function runGate(rows, ...args) { return spawnSync(process.execPath, [GATE, file, ...args], { encoding: 'utf8' }); } -test('a failed rule judgment is soft: passes at the floor', () => { - const result = runGate([...passing(9), row('disc-07', [choice(true), rule(false)])]); +test('a failed rule judgment is informational: counts as passed for the floor and is reported', () => { + const result = runGate([...passing(6), ...[7, 8, 9, 10].map((i) => row(`disc-${i}`, [choice(true), rule(false)]))]); assert.equal(result.status, 0, result.stdout + result.stderr); - assert.match(result.stdout, /9\/10/); + assert.match(result.stdout, /10\/10 passed/); + assert.match(result.stdout, /rule 6\/10 informational/); +}); + +test('a wrong arguable choice counts against the floor, and its rule verdict still reports', () => { + const arguable = [1, 2].map((i) => row(`disc-0${i}`, [choice(false), rule(false)], { arguable: true })); + const result = runGate([...passing(8), ...arguable]); + assert.equal(result.status, 1); + assert.match(result.stdout, /8\/10 passed/); + assert.match(result.stdout, /rule 8\/10 informational/); + assert.match(result.stderr, /pass rate/i); }); test('a failed detection is soft: passes at the floor', () => { @@ -72,14 +82,14 @@ test('a subject error gates because no assert ran', () => { }); test('soft failures below the floor fail the gate', () => { - const soft = [1, 2, 3].map((i) => row(`trans-0${i}`, [rule(false)])); + const soft = [1, 2, 3].map((i) => row(`det-0${i}`, [detect(false)])); const result = runGate([...passing(7), ...soft]); assert.equal(result.status, 1); assert.match(result.stderr, /pass rate/i); }); test('the floor is overridable via argv', () => { - const soft = [1, 2, 3].map((i) => row(`trans-0${i}`, [rule(false)])); + const soft = [1, 2, 3].map((i) => row(`det-0${i}`, [detect(false)])); const result = runGate([...passing(7), ...soft], '0.70'); assert.equal(result.status, 0, result.stdout + result.stderr); }); From 28733dd1a8be29e00a36487dc6e5022fbed4e820 Mon Sep 17 00:00:00 2001 From: Agent Norton Date: Sun, 20 Sep 2026 21:44:00 +0000 Subject: [PATCH 7/7] Let a detection violation name an anchor its quote must contain MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit On the second CI run after the fixture pass det-03's subject flagged the em-dash for the em-dash rule, as on every prior run, and quoted the text on its left: "Same company, same invoice, same artifact —". That shares 15 characters with the key quote, under the overlap matcher's 20-character floor, so the row scored a miss. A violation one character wide can be quoted from either side, and no span heuristic knows which. A violation (or trap) may now name an anchor inside its quote; any subject quote containing the anchor counts as a find of that violation whatever else it covers. The validator requires the anchor to be a non-empty substring of its quote. det-03 anchors on "artifact —". Co-Authored-By: Claude Fable 5.1 --- .../2026-08-29-skill-eval-suite-design.md | 4 ++-- evals/asserts/detection.mjs | 13 ++++++++++- evals/lib/load-evals.mjs | 6 ++++- evals/test/asserts.test.mjs | 23 +++++++++++++++++++ evals/test/load-evals.test.mjs | 14 +++++++++++ home/.agents/skills/prose-register/evals.json | 4 ++-- 6 files changed, 58 insertions(+), 6 deletions(-) diff --git a/docs/superpowers/specs/2026-08-29-skill-eval-suite-design.md b/docs/superpowers/specs/2026-08-29-skill-eval-suite-design.md index 06e78323..323c4f4b 100644 --- a/docs/superpowers/specs/2026-08-29-skill-eval-suite-design.md +++ b/docs/superpowers/specs/2026-08-29-skill-eval-suite-design.md @@ -37,7 +37,7 @@ Phase 1's keyed runs surfaced one flaky grader, and CI planning surfaced an auth 9. **Transformation rubrics list whatever criteria the case defines.** `prose-register` uses `voice_match` where `code-comment-register` uses `placement`; the rubric builder stops hardcoding three keys. 10. **A case can mark its answer key arguable.** `arguable: true` in evals.json, with a grading note saying why, makes a wrong choice count against the pass-rate floor instead of hard-failing the job; a subject that returns no answer still gates. It applies only to discrimination cases. The flag is for a key that cannot gate hard: the author's note concedes another answer is defensible, or the key does not reproduce across environments. Five prose-register cases carry it. Four are flagged on their notes (disc-04, disc-08, disc-09, disc-11). Across ten keyed runs (six local repeats under a bare `$HOME`, four CI runs) the subject picked the key in 6 of 10 for disc-11, 7 of 10 for disc-08, 9 of 10 for disc-04 and 0 of 10 for disc-09; disc-04 is flagged on its note and on the parent PR's report that it flipped once in four earlier runs. disc-12 is flagged on evidence: the CI subject picked the wrong answer on all four bare-runner runs (reading the recap triplet in `before` as a hard line with a breath), while all six local suite runs got it right. See "CI versus local" under After Phase 2. Decision 12 re-cut disc-09 to a two-way pair and edited disc-04, disc-11 and disc-12; all five flags stay until CI runs on the edited cases show the choices hold. 11. **Each skill can set its own pass-rate floor.** `min_pass_rate` at the top of evals.json overrides the 90% default; argv overrides both, and a floor outside (0, 1] is rejected. The values come from CI runs only, at one row below the worst run so far. prose-register: 12, 10, 12 and 10 of 22 (54.5, 45.5, 54.5, 45.5%), so 0.4 (9 of 22). code-comment-register: 15, 15, 13 and 13 of 16 (93.8, 93.8, 81.3, 81.3%), so 0.75 (12 of 16); its 90% default failed two of four runs on judge-rule and detection misses alone. prose-register's two detection cases failed on every run by design, which alone left 20 of 22 rows (decision 12 has them record recall and pass instead). These floors catch a collapse, not a drift, and rest on four runs each. Local runs are not a valid source: they scored 59 to 73% on prose-register. -12. **A key the subject disagrees with on every run is re-cut, not flagged.** A review of five CI runs at the tip of the fixes branch ([2026-09-20-register-eval-review.md](2026-09-20-register-eval-review.md)) attributed 11 of the tip run's 12 failures to the eval side: seven cases failed on every run with the same reasoning, which is a key the model disagrees with, not sampling noise. Three mechanisms follow, all built with zero model sessions. (a) `accepted_rules`, a non-empty string array on discrimination cases, lists the alternatives the rule judge accepts; the rubric shows the judge the reference rule, the skill's rule text and that list, and `grading_note` is human-facing again (disc-07's note asked for a two-rule answer, and the judge read it as a requirement and failed the answers that met it). code-comment-register's discrimination notes explain the key rather than list alternatives, so none migrates, and its rule judge loses that context; its run on the fixture-pass PR measures the effect. (b) `min_recall` on detection cases, a number in [0, 1] defaulting to 1, floors recall; below 1 the case records recall as its score, and a flagged trap still fails it. det-01 and det-02 set 0, since their lists are not exhaustive. The detection grader also matches each key quote against the subject's quotes by overlap rather than by the key's first 60 characters, which scored det-03 a miss on two runs where the subject had dropped the key's first word; one subject quote is one finding, credited to every key it holds whole or else to the one it overlaps most, so a quote that runs from one violation into the next is not a find of both. Detection quotes must appear in the document verbatim, which the validator now enforces. (c) Fixture edits, each recorded in the case's grading note and the provenance block: disc-09 is cut from a three-stage rank to the original-versus-over_tight pair, because the subject ranked the hand-edited stage last on all five CI runs on Lint grounds the skill supports and ranked original above over_tight on every one of them (the six local runs' outputs do not survive; decision 10's 0 of 10 records only that the key ordering never appeared); disc-11 and disc-12 become one-difference pairs by hand edit; disc-04's variants open with their preceding paragraph; trans-06's task states the facts its reference uses, in its own words; trans-02's voice rubric names no required word; disc-08 and disc-10 accept the breath rule, which the subject cited on every run for disc-08 and on four of five for disc-10. The four edited cases (disc-04, disc-09, disc-11, disc-12) stay `arguable` until CI runs on the edited pairs show the choices hold, and disc-08 keeps its flag on its note. The floors from decision 11 stand until five CI runs of the new shape exist; det-01 and det-02 now pass on every run, so the 0.4 floor is two rows looser than when it was calibrated, and it is recalibrated from those runs rather than guessed now. The review predicted 16 to 19 of 22 on the first run with no model or prompt change, and a run at 10 to 12 would have meant the variance is in the model, not the keys. The first run (35495133404, 2026-09-20) scored 15 of 22: choice 12 of 12 with all four edited pairs picking the key, detection 3 of 3, rule 8 of 12, transformation rubric 4 of 7. Every remaining failure is a judged row. code-comment-register scored 14 of 16 with its rule judge at 8 of 10, inside its prior range. +12. **A key the subject disagrees with on every run is re-cut, not flagged.** A review of five CI runs at the tip of the fixes branch ([2026-09-20-register-eval-review.md](2026-09-20-register-eval-review.md)) attributed 11 of the tip run's 12 failures to the eval side: seven cases failed on every run with the same reasoning, which is a key the model disagrees with, not sampling noise. Three mechanisms follow, all built with zero model sessions. (a) `accepted_rules`, a non-empty string array on discrimination cases, lists the alternatives the rule judge accepts; the rubric shows the judge the reference rule, the skill's rule text and that list, and `grading_note` is human-facing again (disc-07's note asked for a two-rule answer, and the judge read it as a requirement and failed the answers that met it). code-comment-register's discrimination notes explain the key rather than list alternatives, so none migrates, and its rule judge loses that context; its run on the fixture-pass PR measures the effect. (b) `min_recall` on detection cases, a number in [0, 1] defaulting to 1, floors recall; below 1 the case records recall as its score, and a flagged trap still fails it. det-01 and det-02 set 0, since their lists are not exhaustive. The detection grader also matches each key quote against the subject's quotes by overlap rather than by the key's first 60 characters, which scored det-03 a miss on two runs where the subject had dropped the key's first word; one subject quote is one finding, credited to every key it holds whole or else to the one it overlaps most, so a quote that runs from one violation into the next is not a find of both. Detection quotes must appear in the document verbatim, which the validator now enforces. On the second run of the new shape (35528316122: 18 of 22 for the floor, 16 strict) det-03's subject quoted the text on the left of the em-dash, 15 characters of overlap with the key, and scored a miss; a violation may now name an `anchor` inside its quote (det-03: "artifact —"), validated as a substring of it, and any subject quote containing the anchor counts as a find. (c) Fixture edits, each recorded in the case's grading note and the provenance block: disc-09 is cut from a three-stage rank to the original-versus-over_tight pair, because the subject ranked the hand-edited stage last on all five CI runs on Lint grounds the skill supports and ranked original above over_tight on every one of them (the six local runs' outputs do not survive; decision 10's 0 of 10 records only that the key ordering never appeared); disc-11 and disc-12 become one-difference pairs by hand edit; disc-04's variants open with their preceding paragraph; trans-06's task states the facts its reference uses, in its own words; trans-02's voice rubric names no required word; disc-08 and disc-10 accept the breath rule, which the subject cited on every run for disc-08 and on four of five for disc-10. The four edited cases (disc-04, disc-09, disc-11, disc-12) stay `arguable` until CI runs on the edited pairs show the choices hold, and disc-08 keeps its flag on its note. The floors from decision 11 stand until five CI runs of the new shape exist; det-01 and det-02 now pass on every run, so the 0.4 floor is two rows looser than when it was calibrated, and it is recalibrated from those runs rather than guessed now. The review predicted 16 to 19 of 22 on the first run with no model or prompt change, and a run at 10 to 12 would have meant the variance is in the model, not the keys. The first run (35495133404, 2026-09-20) scored 15 of 22: choice 12 of 12 with all four edited pairs picking the key, detection 3 of 3, rule 8 of 12, transformation rubric 4 of 7. Every remaining failure is a judged row. code-comment-register scored 14 of 16 with its rule judge at 8 of 10, inside its prior range. 13. **The rule judgment is informational.** On that first run every deterministic assert passed, and four of the seven remaining failures were the rule judge rejecting a rule the subject named that is true of the passage: the "Not X, but Y" move on disc-08, the breath rule on disc-09, "No flex" on disc-10, the reversal rule on disc-11. A passage usually breaks more than one rule, the skill's master rule absorbs cases keyed to others, and widening `accepted_rules` run by run ends with any true rule passing, which the case description itself warns against. The judge still runs and its verdict is recorded per case, and `check-gate` prints the rule count on the summary line, but a row whose only failed asserts are the rule judgment counts as passed for the floor. The choice still gates hard. A judge outage is indistinguishable from a judge disagreement in the results, so it shows as a collapsed rule count on the summary line rather than as a failed gate; read that line. The follow-on is rule identifiers: a stable slug per SKILL.md rule (light tags in the file, or a sidecar the validator checks against the file's text), the RULE line answered with a slug and a sentence, and a deterministic membership check against the key's slug and its accepted set. That removes the judge session, and it makes each case's link to its rule checkable when the rule's wording changes, where today's `rule_quote` goes stale silently; the same slugs name the checks in a generation-side lint. The floor is recalibrated from the first five runs after this decision, since rule-only failures no longer count against it. ## Architecture @@ -79,7 +79,7 @@ The existing heuristics port into per-case asserts the generator emits: - **discrimination** — two asserts. Deterministic: regex-extract the `ANSWER:` line and check the chosen variant against `correct`. Judged: an `llm-rubric` asking whether the `RULE:` line names the same principle as `expected_rule` or one of the case's `accepted_rules`; the case's `grading_note` never reaches the judge (decision 12). Its verdict is recorded and never counts toward the floor (decision 13). (Phase 1 used keyword overlap for the rule; see decision 5.) - **discrimination-structural** — identical to discrimination; the variants are whole openings or closings rather than sentences. - **discrimination-rank** — the prompt lists `stages` and asks for an ordering. Deterministic: the full ordering must equal `correct_ranking`. Judged: the rule is compared against `expected_rule_for_worst`. -- **detection** — deterministic: quote matching with traps as a javascript assert (recall against violations, precision against traps). A key quote counts as found when a subject quote contains it, sits inside it, or overlaps one of its ends by at least 20 characters. A case may set `min_recall` below 1 to record recall as its score instead of failing on it; a flagged trap fails at any floor (decision 12). +- **detection** — deterministic: quote matching with traps as a javascript assert (recall against violations, precision against traps). A key quote counts as found when a subject quote contains it, sits inside it, or overlaps one of its ends by at least 20 characters; a violation may also name an `anchor` inside its quote, and any subject quote containing the anchor counts, for a one-character violation the subject may quote from either side. A case may set `min_recall` below 1 to record recall as its score instead of failing on it; a flagged trap fails at any floor (decision 12). - **transformation** — `llm-rubric` built from whatever rubric fields the case defines (`code-comment-register`: `violation_fixed`, `placement`, `no_new_violation`; `prose-register`: `voice_match` in place of `placement`). **Judge provider.** Every `llm-rubric` names `providers/judge.mjs`, a custom provider spawning `claude -p` with skills and tools disabled and cwd in a scratch directory, so the repo's skills and its `CLAUDE.md` do not reach the judge; the developer's user config still does locally (see After Phase 2). Its model is pinned in the provider. An offline probe (2026-09-19) confirmed promptfoo 0.122.2 accepts a `file://` custom provider as an `llm-rubric` grader and echoes each assertion, `metric` included, into the result's `componentResults`. The judge cannot use `claude --bare`: bare mode reads only `ANTHROPIC_API_KEY`, never an OAuth token. diff --git a/evals/asserts/detection.mjs b/evals/asserts/detection.mjs index 524778f5..f5718d2f 100644 --- a/evals/asserts/detection.mjs +++ b/evals/asserts/detection.mjs @@ -27,12 +27,23 @@ export function overlap(keyQuote, subjectQuote) { export const quotesOverlap = (keyQuote, subjectQuote) => overlap(keyQuote, subjectQuote).length > 0; +// A violation one character wide (det-03's em-dash) can be quoted from either +// side, and a quote that stops at the dash shares too little with the key for +// overlap to count. A target may name an anchor inside its quote that settles +// the match on its own. +function matches(target, quote) { + if (target.anchor && quote.includes(normalize(target.anchor))) { + return { length: normalize(target.quote).length, whole: true }; + } + return overlap(target.quote, quote); +} + // One subject quote is one finding. It is credited to every target it holds // whole and, when it holds none, to the single target it overlaps most; a // quote that runs from one violation into the opening of the next is not a // find of both. function credited(targets, quote) { - const scored = targets.map((target) => ({ target, ...overlap(target.quote, quote) })); + const scored = targets.map((target) => ({ target, ...matches(target, quote) })); const whole = scored.filter((s) => s.whole).map((s) => s.target); if (whole.length) return whole; const best = scored.filter((s) => s.length > 0).sort((a, b) => b.length - a.length)[0]; diff --git a/evals/lib/load-evals.mjs b/evals/lib/load-evals.mjs index 5c63f1e7..20acf0d7 100644 --- a/evals/lib/load-evals.mjs +++ b/evals/lib/load-evals.mjs @@ -151,10 +151,14 @@ export function validateData(data) { // contain verbatim (an ellipsis, a paraphrase) can never be found or // tripped. const document = normalize(item.input_document); - for (const { quote } of [...item.violations, ...item.traps]) { + for (const { quote, anchor } of [...item.violations, ...item.traps]) { if (!document.includes(normalize(quote))) { throw new Error(`${item.id} quote is not in input_document verbatim: "${quote.slice(0, 40)}"`); } + // An anchor settles a match on its own, so it must be part of the quote it stands for. + if (anchor !== undefined && (typeof anchor !== 'string' || anchor === '' || !normalize(quote).includes(normalize(anchor)))) { + throw new Error(`${item.id} anchor must be a non-empty substring of its quote: "${String(anchor).slice(0, 40)}"`); + } } } } diff --git a/evals/test/asserts.test.mjs b/evals/test/asserts.test.mjs index ec02eb01..ecee5e11 100644 --- a/evals/test/asserts.test.mjs +++ b/evals/test/asserts.test.mjs @@ -121,6 +121,29 @@ test('detection counts a quote that starts inside the key and runs past it', () assert.equal(result.score, 1); }); +// det-03's second miss: the subject quoted the text on the left of the dash, +// 15 characters of overlap with the key, under the floor. +const dashQuotedFromTheLeft = '- QUOTE: "Same company, same invoice, same artifact —" | RULE: em-dash'; + +test('detection counts a quote from the far side of a one-character violation when the key names an anchor', () => { + const vars = { + violations: [{ quote: 'same artifact — one rewrite by strangers', anchor: 'artifact —', rule: 'No em-dashes.' }], + traps: [], + }; + const result = assertDetection(dashQuotedFromTheLeft, { vars }); + assert.equal(result.pass, true, result.reason); +}); + +test('detection without an anchor misses a quote that overlaps the key by less than the floor', () => { + const vars = { + violations: [{ quote: 'same artifact — one rewrite by strangers', rule: 'No em-dashes.' }], + traps: [], + }; + const result = assertDetection(dashQuotedFromTheLeft, { vars }); + assert.equal(result.pass, false); + assert.match(result.reason, /0\/1 violations/); +}); + test('detection counts a quote that wraps the key in context on both sides', () => { const output = '- QUOTE: "Then: Larger chunks use more memory. Smaller chunks use more CPU. And so on." | RULE: generic'; const result = assertDetection(output, { vars: { ...detVars, violations: detVars.violations.slice(0, 1) } }); diff --git a/evals/test/load-evals.test.mjs b/evals/test/load-evals.test.mjs index f3da22d4..a2bfb386 100644 --- a/evals/test/load-evals.test.mjs +++ b/evals/test/load-evals.test.mjs @@ -133,6 +133,20 @@ test('validateData rejects a min_recall outside [0, 1]', () => { } }); +test('validateData accepts an anchor that sits inside its quote', () => { + const ok = detectionCase({}); + ok.cases[0].violations = [{ quote: 'quoted line', anchor: 'ed li', rule: 'r' }]; + assert.doesNotThrow(() => validateData(ok)); +}); + +test('validateData rejects an anchor that is empty or not inside its quote', () => { + for (const bad of ['', 'the quoted', 7]) { + const data = detectionCase({}); + data.cases[0].violations = [{ quote: 'quoted line', anchor: bad, rule: 'r' }]; + assert.throws(() => validateData(data), /x1 anchor must be a non-empty substring of its quote/, `anchor ${JSON.stringify(bad)}`); + } +}); + test('validateData rejects a detection quote the document does not contain verbatim', () => { const bad = detectionCase({ input_document: 'Therefore, trust the pilot.', traps: [{ quote: 'Therefore, trust...' }] }); bad.cases[0].violations = [{ quote: 'trust the pilot', rule: 'r' }]; diff --git a/home/.agents/skills/prose-register/evals.json b/home/.agents/skills/prose-register/evals.json index e3270cda..8b7d0fc0 100644 --- a/home/.agents/skills/prose-register/evals.json +++ b/home/.agents/skills/prose-register/evals.json @@ -280,10 +280,10 @@ "source_label": "YAML frontmatter of the essay, not the prose body -- tests whether a model checks metadata fields, not just paragraphs.", "input_document": "description: >-\n Same company, same invoice, same artifact — one rewrite by strangers, one\n by a machine. Why the two binaries don't feel the same, and what\n authorship trust is made of.", "violations": [ - {"quote": "same artifact — one rewrite by strangers", "rule": "No em-dashes. Full stops, semicolons, parentheses, en-dashes. / Lint: \"An em-dash.\"", "fixed_in": "bd0c864 (\"en-dashes in descriptions\") -- one day after this text shipped, and after the commit range this eval was mined from"} + {"quote": "same artifact — one rewrite by strangers", "anchor": "artifact —", "rule": "No em-dashes. Full stops, semicolons, parentheses, en-dashes. / Lint: \"An em-dash.\"", "fixed_in": "bd0c864 (\"en-dashes in descriptions\") -- one day after this text shipped, and after the commit range this eval was mined from"} ], "traps": [], - "grading_note": "At the pinned commit (5b4cbde), this was a live, shipped miss: the em-dash purge (5e09562) predates this description field, which was added later and never caught at the time. It was subsequently fixed on the source repo's main branch in bd0c864, one day after the handoff doc was written -- do not present this case as 'the em-dash is currently live on alannorton.com', since it no longer is as of this file's writing (2026-07-18). Use it as a template for 'check frontmatter/metadata, not just prose,' not as a live bug report. If reusing this eval later, re-verify against the current HEAD of the source repo before citing it as an open issue." + "grading_note": "At the pinned commit (5b4cbde), this was a live, shipped miss: the em-dash purge (5e09562) predates this description field, which was added later and never caught at the time. It was subsequently fixed on the source repo's main branch in bd0c864, one day after the handoff doc was written -- do not present this case as 'the em-dash is currently live on alannorton.com', since it no longer is as of this file's writing (2026-07-18). Use it as a template for 'check frontmatter/metadata, not just prose,' not as a live bug report. If reusing this eval later, re-verify against the current HEAD of the source repo before citing it as an open issue. Anchor (2026-09-20): the violation is one character wide and the subject has quoted the text on either side of it (the left side on the second CI run of the fixture pass, which the overlap matcher alone called a miss), so any subject quote containing 'artifact —' counts." }, { "id": "disc-09",