Skip to content

feat: reins do — delegate a browsing goal to Jev, with benchmark (3/3) - #43

Merged
karngyan merged 99 commits into
mainfrom
feat/reins-do-loop
Sep 28, 2026
Merged

karngyan merged 99 commits into
mainfrom
feat/reins-do-loop

Conversation

@karngyan

@karngyan karngyan commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Part 3 of 3 for reins do (#41 and #42 are merged): the command itself, the benchmark, and every fix from review and live testing.

Why

Driving a site one command at a time costs an agent one turn per click. reins do "<goal>" runs the whole loop inside the daemon, with Jev (TypeSafe's action model) picking each action. It returns only when the goal is met or a decision is needed.

What

  • reins do '<goal>' [--fill name=value]... [--confirm '<label>']... [--continue] [--max-steps 30] [--timeout 60] [--tab] [--browser] [--json]

    • The loop:
      1. Read the page.
      2. Ask Jev once. The request includes speculative heads for the operation, the target, and the text to fill.
      3. Apply the stop rules.
      4. Act and record.
    • Text only ever comes from --fill.
  • Bounded autonomy. Every run ends with a status and, where one applies, a next: line to run:

    • done
    • risky_action
    • needs_text
    • left_site
    • dialog
    • interrupted
    • blocked
    • stuck
    • budget
    • error

    Exit codes: 0 done, 2 stopped, 1 error.

  • Self-check. Every done, and every stuck run, gets a yes/no question to Jev. A low score prints done (unsure: self-check 0.34). In the benchmark, all 17 wrong done results scored below 0.5; 13% of right ones did too.

  • Safety. Risky clicks stop the run. The rule is whole-word, Unicode-aware, and fires on unlabelled buttons. "Accept" on a cookie or consent banner is allowed. Leaving the site stops the run (a leading www. is ignored). Password and hidden values are never sent.

  • Coverage of real pages:

    • SUBMIT_SEARCH (Enter, only in search-like fields)
    • following a new tab
    • open shadow DOM, including names through slots
    • off-screen pagers, and controls whose names match words in the goal
    • waiting for navigation and in-flight data requests
    • typing key by key
    • naming unlabelled controls by their row
    • bounded stale retries
    • retargeting a text field when Jev picks one that fits none of the fills
    • rounded probability ties accepted
  • Browser-wide fixes found along the way:

    • Page focus emulation. Tabs that reins opens had document.hasFocus() false.
    • Links that open a new tab are opened by the extension. A native click made Chrome jump in front of the user's terminal on macOS.
    • reins do --help works.
    • Shell-quoted next: lines.
    • A corrupt key file is reported instead of being treated as missing.
  • Skill. A rewritten reins do section covering when to use it, how to call it, a table of statuses and what to do for each, and its limits.

  • Benchmark. 38 tasks: local fixture pages for tricky widgets plus live sites. A dev set, and an unseen holdout that stayed frozen until its one run. Traces are optional (REINS_JEV_TRACE). There is a report and a /docs/benchmarks page with every limitation listed.

  • CI. Runs the real-browser extension suite against CI's Chrome. REINS_TEST_CHROME is now declared in turbo.json; turbo's strict env had been dropping it.

Benchmark (38 tasks; reins do 5 runs each, Claude Opus 5.5 step by step 3 runs each)

reins do Claude step by step
Dev set (30 tasks, tuned on for five rounds) 124/150 (82.6%) 81/90 (90%)
Unseen holdout (8 tasks, first and only run) 35/40 (87.5%) 21/24 (87.5%)
Model spend ≈ $0.32 for 190 runs $21.83 for 114 runs

On tasks both arms passed, the median was 4.6× faster (range 2.4–22.6×) and 111× cheaper (range 44–535×). The cost figure counts Jev only; the calling agent still spends a turn to issue the command and verify.

Tasks reins do still fails:

  • fx-filters, arxiv, fx-datepicker, fx-orders. Jev says DONE or BLOCKED while the control it needs is visible. The self-check flags every one of these.
  • huggingface, github. Two search boxes, and the choice is a near-tie.
  • xe. The site labels its input "Receiving amount", the same label as its output.

The full report is docs/benchmarks/2026-09-reins-do.md.

The report also records two bugs in the step-by-step commands. They are out of scope here and will be fixed in a follow-up:

  • reins snapshot doesn't read shadow DOM.
  • Some custom controls report "zero size" on click.

Testing

  • pnpm build, pnpm lint and pnpm typecheck are clean.
  • Tests pass: protocol 95, CLI 509, extension 312 (including the real-Chrome suite).
  • Live on the maintainer's Chrome:
    • the benchmark above;
    • new-tab follow;
    • a Finder-frontmost check showing a new-tab link no longer raises Chrome;
    • a Delete-account page stops at risky_action.

🤖 Generated with Claude Code

@karngyan
karngyan added this pull request to stack #44 September 27, 2026 08:49
Base automatically changed from feat/reins-do-extension to main September 28, 2026 16:14
karngyan and others added 25 commits September 28, 2026 21:47
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… only

The goal allowance used substring matching, so "reorder the list" waived
an "Order" button and "paypal login" waived "Pay". It now requires the
whole normalized label bounded by non-alphanumerics. The no-label check
runs before --confirm so --confirm button cannot waive unlabeled buttons,
and non-click actions are never risky.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The boundary class [^a-z0-9] treated non-ASCII letters as boundaries, so
"prépay the bill" waived a "Pay" button and "日本語pay" likewise. The
boundary is now [^\p{L}\p{N}] under the u flag; only regex syntax
characters are escaped, since escaping anything else throws in unicode mode.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Per ruling R15 the --continue loop breaker turns only risky_action,
needs_text, blocked, budget and stuck into stuck + lock; dialog,
interrupted and left_site pass through. It also requires the last
action's outcome to have been observed, so an abort or timeout before
the next read is not mistaken for no change. A dialog result now carries
the current url/title and the last step's pageChanged. needs_text says
when the fills matched nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pand

JSON.stringify gave a double-quoted shell string where $, backticks and
$(...) still expand, so a label like 'Pay $5' never matched --confirm and a
page-controlled label could execute when pasted. Use POSIX single quotes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…estart

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…first, pins the tab once

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… its usage

The done case of nextCommand printed a bare reins snapshot while every
other next line carried --tab/--browser; the usage string now also lists
--browser and --json, and single-quotes the goal and label like the
printed next lines do.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… stricter checkers

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rm, partial results on abort

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… wait

The loop only checked elapsed time between iterations, so a slow jev_act or a
retried Jev call near the deadline overran the CLI's fetch abort and printed a
raw undici error. handleDo now aborts the loop's signal with a TimeoutError
when --timeout elapses; the loop reports that as budget (timed out after Ns),
not interrupted. The CLI waits timeout + 40 s (one full bridge call of slack)
and turns a fetch timeout into one readable line.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y handleDo

- A dialog observation carries no page: stop with `dialog` without touching
  lastFingerprint or the last step's pageChanged, and keep the previous
  url/title when the observation has none.
- A run that began on about:blank / file:// / an error page pins its start
  host on the first http(s) page, so left_site can fire later.
- The select audit label keeps only the field (Country), not the option.
- A missing goal fails before the first jev_observe.
- The busy-tab refusal says a just-stopped run may be finishing its action.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d run memory

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e popup, protocol action type

- jev_observe fills url/title from chrome.tabs.get while a JS dialog blocks
  CDP, so the daemon's result and run state keep a real page.
- The popup's Replace/Cancel paths re-render from live connectivity, like Save.
- jev-snapshot derives its Action type from the protocol's JevAction (type-only
  import; the serialized function stays self-contained).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t notes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
karngyan and others added 26 commits September 28, 2026 21:47
…o settle

The read after an action came two animation frames later. A sort or filter
that navigates or re-renders lands after that, so Jev judged the page
mid-change: DONE right after "Most stars" before the new URL arrived, a
filter whose debounced URL update the read missed, a results page read
while it was still loading.

jevSettle keeps the two frames (and the typed-combobox wait for an option),
then, if the page has started changing since the input, waits until DOM
mutations have been quiet for 300 ms; a page that has not reacted within
150 ms of the frames is read at once. Capped at 1.5 s from the input for
pages that never stop changing. The submit path already worked this way.

Dev set, 3 runs each, against cc8217a (30/45): 33/45 — flights 2→3,
github 0→2, huggingface 1→2; npm 3→2 on the pre-existing race where DONE
follows a suggestion click whose navigation has not started yet (the
baseline's run ended on the same URL and verified only because the check
ran later); every other 3/3 task stayed 3/3, fx-risky risky_action ×3.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… session

`jev_act type` used `Input.insertText`, which fires `input` alone; a field
whose script re-derives its value on keydown/keyup/blur (formatting, a model
written back on blur) never sees the typing and drops it. Type one keyDown
(carrying the character as `text`, so Chromium generates keypress and input)
and one keyUp per character; values over 200 characters are still inserted
whole.

Typing into a form field can bring up a password manager's frame, and Chrome
drops the debugger session under the act. One insertText fell between
commands (observe's re-attach covered it); per-key typing is still sending
when it happens. On a detached/not-attached send failure the act drives the
tab again (polled up to 2 s while the guard clears the frame), reads the
field, and resumes after the characters that landed — or selects all and
starts over when the field holds something else — at most twice per act.

Bench v2 dev, 5 runs: 55/75 vs 52 (fx-form 5/5, no errors).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… options

A form control with no accessible name of its own read as its option text
(a select) or its title alone (an input), so two "Year" inputs in different
table rows, and a month and a day select, were indistinguishable to Jev,
which then clicked "Open Year" instead of typing. Name an unlabelled
input/select/textarea by its context — the first other cell of an
enclosing table row (that cell's own words, controls excluded, up to 60
chars), a fieldset's legend or a labelled group — prefixed to a
title/placeholder-only name ("Start Date: Year"); a select is never named
by its options.

Bench v2 dev: datecalc 0/20 → 8/10 across the round; the pooled 10-run
total is 111/150 vs 108/150, under the round's +6 bar, with the only
pooled drop (github 5→1) traced to an identical request in every arm — a
near-tie the change does not touch. Accepted on that evidence; see
fixloop-log.md, R2.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…s probability

A DONE chosen while a requirement is still open (a filter unapplied, a
query not run, a look-alike link opened) ended the run as done. Put every
DONE to Jev once more on the same observation as one `noul` question —
is every requirement of the goal visibly satisfied on the current page —
with the goal and the rules a filter/sort counts only where shown applied,
a query only once its results show, "open X" only on X. Below P = 0.5 the
DONE is refused: a verdict note enters the history Jev sees, the page is
read again and the loop goes on; after two refusals the third DONE stands.
The accepted probability is `doneConfidence` in the result (`· self-check
0.91` on the done line); the trace records it under `checks`. The check is
a counted, metered Jev call (a DONE costs one more call; at the call cap it
stands unchecked). A refused verdict is not an act: it is no step, and
neither progress nor its absence for the stuck rule, the loop breaker or
the unsubmitted-query rule.

Bench v2 dev, 5 runs: 61/75 vs 57. The check is calibrated — every
verified run scored 0.63–0.94, every done-but-wrong run 0.31–0.45 — though
a refusal did not change any run's outcome (Jev answered DONE again); the
signal is the gain.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The task was reworded when "I Accept" was a risky click; consent-banner
accepts are allowed since ffdb1b9, so the goal is the plain lookup again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TypeSafe returns probabilities rounded to two decimals, so a choice 0.01
under the top (BLOCKED 0.30 vs CLICK 0.31) is a tie as reported, not a
contradiction; the strict arg-max check ended about one run in six with
"TypeSafe returned an unusable answer". validateChoice now accepts a
choice within PROBABILITY_TIE (0.02, the sum check's own tolerance) of
the maximum; everything else stays strict. interpret already validates
only the heads the loop acts on (the operation, the chosen operation's
target, a fill head only when its field is typed into); that rule is now
documented and covered by a test with unusable speculative heads.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Round 2 showed a refused DONE never changed an outcome — Jev answered
DONE again on every re-read — and cost two extra calls per run. The
self-check now runs once per DONE and the DONE stands; doneConfidence is
returned as before. Below 0.5 the human output says so up front
("done (unsure: self-check 0.34) in 4.1s · …") and the next: line's hint
becomes "# self-check says the goal may not be met — verify"; the JSON
is unchanged. The refusal machinery (DONE_CHECK_MIN, DONE_REJECTIONS_MAX,
the verdict note in the history) is gone with it. SKILL.md: a low
self-check means verify, or fall back to step-by-step.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The v1 holdout tasks have been run and discussed, so they no longer
measure generalisation: they are dev now. holdout2 adds 8 unseen tasks,
frozen before round 3 and validated without `reins do` (false at start
and on near-miss states, true on a goal state reached by URL or eval):

- fx-settings: tabbed settings panel with role=switch toggles and Save
- fx-orders: paginated table whose target row sits on page 2
- lit: shadow-DOM search modal on lit.dev (web components)
- musicbrainz: native <select> search type
- openlibrary: web-component "Sort by" menu on search results
- crates: SPA search → crate → Versions route
- iana: find one row in a ~1600-row static table
- osm: disambiguate same-named search results by location

Runner: --set dev|holdout2|all; `all` includes holdout2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Controls and text inside open shadow roots were invisible to `jev_observe`:
`querySelectorAll` and the text walker stop at the boundary, so a site whose
search box or dialog is a web component read as having no such control at
all, and Jev answered BLOCKED at step 0 (live: a documentation site with 0
light-DOM inputs and 9 in shadow roots; bench dev mdn 0/5 → 5/5).

The snapshot now collects controls with a tree walker that descends into
open shadow roots where their hosts sit (document order, so ids stay
positional across trees) and reads text the same way; ancestor walks
(visibility, aria-disabled, the row/group name, consent banners, the guard's
scope) cross the boundary through the host; aria-labelledby/aria-controls
resolve in the control's own tree; jevSubmitFocus compares against the
root's activeElement. actionPoint already hit-tests through
shadowRoot.elementFromPoint. Closed roots stay invisible (spec).

Browser test: a custom element whose open root holds a form[role=search]
input and button — listed (input marked submit), typed into, submitted and
clicked; a closed root's button is not listed.

Dev set (5 runs): 84/110 vs 83 — mdn 0→5, huggingface 0→2; github 3→0,
cambridge 5→3, wolframalpha 5→4, each shown in the fix-loop log to be
independent of this change (byte-identical requests, or no Jev call at
all). Kept over the +3 bar on that evidence; flagged for the maintainer.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…goal

A run can reach its goal and keep acting on the page it is on (re-clicking
the link that led there), and the no-progress rule then reports `stuck`
although the goal is met. Before either stuck (three no-op actions, or
three stale acts) the loop now puts the DONE self-check to the observation
it is on: at STUCK_DONE_MIN (0.8) or above the run is `done` with that
doneConfidence and no reason; below, `stuck` as before, carrying the
probability in the JSON result and on the human line ("stuck: … (self-check
0.06)") for diagnosis. One Jev call, skipped when the budget is spent.

Dev set (5 runs): 87/110 vs 84 (the S1 run). The gain is the usual github
coin and a network-clean cambridge, not this change: no stuck page scored
0.8 this run — the page-only check gives the targeted case 0.21–0.30,
because what the goal asks ("the first story of a section") is in the
history, not on the page. What the change demonstrably adds is the
calibrated value on every stuck result (arxiv 0.06, huggingface 0.05).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The 8 holdout2 tasks have now been run once (on 3c8a364, 16/40) and their
failures discussed, so they no longer measure generalisation. They are
`set: "dev"` (30 dev tasks); a new holdout3 will be built separately.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`name()` read an element's light children, so a web-component option whose
text sits in its own open shadow root, or a button whose caption arrives
through a slot, had no name and was refused as risky ("it has no label").
Read what the element renders instead, the way the accessibility tree does:
an open root's children stand in for the host's own, a slot for its
assigned nodes; aria-hidden subtrees are still skipped and a control with
no words anywhere still reads as its role (and stays risky).

Dev A/B (30 tasks × 5): 102/150 vs 101 — the shadow-DOM search task 0→4;
the drops are the known github coin (byte-identical requests) and noise.
Kept over the +3 bar on that evidence; flagged for the maintainer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rome

A trusted click (Input.dispatchMouseEvent) on <a target=_blank> made Chrome
open a foreground tab and activate itself over whatever app the user was
working in. A ⌘/Ctrl-modified press (tested live on macOS) only moved the
tab behind: the window was still raised. chrome.tabs.create is not.

actionPoint arms a click listener on window next to the pointer probe when
the target is inside a link whose effective target opens a new browsing
context (_blank, a named target that is no frame of the page, or a <base
target> that does; download links excluded). It runs after the page's own
handlers, cancels the link's navigation unless the page already did, and
records the resolved href in the probe. pressAt returns it as newTabUrl and
openLinkTab opens it with chrome.tabs.create next to the opener, active,
with openerTabId — the user sees the new tab as after a normal click.

reins do follows that tab as before (the tabs.onCreated watcher stays for
tabs a page opens from script). reins click reports the additive
ClickResult.openedTabId and prints `ok — opened tab <id> (now active)`.
Script-opened tabs and pages that stop the click's propagation are out of
scope and may still raise Chrome.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ation

The observation is viewport-only, so a target far down a long page or a
pager below the fold is invisible until a scroll happens to land on it,
and Jev answers BLOCKED. List, after the viewport's controls, at most 12
off-screen links/buttons that are either pagers (rel next/prev; a name
like next / previous / page N / load more; a bare number inside a nav
that calls itself pagination — never the aria-current one) or named by a
distinctive word of the goal, marked `offscreen: true` (shown to Jev as a
criterion). `jev_observe` takes the goal's words as `terms` (daemon-side
`goalTerms()`: tokens of 3+ characters minus common words, plus any with
a digit or a dot). Page text stays viewport-only; acting on such a
control scrolls it into view first.

Dev A/B (30 tasks × 5): 110/150 vs 106 — the long-table task 0→5 in one
step per run; no ≥4/5 task below 3/5. Median tokens per call +8%.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A results page can settle with a "Loading…" placeholder while an XHR
fetches the results a second or two later; read at once, Jev saw no
result and re-ran the search, which reloaded the page and started the
wait over. The driven session now enables the Network domain on every act
and observe and counts the tab's XHR/Fetch requests in flight; a settled
document with such requests pending is read again once they land (plus
150 ms for the render), or 1.5 s after the first settled read at the
latest. A page with nothing in flight is read as before.

Dev A/B (30 tasks × 5): 114/150 vs 110 — the async-results task 2→5 with
no repeated submit; no ≥4/5 task below 3/5. Median tokens per call +6%,
median run time +20% (the wait).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… slotted controls

Two halves of one mechanism — run a typed query, then click a
web-component control on the results page — each useless without the
other on the site that showed it:

- The "Open <field>" pseudo-click of the field the last action typed
  into is not offered while the field still holds the typed value: it
  only focuses a field that is focused already (an act that cannot
  change the page), and under a name like "Open Search" Jev took it for
  the way to run the query, three times over. Enter, the page's own
  submit control and the suggestions stay on offer; a field never typed
  into keeps its click.
- actionPoint's hit-test: when the top-most box at the point is light-DOM
  content a host slots into the target (a design-system button's
  caption), elementFromPoint retargets it to the host and the ancestor
  walk stopped there ("covered by <host>"). A host of the target's shadow
  tree standing at the point now counts as the target — the press's
  composed path runs through it.

Dev A/B (30 tasks × 5): 120/150 vs 114 — the web-component search-and-
sort task 0→4; each half alone: 114 and 113. No ≥4/5 task below 3/5.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…asks)

Frozen 2026-09-28 at f1af855 before any `reins do` run on it. Fixtures:
an accordion FAQ whose helpful-vote buttons hide until the answer is
expanded, and a modal wizard whose step 2 is a custom role=radio group.
Live: gutenberg (site search), hn.algolia (facet + sort dropdowns),
debian (two native selects), rustdocs (docs site, method anchor), gopkg
(multi-step navigation), xe (number input + currency comboboxes). Every
checker validated without `reins do`: false at start and on near misses,
true on the goal state. Runner: --set dev|holdout3|all.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The converter's currency comboboxes are searched by typing; reins do takes
text only from --fill, so without from/to every run correctly stopped at
needs_text.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
sameSite strips a leading `www.` from both hosts before the equal-or-
subdomain test, so a run that starts on www.x.org stays on site when the
page moves to packages.x.org (and back). Nothing beyond that: no public-
suffix guessing, a.github.io and b.github.io remain two sites. Live:
a holdout task solved 5/5 but reported `left_site` every run because its
search results live on a sibling host of the www. start page.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fill

When Jev's chosen text field's fill head answers NONE (an output field,
say) while another offered text field's speculative head names a fill
not yet typed this run, type that instead of stopping needs_text: the
candidate with the highest type_text_target probability × fill-head
probability wins and the step is recorded as usual. Only with no such
field does the run stop needs_text as before; a field beyond the first
MAX_FILL_HEADS (no speculative heads) keeps today's stop. An unusable
speculative head is skipped, never thrown on.

Dev A/B with U1 (5 runs): 124/150 vs 122, no ≥4/5 task below 3/5,
fx-risky risky_action ×5; neither mechanism fires on dev. debian ×5:
done 5/5 (was left_site ×5). xe ×5: still needs_text — both of its text
fields carry the site's own "Receiving amount" aria-label, so every head
says NONE (Round 5 in fixloop-log.md).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…come means

A result table (status → what it means → what to do), the self-check as
the number to trust (every wrong done in the benchmark was marked unsure),
and the limits stated plainly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rm by its page check

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… every limitation

Replaces the 4-task write-up. reins do on the 30 dev tasks (e9331fa):
124/150; on the 8 unseen holdout3 tasks (first and only run, f1af855):
35/40. Claude step by step: 81/90 dev (81/87 without fx-newtab, whose
new-tab link its rules forbid), 21/24 holdout3; the three fx-risky runs
are scored by their page check (Claude stopped before deleting). On the
30 tasks both arms pass, the median task is 4.6x faster and 111x cheaper
in model spend (Jev only). Every wrong done (17 of 166) had a self-check
under 0.5; 20 of 149 right ones did too.

Raw data byte-identical to the runner's output in 2026-09-reins-do-v2/
(logs force-added past the *.log ignore), plus four earlier runs behind
the round table. biome now skips JSON anywhere under docs/benchmarks so
those files stay as written. The v1 report shrinks to a closing section;
its JSON files stay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Jev key-set state is a row like a site-permissions row: a green status
  dot, "Ready", the masked key in muted mono, and Replace/Remove on the
  right. No more pale-green text (2.5:1 on white, failed AA).
- Get a key / What's sent to TypeSafe are 11px muted footnote links with
  a soft underline, a ↗ cue, hover to ink, a neutral focus ring and a
  28px tap target.
- Status pill is a neutral chip; only the dot carries state colour.
- Errors (Jev save, unreadable policy) share one tinted alert; the old
  #ef4444 error text was 3.8:1.
- Tokens in oklch with a full prefers-color-scheme dark palette; one
  control height (32px); 11/12/15 type scale; 4/8/12/16 spacing.
- Neutral focus rings everywhere, 0.96 press scale, disabled controls go
  grey instead of half-opacity violet, custom Advanced chevron.
- Width 372px so the subtitle stays on one line in every connection
  state.
- jevReadyText becomes jevMaskedKey ("••••a1b2"); "Ready" is static markup.
- Website popup mock: neutral Connected pill to match.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…dy set

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The code is the record once a feature is built: docs/superpowers (specs,
plans) is removed and git-ignored, and the benchmark keeps its report
only; raw run files are regenerated by bench-do.mjs, not committed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@karngyan
karngyan merged commit f2b0801 into main Sep 28, 2026
2 checks passed
@karngyan
karngyan deleted the feat/reins-do-loop branch September 28, 2026 16:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant