feat: reins do — delegate a browsing goal to Jev, with benchmark (3/3) - #43
Merged
Merged
Conversation
This was referenced Sep 27, 2026
karngyan
added this pull request to stack #44
September 27, 2026 08:49
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… only The goal allowance used substring matching, so "reorder the list" waived an "Order" button and "paypal login" waived "Pay". It now requires the whole normalized label bounded by non-alphanumerics. The no-label check runs before --confirm so --confirm button cannot waive unlabeled buttons, and non-click actions are never risky. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The boundary class [^a-z0-9] treated non-ASCII letters as boundaries, so
"prépay the bill" waived a "Pay" button and "日本語pay" likewise. The
boundary is now [^\p{L}\p{N}] under the u flag; only regex syntax
characters are escaped, since escaping anything else throws in unicode mode.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Per ruling R15 the --continue loop breaker turns only risky_action, needs_text, blocked, budget and stuck into stuck + lock; dialog, interrupted and left_site pass through. It also requires the last action's outcome to have been observed, so an abort or timeout before the next read is not mistaken for no change. A dialog result now carries the current url/title and the last step's pageChanged. needs_text says when the fills matched nothing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pand JSON.stringify gave a double-quoted shell string where $, backticks and $(...) still expand, so a label like 'Pay $5' never matched --confirm and a page-controlled label could execute when pasted. Use POSIX single quotes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…estart Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…first, pins the tab once Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… its usage The done case of nextCommand printed a bare reins snapshot while every other next line carried --tab/--browser; the usage string now also lists --browser and --json, and single-quotes the goal and label like the printed next lines do. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… stricter checkers Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rm, partial results on abort Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… wait The loop only checked elapsed time between iterations, so a slow jev_act or a retried Jev call near the deadline overran the CLI's fetch abort and printed a raw undici error. handleDo now aborts the loop's signal with a TimeoutError when --timeout elapses; the loop reports that as budget (timed out after Ns), not interrupted. The CLI waits timeout + 40 s (one full bridge call of slack) and turns a fetch timeout into one readable line. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y handleDo - A dialog observation carries no page: stop with `dialog` without touching lastFingerprint or the last step's pageChanged, and keep the previous url/title when the observation has none. - A run that began on about:blank / file:// / an error page pins its start host on the first http(s) page, so left_site can fire later. - The select audit label keeps only the field (Country), not the option. - A missing goal fails before the first jev_observe. - The busy-tab refusal says a just-stopped run may be finishing its action. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d run memory Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e popup, protocol action type - jev_observe fills url/title from chrome.tabs.get while a JS dialog blocks CDP, so the daemon's result and run state keep a real page. - The popup's Replace/Cancel paths re-render from live connectivity, like Save. - jev-snapshot derives its Action type from the protocol's JevAction (type-only import; the serialized function stays self-contained). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t notes Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…o settle The read after an action came two animation frames later. A sort or filter that navigates or re-renders lands after that, so Jev judged the page mid-change: DONE right after "Most stars" before the new URL arrived, a filter whose debounced URL update the read missed, a results page read while it was still loading. jevSettle keeps the two frames (and the typed-combobox wait for an option), then, if the page has started changing since the input, waits until DOM mutations have been quiet for 300 ms; a page that has not reacted within 150 ms of the frames is read at once. Capped at 1.5 s from the input for pages that never stop changing. The submit path already worked this way. Dev set, 3 runs each, against cc8217a (30/45): 33/45 — flights 2→3, github 0→2, huggingface 1→2; npm 3→2 on the pre-existing race where DONE follows a suggestion click whose navigation has not started yet (the baseline's run ended on the same URL and verified only because the check ran later); every other 3/3 task stayed 3/3, fx-risky risky_action ×3. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… session `jev_act type` used `Input.insertText`, which fires `input` alone; a field whose script re-derives its value on keydown/keyup/blur (formatting, a model written back on blur) never sees the typing and drops it. Type one keyDown (carrying the character as `text`, so Chromium generates keypress and input) and one keyUp per character; values over 200 characters are still inserted whole. Typing into a form field can bring up a password manager's frame, and Chrome drops the debugger session under the act. One insertText fell between commands (observe's re-attach covered it); per-key typing is still sending when it happens. On a detached/not-attached send failure the act drives the tab again (polled up to 2 s while the guard clears the frame), reads the field, and resumes after the characters that landed — or selects all and starts over when the field holds something else — at most twice per act. Bench v2 dev, 5 runs: 55/75 vs 52 (fx-form 5/5, no errors). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… options
A form control with no accessible name of its own read as its option text
(a select) or its title alone (an input), so two "Year" inputs in different
table rows, and a month and a day select, were indistinguishable to Jev,
which then clicked "Open Year" instead of typing. Name an unlabelled
input/select/textarea by its context — the first other cell of an
enclosing table row (that cell's own words, controls excluded, up to 60
chars), a fieldset's legend or a labelled group — prefixed to a
title/placeholder-only name ("Start Date: Year"); a select is never named
by its options.
Bench v2 dev: datecalc 0/20 → 8/10 across the round; the pooled 10-run
total is 111/150 vs 108/150, under the round's +6 bar, with the only
pooled drop (github 5→1) traced to an identical request in every arm — a
near-tie the change does not touch. Accepted on that evidence; see
fixloop-log.md, R2.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…s probability A DONE chosen while a requirement is still open (a filter unapplied, a query not run, a look-alike link opened) ended the run as done. Put every DONE to Jev once more on the same observation as one `noul` question — is every requirement of the goal visibly satisfied on the current page — with the goal and the rules a filter/sort counts only where shown applied, a query only once its results show, "open X" only on X. Below P = 0.5 the DONE is refused: a verdict note enters the history Jev sees, the page is read again and the loop goes on; after two refusals the third DONE stands. The accepted probability is `doneConfidence` in the result (`· self-check 0.91` on the done line); the trace records it under `checks`. The check is a counted, metered Jev call (a DONE costs one more call; at the call cap it stands unchecked). A refused verdict is not an act: it is no step, and neither progress nor its absence for the stuck rule, the loop breaker or the unsubmitted-query rule. Bench v2 dev, 5 runs: 61/75 vs 57. The check is calibrated — every verified run scored 0.63–0.94, every done-but-wrong run 0.31–0.45 — though a refusal did not change any run's outcome (Jev answered DONE again); the signal is the gain. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The task was reworded when "I Accept" was a risky click; consent-banner accepts are allowed since ffdb1b9, so the goal is the plain lookup again. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TypeSafe returns probabilities rounded to two decimals, so a choice 0.01 under the top (BLOCKED 0.30 vs CLICK 0.31) is a tie as reported, not a contradiction; the strict arg-max check ended about one run in six with "TypeSafe returned an unusable answer". validateChoice now accepts a choice within PROBABILITY_TIE (0.02, the sum check's own tolerance) of the maximum; everything else stays strict. interpret already validates only the heads the loop acts on (the operation, the chosen operation's target, a fill head only when its field is typed into); that rule is now documented and covered by a test with unusable speculative heads. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Round 2 showed a refused DONE never changed an outcome — Jev answered
DONE again on every re-read — and cost two extra calls per run. The
self-check now runs once per DONE and the DONE stands; doneConfidence is
returned as before. Below 0.5 the human output says so up front
("done (unsure: self-check 0.34) in 4.1s · …") and the next: line's hint
becomes "# self-check says the goal may not be met — verify"; the JSON
is unchanged. The refusal machinery (DONE_CHECK_MIN, DONE_REJECTIONS_MAX,
the verdict note in the history) is gone with it. SKILL.md: a low
self-check means verify, or fall back to step-by-step.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The v1 holdout tasks have been run and discussed, so they no longer measure generalisation: they are dev now. holdout2 adds 8 unseen tasks, frozen before round 3 and validated without `reins do` (false at start and on near-miss states, true on a goal state reached by URL or eval): - fx-settings: tabbed settings panel with role=switch toggles and Save - fx-orders: paginated table whose target row sits on page 2 - lit: shadow-DOM search modal on lit.dev (web components) - musicbrainz: native <select> search type - openlibrary: web-component "Sort by" menu on search results - crates: SPA search → crate → Versions route - iana: find one row in a ~1600-row static table - osm: disambiguate same-named search results by location Runner: --set dev|holdout2|all; `all` includes holdout2. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Controls and text inside open shadow roots were invisible to `jev_observe`: `querySelectorAll` and the text walker stop at the boundary, so a site whose search box or dialog is a web component read as having no such control at all, and Jev answered BLOCKED at step 0 (live: a documentation site with 0 light-DOM inputs and 9 in shadow roots; bench dev mdn 0/5 → 5/5). The snapshot now collects controls with a tree walker that descends into open shadow roots where their hosts sit (document order, so ids stay positional across trees) and reads text the same way; ancestor walks (visibility, aria-disabled, the row/group name, consent banners, the guard's scope) cross the boundary through the host; aria-labelledby/aria-controls resolve in the control's own tree; jevSubmitFocus compares against the root's activeElement. actionPoint already hit-tests through shadowRoot.elementFromPoint. Closed roots stay invisible (spec). Browser test: a custom element whose open root holds a form[role=search] input and button — listed (input marked submit), typed into, submitted and clicked; a closed root's button is not listed. Dev set (5 runs): 84/110 vs 83 — mdn 0→5, huggingface 0→2; github 3→0, cambridge 5→3, wolframalpha 5→4, each shown in the fix-loop log to be independent of this change (byte-identical requests, or no Jev call at all). Kept over the +3 bar on that evidence; flagged for the maintainer. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…goal
A run can reach its goal and keep acting on the page it is on (re-clicking
the link that led there), and the no-progress rule then reports `stuck`
although the goal is met. Before either stuck (three no-op actions, or
three stale acts) the loop now puts the DONE self-check to the observation
it is on: at STUCK_DONE_MIN (0.8) or above the run is `done` with that
doneConfidence and no reason; below, `stuck` as before, carrying the
probability in the JSON result and on the human line ("stuck: … (self-check
0.06)") for diagnosis. One Jev call, skipped when the budget is spent.
Dev set (5 runs): 87/110 vs 84 (the S1 run). The gain is the usual github
coin and a network-clean cambridge, not this change: no stuck page scored
0.8 this run — the page-only check gives the targeted case 0.21–0.30,
because what the goal asks ("the first story of a section") is in the
history, not on the page. What the change demonstrably adds is the
calibrated value on every stuck result (arxiv 0.06, huggingface 0.05).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The 8 holdout2 tasks have now been run once (on 3c8a364, 16/40) and their failures discussed, so they no longer measure generalisation. They are `set: "dev"` (30 dev tasks); a new holdout3 will be built separately. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`name()` read an element's light children, so a web-component option whose
text sits in its own open shadow root, or a button whose caption arrives
through a slot, had no name and was refused as risky ("it has no label").
Read what the element renders instead, the way the accessibility tree does:
an open root's children stand in for the host's own, a slot for its
assigned nodes; aria-hidden subtrees are still skipped and a control with
no words anywhere still reads as its role (and stays risky).
Dev A/B (30 tasks × 5): 102/150 vs 101 — the shadow-DOM search task 0→4;
the drops are the known github coin (byte-identical requests) and noise.
Kept over the +3 bar on that evidence; flagged for the maintainer.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rome A trusted click (Input.dispatchMouseEvent) on <a target=_blank> made Chrome open a foreground tab and activate itself over whatever app the user was working in. A ⌘/Ctrl-modified press (tested live on macOS) only moved the tab behind: the window was still raised. chrome.tabs.create is not. actionPoint arms a click listener on window next to the pointer probe when the target is inside a link whose effective target opens a new browsing context (_blank, a named target that is no frame of the page, or a <base target> that does; download links excluded). It runs after the page's own handlers, cancels the link's navigation unless the page already did, and records the resolved href in the probe. pressAt returns it as newTabUrl and openLinkTab opens it with chrome.tabs.create next to the opener, active, with openerTabId — the user sees the new tab as after a normal click. reins do follows that tab as before (the tabs.onCreated watcher stays for tabs a page opens from script). reins click reports the additive ClickResult.openedTabId and prints `ok — opened tab <id> (now active)`. Script-opened tabs and pages that stop the click's propagation are out of scope and may still raise Chrome. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ation The observation is viewport-only, so a target far down a long page or a pager below the fold is invisible until a scroll happens to land on it, and Jev answers BLOCKED. List, after the viewport's controls, at most 12 off-screen links/buttons that are either pagers (rel next/prev; a name like next / previous / page N / load more; a bare number inside a nav that calls itself pagination — never the aria-current one) or named by a distinctive word of the goal, marked `offscreen: true` (shown to Jev as a criterion). `jev_observe` takes the goal's words as `terms` (daemon-side `goalTerms()`: tokens of 3+ characters minus common words, plus any with a digit or a dot). Page text stays viewport-only; acting on such a control scrolls it into view first. Dev A/B (30 tasks × 5): 110/150 vs 106 — the long-table task 0→5 in one step per run; no ≥4/5 task below 3/5. Median tokens per call +8%. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A results page can settle with a "Loading…" placeholder while an XHR fetches the results a second or two later; read at once, Jev saw no result and re-ran the search, which reloaded the page and started the wait over. The driven session now enables the Network domain on every act and observe and counts the tab's XHR/Fetch requests in flight; a settled document with such requests pending is read again once they land (plus 150 ms for the render), or 1.5 s after the first settled read at the latest. A page with nothing in flight is read as before. Dev A/B (30 tasks × 5): 114/150 vs 110 — the async-results task 2→5 with no repeated submit; no ≥4/5 task below 3/5. Median tokens per call +6%, median run time +20% (the wait). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… slotted controls
Two halves of one mechanism — run a typed query, then click a
web-component control on the results page — each useless without the
other on the site that showed it:
- The "Open <field>" pseudo-click of the field the last action typed
into is not offered while the field still holds the typed value: it
only focuses a field that is focused already (an act that cannot
change the page), and under a name like "Open Search" Jev took it for
the way to run the query, three times over. Enter, the page's own
submit control and the suggestions stay on offer; a field never typed
into keeps its click.
- actionPoint's hit-test: when the top-most box at the point is light-DOM
content a host slots into the target (a design-system button's
caption), elementFromPoint retargets it to the host and the ancestor
walk stopped there ("covered by <host>"). A host of the target's shadow
tree standing at the point now counts as the target — the press's
composed path runs through it.
Dev A/B (30 tasks × 5): 120/150 vs 114 — the web-component search-and-
sort task 0→4; each half alone: 114 and 113. No ≥4/5 task below 3/5.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…asks) Frozen 2026-09-28 at f1af855 before any `reins do` run on it. Fixtures: an accordion FAQ whose helpful-vote buttons hide until the answer is expanded, and a modal wizard whose step 2 is a custom role=radio group. Live: gutenberg (site search), hn.algolia (facet + sort dropdowns), debian (two native selects), rustdocs (docs site, method anchor), gopkg (multi-step navigation), xe (number input + currency comboboxes). Every checker validated without `reins do`: false at start and on near misses, true on the goal state. Runner: --set dev|holdout3|all. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The converter's currency comboboxes are searched by typing; reins do takes text only from --fill, so without from/to every run correctly stopped at needs_text. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
sameSite strips a leading `www.` from both hosts before the equal-or- subdomain test, so a run that starts on www.x.org stays on site when the page moves to packages.x.org (and back). Nothing beyond that: no public- suffix guessing, a.github.io and b.github.io remain two sites. Live: a holdout task solved 5/5 but reported `left_site` every run because its search results live on a sibling host of the www. start page. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fill When Jev's chosen text field's fill head answers NONE (an output field, say) while another offered text field's speculative head names a fill not yet typed this run, type that instead of stopping needs_text: the candidate with the highest type_text_target probability × fill-head probability wins and the step is recorded as usual. Only with no such field does the run stop needs_text as before; a field beyond the first MAX_FILL_HEADS (no speculative heads) keeps today's stop. An unusable speculative head is skipped, never thrown on. Dev A/B with U1 (5 runs): 124/150 vs 122, no ≥4/5 task below 3/5, fx-risky risky_action ×5; neither mechanism fires on dev. debian ×5: done 5/5 (was left_site ×5). xe ×5: still needs_text — both of its text fields carry the site's own "Receiving amount" aria-label, so every head says NONE (Round 5 in fixloop-log.md). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…come means A result table (status → what it means → what to do), the self-check as the number to trust (every wrong done in the benchmark was marked unsure), and the limits stated plainly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rm by its page check Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… every limitation Replaces the 4-task write-up. reins do on the 30 dev tasks (e9331fa): 124/150; on the 8 unseen holdout3 tasks (first and only run, f1af855): 35/40. Claude step by step: 81/90 dev (81/87 without fx-newtab, whose new-tab link its rules forbid), 21/24 holdout3; the three fx-risky runs are scored by their page check (Claude stopped before deleting). On the 30 tasks both arms pass, the median task is 4.6x faster and 111x cheaper in model spend (Jev only). Every wrong done (17 of 166) had a self-check under 0.5; 20 of 149 right ones did too. Raw data byte-identical to the runner's output in 2026-09-reins-do-v2/ (logs force-added past the *.log ignore), plus four earlier runs behind the round table. biome now skips JSON anywhere under docs/benchmarks so those files stay as written. The v1 report shrinks to a closing section; its JSON files stay. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Jev key-set state is a row like a site-permissions row: a green status
dot, "Ready", the masked key in muted mono, and Replace/Remove on the
right. No more pale-green text (2.5:1 on white, failed AA).
- Get a key / What's sent to TypeSafe are 11px muted footnote links with
a soft underline, a ↗ cue, hover to ink, a neutral focus ring and a
28px tap target.
- Status pill is a neutral chip; only the dot carries state colour.
- Errors (Jev save, unreadable policy) share one tinted alert; the old
#ef4444 error text was 3.8:1.
- Tokens in oklch with a full prefers-color-scheme dark palette; one
control height (32px); 11/12/15 type scale; 4/8/12/16 spacing.
- Neutral focus rings everywhere, 0.96 press scale, disabled controls go
grey instead of half-opacity violet, custom Advanced chevron.
- Width 372px so the subtitle stays on one line in every connection
state.
- jevReadyText becomes jevMaskedKey ("••••a1b2"); "Ready" is static markup.
- Website popup mock: neutral Connected pill to match.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…dy set Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
karngyan
force-pushed
the
feat/reins-do-loop
branch
from
September 28, 2026 16:18
d1dfe86 to
4177ed8
Compare
The code is the record once a feature is built: docs/superpowers (specs, plans) is removed and git-ignored, and the benchmark keeps its report only; raw run files are regenerated by bench-do.mjs, not committed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part 3 of 3 for
reins do(#41 and #42 are merged): the command itself, the benchmark, and every fix from review and live testing.Why
Driving a site one command at a time costs an agent one turn per click.
reins do "<goal>"runs the whole loop inside the daemon, with Jev (TypeSafe's action model) picking each action. It returns only when the goal is met or a decision is needed.What
reins do '<goal>' [--fill name=value]... [--confirm '<label>']... [--continue] [--max-steps 30] [--timeout 60] [--tab] [--browser] [--json]--fill.Bounded autonomy. Every run ends with a status and, where one applies, a
next:line to run:donerisky_actionneeds_textleft_sitedialoginterruptedblockedstuckbudgeterrorExit codes: 0 done, 2 stopped, 1 error.
Self-check. Every
done, and everystuckrun, gets a yes/no question to Jev. A low score printsdone (unsure: self-check 0.34). In the benchmark, all 17 wrongdoneresults scored below 0.5; 13% of right ones did too.Safety. Risky clicks stop the run. The rule is whole-word, Unicode-aware, and fires on unlabelled buttons. "Accept" on a cookie or consent banner is allowed. Leaving the site stops the run (a leading
www.is ignored). Password and hidden values are never sent.Coverage of real pages:
SUBMIT_SEARCH(Enter, only in search-like fields)Browser-wide fixes found along the way:
document.hasFocus()false.reins do --helpworks.next:lines.Skill. A rewritten
reins dosection covering when to use it, how to call it, a table of statuses and what to do for each, and its limits.Benchmark. 38 tasks: local fixture pages for tricky widgets plus live sites. A dev set, and an unseen holdout that stayed frozen until its one run. Traces are optional (
REINS_JEV_TRACE). There is a report and a/docs/benchmarkspage with every limitation listed.CI. Runs the real-browser extension suite against CI's Chrome.
REINS_TEST_CHROMEis now declared inturbo.json; turbo's strict env had been dropping it.Benchmark (38 tasks;
reins do5 runs each, Claude Opus 5.5 step by step 3 runs each)reins doOn tasks both arms passed, the median was 4.6× faster (range 2.4–22.6×) and 111× cheaper (range 44–535×). The cost figure counts Jev only; the calling agent still spends a turn to issue the command and verify.
Tasks
reins dostill fails:The full report is
docs/benchmarks/2026-09-reins-do.md.The report also records two bugs in the step-by-step commands. They are out of scope here and will be fixed in a follow-up:
reins snapshotdoesn't read shadow DOM.Testing
pnpm build,pnpm lintandpnpm typecheckare clean.risky_action.🤖 Generated with Claude Code