feat: bounded extraction and faster agent workflows - #44
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Browser workflows could lose snapshot delta baselines after intermediate observations, direct-CDP batch fills ignored the documented value alias, and native WebMCP error text could cause a second invocation. This change fixes those paths, adds explicit read settling budgets, preserves code formatting in public-document reads, and makes repeated tool discovery advance.
Recipes gain named, bounded extraction into browser-host artifacts. Reviewed recipes can live in project
.brw/recipes/directories, with optional immutable registry adapters. The initial skill is substantially smaller through linked references. Research documents distinguish measured brw behavior, current competitor capabilities, and proposed algorithms/model orchestration experiments.Validation: full
task checkruns as the pre-push gate; focused real-Chrome and wire-contract regressions cover changed behavior.task agent-eval-verifygraded all eight ordinary/sabotaged runs correctly. Controlled read-settle measurements and the 35-command fixture benchmark are recorded in the October review; no matched cross-product speed claim is made. No live purchases were performed.