feat: thinking, token limits, compaction, skills, tool consent, tool search, subagents, token efficiency (0.0.31) - #50
Merged
Conversation
…search, subagents, token efficiency Shared spec with @dudko.dev/agent-web (same config fields, events and semantics): - Thinking: `thinking` / `stageThinking` → the AI SDK's portable `reasoning` plus provider budgets; thoughts stream as step/final reasoning-delta events. - Limits: run-level token caps (input / output / reasoning / total), per-call output caps per stage, `maxToolCalls`, `maxPlanSteps`; a crossed cap stops at the next boundary, emits budget.exceeded and still answers. - Compaction: auto (history at run start, trace before each call) and manual `agent.compact()`; per-result tool-output caps; stale tool results inside a step's loop become stubs past `clearToolResultsAfterTokens` (sticky, position-keyed). - Skills (agentskills.io SKILL.md): index in the prompt, body on activation by the planner or `load_skill`; `read_skill_file`. - Tool consent: autopilot / ask-writes / ask-all / read-only, glob rules, remember, timeout; switchable live (setToolApprovalMode). - Large catalogues: paginated tools/list, per-server connect timeout, 'auto' tool strategy → `find_tools` search above the threshold. - Prompt caching: run-stable system prompts with an Anthropic breakpoint, OpenAI promptCacheKey, a rolling breakpoint on the newest loop message, tools sorted by name. - Subagents: createSubagentTool in worker_threads or in-process, host tools proxied through the parent's consent gate. - Autonomy rules in the prompts: act, don't ask. - CLI flags/config for all of it; README, CLAUDE.md, env.example; deps bumped. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
- Subagent workers inherit the parent's node options by default; an explicit
execArgv is passed only to strip entry-point flags (-e / -p /
--input-type). Passing process.execArgv explicitly failed on Node 24 under
`node --test` ("Initiated Worker with invalid execArgv flags").
- The real-model gate (Qwen2.5-3B) regressed: the rewritten planner rule put
the canned "Answer the user directly" next to the lookup rule, extra
executor rules made the 3B echo a guessed answer alongside the real lookup,
and sorting tools by name reordered what the server declared. Reproduced
locally with the pinned llama.cpp image and weights, fresh server, two
attempts (as CI): planner rules restored with the autonomy rule added as
rule 6, executor prompt as measured, tools in declaration order (servers
already mount in declaration order, so the cached prefix stays stable).
Now 2/2 on a fresh server.
- Version 0.0.31.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
siarheidudko
marked this pull request as ready for review
October 3, 2026 00:26
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR implements the agent spec it shares with
@dudko.dev/agent-web(browser): the same config fields, events and semantics. Full documentation is in the README. Releases as 0.0.31.What's in it
Thinking
thinking/stageThinkingmap to the AI SDK's portablereasoninglevel, plus provider budgets.Limits
maxToolCallsandmaxPlanSteps.budget.exceeded, and still produces an answer.Compaction
agent.compact(), and/compactin the CLI.clearToolResultsAfterTokens. Clearing is sticky and keyed by position.Skills
SKILL.mdformat from agentskills.io.load_skill/read_skill_file.Tool consent
autopilot,ask-writes,ask-all,read-only.--approvaland/autopilot.Large MCP catalogues
tools/listis read to the end (pagination).connectTimeoutMs.'auto'strategy switches tofind_toolssearch above the threshold.Prompt caching
promptCacheKey.Subagents
createSubagentToolruns inworker_threads(dist/subagent-worker.js) or in-process.Autonomy
[BLOCKER]→ replanner.Dependencies
ai7.0.127,@modelcontextprotocol/sdk1.32, and the@ai-sdk/*providers.typescriptand@types/nodestay pinned, per the org ignore list.CI fixes in the second commit
process.execArgv. Node 24 rejects the per-process flagsnode --testadds; workers inherit the options themselves. Reproduced onnode:24in Docker and fixed: 256/256.mainpasses, the branch failed.Checks
npm run typecheck,npm run format:check,npm run buildandnpm testall pass on Node 22 and Node 24 (256 passed, live test skipped locally).🤖 Generated with Claude Code
https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5