Skip to content

feat: thinking, limits, compaction, skills, tool consent, tool search, subagents, attachments, VFS, token efficiency (0.0.20) - #18

Merged
siarheidudko merged 6 commits into
mainfrom
claude/agents-review-improvements-2m2406
Oct 3, 2026
Merged

siarheidudko merged 6 commits into
mainfrom
claude/agents-review-improvements-2m2406

Conversation

@siarheidudko

@siarheidudko siarheidudko commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

This PR implements the agent spec it shares with @dudko.dev/agent (Node; released as 0.0.31): the same config fields, events and semantics. Full documentation is in docs/capabilities.md. Releases as 0.0.20.

What's in it

Thinking

  • thinking / stageThinking map to the AI SDK's portable reasoning level, plus provider budgets.
  • Thoughts stream as step.reasoning-delta / final.reasoning-delta.

Limits

  • Token caps per run, by kind: input / output / reasoning / total.
  • Per-call output caps per stage.
  • maxToolCalls and maxPlanSteps.
  • When a cap is crossed, the run stops at the next boundary, emits budget.exceeded, and still produces an answer.

Compaction

  • Automatic: the history before planning, the trace between steps.
  • Manual: agent.compact().
  • Every tool result the model sees is capped.
  • Context editing: inside a step's tool loop, past compaction.clearToolResultsAfterTokens, the oldest tool results become one-line stubs. Clearing is sticky and keyed by position.

Skills

  • SKILL.md format from agentskills.io.
  • Only the index sits in the prompt; the body loads on activation (by the planner or via load_skill).
  • read_skill_file reads bundled files.

Tool consent

  • Modes: autopilot, ask-writes, ask-all, read-only.
  • Glob rules, "remember", timeout.
  • setToolApprovalMode switches the mode live.

Large catalogues

  • MCP tool lists are read to the end (pagination).
  • Per-server connect timeout.
  • The default toolSelectionStrategy: 'auto' switches to find_tools search above toolSearchThreshold.

Prompt caching

  • System prompts are run-stable and carry an Anthropic breakpoint.
  • OpenAI calls carry a promptCacheKey.
  • A rolling breakpoint on the newest message of every tool-loop round.
  • Tools keep declaration order. Sorting them was measured with a local Qwen2.5-3B: the order swings a 3B either way, and it buys no cache stability.

Subagents

  • createSubagentTool runs in-process or in a Web Worker (serveSubagentWorker).
  • Host tools are proxied through the parent's consent gate.

Attachments

  • Images, PDFs, files and URLs as run attachments (file parts).
  • Capability detection; AttachmentsNotSupportedError gives an actionable message instead of a raw provider error.

Virtual file system

  • VirtualFileSystem (IndexedDB) plus createFileTools.

Autonomy

  • The prompts tell the agent to act and state its assumptions rather than ask; permission is the system's job.

Merge order

  1. dudko-dev/agent — merged and released as 0.0.31.
  2. This PR, released as 0.0.20.
  3. dudko-dev/agent-web-react — next, once 0.0.20 is on npm.

Checks

  • npm run typecheck, npm run format:check, npm run build and npm test all pass (144 tests). CI is green.
  • Ran the Node gate's real-model scenario ("what is the secret code?") against this package on a local Qwen2.5-3B: the tool is driven and the real value reported.
  • Smoke-tested against real Gemini through the React demo.

🤖 Generated with Claude Code

https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5

claude added 6 commits October 2, 2026 22:30
…search, subagents

Agent loop:
- thinking / stageThinking: portable reasoning level or exact budget,
  streamed thoughts (step/final.reasoning-delta), reasoning tokens in usage
- limits: run caps for input / output / reasoning / total tokens (checked
  between steps and at every tool round) + per-call output caps;
  maxToolCalls, maxPlanSteps; budget.exceeded event, the run still answers
- context: step results and tool findings now reach later steps and the
  synthesizer; auto + manual compaction (history, run trace) and a
  model-facing tool-output cap; agent.compact()
- prompt caching: run-stable system prompts, Anthropic cache breakpoints,
  OpenAI promptCacheKey; cached tokens in usage
- skills: SKILL.md parsing/loading, planner selection, load_skill /
  read_skill_file
- tool consent: autopilot / ask-writes / ask-all / read-only, glob rules,
  "always allow", live setToolApprovalMode(); gate inside execute so native,
  prompted, MCP and subagent host-tool calls are all covered
- large catalogues: 'search' / 'auto' tool selection with find_tools,
  condensed planner catalogue, schema-derived hints for prompted catalogues
- subagents: createSubagentTool (in-process or Web Worker) +
  serveSubagentWorker, proxied host tools under the parent's gate
- autonomy: prompts never ask the user; data questions are planned as tool work
- native tool loop counts the usage of every step (was the last step only)

MCP connector: paginated tools/list, per-server connectTimeoutMs,
readOnlyHint → read-only tools.

Deps: ai 7.0.127, @modelcontextprotocol/sdk 1.32.0, @ai-sdk/* and
@browser-ai/* latest, typescript 6 (ignoreDeprecations for tsup's dts).

Docs: docs/capabilities.md, README, design.md, tasks.md, CLAUDE.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
…le system

- run(goal, { images }) sends images to the planner, executor and
  synthesizer as image parts; agent.capabilities.images / supportsImages()
  / `vision` decide whether a model can take them
- a text-only or prompted-mode model ends an image run at once (no tokens)
  with ImagesNotSupportedError's message; a provider refusal of image input
  is turned into the same clear error instead of a raw API error
- memory keeps a text note of attached images, never their bytes
- VirtualFileSystem: an IndexedDB (or in-memory) workspace with namespaces,
  data-URL helpers and change events; IndexedDB schema v2 adds a `files`
  store (v1 databases upgrade in place)
- createFileTools: fs_list / fs_read (read-only) and fs_write / fs_delete

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
run(goal, { files }) alongside images: PDFs and other files go to the model
as file parts, http(s) URLs as links (providers that accept links fetch them,
the SDK downloads for the rest). Images use v7 file parts too (the image part
is deprecated). agent.capabilities reports images / pdf / files; `inputs`
overrides per kind. A kind the model is known not to take ends the run before
any call with AttachmentsNotSupportedError's message (PDFs: convert to text
or switch model); a provider refusal is mapped to the same message.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
…ng, sorted tools

- Prompt caching now also moves an Anthropic breakpoint to the newest message
  of every tool-loop round (stale message breakpoints removed, so a request
  carries at most two), so each round reads the earlier rounds from cache.
- Context editing inside a step's tool loop: past
  compaction.clearToolResultsAfterTokens the oldest tool results become
  one-line stubs (sticky), keeping the newest keepToolResults; reported as
  context.compacted with scope 'tool-results'.
- Tools are sorted by name, so the cached prefix is identical regardless of
  MCP connection order.
- docs/capabilities.md: a "Token efficiency" section mapping the practices of
  Claude Code / Copilot to the agent's knobs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
Some servers repeat tool call ids across rounds; keying the sticky set by
the result's position (the loop only appends) clears exactly the oldest.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
Sorting tools by name bought no cache stability (hosts merge their tool sets
in a fixed order, useMcpServers in list order) and reorders what a server
declared. Measured with a local Qwen2.5-3B (the Node sibling's release gate):
tool order swings a 3B either way, so the agent no longer imposes one.
sortTools is removed (it was never released).

Version 0.0.20.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
@siarheidudko siarheidudko changed the title feat: thinking, limits, compaction, skills, tool consent, tool search, subagents, attachments, VFS, token efficiency feat: thinking, limits, compaction, skills, tool consent, tool search, subagents, attachments, VFS, token efficiency (0.0.20) Oct 3, 2026
@siarheidudko
siarheidudko marked this pull request as ready for review October 3, 2026 00:39
@siarheidudko
siarheidudko merged commit bcd6ec4 into main Oct 3, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants