Skip to content

feat: thinking, token limits, compaction, skills, tool consent, tool search, subagents, token efficiency (0.0.31) - #50

Merged
siarheidudko merged 2 commits into
mainfrom
claude/agents-review-improvements-2m2406
Oct 3, 2026
Merged

siarheidudko merged 2 commits into
mainfrom
claude/agents-review-improvements-2m2406

Conversation

@siarheidudko

@siarheidudko siarheidudko commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

This PR implements the agent spec it shares with @dudko.dev/agent-web (browser): the same config fields, events and semantics. Full documentation is in the README. Releases as 0.0.31.

What's in it

Thinking

  • thinking / stageThinking map to the AI SDK's portable reasoning level, plus provider budgets.
  • Thoughts stream as reasoning-delta events.

Limits

  • Token caps per run, by kind: input / output / reasoning / total.
  • Per-call output caps per stage.
  • maxToolCalls and maxPlanSteps.
  • When a cap is crossed, the run stops at the next boundary, emits budget.exceeded, and still produces an answer.

Compaction

  • Automatic: the history at run start, the trace before each call.
  • Manual: agent.compact(), and /compact in the CLI.
  • Every tool result the model sees is capped.
  • Context editing: inside a step's tool loop, the oldest tool results become stubs past clearToolResultsAfterTokens. Clearing is sticky and keyed by position.

Skills

  • SKILL.md format from agentskills.io.
  • Only the index sits in the prompt; the body loads on activation.
  • Built-in load_skill / read_skill_file.

Tool consent

  • Modes: autopilot, ask-writes, ask-all, read-only.
  • Glob rules, "remember", timeout.
  • Switchable live; the CLI has --approval and /autopilot.

Large MCP catalogues

  • tools/list is read to the end (pagination).
  • Per-server connectTimeoutMs.
  • The 'auto' strategy switches to find_tools search above the threshold.

Prompt caching

  • System prompts are run-stable and carry an Anthropic breakpoint.
  • OpenAI calls carry a promptCacheKey.
  • A rolling breakpoint on the newest message of every tool-loop round.
  • Tools keep declaration order, which is already deterministic.

Subagents

  • createSubagentTool runs in worker_threads (dist/subagent-worker.js) or in-process.
  • Host tools are proxied through the parent's consent gate.

Autonomy

  • Planner, replanner and synthesizer rules: act, state assumptions, never ask the user.
  • The executor prompt stays as measured; a blocked step goes [BLOCKER] → replanner.

Dependencies

  • ai 7.0.127, @modelcontextprotocol/sdk 1.32, and the @ai-sdk/* providers.
  • typescript and @types/node stay pinned, per the org ignore list.

CI fixes in the second commit

  • Node 24: subagent workers no longer receive an explicit process.execArgv. Node 24 rejects the per-process flags node --test adds; workers inherit the options themselves. Reproduced on node:24 in Docker and fixed: 256/256.
  • Real model smoke was a real regression, caused by the rewritten planner/executor wording and by sorting tools.
    • Reproduced locally with the pinned llama.cpp image and the same Qwen2.5-3B: main passes, the branch failed.
    • Fix: the measured wording is restored, the autonomy rule is kept in the planner, and tools keep declaration order.
    • Now 2/2 on a fresh server, the same way CI runs it.

Checks

  • npm run typecheck, npm run format:check, npm run build and npm test all pass on Node 22 and Node 24 (256 passed, live test skipped locally).
  • "Real model smoke" is green.

🤖 Generated with Claude Code

https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5

claude added 2 commits October 2, 2026 23:51
…search, subagents, token efficiency

Shared spec with @dudko.dev/agent-web (same config fields, events and
semantics):

- Thinking: `thinking` / `stageThinking` → the AI SDK's portable `reasoning`
  plus provider budgets; thoughts stream as step/final reasoning-delta events.
- Limits: run-level token caps (input / output / reasoning / total), per-call
  output caps per stage, `maxToolCalls`, `maxPlanSteps`; a crossed cap stops
  at the next boundary, emits budget.exceeded and still answers.
- Compaction: auto (history at run start, trace before each call) and manual
  `agent.compact()`; per-result tool-output caps; stale tool results inside a
  step's loop become stubs past `clearToolResultsAfterTokens` (sticky,
  position-keyed).
- Skills (agentskills.io SKILL.md): index in the prompt, body on activation
  by the planner or `load_skill`; `read_skill_file`.
- Tool consent: autopilot / ask-writes / ask-all / read-only, glob rules,
  remember, timeout; switchable live (setToolApprovalMode).
- Large catalogues: paginated tools/list, per-server connect timeout,
  'auto' tool strategy → `find_tools` search above the threshold.
- Prompt caching: run-stable system prompts with an Anthropic breakpoint,
  OpenAI promptCacheKey, a rolling breakpoint on the newest loop message,
  tools sorted by name.
- Subagents: createSubagentTool in worker_threads or in-process, host tools
  proxied through the parent's consent gate.
- Autonomy rules in the prompts: act, don't ask.
- CLI flags/config for all of it; README, CLAUDE.md, env.example; deps bumped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
- Subagent workers inherit the parent's node options by default; an explicit
  execArgv is passed only to strip entry-point flags (-e / -p /
  --input-type). Passing process.execArgv explicitly failed on Node 24 under
  `node --test` ("Initiated Worker with invalid execArgv flags").
- The real-model gate (Qwen2.5-3B) regressed: the rewritten planner rule put
  the canned "Answer the user directly" next to the lookup rule, extra
  executor rules made the 3B echo a guessed answer alongside the real lookup,
  and sorting tools by name reordered what the server declared. Reproduced
  locally with the pinned llama.cpp image and weights, fresh server, two
  attempts (as CI): planner rules restored with the autonomy rule added as
  rule 6, executor prompt as measured, tools in declaration order (servers
  already mount in declaration order, so the cached prefix stays stable).
  Now 2/2 on a fresh server.
- Version 0.0.31.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
@siarheidudko siarheidudko changed the title feat: thinking, token limits, compaction, skills, tool consent, tool search, subagents, token efficiency feat: thinking, token limits, compaction, skills, tool consent, tool search, subagents, token efficiency (0.0.31) Oct 3, 2026
@siarheidudko
siarheidudko marked this pull request as ready for review October 3, 2026 00:26
@siarheidudko
siarheidudko merged commit 2bbba82 into main Oct 3, 2026
9 checks passed
@siarheidudko
siarheidudko deleted the claude/agents-review-improvements-2m2406 branch October 3, 2026 00:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants