Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 18 additions & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,14 +7,28 @@
## Layout

- `src/` — library + CLI source (`.ts`, imported with explicit `.ts` extensions)
- `agent.ts` (`createAgent`, `IAgent`), `runner.ts` (the loop), `planner.ts` /
`executor.ts` / `replanner.ts` / `synthesizer.ts`, `prompts.ts` (system =
run-stable, user = dynamic), `call-options.ts` (per-stage SDK options).
- Feature modules: `thinking.ts`, `limits.ts`, `compaction.ts`, `skills.ts`,
`approval.ts`, `tool-wrap.ts` (approval gate + tool-call budget + output
cap, wrapped around every tool's `execute`), `tool-search.ts`,
`caching.ts` (system breakpoint + rolling message breakpoint),
`context-editing.ts` (stale tool results → stubs inside a step's loop),
`subagent.ts` + `subagent-worker.ts` (worker entry, built to
`dist/subagent-worker.js`).
- `index.ts` is exports only.
- `tests/` — `node:test` suites, run via `node --experimental-strip-types`
- `dist/` — `tsup` build output (do not edit)

## Test layers

- `tests/*.test.ts` — units plus `integration.test.ts`, which runs the whole
loop over real sockets against a scripted OpenAI-compatible endpoint, a real
MCP server and a real authorization server (`tests/helpers/`). No key, ~5s.
MCP server and a real authorization server (`tests/helpers/`), and
`features-integration.test.ts`, which drives thinking, limits, compaction,
skills, approval, tool search and subagents (worker + in-process) the same
way. No key, ~7s.
- `tests/live-model.test.ts` — the same loop against a REAL model. Skipped
unless `AGENT_LIVE_MODEL_URL` points at an OpenAI-compatible endpoint; CI
starts one via `live-model.yml`. Asserts mechanics only (the loop finished, a
Expand All @@ -41,6 +55,9 @@ If `npm install` is needed (e.g. lockfile changed), run it with `--no-audit --no
- Use `.ts` extensions in relative imports (project relies on `--experimental-strip-types`).
- Zod v4 is used; when passing heterogeneous schemas through a shared array/iterable, type the collection as `z.ZodType` to avoid union-narrowing errors.
- Anthropic's native structured output rejects `maxItems` on arrays — never add `.max()` to Zod arrays that flow into structured output. The guard test in `tests/anthropic-schema-compat.test.ts` enforces this.
- Keep system prompts run-stable (prompt caching): anything that changes per call (history, request, plan, trace) goes in the user prompt. Tests match stages on the role phrases "You are the Planner / Executor / Replanner / Synthesizer" — keep them.
- Tools reach the run through `runContext` (AsyncLocalStorage): emit, current step and the per-run state. Gate logic belongs inside `execute` (see `tool-wrap.ts`), not in the executor's stream loop.
- The public API, config fields, event names and semantics are shared with the browser sibling `@dudko.dev/agent-web`; keep names aligned when changing them.

## Boundaries

Expand Down
283 changes: 274 additions & 9 deletions README.md

Large diffs are not rendered by default.

54 changes: 46 additions & 8 deletions env.example
Original file line number Diff line number Diff line change
Expand Up @@ -105,19 +105,57 @@ AGENT_MAX_STEPS_PER_TASK=8
# Prevents the LLM from looping in "revise -> execute -> revise". Default 2.
# AGENT_MAX_REVISIONS=2

# Hard cap on total tokens per run (input + output). When crossed the agent
# exits the execution loop and proceeds to synthesize the final answer.
# AGENT_MAX_TOTAL_TOKENS=200000
# Per-run token caps. When one is crossed (checked between steps AND at every
# LLM step inside an executor step) the agent stops executing steps and
# synthesizes the final answer from what it has.
# AGENT_MAX_TOTAL_TOKENS=200000 # input + output
# AGENT_MAX_INPUT_TOKENS=150000
# AGENT_MAX_OUTPUT_TOKENS=30000 # includes reasoning
# AGENT_MAX_REASONING_TOKENS=20000

# Cap on tool calls per run (across steps). Further calls fail with
# "tool-call budget exhausted" and the run goes to the final answer.
# AGENT_MAX_TOOL_CALLS=40

# Thinking / reasoning for every stage:
# off explicitly disabled
# minimal|low|medium|high|xhigh portable effort level
# provider-default the provider's default level
# <number> an exact budget in tokens (Anthropic / Google)
# Unset: nothing is sent (the provider's default behaviour).
# AGENT_THINKING=medium

# Tool approval (consent) mode:
# autopilot (default) every tool call runs
# ask-writes read-only tools run, others ask at the REPL prompt
# ask-all every call asks
# read-only read-only tools run, others are denied
# Answer y / n / a(lways) at the prompt; /autopilot toggles at runtime and
# /approval <mode> switches modes.
# AGENT_TOOL_APPROVAL=ask-writes

# Folder of agentskills.io-style skills: <dir>/<skill>/SKILL.md (+ bundled
# text files). The planner picks the ones that apply; the executor can load
# any with load_skill.
# AGENT_SKILLS_DIR=./skills

# Compaction: history / trace summaries kick in at 50% of the context window.
# AGENT_COMPACTION=off disables the automatic part (/compact still works).
# AGENT_CONTEXT_WINDOW_TOKENS=128000
# AGENT_COMPACTION=on

# Tool selection strategy for the executor:
# - all (default) every step sees the full filtered ToolSet.
# Works well with <=50 tools.
# - auto (default) 'all' up to 40 tools, 'search' above.
# - all every step sees the full filtered ToolSet.
# Works well with <=40 tools.
# - plan-narrowed executor gets only the tools listed in
# step.suggestedTools. The planner MUST populate
# suggestedTools for every step that needs tools.
# Cuts context tokens significantly when the catalog
# has 100+ tools.
# AGENT_TOOL_SELECTION_STRATEGY=all
# - search each step starts with the suggested tools and the
# ones found earlier, plus find_tools, which searches
# the whole catalogue and activates matches. For
# catalogues of hundreds of tools.
# AGENT_TOOL_SELECTION_STRATEGY=auto

# Fine-grained MCP tool control. Names use the `serverName__toolName` format.
# availableTools wins over excludedTools.
Expand Down
Loading
Loading