agent6 treats the LLM as untrusted. Concrete claims below, layer by layer, each with what it means and where it stops.
Open a GitHub issue prefixed [security]. Include: agent6 version
(agent6 --version), kernel + distro (uname -a, /etc/os-release),
agent6 check sandbox output, and a minimal repro (ideally a failing test under
tests/security/).
Adversary: a fully malicious worker model, or an honest model that has been prompt-injected by a file in the workspace.
We assume the adversary controls:
- the text of every LLM response
- the choice of tool calls and their arguments (within the published JSON schema)
- the content of any file the agent reads during the run
We do NOT assume the adversary controls:
- the kernel
- the agent6 binary
- the provider endpoints
Under that adversary, agent6 aims to hold:
-
No writes outside the workspace.
-
No reads outside the workspace and a read-only system set.
- The system set (
/usr /bin /sbin /lib /lib64 /etc /dev /proc /tmp) exists so installed toolchains resolve. hardenedalso exposes$HOME+/run(Landlock can't carve them out);sandbox.extra_read_pathsadds more.
- The system set (
-
The agent process's own egress is NOT bounded. agent6 talks to the configured providers; nothing stops the PROCESS reaching elsewhere.
- There was once an empty netns plus a broker proxying provider calls. It was
deleted: under
strictthe agent process has no filesystem confinement, so code execution there could write~/.ssh/authorized_keysor a cron entry and exfiltrate on its own schedule. Blocking the socket while leaving that open is a partial mitigation that reads as a guarantee, and it cost four special cases plus a finding of its own. - What IS bounded is what a COMMAND reaches:
sandbox.tool_network(defaultauto; §8). That boundary is per-jailed-child and unchanged.
- There was once an empty netns plus a broker proxying provider calls. It was
deleted: under
-
agent6's own git never pushes,
--forces, rewrites history, orreset --hards (§5).- This does NOT bind a
gitthe model runs viarun_command; that path is bounded by the sandbox (protect_gitkeeps.gitunwritable understrict; push needs egress).
- This does NOT bind a
-
No persistence after the run: no daemon, cron,
.bashrcwrite, or setuid binary.- No setuid/setgid bit can be set.
chmod/fchmod/fchmodat/fchmodat2are denied by seccomp when the mode carriesS_ISUID/S_ISGID(ordinary chmod is untouched).fchmodat2is named separately because it SUPERSEDEDfchmodaton Linux 6.6+: a filter listing only the older spelling is open on any current kernel, andchmodnever reaches it, so only a direct syscall test finds that. The bit would land on the HOST inode and outlive the jail, and undersudo agent6 --allow-rootthe uid_map makes the jailed child real root -- a setuid-root binary left in the workspace is local root for anyone who runs it. Mountnosuiddoes not cover this: it stops the JAIL honouring the bit, not the host. Every mount carriesnosuidandnodevanyway -- the workspace, the system binds, the protect and read-only binds, the writable grants, and the private/tmp. Notnoexecon/tmp: HOME lives there and toolchains run helpers from it, and a child that can already execute from the workspace gains nothing from being stopped there. - Children can only write inside the jail's mount namespace (strict) or the Landlock write grants (hardened).
- Nothing a command starts outlives it. strict's PID namespace takes the
whole tree down. hardened has none, so the agent makes itself a child
subreaper: a
setsiddaemon that escapes the launcher's process group reparents to the agent, which kills it when the command returns. - The bound is the command's own descendants. On hardened a command can
still hand work to a user daemon that was ALREADY running (a tmux server,
systemd --user):AF_UNIXconnect has no Landlock hook, and without a mount namespace those sockets are nameable. strict does not expose them.
- No setuid/setgid bit can be set.
Applied at run/resume start, before any network object. Restricts the Python
process irrevocably, inherited by every child:
| Landlock rule | Allowed |
|---|---|
| FS read+exec | cwd, $HOME, /usr, /etc, /tmp, /bin /sbin /lib /lib64 /dev, /run + /proc when present |
| FS write | cwd, /tmp, the /dev char devices, /proc when present |
strictskips this layer.- Its per-command namespaces are stronger, and this would
break the jail's
pivot_root/mountat Landlock ABI ≥ 7.
- Its per-command namespaces are stronger, and this would
break the jail's
- The read+exec set mirrors the jail child's roots.
- The launcher opens each from here to grant the child, so a missing one
(e.g.
/devon merged-/usr) makes the child's execve fail EACCES.
- The launcher opens each from here to grant the child, so a missing one
(e.g.
- Neither level bounds the AGENT's egress. Both say so rather than pretending.
- Landlock filters connects by PORT, not host, so the only rule available here was "any host on the provider ports". That stops nothing worth stopping -- an exfiltrating agent needs one HTTPS endpoint, and every host offers one -- while breaking a legitimate tool on another port. It was removed: a security claim nobody can rely on is worse than no claim.
- Bind/listen is not denied either. Blocking inbound while outbound stays open is the same non-claim, and it cost a dev server the model runs its listening socket. Nothing it binds outlives the command (§5).
apply_edit is in-process; every run_verify_command/run_command runs in
agent6-jail, as does every run_background. Under strict a RUN's commands
are served by ONE launcher process, so they share its netns, PID namespace and
private /tmp: a server run_background starts is reachable by the next
command on loopback, and closing the run's request channel takes the PID
namespace and everything in it down. Its confinement is fixed when it opens, so
the policy is the run's rather than the first command's, and it grants the
background log root (<session>/shells/logs) up front -- a run's background
commands can therefore write each other's logs, but not their own exit code or
name, which live outside that root. Anywhere else (hardened, a dispatcher with
no run) each command gets its own launcher. Under strict it:
- Forks a new user/mount/PID/IPC/UTS/net namespace.
pivot_roots into a minimal bind-mount rootfs on a fresh tmpfs: cwd + private/tmpwritable, system paths read-only,extra_read_pathsgrants read+exec at their real paths (extra_rwgrants writable at theirs). Operator-tool dirs join as read+exectool_pathsmounts (standard bin dirs that exist, the real dirs their symlinks resolve to, uv-managed CPythons), derived bysandbox.jail.operator_tool_paths-- which never mounts agent6's OWN config, state, data or cache dirs however a tool symlink resolves into them, sosecrets.tomland the run history stay outside the jail by construction rather than by directory layout. A tool dir whose read-only remount fails is detached, and a failed detach refuses the run: best-effort means unreachable, never writable. run_command/verify jails and machine tool jails share that one computation, andmachine checkprobes the same PATH.- Exposes curated
/dev(null zero urandom random full); omits/dev/tty(it would let a child write escape sequences to the parent's terminal). - Mounts a fresh private
/proc; if that fails, leaves/procempty (never the host's, which would leak process info).- The launcher is PID 1 of that namespace, so
/proc/1/environis readable by the jailed command. It is spawned with an EMPTY environment for that reason: inheriting agent6's put a provider key supplied viaapi_key_envone file read away from the model. The launcher needs none -- its policy arrives on stdin and the child's env is set explicitly in it.
- The launcher is PID 1 of that namespace, so
- Applies Landlock FS rules (net confinement is the namespace); best-effort: a kernel without Landlock skips this layer, warned loudly at run entry.
- Installs a seccomp deny-list: dangerous syscalls (ptrace, mount, setns,
unshare, kexec, bpf, perf, keyctl, module loading, reboot, clock-set, …)
return
EPERM, the rest allowed. - Sets
NO_NEW_PRIVS, so the kernel ignores setuid bits (sudo/setuid can't escalate). execves the binary and SIGKILLs the group at the wall-clock timeout.
Notes:
- The memory cap is operational, not a threat-model control. A per-process
RLIMIT_DATA([sandbox].memory_limit_mb, default 4096,0off; notRLIMIT_AS, so V8/JVM/ASAN keep working) stops one runaway allocation, nothing more. - The seccomp layer is a deny-list (defense-in-depth), not a boundary. It
enumerates known-dangerous syscalls, so by construction it does not catch
everything: namespace creation via
clone/clone3(CLONE_NEWUSER|CLONE_NEWNS) is not blocked, onlyunshareis. Accepted, because a nested namespace grants no new access against the host — the whole mount family (mount,mount_setattr,open_tree,move_mount,fsopen/fsconfig/fsmount/fspick,pivot_root,umount2) is denied, Landlock is inherited and irrevocable, andmknodchecks caps against the initial user namespace. Filteringclone3is declined on purpose: seccomp cannot read its flags (they sit behind a struct pointer), and denying it outright would break glibc/Go process spawning (they fall back tocloneonly onENOSYS). The real boundaries are the namespaces, Landlock, and the mount-syscall denials — not this list. - No
capset.strictmaps namespaced-root to your uid;hardenedkeeps the caller's caps (none for a normal user). hardeneddrops the namespaces + rootfs; Landlock, seccomp,NO_NEW_PRIVS, and the timeout remain.- The policy arrives as JSON on stdin from
run_in_jail; the Rust side validates it against a strict schema and refuses unknown fields.
The jail is one-way: the agent works within the environment you give it and can't expand it.
sudocan't escalate, even passwordless.NO_NEW_PRIVSvoids setuid, so jailedsudofails regardless of anyNOPASSWDrule.- Package installs are impossible.
apt/dnf/apkneed all three of root (blocked), mirror network (provider-only egress), and/usr//varwrites (denied). - Compiling and running host-installed toolchains works.
- Every command tool --
run_command,run_verify_command,run_background,stop_background-- answers torun_commandsand runs jailed. They just can't install new tools, and a networked build step needstool_networkloosened. - The verify gate is a command like any other:
run_commands = "no"withholds it too, and such a run starts gateless rather than chasing a green it can never reach.
- Every command tool --
- Provisioning is operator-first. Install toolchains, venvs, and deps
yourself before/outside agent6; widen access via config, never sudo
(
extra_read_paths,tool_network,[providers.*].base_url, all inconfig show). - Running agent6 as root (
--allow-root/AGENT6_ALLOW_ROOT=1) weakens the boundary.strictmaps inside-root to real root, so jailed children run as real root under only Landlock + seccomp +NO_NEW_PRIVS.- Still no writes outside the workspace and no egress beyond providers, but
the allowed reads now include root-only files (
/etc/shadowunderhardened;strict's rootfs hides them). Run as your normal user.
Everything the model can influence runs through run_in_jail (§2). A fixed set
of modules also shells out directly with subprocess.run/Popen; each has
fixed argv depending only on operator input, never LLM output.
tests/security/test_subprocess_allowlist.py pins the file list; audit with
rg 'subprocess\.|os\.(system|exec|posix_spawn)' src/agent6/.
-
git_ops.py: agent6's own git operations (§5). -
sandbox/detect.py: probes the host's sandboxing capabilities. -
sandbox/exec_confined.py: confines itself, thenexecvps the argv after--. For a long-lived child agent6 spawns but does not drive (a configured MCP server): the jail launcher captures stdio and owns the process to completion, which cannot host a live MCP pipe. Both mechanisms are restrict-self-then-exec and inherited, so the server and everything it spawns get them: Landlock's domain is irrevocable acrossexecve, and[mcp.servers.<name>.sandbox].network = "none"additionallyunshares a user + network namespace, which the stdio pipes (file descriptors) survive untouched. Either mechanism stands alone: a block naming no paths applies no Landlock domain (an empty one would grant nothing, to the shim and to whatever it becomes), and a shim asked for neither refuses rather than exec unconfined while looking confined. The namespace work runs BEFORE Landlock, because it writes/proc/self/{setgroups,uid_map,gid_map}(mapping the operator's uid straight through, so the server does not becomenobody) and opens a socket to bringloup -- none of which the domain grants. A host whose kernel forbids unprivileged user namespaces makes the server REFUSE rather than start connected. The paths and the argv are the operator's, from config; no LLM input reaches it, andpreexec_fn-- the obvious alternative -- is unsafe in a threaded process, which the MCP client is. -
sandbox/jail.py: the jail launcher itself. -
tools/lsp.py: thetylanguage server, exe resolved from PATH. -
tools/mcp_client.py: operator-configured[mcp.servers.*]server commands. -
providers/token_command.py: the operator-configured[providers.*].token_commandthat mints a provider bearer; argv from config. -
sessions/ipc.py:ps -p <pid> -o lstart=on hosts without /proc (macOS) for the worker.pid start-time identity; fixed argv over a pid agent6 itself recorded. -
ui/cli/_btw.py: spawnsagent6 askdetached for/btw, so the side question keeps provider egress when the run itself is confined. argv is the agent6 exe plus the question the OPERATOR typed at the pause menu -- never LLM output -- with--before it so a question starting with a dash cannot read as a flag. -
ui/spawn.py: the shared front-end spawn helper; spawns the agent6 CLI detached for run/machine launches and capturessessions merge/prune/config set; argv is the agent6 exe plus operator-chosen args. -
ui/notify.py: firesnotify-sendwith fixed argv (exe,--end-of-options, two positional data args, no shell) for the device-present machine notification; the message is inert data, never a command or an option. -
ui/cli/helpers:$EDITORfor plan, notes and steer editing.git diff/logfor the review subcommand and thesessions/askdiff views; argv from the run manifest the CLI wrote outside the jail.rgfor history search.- The fixed-argv
python -m agent6.ui.tuico-process behindrun --tui. ui/cli/system_cmds.py:cp/rm/apparmor_parservia sudo with fixed argv foragent6 system apparmor(operator host setup).
-
app/helpers:app/finalize.py: the operator[notify].on_completehook fired at run end; argv from config, env fromhook_env(a minimal base plusAGENT6_SESSION_*, never the provider keys in the operator environment).app/machine/_scriptcheck.py: ruff/ty with fixed argv to statically read generated scripts, which only ever execute viarun_in_jail.- The
machine runsupervisor (app/machine_agent.py): spawns each agent state as a fixed-argvpython -m agent6.ui.cli.machine_agentsubprocess whose request travels in a temp file, never on argv; its operator[machine.notify].on_eventhook (argv from config, fired fromapp/machine/_preflight.py) runs on the host with the same minimalhook_envbase plusAGENT6_MACHINE_*, mirroring[notify].on_complete. ui/cli/skills_cmds.py:git clone --depth 1 -- <url>with fixed argv foragent6 skills install; the URL is operator-supplied on the CLI and nothing fetched is ever executed.
-
ui/tui/clipboard.py: fixed-argvtmux set-buffer -wwith the copied transcript text as one inert data argument. -
ui/tui/conversation.py: the operator's$PAGER, argv from the environment, transcript text on stdin.
You set sandbox.isolation; it resolves against the host to the effective
isolation level. No silent downgrade: a request the host can't meet is refused, and
auto reaches none only when the host offers no confinement mechanism at
all (non-Linux, or a Linux kernel with neither userns nor Landlock) -- always
loudly. Capabilities are probed (unshare for userns, the Landlock ABI
syscall for hardened), never guessed from the kernel version.
sandbox.isolation |
Host | Effective |
|---|---|---|
auto (default) |
Linux + user namespaces | strict |
auto |
Linux, no userns, Landlock | hardened |
auto |
Linux, no userns, no Landlock | none (loud warning) |
auto |
non-Linux | none |
strict |
Linux + user namespaces | strict |
strict |
else | ⛔ refuse |
hardened |
Linux + Landlock | hardened |
hardened |
else | ⛔ refuse (Landlock is hardened's only FS boundary) |
none (opt-out) |
any | none (the environment is the boundary) |
-
strict: full namespaces +
pivot_root+ Landlock + seccomp +NO_NEW_PRIVS.- On a kernel without Landlock the jail's in-rootfs Landlock layer is skipped (namespaces, read-only binds, and seccomp still confine) with a loud once-per-run warning; hardened has no mounts to fall back on, so it refuses instead.
-
hardened: Landlock + seccomp +
NO_NEW_PRIVS, no namespaces.- Works in default-seccomp Docker (the container blocks the inner
clone(CLONE_NEW*)); the container is the blast radius. - With no PID namespace, teardown is the agent's job (§5): it holds
PR_SET_CHILD_SUBREAPER, and each command kills every process that appeared during it from outside the agent's session. A survivor the sweep cannot kill fails the command rather than passing silently.
- Works in default-seccomp Docker (the container blocks the inner
-
none: unsandboxed, always with a loud warning.
-
Unsandboxing is explicit and self-authorizing.
isolation = "none",--dangerously-disable-sandbox, orAGENT6_DANGEROUSLY_DISABLE_SANDBOX=1. The LLM can't reach argv/env, so setting one is the consent. -
Sandbox-off + auto-approved
run_commandadds a one-time gate. For that combination only:Continue? [y/N]interactively, a warning in CI/machine run. -
CI should set
strictto fail loud if the sandbox is weaker than expected.
fetchis the model's only egress, and it is narrow. One https URL, GET, no redirects followed, no credential, text only, 1 MiB. Hosts onsandbox.fetch_hostsare read without asking; any other host prompts, and an absent operator is a no. It exists because a jailed command has no network, so it is hidden entirely whentool_network = "allow". A GET can still carry data out in its path -- the allow-list is empty by default for that reason.- The LLM only sees the fixed set in
src/agent6/tools/schema.py.- Structured edits, read-only navigation, fixed-argv verify/metric commands,
finish_session,ask_user, a curator task notepad, a cross-run memory notepad, and capability-gatedrun_command. - No
shell, nowrite_file(writes go throughapply_edit, which refuses paths outside cwd), noweb_fetch, noeval. - Adding a tool needs a security review note (AGENTS.md).
- Structured edits, read-only navigation, fixed-argv verify/metric commands,
- The memory notepad and notes scratchpad are prompt-injection persistence
channels.
add_memory/invalidate_memory(run mode) write fixed markdown under<state-dir>/<repo-id>/memories/(code picks the path; the model supplies only a schema-validated scope + text); active notes join later runs' system prompt on the same repo.write_notesreplaces one fixed file,<state-dir>/<repo-id>/notes.md, on the same terms: code owns the path, the model supplies only a length-bounded string.- Mitigated: both are inert data (never executed), the injected blocks are
size-capped and framed as untrusted, and both stores are operator-auditable
(
agent6 memory list --all,agent6 memory invalidatekeeps the trail; notes are one readable markdown file). Neither is ever mounted into the jail, so a jailed command cannot reach or rewrite them. - Notes are whole-file replace, so a hostile write can erase earlier notes. That is the same authority the agent already has over its own context and costs nothing outside it; memories, which carry the audit trail, stay append-only precisely because they must survive this.
- Neither weakens a boundary here: sandbox/egress/git policy come from config, not prompt content.
- agent6's own git refuses the destructive ops, by construction.
git_ops.pyis the only module through which agent6 invokes git; it wraps the safe ops (status, add, commit, diff, branch, checkout) and refusespush,reset --hard,commit --amend,rebase,filter-branch/filter-repo,branch -D/--force, and any--force/-fon a destructive verb.
- A
gitthe model runs viarun_commandis bounded by the sandbox, not this list, and its argv is NOT screened.protect_git(default on) keeps.gitunwritable understrict, which re-binds it read-only, recursively: a mount nested under it (e.g..git/objectson its own bind) stays visible and read-only rather than shadowed. A rewrite fails andpushhas no egress. It is STRICT-ONLY: see below. Onhardenedthe default degrades with a warning and an explicitly-settruerefuses to run.- The protected scope is the project's own
.git: agent6's operational state, the repository it commits to each turn. A nested.git(a vendored repo's, a submodule's) is content, like any other file in the workspace -- tracked by the root repo or untracked, with no guarantee offered either way. Naming it would close nothing: a planted git config is one host-execution vector among many in repo content (.envrc, aMakefile,conftest.py). - agent6 used to refuse mutating git subcommands (plus the
-c alias.*injection that dodged them) inrun_commandargv. Removed: a blocklist enumerates badness, and a model that writes a shell script and runs it walks past it, so it bought complexity and a false sense of a boundary the jail already owns.
- git_ops neutralizes repo-controlled host code in a poisoned
.git/config.core.fsmonitoranddiff.externalare always off;.git/hooks/*run only undergit.run_repo_hooks = true(default false;core.hooksPathpoints away so a hook can't fire on agent6's auto-commit).- Defense in depth on top of
protect_git: those settings bound what a poisoned.git/configcould do, andprotect_gitstops the model writing one in the first place. protect_gitis strict-only, and hardened leaves.gitwritable. The threat is real there: a jailed command can plant afilter.<n>.cleanplus a.gitattributes, and agent6's own auto-commit --git add -Aon the HOST, outside the jail -- then runs it, reaching$HOMEand the network. Hardened has no mount namespace, so the only tool is Landlock, and Landlock cannot express this. Two of its properties close the door together: a grant on a directory is RECURSIVE (no "this directory only"), and stacked rulesets INTERSECT (an access needs every layer to allow it). To deny.gitsome layer must not grant it, which by recursion means that layer cannot grant the workspace root either -- so the root becomes unwritable overall and nothing new can be created in it. Measured, both shapes: granting the root allows.gittoo; granting only its children denies both. A second layer cannot subtract what a first one granted. agent6 shipped the children-only carve for a while. Its cost was that no NEW top-level entry could be created --touch,mkdir,mkfifoall failed at the workspace root, and coreutils reported "File exists", sending operators looking in the wrong place. That is too much to pay for a protection the operator can have properly by usingstrict. The in-process edit tools (apply_edit/apply_patch) refuse, on both isolation levels, a write into the project's own.git, raw or symlink-resolved. That guard covers only in-process edits, not jailed commands.
- The edit tools refuse writes into an in-repo venv or
site-packages.- A
pyvenv.cfgdir orsite-packagesancestor: a run rewriting an editable-install.pthwould silently corrupt the venv, invisible inruns diff/merge since venvs are gitignored. Reads stay allowed. - Related limit: an editable install records the host path in its
.pth, absent under the jail's/workspace, so averify_commandimporting the project canModuleNotFoundError. Fix with pytestpythonpath, aconftest.py, or a non-editable install.
- A
- Provider keys are
0600, owner-only, and never leave agent6's process.- In
$XDG_CONFIG_HOME/agent6/secrets.toml(refused if group/other-readable or foreign-owned, like an SSH key), or from[providers.<name>].api_key_env(env wins). Never in transcripts, never inconfig show(redacted), never mounted into the jail.
- In
agent6 connectnever executes remote input.- It only prompts locally (
getpass) and writes config/secrets. It makes one read-onlyGETto the provider's key endpoint to confirm auth (status only;--no-verifyto skip). - During a run agent6 opens no listening socket (an MCP server is spawned on
stdio or DIALLED at an operator-set
url-- outbound either way, never a listener; the web UI is private unix socket); the only accept-side socket is opt-inagent6 web(§7).
- It only prompts locally (
- Running as root is refused without an explicit opt-in.
--allow-root/AGENT6_ALLOW_ROOT=1(+ a banner). Undersudo, agent6 reads the real user's config/secrets (fromSUDO_UID/SUDO_USER), not root's, and chowns state-dir writes back. It doesn't drop privileges in-process: the jail, not the uid, is the boundary.
- An in-process
GraphCuratorowns the task graph.- It validates every mutation against a pydantic schema before writing, and holds a per-mutation flock on the session dir. A write-path fault after the in-memory update reloads from disk before surfacing, so a later read never observes a node that was never persisted.
- The session directory is safe because of its location, not any single writer.
- Per-repo state lives at
$XDG_STATE_HOME/agent6/<repo-id>/(override with[agent6].state_dir), outside the cwd jailed commands run on.
- Per-repo state lives at
- The config write lock is a concurrency optimization, not a boundary.
- Publishes are atomic, so a torn config is impossible with or without it;
the lock only serializes read-modify-write cycles. It FAILS OPEN by
design (a planted symlink is refused
O_NOFOLLOW; a stale root-owned lock is ignored), so it is never a way to block or redirect a write. Without the lock a rollback could erase a concurrent writer's update, so the write is kept and the error says "kept as written" (docs/config.md).
- Publishes are atomic, so a torn config is impossible with or without it;
the lock only serializes read-modify-write cycles. It FAILS OPEN by
design (a planted symlink is refused
agent6 run --parallel, agent6 sessions compare, and a live run's /parallel
steer directive (§ architecture.md)
each spawn subordinate work. Nothing here loosens the sandbox:
- Every lane is an ordinary run. A lane is a plain detached
agent6 runon its own clone: its own jail persandbox.isolation, its ownrun_commandspolicy. Nothing shares a sandbox socket across lanes or with the parent run. - Recursion is blocked by an env guard, not policy. Every spawned lane
carries
AGENT6_SUBRUN=1; both the--parallelflag and the coordinator'slane_spawnerwiring refuse when it is set, so a lane can never itself fan out or dispatch (depth 1 by construction). - A lane's config carries key references, never secret values. The
orchestrator writes each lane a
--configfile viamaterialize(), a dump of the resolvedConfigmodel (providerbase_url,api_key_envnames, etc.) --Confignever holds a raw API key. The lane's own process reads the samesecrets.toml/ provider env var as any other run, same user, same host. - No new subprocess call site.
workflows/subrun.py,app/parallel.py, andui/cli/parallel.pyadd no directsubprocessuse; lane git plumbing (clone/fetch/merge) goes throughgit_ops.pyand lane spawning goes throughui/spawn.py, both already on the §2b allowlist. Thetests/security/test_subprocess_allowlist.pypin needed no new entry. - Dirty-tree refusal, not auto-stash. A lane clones committed HEAD only,
so
--parallelrefuses a dirty origin undergit.require_clean_worktree(the same policy and message shapeagent6 runuses) rather than carrying uncommitted work into a lane it cannot see.
- The loop opens no accept-side socket.
- Only outbound HTTPS to the provider; the task graph is an in-process curator, no socket.
agent6 webis the one accept-side socket, and only when you start it.- Loopback (
127.0.0.1) by default, no app auth (run behindtailscale serve; the tailnet identity is the access control, see the web UI). - A non-loopback bind is refused unless opted in:
[web].hostneeds[web].allow_non_loopback = true,--hostneeds--allow-non-loopback.
- Loopback (
- The server renders folded state and drives typed contracts only; it executes
nothing.
- New-work spawns fixed argv with the task behind
--; machine-run is allow-listed to authored files; answers write only the addressed run's answer files (session id, answer id, machine target state dir each validated to one path component); merge/prune/config-set are fixed agent6 subcommands.
- New-work spawns fixed argv with the task behind
- State-changing POSTs carry a CSRF guard.
- Body must be
Content-Type: application/json(a cross-sitefetchwith it triggers a preflight the server never answers) and anyOriginmust matchHost. Holds on loopback and behindtailscale serve. - It does NOT cover DNS rebinding (that needs a Host allow-list incompatible with the tailnet name).
- Body must be
- Request framing is bounded. 1 MiB body cap (413), chunked refused (411), any unread-body refusal closes the connection.
- The machine write surface (
POST /api/machine/<name>/{poke,answer,approve,steer}) uses the same guards.pokewrites only the instance signal file (inert JSON the nexttoolreads); the others write only the current agent state's per-state dir. PWA assets are static; the service worker is a no-op passthrough (no Web Push/VAPID). - No telemetry, no auto-update, no remote control plane.
machine runis a supervisor that makes no network calls.- Each
toolstate is jailed, so a per-toolallow_networksets its netns independently. - This lets a machine keep agents on the provider API while one reviewed,
fixed-argv
toolreaches the network: unlikerun_command(LLM-chosen argv), atoolisn't a free exfil channel.
- Each
Egress = tool_network × per-tool allow_network; the
effective isolation level decides what's enforceable. "offline" = no egress.
Agent egress is unconfined at every isolation level (claim 3). Only jailed commands have a network boundary.
Jailed-command egress (run_command, machine tool) by tool_network
(cells = strict, where a per-child netns makes "offline" real):
| jailed command | auto (def) |
block |
only_explicit_states |
allow |
|---|---|---|---|---|
run_command |
offline | offline | offline | host network |
tool, allow_network auto(def)/block |
offline | offline | offline | offline |
tool, allow_network = allow |
⛔ refuse | ⛔ refuse | host network | host network |
auto is the secure default that runs everywhere (see AGENTS.md "Secure by
default, degrade or refuse"): on strict it is offline above; on hardened
(no netns) it cannot be offline, so a jailed child inherits the agent process's
network and a once-per-run warning says so. block is
the ENFORCE form — it refuses on hardened rather than run under-confined. (On
none nothing is enforced or refused: it is the explicit unsandboxed opt-out
with its own loud warning, below.)
Refusals (fail-closed):
| Configuration | When |
|---|---|
a tool sets allow_network = allow under tool_network auto/block |
machine start |
tool_network = only_explicit_states, or explicit tool_network = block |
run start, hardened ¹ |
a machine with tool states, or a tool with allow_network = block, under tool_network auto/block |
machine start, hardened ¹ |
-
⚠
none(non-Linux, or explicit opt-out) is unsandboxed: nothing enforced, nothing refused, loud warning. -
¹ per-command isolation needs a netns, so it's
strict-only. Onhardeneda jailed child inherits the agent's Landlock (host-agnostic, per provider port; UDP unconfined). The secure defaultautodegrades there with a warning; an EXPLICIT enforce (block,only_explicit_states) refuses rather than run silently under-confined.
More fail-closed properties:
- Operator-gated policy.
tool_networkis read only from the operator's config; a machine's[config]overlay is rejected at load if it declares[providers.*],[sandbox.*],[presets.*], orgit.run_repo_hooks.- Otherwise a strategy preset or a host
[machine.notify]argv could splice into the resolved config, andrun_repo_hookswould run repo.git/hookson the host on amode="run"commit. Atoolonly declaresallow_network; honoringallowis the operator's call, and every conflict is refused at startup naming the state.
- Otherwise a strategy preset or a host
- Bundle confinement. Scripts live in a reviewed
scripts/beside the.asm.toml;machine checkverifies every entry and static reference resolves inside the bundle (escaping symlinks rejected).- Scripts are operator-authored and committed, never fetched/generated at run
time, and the
.asm.toml+scripts/are RO in every jail during a run, so a state can't rewrite its own logic or add anallow_networkflag.
- Scripts are operator-authored and committed, never fetched/generated at run
time, and the
- Notifications don't widen the agent's surface. Front-ends render
machine.notifyas an overlay, andattach/TUI callnotify-sendwith a FIXED argv (no shell), so a model message is inert data.- The out-of-band hook
[machine.notify].on_eventruns an operator argv on the host with a minimal env -- PATH/HOME/locale/desktop-bus plusAGENT6_MACHINE_*(hook_envinapp/finalize.py), never the full environment with its provider keys (mirrors[notify].on_complete); a[config]overlay setting[machine.notify]is rejected at load. No Web Push/VAPID.
- The out-of-band hook
- A skill is config, not repo content: install only from trusted sources.
agent6 skills install <url>is an operator-initiated CLI fetch (same trust class asconnect); what it installs enters the system prompt/tool results verbatim.
- Nothing in a skill runs at install or load.
- Its scripts run only if the model runs them through the jailed command path,
subject to
run_commands.
- Its scripts run only if the model runs them through the jailed command path,
subject to
use_skillis read-only and path-contained. Serves the skill's own dir only (symlinks/..resolved first), never the repo or network. Skill dirs aren't mounted into the jail; content reaches the model engine-side.- Repo-local
.claude/skills/are deliberately NOT discovered. Third-party repo content must not enter the prompt; only the installed dir +[skills].extra_dirsare scanned.
tests/security/test_prompt_injection.py
runs an adversarial corpus through the planner/worker/reviewer prompts and
asserts no exfiltration, no out-of-policy tool calls, and no following embedded
instructions to weaken constraints. It's a smoke test, not a proof: the
structural defenses above are the real mitigation; the corpus catches prompt
regressions.
-
Landlock TCP rules need Linux ≥ 6.7 (ABI ≥ 4); older kernels leave the agent process net-unconfined (children stay net-isolated in
strictvia the empty netns). -
User namespaces must be enabled; some distros disable them, and agent6 refuses
strictthere. -
AppArmor userns (Ubuntu 24.04+) blocks unprivileged userns without a profile.
- agent6 ships one scoped to the launcher (
agent6 system apparmor install); with it, per-command jailing isstrict, without ithardened.
- agent6 ships one scoped to the launcher (
-
seccomp is required; kernels that block it from unprivileged callers make the jail fail closed.
-
Devcontainers get
hardened; the container is the FS blast radius, network still Landlock-confined when supported. The XDG state base is ephemeral (lost on rebuild), so mount a volume at the state dir or set[agent6].state_dirto persist runs. -
agent6 installed inside the project it works on (pip into the project's own venv) puts the running agent's code in the jail's writable workspace: a jailed command can rewrite it, and the next tool call runs the rewrite as you, outside the jail. Install agent6 outside the tree (pipx /
uv tool); agent6 warns at run entry when it detects this shape. -
Side channels: no claim about timing/cache/speculative side channels; don't co-locate agent6 with secrets if Spectre-class attacks are in your model.
-
Supply chain: pin your install. Runtime deps
pydantic,httpx2,argcomplete, thetree-sitterpair,textual,ruff,ty; build-dephatchling; the jail's Rust cratesnix,libc,landlock,seccompiler,serde,serde_json.