Skip to content

Bi-directional A2App: make the app→agent direction actually happen - #11

Merged
CraftOS-dev merged 9 commits into
mainfrom
korivi-a2app-bidirectional
Sep 29, 2026
Merged

CraftOS-dev merged 9 commits into
mainfrom
korivi-a2app-bidirectional

Conversation

@korivi-CraftOS

@korivi-CraftOS korivi-CraftOS commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

The gap

A2App already carries the app→agent plane: an app emits an event, a task lands in its queue, and a2app <app> tasks can claim it. Nothing made an agent turn up to take it, and nothing let an app put work there in the first place.

Both halves are here.


The consumer half: the ladder

agent-app <app> bridge is a per-app, opt-in service that watches one app's queue and triggers a harness by the deepest route it offers. It reports which rung it landed on, every time:

Rung Route When
1 inbound the harness already serves an HTTP endpoint that starts a run — deepest, and rare
2 headless its one-shot CLI (claude -p, codex exec, gemini -p, aider --message). Universal, and the default: the built-in profiles describe it, so most machines land here with nothing configured
3 gateway a local gateway that is not up yet; the bridge starts it and posts to it
4 subscribe the harness cannot be triggered, only poll. bridge start hands over the listen command instead of daemonizing
5 — none of the above. It says bi-directional operation is not supported here and names what would change it, rather than starting and silently delivering nothing
agent-app <dir> bridge                  # which rung, what is running, how much is waiting
agent-app <dir> bridge start            # background service (pid + log in .a2app/)
agent-app <dir> bridge start --once     # drain what is claimable now and exit (cron, CI)
agent-app <dir> bridge start --dry-run  # the exact prompts a run would produce; claims nothing
agent-app <dir> bridge stop

Profiles live in ~/.a2app/harnesses.json; a configured entry replaces a same-id built-in rather than merging into it. tokenEnv names the variable holding a token, never the token itself.

Rung 4's half is a2app <app> tasks next [--wait <ms>]: block until a task arrives, claim it, print it, exit. An idle queue is exit 0 with no task — a listen loop is idle almost all the time, and reporting that as a failure would make it indistinguishable from exit 3, which is the one that needs a human.

What it holds to

  • One run per task. The claim comes before delivery, so two bridges — or a bridge and a harness polling — can watch one queue and each task still runs once. The one hole is the app's own safety net: a task handed over HTTP stays working, and a harness that never reports has it returned to the queue after 60s and delivered again. Suppressing that would strand every task a dead agent was holding, so it is allowed — but the bridge remembers what it handed over and says so when one comes back, naming the fix. A duplicate run nobody is told about is the failure worth preventing.
  • The claim stays alive. The adapter sweeps a working task back after 60s without an update. A real run takes minutes, so progress heartbeats go out while a headless run is in flight — without them the app would redeliver a task still being worked on, the sweeper doing exactly its job, into a duplicate. A heartbeat the app refuses means the claim was lost, and the run is killed rather than allowed to finish against a task somebody else now owns.
  • Tasks are closed only where that is knowable. A headless run ends when the process exits, so its exit code closes the task — unless the harness closed it first, in which case its own result stands, input-required included. An HTTP 2xx is an acknowledgement, not a completion, so those are left open for the harness to close, and the prompt tells that harness it has no safety net. A refused trigger does fail the task: nothing started.
  • The payload is data. Fenced with a per-delivery nonce so it cannot close its own fence, labelled as data, capped with a pointer to tasks get <id>, and the harness is spawned with no shell and an argument array — a record whose title is a shell command arrives as a title.
  • It takes work only from the app addressed. The port is asked who it is, the same discipline serve and stop apply.
  • Nothing starts on its own, and --dry-run claims nothing.
  • It is quiet enough to leave running. .a2app/bridge.log is not rotated, so repetition is the enemy: a task filtered out by --capability is announced once rather than every poll, and a down app is complained about once (with one line when it returns) rather than every few seconds. Verified: a bridge idling against a live app writes one line.

The producer half: letting an app queue work

The delivery half had no producer. The creator skill told authors to "declare it in the trigger manifest"; every blueprint replied that no such manifest exists and to handle events "with plain code"; and the mechanism that is actually real — the adapter's trigger() — was described as a thing to consult directly or not use. An author following that guidance would never queue agent work, so the queue would always be empty.

blueprint-react-node now carries the seam. server.mjs is system-owned, so the adapter handle lives where an author may not edit; they get trigger in an operation runner's toolbox beside db and persist, and declare their event types in a2app.schema.mjs:

events: [{ type: "task.needs_triage" }],

"request-triage": (args, _ctx, { db, trigger }) => {
  const task = db.tasks?.[args?.task];
  if (!task) return { ok: false, reason: "no such task" };
  const { taskId } = trigger("task.needs_triage", { task: task.id }, "triage");
  return { ok: true, queued: taskId };
},

server.mjs passes that list to the adapter, so an undeclared type is refused at fire time — what an app can ever ask for is fixed by its author, in code they own. There is still no manifest of instructions, deliberately: an app names a capability, never a command, which is the property the old text was reaching for.

The docs now say what is true per stack — react-node and python-fastapi have the queue and how to reach it; pocketbase does not and says so, instead of implying the feature is missing everywhere. Both skills now also say that a queued task moves only when a bridge or a polling harness is listening, so a feature that queues work does not look broken.


Also in here

  • claimTask / claim_task take an optional credential in both SDKs, and tasks claim an optional --as: a claim can only ever name the caller's own credential, so requiring it made an agent look up a value that could not change the outcome.
  • flagAll for repeatable flags — honouring only the first --capability would narrow a filter the caller widened.
  • Delivered prompts got shorter: a headless harness is spawned in the app's directory, so the app is addressed as . rather than by an absolute path repeated four times. Routes that run elsewhere still get the full path.
  • agent-app list marks an app whose bridge is alive as bridge:<harness>/<route>. A bridge is a standing capability — while it runs, that app can start agent runs on this machine — and one started weeks ago was visible nowhere.
  • A lock broken after its holder died no longer fails the command when the tidy-up delete loses a race with an open handle on Windows.

Verification

  • pnpm -r build and pnpm -r typecheck clean; CI green on Ubuntu across all three jobs.
  • Conformance 82/82 (was 80). Class C gains two checks for tasks next against the real adapter.
  • framework/cli/test/bridge.test.mjs — all five rungs, config that is never guessed at, claim-before-deliver, a payload reaching the harness as one argument with no shell, the handoff and terminal-state rules, input-required left parked, --dry-run claiming nothing, and the detached daemon's start/status/stop.
  • toolkits/blueprint-react-node/test/app-to-agent.test.mjs — the seam exists, every fired event type is declared, and the worked example sends ids rather than prose.
  • End to end on a scaffolded app, not just in unit tests: scaffold → serve → request-triage → the task appears in the queue → bridge start --once claims it, runs the harness with the fenced prompt, and completes it. Firing the identical trigger twice returned the same task, which is the dedup rule doing its job.

Bugs found by writing the tests, each verified to fail against the unfixed code:

  1. The detached child resolved a relative app path against its own working directory — a bridge looking for an app inside the app.
  2. An HTTP trigger was recorded as a completed task, when 2xx only means received.
  3. input-required — an agent parking a task for a human — was being overwritten with "completed".
  4. Rung 3 was implemented but never exercised; covering it turned up a gateway that --once started and nothing could then stop.
  5. Three faults that only appear on the timescale a service runs: a filtered task announced every poll forever, a down app complained about every poll forever, and a redelivered handoff that silently broke "one run per task" for the deepest rung.

Not in scope

  • The bridge is local-only: agent-app acts on an app's files, so a URL is refused there by construction. A connected remote app is covered by a2app <url> tasks next --wait, which needs no files.
  • One task at a time. Two harness runs writing into one app concurrently is a race the app's guard cannot see, since both are valid writes.
  • blueprint-pocketbase-react has no tasks/events surface; that is documented rather than papered over.
  • No spec change: this adds no wire fields.

A2App already carries the app->agent plane: an app emits an event, a task
lands in its queue, and `a2app <app> tasks` can claim it. Nothing made an
agent turn up to take it, so a submitted task sat in `submitted` until
someone happened to look. Triggering is the one part of this that depends
on the harness rather than on the app, so it belongs to the framework.

Adds `agent-app <app> bridge` — a per-app, opt-in service that watches one
app's queue and triggers a harness by the deepest route that harness
offers, and reports which rung it landed on:

  1 inbound    an HTTP endpoint the harness already serves (rare)
  2 headless   its one-shot CLI (claude -p, codex exec, gemini -p,
               aider --message). Universal, and the default: the built-in
               profiles describe it, so most machines land here with
               nothing configured
  3 gateway    a local gateway that is not up yet; the bridge starts it
  4 subscribe  the harness cannot be triggered, only poll — `bridge start`
               hands over `a2app <app> tasks next --wait` instead of
               daemonizing
  5 none       none of the above. It says bi-directional operation is not
               supported here and names what would change it, rather than
               starting and delivering nothing

Harness profiles live in ~/.a2app/harnesses.json; a configured entry
replaces a same-id built-in outright rather than merging into it.

`a2app <app> tasks next [--wait]` is rung 4's half: block until a task
arrives, claim it, print it, exit. An idle queue is exit 0 with no task —
a listen loop is idle almost all the time, and reporting that as a failure
would make it indistinguishable from exit 3.

What the bridge holds to:

- One run per task. The claim comes before delivery, so two bridges, or a
  bridge and a harness polling, can watch one queue and each task still
  runs once.
- The claim stays alive. The adapter sweeps a `working` task back after
  60s without an update; a real run takes minutes, so progress heartbeats
  go out while a headless run is in flight. A heartbeat the app refuses
  means the claim was lost, and the run is killed rather than allowed to
  finish against a task someone else now owns.
- Tasks are closed only where that is knowable. A headless run ends when
  the process exits, so its exit code closes the task — unless the harness
  closed it first, in which case its own result stands, `input-required`
  included. An HTTP 2xx is an acknowledgement, not a completion, so those
  tasks are left open for the harness to close. A refused trigger does
  fail the task: nothing started.
- The payload is data. Fenced with a per-delivery nonce so it cannot close
  its own fence, labelled as data rather than instructions, capped with a
  pointer to `tasks get <id>`, and the harness is spawned with no shell and
  an argument array.
- Nothing starts on its own, and `--dry-run` claims nothing.

Also: `claimTask`/`claim_task` take an optional credential in both SDKs and
`tasks claim` an optional `--as`, since a claim can only ever name the
caller's own credential; `flagAll` for repeatable flags; and a lock broken
after its holder died no longer fails the command if the tidy-up delete
loses a race with an open handle on Windows.

Conformance class C grows two checks for `tasks next` against the real
adapter (82/82), and the bridge has its own suite covering the ladder,
claim-before-deliver, no-shell delivery, the handoff and terminal-state
rules, and the detached daemon's start/status/stop.
Two things found reading it back.

A stop was noticed only when the poll interval elapsed, because the gap
between passes was one long timer. That is a Ctrl-C that appears to hang,
and a `bridge stop` that falls through to a force-kill because the process
did not exit in time. The wait is now sliced and checks the stop flag, so
stopping is immediate and a graceful exit is the normal path.

And a file on PATH that is not executable was reported as an installed
harness, which passes the ladder's availability check and then fails at
spawn. Rung 5 exists to be told up front, so POSIX now checks X_OK;
Windows has no execute bit and keeps using the PATHEXT suffix.
The gateway rung was implemented but never exercised: `mode: "gateway"`
appeared nowhere in the tests, so "the framework starts the gateway for
you" was a claim rather than a checked fact.

It now has a test that can only pass if the framework really starts it —
the gateway writes its pid on boot, the test asserts that file is absent
beforehand and present after, and the trigger is delivered to a port that
exists only because the bridge brought it up. It also pins the two rules
that go with it: an already-running gateway is reused rather than started
a second time, and an acknowledgement is a handoff, not a completion.

Writing it turned up a wart. A gateway outlives the pass that started it,
but only a long-running `bridge start` records its pid for `bridge stop`.
A `--once` pass cleared that record on the way out, leaving a process
nothing could stop. Reuse is the right behaviour for a service that a
cron'd pass hits repeatedly, so the pass now names the pid it started
instead of pretending it owns nothing.
The delivery half had no producer. `agent-app bridge` takes tasks off an
app's queue, but nothing put them there: the creator skill told authors to
"declare it in the trigger manifest", every blueprint replied that no such
manifest exists and to handle events "with plain code", and the mechanism
that is actually real — the adapter's `trigger()` — was described as a
thing to consult directly or not use. An author following that guidance
would never queue agent work, so the queue would always be empty.

blueprint-react-node now carries the seam. `server.mjs` is system-owned,
so the adapter handle lives where an author may not edit; they get
`trigger` in an operation runner's toolbox beside `db` and `persist`, and
declare their event types in `a2app.schema.mjs` — which `server.mjs`
passes to the adapter, so an undeclared type is refused at fire time. A
worked `request-triage` operation shows the shape: declared type, a
capability, and the record's ID rather than a copy or prose.

The docs now say what is true per stack: react-node and python-fastapi
have the queue and how to reach it, pocketbase does not and says so
instead of implying the feature is missing everywhere. The creator skill
describes the mechanism that exists and keeps the property the old text
was reaching for — an app names a capability, never an instruction, so a
compromised app cannot steer an agent — and both skills now say a queued
task moves only when a bridge or a polling harness is listening.

Also shortens every delivered prompt. A headless harness is spawned in the
app's directory, so the app is addressed as `.` rather than by an absolute
path repeated four times, which on a deeply nested app was most of what
the agent read. Routes that run elsewhere still get the full path.

Verified end to end on a scaffolded app, not just in unit tests: scaffold →
serve → `request-triage` → the task appears in the queue → `bridge start
--once` claims it, runs the harness with the fenced prompt, and completes
it. Firing the identical trigger twice returned the same task, which is
the dedup rule doing its job.
Three faults that only show up on the timescale a service actually runs,
none of which a test that finishes in a second would ever see.

A task filtered out by --capability is never claimed, so it stays in the
queue and was announced on every pass — a line every five seconds, forever,
into a log nothing rotates. An app that goes down was complained about just
as often. Both are now said once: the filtered task per id, the poll failure
until the message changes, with one line when the app answers again.

The third is the one that matters. A task handed over an HTTP route stays
`working` by design, because a 2xx is an acknowledgement and only the
harness knows when the work is done. If that harness never reports, the app
returns the task to the queue after 60s and the bridge delivers it again —
so "one run per task", the promise at the top of this file, quietly stopped
holding for the deepest rung. Suppressing the redelivery would strand every
task a dead agent was holding, so it still happens; what changes is that the
bridge remembers what it handed over, counts it, and says so, naming the fix
(the harness must call `tasks progress`). A duplicate run nobody is told
about is the failure worth preventing, not the duplicate itself.

`agent-app list` also marks an app whose bridge is alive. A bridge is a
standing capability — while it runs, that app can start agent runs on this
machine — and one started weeks ago was visible nowhere.
…a port

harness-plugins/README.md had a ladder for getting a launched app in front
of a person, and nothing for the direction that now exists: getting an agent
to pick up work an app queued. Rungs 1 and 3 are exactly where a plugin is
the only thing that knows its harness's API, so the contract belongs there.

It is one file. A plugin whose harness offers an endpoint or a gateway
writes that route into harnesses.json at install time and every app on the
machine can use it; the same entry is how a plugin corrects a headless
invocation whose flags changed. A plugin that does nothing here is not
broken — its users land on rung 2 or 4, which need no plugin code — and no
plugin may start a bridge on a user's behalf, because a bridge lets an app
start agent runs.

Also closes a race in the gateway test. It picks a port by binding and
releasing, which leaves a window for another process on a loaded machine to
take it first; the stand-in gateway now retries the bind instead of dying.
A test that is flaky under load teaches people to re-run rather than read.
@ahmad-ajmal
ahmad-ajmal force-pushed the korivi-a2app-bidirectional branch from 21d5384 to 9d95c80 Compare September 28, 2026 11:56
@CraftOS-dev
CraftOS-dev merged commit de0f613 into main Sep 29, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants