Repository navigation
Bi-directional A2App: make the app→agent direction actually happen - #11
Merged
Merged
Conversation
A2App already carries the app->agent plane: an app emits an event, a task
lands in its queue, and `a2app <app> tasks` can claim it. Nothing made an
agent turn up to take it, so a submitted task sat in `submitted` until
someone happened to look. Triggering is the one part of this that depends
on the harness rather than on the app, so it belongs to the framework.
Adds `agent-app <app> bridge` — a per-app, opt-in service that watches one
app's queue and triggers a harness by the deepest route that harness
offers, and reports which rung it landed on:
1 inbound an HTTP endpoint the harness already serves (rare)
2 headless its one-shot CLI (claude -p, codex exec, gemini -p,
aider --message). Universal, and the default: the built-in
profiles describe it, so most machines land here with
nothing configured
3 gateway a local gateway that is not up yet; the bridge starts it
4 subscribe the harness cannot be triggered, only poll — `bridge start`
hands over `a2app <app> tasks next --wait` instead of
daemonizing
5 none none of the above. It says bi-directional operation is not
supported here and names what would change it, rather than
starting and delivering nothing
Harness profiles live in ~/.a2app/harnesses.json; a configured entry
replaces a same-id built-in outright rather than merging into it.
`a2app <app> tasks next [--wait]` is rung 4's half: block until a task
arrives, claim it, print it, exit. An idle queue is exit 0 with no task —
a listen loop is idle almost all the time, and reporting that as a failure
would make it indistinguishable from exit 3.
What the bridge holds to:
- One run per task. The claim comes before delivery, so two bridges, or a
bridge and a harness polling, can watch one queue and each task still
runs once.
- The claim stays alive. The adapter sweeps a `working` task back after
60s without an update; a real run takes minutes, so progress heartbeats
go out while a headless run is in flight. A heartbeat the app refuses
means the claim was lost, and the run is killed rather than allowed to
finish against a task someone else now owns.
- Tasks are closed only where that is knowable. A headless run ends when
the process exits, so its exit code closes the task — unless the harness
closed it first, in which case its own result stands, `input-required`
included. An HTTP 2xx is an acknowledgement, not a completion, so those
tasks are left open for the harness to close. A refused trigger does
fail the task: nothing started.
- The payload is data. Fenced with a per-delivery nonce so it cannot close
its own fence, labelled as data rather than instructions, capped with a
pointer to `tasks get <id>`, and the harness is spawned with no shell and
an argument array.
- Nothing starts on its own, and `--dry-run` claims nothing.
Also: `claimTask`/`claim_task` take an optional credential in both SDKs and
`tasks claim` an optional `--as`, since a claim can only ever name the
caller's own credential; `flagAll` for repeatable flags; and a lock broken
after its holder died no longer fails the command if the tidy-up delete
loses a race with an open handle on Windows.
Conformance class C grows two checks for `tasks next` against the real
adapter (82/82), and the bridge has its own suite covering the ladder,
claim-before-deliver, no-shell delivery, the handoff and terminal-state
rules, and the detached daemon's start/status/stop.
Two things found reading it back. A stop was noticed only when the poll interval elapsed, because the gap between passes was one long timer. That is a Ctrl-C that appears to hang, and a `bridge stop` that falls through to a force-kill because the process did not exit in time. The wait is now sliced and checks the stop flag, so stopping is immediate and a graceful exit is the normal path. And a file on PATH that is not executable was reported as an installed harness, which passes the ladder's availability check and then fails at spawn. Rung 5 exists to be told up front, so POSIX now checks X_OK; Windows has no execute bit and keeps using the PATHEXT suffix.
The gateway rung was implemented but never exercised: `mode: "gateway"` appeared nowhere in the tests, so "the framework starts the gateway for you" was a claim rather than a checked fact. It now has a test that can only pass if the framework really starts it — the gateway writes its pid on boot, the test asserts that file is absent beforehand and present after, and the trigger is delivered to a port that exists only because the bridge brought it up. It also pins the two rules that go with it: an already-running gateway is reused rather than started a second time, and an acknowledgement is a handoff, not a completion. Writing it turned up a wart. A gateway outlives the pass that started it, but only a long-running `bridge start` records its pid for `bridge stop`. A `--once` pass cleared that record on the way out, leaving a process nothing could stop. Reuse is the right behaviour for a service that a cron'd pass hits repeatedly, so the pass now names the pid it started instead of pretending it owns nothing.
The delivery half had no producer. `agent-app bridge` takes tasks off an app's queue, but nothing put them there: the creator skill told authors to "declare it in the trigger manifest", every blueprint replied that no such manifest exists and to handle events "with plain code", and the mechanism that is actually real — the adapter's `trigger()` — was described as a thing to consult directly or not use. An author following that guidance would never queue agent work, so the queue would always be empty. blueprint-react-node now carries the seam. `server.mjs` is system-owned, so the adapter handle lives where an author may not edit; they get `trigger` in an operation runner's toolbox beside `db` and `persist`, and declare their event types in `a2app.schema.mjs` — which `server.mjs` passes to the adapter, so an undeclared type is refused at fire time. A worked `request-triage` operation shows the shape: declared type, a capability, and the record's ID rather than a copy or prose. The docs now say what is true per stack: react-node and python-fastapi have the queue and how to reach it, pocketbase does not and says so instead of implying the feature is missing everywhere. The creator skill describes the mechanism that exists and keeps the property the old text was reaching for — an app names a capability, never an instruction, so a compromised app cannot steer an agent — and both skills now say a queued task moves only when a bridge or a polling harness is listening. Also shortens every delivered prompt. A headless harness is spawned in the app's directory, so the app is addressed as `.` rather than by an absolute path repeated four times, which on a deeply nested app was most of what the agent read. Routes that run elsewhere still get the full path. Verified end to end on a scaffolded app, not just in unit tests: scaffold → serve → `request-triage` → the task appears in the queue → `bridge start --once` claims it, runs the harness with the fenced prompt, and completes it. Firing the identical trigger twice returned the same task, which is the dedup rule doing its job.
Three faults that only show up on the timescale a service actually runs, none of which a test that finishes in a second would ever see. A task filtered out by --capability is never claimed, so it stays in the queue and was announced on every pass — a line every five seconds, forever, into a log nothing rotates. An app that goes down was complained about just as often. Both are now said once: the filtered task per id, the poll failure until the message changes, with one line when the app answers again. The third is the one that matters. A task handed over an HTTP route stays `working` by design, because a 2xx is an acknowledgement and only the harness knows when the work is done. If that harness never reports, the app returns the task to the queue after 60s and the bridge delivers it again — so "one run per task", the promise at the top of this file, quietly stopped holding for the deepest rung. Suppressing the redelivery would strand every task a dead agent was holding, so it still happens; what changes is that the bridge remembers what it handed over, counts it, and says so, naming the fix (the harness must call `tasks progress`). A duplicate run nobody is told about is the failure worth preventing, not the duplicate itself. `agent-app list` also marks an app whose bridge is alive. A bridge is a standing capability — while it runs, that app can start agent runs on this machine — and one started weeks ago was visible nowhere.
…a port harness-plugins/README.md had a ladder for getting a launched app in front of a person, and nothing for the direction that now exists: getting an agent to pick up work an app queued. Rungs 1 and 3 are exactly where a plugin is the only thing that knows its harness's API, so the contract belongs there. It is one file. A plugin whose harness offers an endpoint or a gateway writes that route into harnesses.json at install time and every app on the machine can use it; the same entry is how a plugin corrects a headless invocation whose flags changed. A plugin that does nothing here is not broken — its users land on rung 2 or 4, which need no plugin code — and no plugin may start a bridge on a user's behalf, because a bridge lets an app start agent runs. Also closes a race in the gateway test. It picks a port by binding and releasing, which leaves a window for another process on a loaded machine to take it first; the stand-in gateway now retries the bind instead of dying. A test that is flaky under load teaches people to re-run rather than read.
ahmad-ajmal
force-pushed
the
korivi-a2app-bidirectional
branch
from
September 28, 2026 11:56
21d5384 to
9d95c80
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The gap
A2App already carries the app→agent plane: an app emits an event, a task lands in its queue, and
a2app <app> taskscan claim it. Nothing made an agent turn up to take it, and nothing let an app put work there in the first place.Both halves are here.
The consumer half: the ladder
agent-app <app> bridgeis a per-app, opt-in service that watches one app's queue and triggers a harness by the deepest route it offers. It reports which rung it landed on, every time:inboundheadlessclaude -p,codex exec,gemini -p,aider --message). Universal, and the default: the built-in profiles describe it, so most machines land here with nothing configuredgatewaysubscribebridge starthands over the listen command instead of daemonizingProfiles live in
~/.a2app/harnesses.json; a configured entry replaces a same-id built-in rather than merging into it.tokenEnvnames the variable holding a token, never the token itself.Rung 4's half is
a2app <app> tasks next [--wait <ms>]: block until a task arrives, claim it, print it, exit. An idle queue is exit 0 with no task — a listen loop is idle almost all the time, and reporting that as a failure would make it indistinguishable from exit 3, which is the one that needs a human.What it holds to
working, and a harness that never reports has it returned to the queue after 60s and delivered again. Suppressing that would strand every task a dead agent was holding, so it is allowed — but the bridge remembers what it handed over and says so when one comes back, naming the fix. A duplicate run nobody is told about is the failure worth preventing.workingtask back after 60s without an update. A real run takes minutes, so progress heartbeats go out while a headless run is in flight — without them the app would redeliver a task still being worked on, the sweeper doing exactly its job, into a duplicate. A heartbeat the app refuses means the claim was lost, and the run is killed rather than allowed to finish against a task somebody else now owns.input-requiredincluded. An HTTP 2xx is an acknowledgement, not a completion, so those are left open for the harness to close, and the prompt tells that harness it has no safety net. A refused trigger does fail the task: nothing started.tasks get <id>, and the harness is spawned with no shell and an argument array — a record whose title is a shell command arrives as a title.serveandstopapply.--dry-runclaims nothing..a2app/bridge.logis not rotated, so repetition is the enemy: a task filtered out by--capabilityis announced once rather than every poll, and a down app is complained about once (with one line when it returns) rather than every few seconds. Verified: a bridge idling against a live app writes one line.The producer half: letting an app queue work
The delivery half had no producer. The creator skill told authors to "declare it in the trigger manifest"; every blueprint replied that no such manifest exists and to handle events "with plain code"; and the mechanism that is actually real — the adapter's
trigger()— was described as a thing to consult directly or not use. An author following that guidance would never queue agent work, so the queue would always be empty.blueprint-react-nodenow carries the seam.server.mjsis system-owned, so the adapter handle lives where an author may not edit; they gettriggerin an operation runner's toolbox besidedbandpersist, and declare their event types ina2app.schema.mjs:server.mjspasses that list to the adapter, so an undeclared type is refused at fire time — what an app can ever ask for is fixed by its author, in code they own. There is still no manifest of instructions, deliberately: an app names a capability, never a command, which is the property the old text was reaching for.The docs now say what is true per stack — react-node and python-fastapi have the queue and how to reach it; pocketbase does not and says so, instead of implying the feature is missing everywhere. Both skills now also say that a queued task moves only when a bridge or a polling harness is listening, so a feature that queues work does not look broken.
Also in here
claimTask/claim_tasktake an optional credential in both SDKs, andtasks claiman optional--as: a claim can only ever name the caller's own credential, so requiring it made an agent look up a value that could not change the outcome.flagAllfor repeatable flags — honouring only the first--capabilitywould narrow a filter the caller widened..rather than by an absolute path repeated four times. Routes that run elsewhere still get the full path.agent-app listmarks an app whose bridge is alive asbridge:<harness>/<route>. A bridge is a standing capability — while it runs, that app can start agent runs on this machine — and one started weeks ago was visible nowhere.Verification
pnpm -r buildandpnpm -r typecheckclean; CI green on Ubuntu across all three jobs.tasks nextagainst the real adapter.framework/cli/test/bridge.test.mjs— all five rungs, config that is never guessed at, claim-before-deliver, a payload reaching the harness as one argument with no shell, the handoff and terminal-state rules,input-requiredleft parked,--dry-runclaiming nothing, and the detached daemon's start/status/stop.toolkits/blueprint-react-node/test/app-to-agent.test.mjs— the seam exists, every fired event type is declared, and the worked example sends ids rather than prose.request-triage→ the task appears in the queue →bridge start --onceclaims it, runs the harness with the fenced prompt, and completes it. Firing the identical trigger twice returned the same task, which is the dedup rule doing its job.Bugs found by writing the tests, each verified to fail against the unfixed code:
input-required— an agent parking a task for a human — was being overwritten with "completed".--oncestarted and nothing could then stop.Not in scope
agent-appacts on an app's files, so a URL is refused there by construction. A connected remote app is covered bya2app <url> tasks next --wait, which needs no files.blueprint-pocketbase-reacthas no tasks/events surface; that is documented rather than papered over.