Skip to content
View wippa-studios's full-sized avatar
🙂
ready-to-work
🙂
ready-to-work

Block or report wippa-studios

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
wippa-studios/README.md
wippa studios

Build an agent. Let it talk to other agents. Then check the protocol underneath isn't lying to you.


Projects Public repos with CI Upstream PRs merged Runtime deps Defects found in my own protocol License


🧭 The stack

  ┌─────────────────┐
  │   agent-core    │   build one agent that can plan, act, and self-budget
  └────────┬────────┘
           ▼
  ┌─────────────────┐
  │  agent-protect  │   run agents you didn't write, without trusting them
  └────────┬────────┘
           ▼
  ┌─────────────────┐
  │   wippa-uacp    │   the wire protocol, so any two agents can talk
  └────────┬────────┘
           ▼
  ┌─────────────────┐
  │  uacp-interop   │   check that the protocol agrees with its own code
  └─────────────────┘

Four separate projects, no dependency between them. The arrows mean "this is what you reach for next", not "this imports that". The one real link is called out below.


📡 The protocol work

A conformance harness that tests the spec against itself.

CI PyPI Python

Probes the spec against its own schemas, both reference implementations, live servers, and its own docs, then reports where they disagree.

96 tests · runs offline

uvx uacp-interop

An HTTP-shaped wire protocol for agent-to-agent traffic.

CI

JSON Schema spec, implemented twice (TypeScript + Python) against one cross-language suite. Dot-notation capability namespaces, correlation IDs, streaming, cancellation, discovery, auth envelope. A2A and MCP adapters.

Zero runtime dependencies · 161 Python checks · 96 TS assertions Node 18/20/22 · Python 3.10–3.12

🐛 What the harness found in a protocol I wrote

A conformance checklist proves your implementation matches your own reading of a spec. It says nothing about whether your reading and your code agree. uacp-interop exists to close that gap, and it found eleven defects, all now fixed. The greatest hits:

# Finding Why it looked fine What it actually did
🔓 An authorization check nobody called Both implementations shipped a correct, unit-tested check_authorization. The spec said a bus "can enforce" allow-lists. Bus.send never invoked it. A caller the policy denied reached the protected capability.
👤 metadata.auth missing from the Python model It existed in the TypeScript one. An authenticated message decoded as anonymous, with no error.
⏳ async defaulted to true Schema and prose were consistent. A plain request/response capability was documented as streaming; the consumer waited for a stream.end that never came. A hang, not an error.
🔑 Audit trail stored bearer tokens in plaintext It was "just logging". Credentials at rest, in the log.
🏷️ capability globally required Tidy and uniform. heartbeat and register invented a placeholder; both implementations sent the literal "_internal".
➖ Hyphens banned in capability names A deliberate, reasoned exclusion. resolve-library-id and query-docs couldn't be named without renaming them, and a renamed tool is a different tool to any client that calls it by name.
🧠 The lessons behind the table (click to expand)

"Can" reads as a capability, not an obligation. No schema comparison could have found the unwired authorization check, because the schema and the prose were both fine. The prose under-committed. Only reading the code could catch it.

The check that guards it is source probing, not behavioural testing. It doesn't execute the bus. It probes the source, per language, for the call itself, so it catches the wiring being removed. The first version recorded enforced: true by hand, which meant it was asserting my own reading back at me. Had the call been deleted later it would have reported clean forever, and a green test count would have been a true statement about nothing.

The hyphen bug is the one I'd underline. Two of six live MCP servers announce a name containing a hyphen and couldn't supply a conformant agent id at all, so the protocol couldn't describe a third of what it exists to interoperate with. I'd excluded hyphens on purpose, arguing an agent is an identity rather than a namespace. But the separator is the dot, so nothing technical turned on it. It was tidied-up symmetry, which is precisely the failure the guarding test predicted and couldn't prevent. No amount of self-consistent testing would have found it; only measuring other people's servers did.

One finding that isn't a bug: nothing bounds a pipeline as a whole. Two conforming agents calling each other ran at ~20k round trips/second, limited only by the host recursion limit, and the caller received a success response with no sign the loop had been cut short.

🪞 …and what it found back in the protocol repo

Three defects that all looked fine:

1. A test suite reporting "7 errors" while executing zero assertions

The cross-language suite was written as a standalone script; under pytest every test function requested a fixture nothing provided. Its TypeScript peer ran 54 real assertions the whole time, which made the broken side look healthy. With it running, the async default fix the spec had already merged turned out never to have been applied to the Python implementation.

2. ACLs written as deploy.* matched nothing

check_authorization compared for exact equality plus a bare *, so every ACL with a trailing .* (the form the spec and tests both use) matched nothing, and denied=["deploy.*"] did not deny deploy.*. The four assertions that would have caught it lived in a main() nobody ran.

3. A multi-agent example where no agent ever called another

It registered three agents, then drove them from the script with from_="pipeline", a bare string rather than an agent. Every hop went script → agent → script. The output was correct throughout, which is exactly why it survived.

And the fourth, still partly open: a permissive authorization default. A bus with no ACLs routes everything, which is defensible only when every caller is a function in the same process. Inverting it outright would break every single-author use while protecting nothing, so Bus now takes trust='trusted' | 'untrusted' and refuses to construct an untrusted bus with no policy. Reaching a networked bus is now a decision rather than an omission.


🤖 The agent work

The framework.

CI

  • ⚡ Metabolic budgeting: every step spends from a budget, so a runaway plan starves itself out instead of the budget
  • 🏛️ A Governor critiques a plan before an LLM proposer may execute it
  • 🕸️ Graph memory, a DSL executor, an Express dashboard

184 tests across 17 files

🛡️ @wippa/core

The sandbox (agent-protect).

CI

Found a CrewAI or LangChain repo and want to run it without trusting it?

clone → scan for secrets & high-entropy strings → build → execute in Docker with egress filtering and secret masking.

97 tests · one runtime dependency

Also: wippa-automations, an AI-first Zapier alternative, and wippa-opencode-agents, a portable multi-agent OpenCode configuration.


🔗 Where the four connect

They don't share code, so it's worth being precise about the one real link:

uacp-interop vendors wippa-uacp's JSON Schemas and records the exact upstream commit in PROVENANCE.json, re-checked for drift on every run.

It deliberately does not pip install the protocol, so a conformance report can test a released spec without silently testing something newer. Tracking main is how a report ends up certifying a spec that was never published. (It's also why a fixed defect keeps showing up as a live finding until the schema is re-vendored.)

The trust boundary: part design, part gap

agent-protect exists because untrusted agents are dangerous; UACP exists to put agents on a shared bus.

  • ✅ Designed: a bus carrying an agent the operator doesn't trust is enforceable. Bus takes an ACL set, rejects unauthorized calls with AUTHORIZATION before routing, and won't construct at all if told it's untrusted without a policy.
  • ⚠️ Still missing: a cross-pipeline cost ceiling. Two well-behaved agents in a loop are bounded by nothing in the protocol, and the caller can't tell a completed pipeline from a truncated one. That's a specification gap, not an implementation bug: no amount of better code makes it conformant, because there's no claim to have failed.

Both are written up in CONFORMANCE-TESTING.md and §15.8 rather than left implicit.


🧪 Tools & experiments

Project What it is
collab-vLLM Run Qwen3 on a free Colab GPU and point OpenCode at it as an OpenAI-compatible provider
OpenBookmaker Open-source Betfair-style exchange: back/lay order book, tick-ladder pricing, matching engine, cash-out
wippa-bet-lab Backtest and paper-trade strategies. Research in greyhounds · tennis-pro · nba-data
wippa-bet-agent Real-time Betfair Exchange paper-trading agent. Push-driven Stream API, a rule DSL, and a generator that searches recorded data for systems worth promoting. Paper money only — no order-placement path exists in the stack
wippa-project-manager Project management where an agent can be the assignee. A task declares capability shapes, an agent publishes capability shapes, and "done" is a signed run report — a task cannot be completed by anyone asserting a status. Zero runtime dependencies · PyPI
wippa-sky-v3 A city builder with a one-point perspective renderer. Earlier: wippa-sky, wippa-build

🌍 Contributing upstream

I fix things in other people's projects. Three have merged.

✅ Merged: flipt#6605

A stale-read race in updateSnapshot. It published a new evaluation snapshot to subscribers before swapping it into the store. A client that refetched on the hint read the previous snapshot (or got a not-found for a namespace that had just been created), and because the stream had already recorded the digest, no second event followed. It stayed stale until the next flag change. One reorder fixes it. A reviewer asked for the regression test to move into the existing test file; a maintainer then extended the fix and merged v2 in. 48694cbd

✅ Merged: gala#598

uacp-interop's own methodology, pointed at something else. The maintainer of gala asked for a check that its Nix-built stdlib byte-matches the Bazel one, and supplied the shell commands. Nothing compared them, so a codegen change that genuinely moved stdlib output would leave the Nix build succeeding with a stale embedded stdlib and no signal. The prerequisite had already landed the day before, so the lane was the whole of it. 0421b5c

I could not run it — no Nix, no Bazel where I wrote it — so the body says so plainly and names the one thing a reviewer should check instead. martianoff merged it on the description, not on my verification.

✅ Merged: serverless#13901

A TypeError: functionArn.split is not a function crash in API Gateway authorizer validation: any CloudFormation-intrinsic arn without a name reached a string-only ARN parser. Replaced with an actionable validation error, and covered the error path, which had no test at all. czubocha merged it. fe70b85c

🔍 Diagnosed: starnet#18

Interactive replies in a conversation that also held scheduled-routine results couldn't be rated, because eligibility was read from the conversation's origin instead of the individual run's. I traced the fix commit and confirmed it was an ancestor of the shipping branch; the maintainer shipped it in v0.12.4 and closed the issue thanking me for the flag. My contribution there was the diagnosis, not a patch.

❌ Closed: cli/cli#14544 — the best bug I found, closed for process

gh pr merge --delete-branch deletes someone else's work. It matches the head branch by name in whatever repo your shell happens to be in. If an unrelated repo has a worktree on a same-named branch holding unpushed commits, gh removes the worktree and runs git branch -D on it — and since git branch -D bypasses git's own merged check, those commits exist only in a reflog. I reproduced it, wrote a scope guard plus an ancestry check, and got 7/7 green CI with 8 of 13 new tests failing against pristine source.

It was still the wrong contribution, and I closed it myself.

cli/cli takes external PRs only for help wanted issues with explicit Acceptance Criteria, and their AGENTS.md asks you to verify that before implementing — and to not automatically publish a comment asking for the label. The issue had needs-triage. Their own bot labelled the PR unmet-requirements twenty minutes after it opened.

The mistake that actually mattered wasn't opening it. It was that, having found the help wanted rule in CONTRIBUTING.md, I recommended a public comment as the path their docs prescribe — without having opened AGENTS.md, which forbids exactly that. Incomplete file list, not faulty reasoning. Those fail differently: faulty reasoning gets fixed, an incomplete checklist gets repeated. AGENTS.md and CLAUDE.md now get read before CONTRIBUTING.md, because an agent-facing file governs permission and CONTRIBUTING only governs the bar.

The branch is still up. If #14537 ever gets triaged into scope it's one reopen away.

Closing the case martianoff couldn't reproduce. They'd reported that a local type in their own sample shadowed an import, and it didn't happen for them. It was a self-hosting resolver that only treated a package-qualified path as a local name, so http.Get shadowed the imported http while net/http.Get did not — and their sample used the qualified form. One file, +189/−2, with the 14/14 regression test.

There's a second defect in that resolver I did not fix: a bare name with no local declaration is still handled unsoundly at step 5, and two committed examples depend on it. I flagged it as its own issue rather than quietly widening a PR that was already a maintainer request.

Transaction support, closing a help wanted that had been open four years. The mechanism existed three times over in that codebase and was not on the path a request takes: a package-level OpenTransactions map nothing read, an implemented-but-unused TransactionManager, and Insert/Update/DeleteWithTransaction already declared on the adapter interface. This is the same shape as the unwired authorization check above — correct, unit-tested, unreachable.

Wiring it was the easy half. Two defects in my own first cut made it non-functional, and I only found them by reading the bot's review rather than declaring it done: the transaction was opened on the request context, which database/sql rolls back the moment that request returns, and the commit path read its id with r.PathValue, which gorilla/mux v1.8.1 leaves empty — so every commit returned 410. My tests had missed both because they called the handlers directly instead of going through the router, and the stub ignored the context. Test and code shared one wrong assumption and so agreed with each other. Each fix has a test that fails without it.

🚧 Also open: kaneo#1851

A pending-invitation list that was permanently blank, because the page read invitations off a full-workspace payload that carries no such field, and now reads the dedicated query the invite mutations already invalidate. The regression test fails on a pristine main and the failure is the blank list itself.

🚧 Also open: IBM#7001

Covers a response.json() failure on a non-standard 2xx — a branch that existed in the grammar and in the transpiler but had no test reaching it, so _handle_json_parse_error was unguarded. Test-only.

📬 All open PRs (9)

gala#610 · prest#1046 · kaneo#1851 · flumine#864 · OpenSandbox#2012 · crawl4ai#2299 · starnet#45 · tirith#267 · IBM#7001


green means CI actually ran on the head commit and passed — gala#610 (25 checks) and kaneo#1851 (21 checks). The rest are genuinely silent: first-time forks sit unapproved until a maintainer clicks, so nothing has run and nothing has been asserted.

🔇 I measured the wrong ref and published the result

I read check-runs on every open PR's merge ref, reasoned that branch protection evaluates the merge ref so that must be the right thing to query, and got zero for all of them. Zero is exactly what a first-time fork's unapproved workflows also return, so "nothing has run" and "I looked in the wrong place" were indistinguishable. I wrote it up as a finding and put it on this page.

Then I needed to answer a different question — does this repo ever run CI for someone who isn't a collaborator? — and built a probe that reads the head SHA. Re-running the same measurement there:

PR head SHA merge ref actually
gala#610 25 checks, 0 failing 0 green
kaneo#1851 21 checks, 0 failing 0 green
prest#1046, OpenSandbox#2012, IBM#7001 2, 2, 1 — all passing 0 green
tirith#267, flumine#864, starnet#45, crawl4ai#2299 0 0 genuinely silent

Two fully green PRs, including the maintainer-endorsed one, and I'd written off all of them. The mistake wasn't carelessness in the reading — it was accepting a plausible justification for a measurement in place of validating it. Branch protection does evaluate the merge ref; that is true, and it is not an argument for querying it to ask "has CI run on my code". Four PRs really are silent, and that part stands.

The probe also had to be thrown away once. Its first version counted dependabot[bot] as evidence of outside contributors, and since dependabot is auto-approved in essentially every repo, it reported cli/cli, gala and kaneo all as "CI runs for forks" — every repository passing, on the strongest possible false evidence. Bots are excluded now. A gate that cannot fail is worse than no gate, because you stop looking.

Three gates now sit in front of scoring, all cheap, all of which were tested against a real failure of my own:

  • Does this repo run CI for a first-time fork? Not "does it have CI" — whether a stranger's PR gets a verdict. Measured empirically against recent human outside-contributor merges, with bots excluded. A repo that won't, can't merge you.
  • Is anyone already on the issue? closed_by_pull_requests alone isn't enough — on gVisor#14966 it reads zero while the issue body names an open PR fixing it. Four signals have to agree.
  • Is a contribution even admissible? AGENTS.md and CLAUDE.md first, then CONTRIBUTING.md, then the PR's own labels. A maintainer bot's unmet-requirements is a recorded verdict, and it outranks my reading of their docs.

The third gate exists because the first two passed and I still got it wrong — see the cli/cli entry above. The one repo that did run my CI is also the one with the strictest entry policy, and passing a CI check is not the same as being allowed in.

And the objective itself changed. Thirteen open PRs with zero merges is a failure state that looks like progress, so fame now carries weight 0.0 in the scorer. A merged PR in a mid-size repo beats an unmerged one in a famous repo, and a maintainer-endorsed task beats both. All three merges so far came from a maintainer's backlog, not an issue feed — which is also why a fresh-sorted sweep found nothing available in roughly 12,800 issues across six attempts. Fresh means a competent maintainer just filed it and is already on it.


Everything here is MIT licensed. Happy to take a bug report on any of it.

Pinned Loading

  1. wippa-uacp wippa-uacp Public

    Universal Agent Communication Protocol — HTTP for AI agents

    Python 1

  2. uacp-interop uacp-interop Public

    Conformance harness for the UACP protocol. Finds real bugs by probing the spec against its own code and against live MCP servers — 6 defects found and fixed, 1 open. pip install uacp-interop

    Python 1

  3. wippa-agent-core wippa-agent-core Public

    An agent framework with metabolic budgeting, graph memory, Governor plan critic, LLM proposer, DSL executor, and Express dashboard

    TypeScript 1

  4. wippa-agent-protect wippa-agent-protect Public

    Sandboxed AI agent execution: clone, scan, build and run untrusted agent repos in Docker with egress filtering and secret masking. Publishes @wippa/core.

    TypeScript 1

  5. wippa-project-manager wippa-project-manager Public

    Agent-native project management: agents are first-class assignees, and "done" is a signed run report rather than a checkbox.

    TypeScript 1