Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 7 additions & 6 deletions CITATION.cff
Original file line number Diff line number Diff line change
Expand Up @@ -85,18 +85,19 @@ references:
description: "arXiv preprint identifier"
url: "https://arxiv.org/abs/2402.17753"
- type: article
title: "Governed Memory: Human-in-the-Loop Co-memorize for Long-Lived Agent Memory"
title: "Governed Memory: A Production Architecture for Multi-Agent Workflows"
authors:
- name: "Governed Memory authors (see arXiv listing)"
- family-names: "Taheri"
given-names: "Hamed"
year: 2026
identifiers:
- type: other
value: "arXiv:2603.17787"
description: "arXiv preprint identifier (per docs/kickoff/FINDINGS-CONTEXT.md; verify on arXiv before publication)"
description: "arXiv preprint identifier"
notes: >-
Citation transcribed from project kickoff notes. The diff-and-approve
workflow in memorywire's governance channel draws on the Co-memorize HITL
pattern surfaced in this paper.
memorywire's human-in-the-loop diff-and-approve governance contrasts with
this line of work, which enforces memory governance automatically for
autonomous multi-agent workflows rather than through a human-approval gate.
- type: software
title: "Model Context Protocol (MCP)"
authors:
Expand Down
6 changes: 5 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -383,7 +383,11 @@ uv pip install pytest pytest-asyncio pytest-cov ruff mypy

Modelled on [MCP](https://modelcontextprotocol.io) (cross-vendor protocol shape), informed by the [LongMemEval](https://arxiv.org/abs/2410.10813), [LoCoMo](https://arxiv.org/abs/2402.17753), and Governed Memory papers, and by the published architecture writeups of mem0, Letta, Cognee, Zep/Graphiti.

The diff-and-approve workflow draws on the Co-memorize HITL pattern surfaced in the *Governed Memory* literature.
The diff-and-approve workflow mirrors code review and change-management gates;
what memorywire adds is its standardization at the wire-format layer. This contrasts
with automated governance for autonomous agents such as
[Governed Memory](https://arxiv.org/abs/2603.17787), which enforces write policy
without a human gate.

## Prior work and naming

Expand Down
4 changes: 2 additions & 2 deletions docs/MCP-RELATIONSHIP.md
Original file line number Diff line number Diff line change
Expand Up @@ -348,5 +348,5 @@ references.
- Cloudflare Web Bot Auth Internet-Draft precedent —
`draft-meunier-web-bot-auth-architecture` (IETF). The
ship-spec-then-propose-upstream pattern memorywire follows.
- "Governed Memory" (arXiv 2603.17787) — the Co-memorize HITL
pattern that informs memorywire's governance channel.
- "Governed Memory" (arXiv 2603.17787) — automated governance for autonomous agents,
the human-gated contrast to memorywire's governance channel.
2 changes: 1 addition & 1 deletion docs/launch-post.md
Original file line number Diff line number Diff line change
Expand Up @@ -64,7 +64,7 @@ This is the piece I most want operators to test, because it has the strongest st

A `remember` with `approval_required: true` stages the row behind a sentinel (`PENDING_APPROVAL_DELETED_AT = -1` in `memories.deleted_at`) so it cannot influence recall until a human approves. The Pro-tier UI (Starlette + HTMX, source-available under FSL 1.1, auto-converts to Apache-2.0 after 2 years) renders the pending memory with a structured diff against the closest live counterpart; the reviewer approves or rejects; the decision is journaled into an append-only audit log keyed by `approved_by`. The same flow covers `forget` and `merge`. An approval-learning loop sits on top: track which patterns the reviewer always approves or rejects and auto-allow after N consistent decisions.

This matters because *memory governance is becoming a compliance problem*, not just a quality-of-life one. EU AI Act transparency / human-oversight, HIPAA audit-trail expectations on systems that persist patient context, SOC2 data-handling — all assume you can answer "what does this system remember, who decided, and when?" Today, with no protocol-level governance primitive across memory frameworks, the answer is "instrument each framework separately." memorywire's channel is one place to instrument and one log to query. The closest published precedent is the Co-memorize HITL pattern from the *Governed Memory* paper (arXiv 2603.17787); memorywire's contribution is shipping it as protocol surface and a working UI.
This matters because *memory governance is becoming a compliance problem*, not just a quality-of-life one. EU AI Act transparency / human-oversight, HIPAA audit-trail expectations on systems that persist patient context, SOC2 data-handling — all assume you can answer "what does this system remember, who decided, and when?" Today, with no protocol-level governance primitive across memory frameworks, the answer is "instrument each framework separately." memorywire's channel is one place to instrument and one log to query. The diff-and-approve discipline itself mirrors code review; memorywire's contribution is shipping it as protocol surface and a working UI. This contrasts with automated memory governance for autonomous agents such as *Governed Memory* (arXiv 2603.17787), which enforces write policy without a human gate.

## Security in one paragraph

Expand Down
2 changes: 1 addition & 1 deletion docs/launch/01-tweet-thread.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,7 @@ Empirical numbers in the paper:
## 4/6

```
The piece I leaned on most: every remember() can stage behind a human-review queue with a structured diff against current memory state. The Co-memorize / Governed Memory pattern, productized.
The piece I leaned on most: every remember() can stage behind a human-review queue with a structured diff against current memory state. Human-in-the-loop governance, productized as protocol surface — unlike automated approaches such as Governed Memory (arXiv 2603.17787), the human stays in the gate.

Full threat model in the paper §6 — 6 adversaries mapped to OWASP + CWE.
```
Expand Down
Binary file modified docs/paper/arxiv-submission.tar.gz
Binary file not shown.
14 changes: 7 additions & 7 deletions docs/paper/arxiv-submission/memorywire-paper.tex
Original file line number Diff line number Diff line change
Expand Up @@ -72,7 +72,7 @@ \subsection{The problem: islanded memory frameworks}

Agent runtimes that maintain memory across sessions are now a category. Open-source frameworks include mem0, Letta (formerly MemGPT), Cognee, Zep/Graphiti, MemoryOS, and MemTensor MemOS; closed commercial offerings include Oracle's AI Agent Memory and the memory layers shipped inside major hosted-agent platforms. What the category has not produced is a shared wire format. Each framework defines its own SDK surface, JSON shape for memory records, embedding-provider integration, taxonomy (or absence of taxonomy) for memory types, and implicit lifecycle for record creation and deletion. The heterogeneity is non-trivial to bridge: mem0 stores records under a \texttt{memories[]} list keyed by \texttt{user\_id} with a heterogeneous \texttt{created\_at} representation; Letta stores archival memory keyed by \texttt{agent\_id} and exposes a \texttt{tags} list as the only structured-metadata sink; Cognee mints internal \texttt{data\_id} UUIDs that are not surfaced through its public \texttt{add} API, making per-record deletion impossible from outside the pipeline; sqlite-vec stores tables keyed by a stable ULID-shaped string; pgvector exposes records through an application-chosen SQL schema. Re-platforming an agent from one framework to another therefore requires a bespoke migrator and field-level losses where the source framework encodes more state than the target's data model holds.

The same heterogeneity means there is no shared \emph{governance} surface. Each framework provides a write API and a read API; none mediate the write with a ``diff against current state, present to a human, commit only on approval'' workflow. The Co-memorize human-in-the-loop pattern, formalized in the \emph{Governed Memory} line of work~\cite{taheri2026governed}, has no production implementation an off-the-shelf agent can drop in. Operators who want auditability over what enters long-term memory must build it themselves and accept that the framework can bypass them.
The same heterogeneity means there is no shared \emph{governance} surface. Each framework provides a write API and a read API; none mediate the write with a ``diff against current state, present to a human, commit only on approval'' workflow. Production governance work such as Taheri's \emph{Governed Memory}~\cite{taheri2026governed} enforces write policy automatically across autonomous agents rather than through such a human-approval gate, so this human-in-the-loop workflow has no production implementation an off-the-shelf agent can drop in. Operators who want auditability over what enters long-term memory must build it themselves and accept that the framework can bypass them.

This is the gap memorywire addresses. It is not ``we need a better retrieval algorithm'' --- the algorithms in the category (vector search, hybrid lexical-semantic RRF fusion, graph hop boosts, FSM-encoded procedures, STM/LTM consolidation) are well understood. It is ``we need a shared protocol so any client can talk to any backend, any agent can carry its memory across runtimes, and any write can be diffed and approved.'' Structurally it is the gap MCP closed for \emph{tool use}, applied to \emph{memory}.

Expand All @@ -84,7 +84,7 @@ \subsection{Contributions}
\begin{itemize}[leftmargin=*]
\item \textbf{C1.} A wire format for five memory operations over four memory types, expressed as JSON Schema 2020-12~\cite{jsonschema_2020_12} (\texttt{docs/spec/v0.md}). The operations are \texttt{remember}, \texttt{recall}, \texttt{forget}, \texttt{merge}, \texttt{expire}; the types are \texttt{semantic}, \texttt{episodic}, \texttt{procedural}, \texttt{emotional}. The schemas are vendor-neutral, transport-agnostic, and explicitly versioned with a breaking-change policy through v0.5.
\item \textbf{C2.} A reference implementation in Python 3.11+ with five production-backend adapters (sqlite-vec, mem0, Letta, Cognee, pgvector), all implementing a single \texttt{MemoryStore} Protocol. The reference includes a memory router that fans operations across $N$ stores in parallel and fuses recall results via Reciprocal Rank Fusion ($k=60$)~\cite{cormack2009rrf} with an optional one-hop graph boost, plus a tolerant partial-failure model where a single rogue or unavailable backend cannot crash the operation.
\item \textbf{C3.} A governance UI implementing the Co-memorize diff-and-approve pattern over \texttt{remember}, \texttt{forget}, and \texttt{merge}. Writes flagged \texttt{approval\_required} are staged behind a \texttt{PENDING\_APPROVAL\_DELETED\_AT = -1} sentinel and remain invisible to \texttt{recall} until a reviewer commits or rejects them through the UI. The same audit log is the single source of truth for all governance and mutation events.
\item \textbf{C3.} A governance UI implementing a human-in-the-loop diff-and-approve workflow over \texttt{remember}, \texttt{forget}, and \texttt{merge}. Writes flagged \texttt{approval\_required} are staged behind a \texttt{PENDING\_APPROVAL\_DELETED\_AT = -1} sentinel and remain invisible to \texttt{recall} until a reviewer commits or rejects them through the UI. The same audit log is the single source of truth for all governance and mutation events.
\item \textbf{C4.} An empirical evaluation comprising (a)~a microbenchmark on 100 hand-authored facts $\times$ 50 labelled queries against a real sentence-transformer embedder, (b)~an adversarial-fusion experiment that sweeps a 1-of-$N$ rank-0 injection attack across three fusion algorithms (RRF, MAX, weighted), and (c)~a cross-adapter conformance suite of 16 protocol-invariant scenarios run against all five shipped adapters (68 PASS / 12 SKIP / 0 FAIL out of 80 cells).
\item \textbf{C5.} A six-adversary threat model with line-level mitigation citations into the reference implementation, plus an open-data artifact (\texttt{docs/adversarial-results.\{rrf,max,weighted\}.json}, the labelled microbench corpus, the conformance scenario list) sufficient to reproduce every empirical claim without re-running paid evaluators.
\end{itemize}
Expand All @@ -97,7 +97,7 @@ \subsection{Limitations, front-loaded}
\subsection{Paper roadmap}
\label{sec:intro-roadmap}

\Cref{sec:related} places memorywire against prior work in agent memory frameworks, cross-vendor protocols (particularly MCP), and the Governed Memory line. \Cref{sec:spec} specifies the wire format: operations, types, the \texttt{MemoryStore} Protocol, router semantics, and the governance channel. \Cref{sec:impl} describes the reference implementation including the five backend adapters, the procedural-memory FSM backend, the STM$\leftrightarrow$LTM transformer, and the governance UI. \Cref{sec:eval} reports the empirical evaluation. \Cref{sec:threats} is the threat model. \Cref{sec:mcp} details the relationship to MCP. \Cref{sec:future} and \Cref{sec:conclusion} lay out future work and the bet we are making.
\Cref{sec:related} places memorywire against prior work in agent memory frameworks, cross-vendor protocols (particularly MCP), and prior production work on memory governance. \Cref{sec:spec} specifies the wire format: operations, types, the \texttt{MemoryStore} Protocol, router semantics, and the governance channel. \Cref{sec:impl} describes the reference implementation including the five backend adapters, the procedural-memory FSM backend, the STM$\leftrightarrow$LTM transformer, and the governance UI. \Cref{sec:eval} reports the empirical evaluation. \Cref{sec:threats} is the threat model. \Cref{sec:mcp} details the relationship to MCP. \Cref{sec:future} and \Cref{sec:conclusion} lay out future work and the bet we are making.

\section{Background and Related Work}
\label{sec:related}
Expand Down Expand Up @@ -133,9 +133,9 @@ \subsection{Human-memory taxonomy}

\subsection{Human-in-the-loop approval for agent actions}

The governance channel in memorywire implements the Co-memorize diff-and-approve pattern formalized in the Governed Memory line of work~\cite{taheri2026governed}. The pattern is: when an agent proposes to write a memory, the system computes a structured diff between the proposed write and the current state, presents the diff to a human reviewer, and commits the write only on approval. The pattern generalizes to any state mutation; memorywire applies it to \texttt{remember}, \texttt{forget}, and \texttt{merge}, and excludes \texttt{recall} and \texttt{expire} from the default approval surface (with \texttt{recall} flagged for v0.2 reconsideration; see \Cref{sec:threats}).
The governance channel in memorywire implements a human-in-the-loop diff-and-approve workflow: when an agent proposes to write a memory, the system computes a structured diff between the proposed write and the current state, presents the diff to a human reviewer, and commits the write only on approval. This contrasts with production governance for autonomous agents such as Taheri's \emph{Governed Memory}~\cite{taheri2026governed}, which enforces write policy automatically rather than routing writes through a human gate. The workflow generalizes to any state mutation; memorywire applies it to \texttt{remember}, \texttt{forget}, and \texttt{merge}, and excludes \texttt{recall} and \texttt{expire} from the default approval surface (with \texttt{recall} flagged for v0.2 reconsideration; see \Cref{sec:threats}).

The Co-memorize pattern is not novel to this paper. What is new is its standardization at the wire-format layer: memorywire defines a \texttt{governance} JSON schema for the diff-and-approve message and ships a reference UI that any backend adapter inherits transparently. An agent calling \texttt{remember(content="\ldots", approval\_required=true)} gets the governance flow regardless of which of the five backends actually stores the row.
Diff-and-approve as a review discipline is not itself novel---it mirrors code review and change-management gates. What is new here is its standardization at the wire-format layer: memorywire defines a \texttt{governance} JSON schema for the diff-and-approve message and ships a reference UI that any backend adapter inherits transparently. An agent calling \texttt{remember(content="\ldots", approval\_required=true)} gets the governance flow regardless of which of the five backends actually stores the row.

\section{The memorywire Wire Format}
\label{sec:spec}
Expand Down Expand Up @@ -307,7 +307,7 @@ \subsection{The STM$\leftrightarrow$LTM transformer}
\subsection{The governance UI}
\label{sec:impl-ui}

The governance UI (\texttt{ui/src/memorywire\_ui/}) is a Starlette server with HTMX-driven templates. It shares the sqlite-vec adapter's SQLite database, so a reviewer sees pending writes as rows with the \texttt{PENDING\_APPROVAL\_DELETED\_AT = -1} sentinel in \texttt{deleted\_at} and a \texttt{pending: yes} badge in the UI. Reviewers can approve (clear the sentinel; the row becomes live), reject (hard-delete or soft-delete depending on the policy), or apply a Co-memorize transformation --- typically a \texttt{merge} against an existing canonical row or a \texttt{forget} of a similar row that the new write supersedes.
The governance UI (\texttt{ui/src/memorywire\_ui/}) is a Starlette server with HTMX-driven templates. It shares the sqlite-vec adapter's SQLite database, so a reviewer sees pending writes as rows with the \texttt{PENDING\_APPROVAL\_DELETED\_AT = -1} sentinel in \texttt{deleted\_at} and a \texttt{pending: yes} badge in the UI. Reviewers can approve (clear the sentinel; the row becomes live), reject (hard-delete or soft-delete depending on the policy), or apply a reconciling transformation --- typically a \texttt{merge} against an existing canonical row or a \texttt{forget} of a similar row that the new write supersedes.

Authentication is opt-in via the \texttt{MEMORYWIRE\_UI\_TOKEN} environment variable. When set, every request requires either an \texttt{Authorization: Bearer <token>} header or an \texttt{memorywire\_ui\_session} cookie; comparison uses \texttt{hmac.compare\_digest} for constant-time matching. CSRF is enforced through a double-submit-cookie pattern signed with HMAC-SHA256 over \texttt{nonce.ts} with a 24-hour TTL. Without \texttt{MEMORYWIRE\_UI\_TOKEN} the UI is unauthenticated; binding a non-loopback host without a token fires a stderr warning on boot (\texttt{ui/src/memorywire\_ui/middleware.py:55--66}).

Expand Down Expand Up @@ -600,7 +600,7 @@ \section{Conclusion}

\section*{Acknowledgments}

We acknowledge the open-source projects memorywire composes with: mem0, Letta (formerly MemGPT), Cognee, Zep/Graphiti, MemoryOS, MemTensor MemOS, sqlite-vec, pgvector, \texttt{pytransitions}, Starlette, HTMX, sentence-transformers, and the Model Context Protocol community. The Co-memorize diff-and-approve pattern draws on the Governed Memory line of work. RRF as a fusion primitive is from Cormack, Clarke, and B{\"u}ttcher. The human-memory taxonomy mapping follows the cognitive-science literature established by Tulving and Squire.
We acknowledge the open-source projects memorywire composes with: mem0, Letta (formerly MemGPT), Cognee, Zep/Graphiti, MemoryOS, MemTensor MemOS, sqlite-vec, pgvector, \texttt{pytransitions}, Starlette, HTMX, sentence-transformers, and the Model Context Protocol community. RRF as a fusion primitive is from Cormack, Clarke, and B{\"u}ttcher. The human-memory taxonomy mapping follows the cognitive-science literature established by Tulving and Squire.

\bibliographystyle{plain}
\bibliography{memorywire}
Expand Down
Loading
Loading