diff --git a/CITATION.cff b/CITATION.cff index 6873303..0228de9 100644 --- a/CITATION.cff +++ b/CITATION.cff @@ -85,18 +85,19 @@ references: description: "arXiv preprint identifier" url: "https://arxiv.org/abs/2402.17753" - type: article - title: "Governed Memory: Human-in-the-Loop Co-memorize for Long-Lived Agent Memory" + title: "Governed Memory: A Production Architecture for Multi-Agent Workflows" authors: - - name: "Governed Memory authors (see arXiv listing)" + - family-names: "Taheri" + given-names: "Hamed" year: 2026 identifiers: - type: other value: "arXiv:2603.17787" - description: "arXiv preprint identifier (per docs/kickoff/FINDINGS-CONTEXT.md; verify on arXiv before publication)" + description: "arXiv preprint identifier" notes: >- - Citation transcribed from project kickoff notes. The diff-and-approve - workflow in memorywire's governance channel draws on the Co-memorize HITL - pattern surfaced in this paper. + memorywire's human-in-the-loop diff-and-approve governance contrasts with + this line of work, which enforces memory governance automatically for + autonomous multi-agent workflows rather than through a human-approval gate. - type: software title: "Model Context Protocol (MCP)" authors: diff --git a/README.md b/README.md index ad859c6..9ca550b 100644 --- a/README.md +++ b/README.md @@ -383,7 +383,11 @@ uv pip install pytest pytest-asyncio pytest-cov ruff mypy Modelled on [MCP](https://modelcontextprotocol.io) (cross-vendor protocol shape), informed by the [LongMemEval](https://arxiv.org/abs/2410.10813), [LoCoMo](https://arxiv.org/abs/2402.17753), and Governed Memory papers, and by the published architecture writeups of mem0, Letta, Cognee, Zep/Graphiti. -The diff-and-approve workflow draws on the Co-memorize HITL pattern surfaced in the *Governed Memory* literature. +The diff-and-approve workflow mirrors code review and change-management gates; +what memorywire adds is its standardization at the wire-format layer. This contrasts +with automated governance for autonomous agents such as +[Governed Memory](https://arxiv.org/abs/2603.17787), which enforces write policy +without a human gate. ## Prior work and naming diff --git a/docs/MCP-RELATIONSHIP.md b/docs/MCP-RELATIONSHIP.md index b3cbdcf..ce8d7f8 100644 --- a/docs/MCP-RELATIONSHIP.md +++ b/docs/MCP-RELATIONSHIP.md @@ -348,5 +348,5 @@ references. - Cloudflare Web Bot Auth Internet-Draft precedent — `draft-meunier-web-bot-auth-architecture` (IETF). The ship-spec-then-propose-upstream pattern memorywire follows. -- "Governed Memory" (arXiv 2603.17787) — the Co-memorize HITL - pattern that informs memorywire's governance channel. +- "Governed Memory" (arXiv 2603.17787) — automated governance for autonomous agents, + the human-gated contrast to memorywire's governance channel. diff --git a/docs/launch-post.md b/docs/launch-post.md index 8f08c27..88dd003 100644 --- a/docs/launch-post.md +++ b/docs/launch-post.md @@ -64,7 +64,7 @@ This is the piece I most want operators to test, because it has the strongest st A `remember` with `approval_required: true` stages the row behind a sentinel (`PENDING_APPROVAL_DELETED_AT = -1` in `memories.deleted_at`) so it cannot influence recall until a human approves. The Pro-tier UI (Starlette + HTMX, source-available under FSL 1.1, auto-converts to Apache-2.0 after 2 years) renders the pending memory with a structured diff against the closest live counterpart; the reviewer approves or rejects; the decision is journaled into an append-only audit log keyed by `approved_by`. The same flow covers `forget` and `merge`. An approval-learning loop sits on top: track which patterns the reviewer always approves or rejects and auto-allow after N consistent decisions. -This matters because *memory governance is becoming a compliance problem*, not just a quality-of-life one. EU AI Act transparency / human-oversight, HIPAA audit-trail expectations on systems that persist patient context, SOC2 data-handling — all assume you can answer "what does this system remember, who decided, and when?" Today, with no protocol-level governance primitive across memory frameworks, the answer is "instrument each framework separately." memorywire's channel is one place to instrument and one log to query. The closest published precedent is the Co-memorize HITL pattern from the *Governed Memory* paper (arXiv 2603.17787); memorywire's contribution is shipping it as protocol surface and a working UI. +This matters because *memory governance is becoming a compliance problem*, not just a quality-of-life one. EU AI Act transparency / human-oversight, HIPAA audit-trail expectations on systems that persist patient context, SOC2 data-handling — all assume you can answer "what does this system remember, who decided, and when?" Today, with no protocol-level governance primitive across memory frameworks, the answer is "instrument each framework separately." memorywire's channel is one place to instrument and one log to query. The diff-and-approve discipline itself mirrors code review; memorywire's contribution is shipping it as protocol surface and a working UI. This contrasts with automated memory governance for autonomous agents such as *Governed Memory* (arXiv 2603.17787), which enforces write policy without a human gate. ## Security in one paragraph diff --git a/docs/launch/01-tweet-thread.md b/docs/launch/01-tweet-thread.md index 4c023bd..049c952 100644 --- a/docs/launch/01-tweet-thread.md +++ b/docs/launch/01-tweet-thread.md @@ -44,7 +44,7 @@ Empirical numbers in the paper: ## 4/6 ``` -The piece I leaned on most: every remember() can stage behind a human-review queue with a structured diff against current memory state. The Co-memorize / Governed Memory pattern, productized. +The piece I leaned on most: every remember() can stage behind a human-review queue with a structured diff against current memory state. Human-in-the-loop governance, productized as protocol surface — unlike automated approaches such as Governed Memory (arXiv 2603.17787), the human stays in the gate. Full threat model in the paper §6 — 6 adversaries mapped to OWASP + CWE. ``` diff --git a/docs/paper/arxiv-submission.tar.gz b/docs/paper/arxiv-submission.tar.gz index b5e8e51..94321e5 100644 Binary files a/docs/paper/arxiv-submission.tar.gz and b/docs/paper/arxiv-submission.tar.gz differ diff --git a/docs/paper/arxiv-submission/memorywire-paper.tex b/docs/paper/arxiv-submission/memorywire-paper.tex index 30b0615..19be98d 100644 --- a/docs/paper/arxiv-submission/memorywire-paper.tex +++ b/docs/paper/arxiv-submission/memorywire-paper.tex @@ -72,7 +72,7 @@ \subsection{The problem: islanded memory frameworks} Agent runtimes that maintain memory across sessions are now a category. Open-source frameworks include mem0, Letta (formerly MemGPT), Cognee, Zep/Graphiti, MemoryOS, and MemTensor MemOS; closed commercial offerings include Oracle's AI Agent Memory and the memory layers shipped inside major hosted-agent platforms. What the category has not produced is a shared wire format. Each framework defines its own SDK surface, JSON shape for memory records, embedding-provider integration, taxonomy (or absence of taxonomy) for memory types, and implicit lifecycle for record creation and deletion. The heterogeneity is non-trivial to bridge: mem0 stores records under a \texttt{memories[]} list keyed by \texttt{user\_id} with a heterogeneous \texttt{created\_at} representation; Letta stores archival memory keyed by \texttt{agent\_id} and exposes a \texttt{tags} list as the only structured-metadata sink; Cognee mints internal \texttt{data\_id} UUIDs that are not surfaced through its public \texttt{add} API, making per-record deletion impossible from outside the pipeline; sqlite-vec stores tables keyed by a stable ULID-shaped string; pgvector exposes records through an application-chosen SQL schema. Re-platforming an agent from one framework to another therefore requires a bespoke migrator and field-level losses where the source framework encodes more state than the target's data model holds. -The same heterogeneity means there is no shared \emph{governance} surface. Each framework provides a write API and a read API; none mediate the write with a ``diff against current state, present to a human, commit only on approval'' workflow. The Co-memorize human-in-the-loop pattern, formalized in the \emph{Governed Memory} line of work~\cite{taheri2026governed}, has no production implementation an off-the-shelf agent can drop in. Operators who want auditability over what enters long-term memory must build it themselves and accept that the framework can bypass them. +The same heterogeneity means there is no shared \emph{governance} surface. Each framework provides a write API and a read API; none mediate the write with a ``diff against current state, present to a human, commit only on approval'' workflow. Production governance work such as Taheri's \emph{Governed Memory}~\cite{taheri2026governed} enforces write policy automatically across autonomous agents rather than through such a human-approval gate, so this human-in-the-loop workflow has no production implementation an off-the-shelf agent can drop in. Operators who want auditability over what enters long-term memory must build it themselves and accept that the framework can bypass them. This is the gap memorywire addresses. It is not ``we need a better retrieval algorithm'' --- the algorithms in the category (vector search, hybrid lexical-semantic RRF fusion, graph hop boosts, FSM-encoded procedures, STM/LTM consolidation) are well understood. It is ``we need a shared protocol so any client can talk to any backend, any agent can carry its memory across runtimes, and any write can be diffed and approved.'' Structurally it is the gap MCP closed for \emph{tool use}, applied to \emph{memory}. @@ -84,7 +84,7 @@ \subsection{Contributions} \begin{itemize}[leftmargin=*] \item \textbf{C1.} A wire format for five memory operations over four memory types, expressed as JSON Schema 2020-12~\cite{jsonschema_2020_12} (\texttt{docs/spec/v0.md}). The operations are \texttt{remember}, \texttt{recall}, \texttt{forget}, \texttt{merge}, \texttt{expire}; the types are \texttt{semantic}, \texttt{episodic}, \texttt{procedural}, \texttt{emotional}. The schemas are vendor-neutral, transport-agnostic, and explicitly versioned with a breaking-change policy through v0.5. \item \textbf{C2.} A reference implementation in Python 3.11+ with five production-backend adapters (sqlite-vec, mem0, Letta, Cognee, pgvector), all implementing a single \texttt{MemoryStore} Protocol. The reference includes a memory router that fans operations across $N$ stores in parallel and fuses recall results via Reciprocal Rank Fusion ($k=60$)~\cite{cormack2009rrf} with an optional one-hop graph boost, plus a tolerant partial-failure model where a single rogue or unavailable backend cannot crash the operation. - \item \textbf{C3.} A governance UI implementing the Co-memorize diff-and-approve pattern over \texttt{remember}, \texttt{forget}, and \texttt{merge}. Writes flagged \texttt{approval\_required} are staged behind a \texttt{PENDING\_APPROVAL\_DELETED\_AT = -1} sentinel and remain invisible to \texttt{recall} until a reviewer commits or rejects them through the UI. The same audit log is the single source of truth for all governance and mutation events. + \item \textbf{C3.} A governance UI implementing a human-in-the-loop diff-and-approve workflow over \texttt{remember}, \texttt{forget}, and \texttt{merge}. Writes flagged \texttt{approval\_required} are staged behind a \texttt{PENDING\_APPROVAL\_DELETED\_AT = -1} sentinel and remain invisible to \texttt{recall} until a reviewer commits or rejects them through the UI. The same audit log is the single source of truth for all governance and mutation events. \item \textbf{C4.} An empirical evaluation comprising (a)~a microbenchmark on 100 hand-authored facts $\times$ 50 labelled queries against a real sentence-transformer embedder, (b)~an adversarial-fusion experiment that sweeps a 1-of-$N$ rank-0 injection attack across three fusion algorithms (RRF, MAX, weighted), and (c)~a cross-adapter conformance suite of 16 protocol-invariant scenarios run against all five shipped adapters (68 PASS / 12 SKIP / 0 FAIL out of 80 cells). \item \textbf{C5.} A six-adversary threat model with line-level mitigation citations into the reference implementation, plus an open-data artifact (\texttt{docs/adversarial-results.\{rrf,max,weighted\}.json}, the labelled microbench corpus, the conformance scenario list) sufficient to reproduce every empirical claim without re-running paid evaluators. \end{itemize} @@ -97,7 +97,7 @@ \subsection{Limitations, front-loaded} \subsection{Paper roadmap} \label{sec:intro-roadmap} -\Cref{sec:related} places memorywire against prior work in agent memory frameworks, cross-vendor protocols (particularly MCP), and the Governed Memory line. \Cref{sec:spec} specifies the wire format: operations, types, the \texttt{MemoryStore} Protocol, router semantics, and the governance channel. \Cref{sec:impl} describes the reference implementation including the five backend adapters, the procedural-memory FSM backend, the STM$\leftrightarrow$LTM transformer, and the governance UI. \Cref{sec:eval} reports the empirical evaluation. \Cref{sec:threats} is the threat model. \Cref{sec:mcp} details the relationship to MCP. \Cref{sec:future} and \Cref{sec:conclusion} lay out future work and the bet we are making. +\Cref{sec:related} places memorywire against prior work in agent memory frameworks, cross-vendor protocols (particularly MCP), and prior production work on memory governance. \Cref{sec:spec} specifies the wire format: operations, types, the \texttt{MemoryStore} Protocol, router semantics, and the governance channel. \Cref{sec:impl} describes the reference implementation including the five backend adapters, the procedural-memory FSM backend, the STM$\leftrightarrow$LTM transformer, and the governance UI. \Cref{sec:eval} reports the empirical evaluation. \Cref{sec:threats} is the threat model. \Cref{sec:mcp} details the relationship to MCP. \Cref{sec:future} and \Cref{sec:conclusion} lay out future work and the bet we are making. \section{Background and Related Work} \label{sec:related} @@ -133,9 +133,9 @@ \subsection{Human-memory taxonomy} \subsection{Human-in-the-loop approval for agent actions} -The governance channel in memorywire implements the Co-memorize diff-and-approve pattern formalized in the Governed Memory line of work~\cite{taheri2026governed}. The pattern is: when an agent proposes to write a memory, the system computes a structured diff between the proposed write and the current state, presents the diff to a human reviewer, and commits the write only on approval. The pattern generalizes to any state mutation; memorywire applies it to \texttt{remember}, \texttt{forget}, and \texttt{merge}, and excludes \texttt{recall} and \texttt{expire} from the default approval surface (with \texttt{recall} flagged for v0.2 reconsideration; see \Cref{sec:threats}). +The governance channel in memorywire implements a human-in-the-loop diff-and-approve workflow: when an agent proposes to write a memory, the system computes a structured diff between the proposed write and the current state, presents the diff to a human reviewer, and commits the write only on approval. This contrasts with production governance for autonomous agents such as Taheri's \emph{Governed Memory}~\cite{taheri2026governed}, which enforces write policy automatically rather than routing writes through a human gate. The workflow generalizes to any state mutation; memorywire applies it to \texttt{remember}, \texttt{forget}, and \texttt{merge}, and excludes \texttt{recall} and \texttt{expire} from the default approval surface (with \texttt{recall} flagged for v0.2 reconsideration; see \Cref{sec:threats}). -The Co-memorize pattern is not novel to this paper. What is new is its standardization at the wire-format layer: memorywire defines a \texttt{governance} JSON schema for the diff-and-approve message and ships a reference UI that any backend adapter inherits transparently. An agent calling \texttt{remember(content="\ldots", approval\_required=true)} gets the governance flow regardless of which of the five backends actually stores the row. +Diff-and-approve as a review discipline is not itself novel---it mirrors code review and change-management gates. What is new here is its standardization at the wire-format layer: memorywire defines a \texttt{governance} JSON schema for the diff-and-approve message and ships a reference UI that any backend adapter inherits transparently. An agent calling \texttt{remember(content="\ldots", approval\_required=true)} gets the governance flow regardless of which of the five backends actually stores the row. \section{The memorywire Wire Format} \label{sec:spec} @@ -307,7 +307,7 @@ \subsection{The STM$\leftrightarrow$LTM transformer} \subsection{The governance UI} \label{sec:impl-ui} -The governance UI (\texttt{ui/src/memorywire\_ui/}) is a Starlette server with HTMX-driven templates. It shares the sqlite-vec adapter's SQLite database, so a reviewer sees pending writes as rows with the \texttt{PENDING\_APPROVAL\_DELETED\_AT = -1} sentinel in \texttt{deleted\_at} and a \texttt{pending: yes} badge in the UI. Reviewers can approve (clear the sentinel; the row becomes live), reject (hard-delete or soft-delete depending on the policy), or apply a Co-memorize transformation --- typically a \texttt{merge} against an existing canonical row or a \texttt{forget} of a similar row that the new write supersedes. +The governance UI (\texttt{ui/src/memorywire\_ui/}) is a Starlette server with HTMX-driven templates. It shares the sqlite-vec adapter's SQLite database, so a reviewer sees pending writes as rows with the \texttt{PENDING\_APPROVAL\_DELETED\_AT = -1} sentinel in \texttt{deleted\_at} and a \texttt{pending: yes} badge in the UI. Reviewers can approve (clear the sentinel; the row becomes live), reject (hard-delete or soft-delete depending on the policy), or apply a reconciling transformation --- typically a \texttt{merge} against an existing canonical row or a \texttt{forget} of a similar row that the new write supersedes. Authentication is opt-in via the \texttt{MEMORYWIRE\_UI\_TOKEN} environment variable. When set, every request requires either an \texttt{Authorization: Bearer } header or an \texttt{memorywire\_ui\_session} cookie; comparison uses \texttt{hmac.compare\_digest} for constant-time matching. CSRF is enforced through a double-submit-cookie pattern signed with HMAC-SHA256 over \texttt{nonce.ts} with a 24-hour TTL. Without \texttt{MEMORYWIRE\_UI\_TOKEN} the UI is unauthenticated; binding a non-loopback host without a token fires a stderr warning on boot (\texttt{ui/src/memorywire\_ui/middleware.py:55--66}). @@ -600,7 +600,7 @@ \section{Conclusion} \section*{Acknowledgments} -We acknowledge the open-source projects memorywire composes with: mem0, Letta (formerly MemGPT), Cognee, Zep/Graphiti, MemoryOS, MemTensor MemOS, sqlite-vec, pgvector, \texttt{pytransitions}, Starlette, HTMX, sentence-transformers, and the Model Context Protocol community. The Co-memorize diff-and-approve pattern draws on the Governed Memory line of work. RRF as a fusion primitive is from Cormack, Clarke, and B{\"u}ttcher. The human-memory taxonomy mapping follows the cognitive-science literature established by Tulving and Squire. +We acknowledge the open-source projects memorywire composes with: mem0, Letta (formerly MemGPT), Cognee, Zep/Graphiti, MemoryOS, MemTensor MemOS, sqlite-vec, pgvector, \texttt{pytransitions}, Starlette, HTMX, sentence-transformers, and the Model Context Protocol community. RRF as a fusion primitive is from Cormack, Clarke, and B{\"u}ttcher. The human-memory taxonomy mapping follows the cognitive-science literature established by Tulving and Squire. \bibliographystyle{plain} \bibliography{memorywire} diff --git a/docs/paper/arxiv-submission/memorywire.bib b/docs/paper/arxiv-submission/memorywire.bib index 589285f..68f334d 100644 --- a/docs/paper/arxiv-submission/memorywire.bib +++ b/docs/paper/arxiv-submission/memorywire.bib @@ -39,7 +39,8 @@ @misc{packer2023memgpt year = {2023}, eprint = {2310.08560}, archivePrefix = {arXiv}, - primaryClass = {cs.AI} + primaryClass = {cs.AI}, + note = {arXiv:2310.08560} } @misc{wu2024longmemeval, @@ -48,7 +49,8 @@ @misc{wu2024longmemeval year = {2024}, eprint = {2410.10813}, archivePrefix = {arXiv}, - primaryClass = {cs.CL} + primaryClass = {cs.CL}, + note = {arXiv:2410.10813} } @misc{maharana2024locomo, @@ -57,21 +59,22 @@ @misc{maharana2024locomo year = {2024}, eprint = {2402.17753}, archivePrefix = {arXiv}, - primaryClass = {cs.CL} + primaryClass = {cs.CL}, + note = {arXiv:2402.17753} } % Verified via arXiv listing 2026-05-28: arXiv:2603.17787 resolves to a real % paper by Hamed Taheri titled "Governed Memory: A Production Architecture % for Multi-Agent Workflows" (submitted 2026-03-18). The 2603 YYMM prefix % (March 2026) is unusual but valid for an arXiv submission of that month. -% TODO verify arXiv ID before submission (re-check listing close to camera-ready). @misc{taheri2026governed, author = {Taheri, Hamed}, title = {Governed Memory: A Production Architecture for Multi-Agent Workflows}, year = {2026}, eprint = {2603.17787}, archivePrefix = {arXiv}, - primaryClass = {cs.MA} + primaryClass = {cs.AI}, + note = {arXiv:2603.17787} } @misc{chhikara2025mem0, @@ -80,7 +83,8 @@ @misc{chhikara2025mem0 year = {2025}, eprint = {2504.19413}, archivePrefix = {arXiv}, - primaryClass = {cs.AI} + primaryClass = {cs.AI}, + note = {arXiv:2504.19413} } @misc{mcp_spec_2025, diff --git a/docs/paper/memorywire-paper.md b/docs/paper/memorywire-paper.md index f3d263b..17210f1 100644 --- a/docs/paper/memorywire-paper.md +++ b/docs/paper/memorywire-paper.md @@ -19,7 +19,7 @@ Agent-memory frameworks --- mem0, Letta/MemGPT, Cognee, Zep/Graphiti, MemoryOS, Agent runtimes that maintain memory across sessions are now a category. Open-source frameworks include mem0, Letta (formerly MemGPT), Cognee, Zep/Graphiti, MemoryOS, and MemTensor MemOS; closed commercial offerings include Oracle's AI Agent Memory and the memory layers shipped inside major hosted-agent platforms. What the category has not produced is a shared wire format. Each framework defines its own SDK surface, JSON shape for memory records, embedding-provider integration, taxonomy (or absence of taxonomy) for memory types, and implicit lifecycle for record creation and deletion. The heterogeneity is non-trivial to bridge: mem0 stores records under a `memories[]` list keyed by `user_id` with a heterogeneous `created_at` representation; Letta stores archival memory keyed by `agent_id` and exposes a `tags` list as the only structured-metadata sink; Cognee mints internal `data_id` UUIDs that are not surfaced through its public `add` API, making per-record deletion impossible from outside the pipeline; sqlite-vec stores tables keyed by a stable ULID-shaped string; pgvector exposes records through an application-chosen SQL schema. Re-platforming an agent from one framework to another therefore requires a bespoke migrator and field-level losses where the source framework encodes more state than the target's data model holds. -The same heterogeneity means there is no shared *governance* surface. Each framework provides a write API and a read API; none mediate the write with a "diff against current state, present to a human, commit only on approval" workflow. The Co-memorize human-in-the-loop pattern, formalized in the *Governed Memory* line of work, has no production implementation an off-the-shelf agent can drop in. Operators who want auditability over what enters long-term memory must build it themselves and accept that the framework can bypass them. +The same heterogeneity means there is no shared *governance* surface. Each framework provides a write API and a read API; none mediate the write with a "diff against current state, present to a human, commit only on approval" workflow. Production governance work such as Taheri's *Governed Memory* (arXiv:2603.17787) enforces write policy automatically across autonomous agents rather than through such a human-approval gate, so this human-in-the-loop workflow has no production implementation an off-the-shelf agent can drop in. Operators who want auditability over what enters long-term memory must build it themselves and accept that the framework can bypass them. This is the gap memorywire addresses. It is not "we need a better retrieval algorithm" — the algorithms in the category (vector search, hybrid lexical-semantic RRF fusion, graph hop boosts, FSM-encoded procedures, STM/LTM consolidation) are well understood. It is "we need a shared protocol so any client can talk to any backend, any agent can carry its memory across runtimes, and any write can be diffed and approved." Structurally it is the gap MCP closed for *tool use*, applied to *memory*. @@ -29,7 +29,7 @@ This paper makes five contributions: - **C1.** A wire format for five memory operations over four memory types, expressed as JSON Schema 2020-12 (`docs/spec/v0.md`). The operations are `remember`, `recall`, `forget`, `merge`, `expire`; the types are `semantic`, `episodic`, `procedural`, `emotional`. The schemas are vendor-neutral, transport-agnostic (REST-friendly request/response idiom that translates mechanically to JSON-RPC), and explicitly versioned with a breaking-change policy through v0.5. - **C2.** A reference implementation in Python 3.11+ with five production-backend adapters (sqlite-vec, mem0, Letta, Cognee, pgvector), all implementing a single `MemoryStore` Protocol. The reference includes a memory router that fans operations across N stores in parallel and fuses recall results via Reciprocal Rank Fusion (k=60) with an optional one-hop graph boost, plus a tolerant partial-failure model where a single rogue or unavailable backend cannot crash the operation. -- **C3.** A governance UI implementing the Co-memorize diff-and-approve pattern over `remember`, `forget`, and `merge`. Writes flagged `approval_required` are staged behind a `PENDING_APPROVAL_DELETED_AT = -1` sentinel and remain invisible to `recall` until a reviewer commits or rejects them through the UI. The same audit log is the single source of truth for all governance and mutation events. +- **C3.** A governance UI implementing a human-in-the-loop diff-and-approve workflow over `remember`, `forget`, and `merge`. Writes flagged `approval_required` are staged behind a `PENDING_APPROVAL_DELETED_AT = -1` sentinel and remain invisible to `recall` until a reviewer commits or rejects them through the UI. The same audit log is the single source of truth for all governance and mutation events. - **C4.** An empirical evaluation comprising (a) a microbenchmark on 100 hand-authored facts × 50 labelled queries against a real sentence-transformer embedder, (b) an adversarial-fusion experiment that sweeps a 1-of-N rank-0 injection attack across three fusion algorithms (RRF, MAX, weighted), and (c) a cross-adapter conformance suite of 16 protocol-invariant scenarios run against all five shipped adapters (68 PASS / 12 SKIP / 0 FAIL out of 80 cells). - **C5.** A six-adversary threat model with line-level mitigation citations into the reference implementation, plus an open-data artifact (`docs/adversarial-results.{rrf,max,weighted}.json`, the labelled microbench corpus, the conformance scenario list) sufficient to reproduce every empirical claim without re-running paid evaluators. @@ -39,7 +39,7 @@ We state up front what this paper is *not*. It is not a new algorithm: RRF is fr ### 1.4 Paper roadmap -Section 2 places memorywire against prior work in agent memory frameworks, cross-vendor protocols (particularly MCP), and the Governed Memory line. Section 3 specifies the wire format: operations, types, the `MemoryStore` Protocol, router semantics, and the governance channel. Section 4 describes the reference implementation including the five backend adapters, the procedural-memory FSM backend, the STM↔LTM transformer, and the governance UI. Section 5 reports the empirical evaluation. Section 6 is the threat model. Section 7 details the relationship to MCP. Sections 8 and 9 lay out future work and the bet we are making. +Section 2 places memorywire against prior work in agent memory frameworks, cross-vendor protocols (particularly MCP), and prior production work on memory governance. Section 3 specifies the wire format: operations, types, the `MemoryStore` Protocol, router semantics, and the governance channel. Section 4 describes the reference implementation including the five backend adapters, the procedural-memory FSM backend, the STM↔LTM transformer, and the governance UI. Section 5 reports the empirical evaluation. Section 6 is the threat model. Section 7 details the relationship to MCP. Sections 8 and 9 lay out future work and the bet we are making. ## 2. Background and Related Work @@ -75,9 +75,9 @@ We do not claim that the four-type taxonomy is the *correct* one in any deep sen ### 2.5 Human-in-the-loop approval for agent actions -The governance channel in memorywire implements the Co-memorize diff-and-approve pattern formalized in the Governed Memory line of work (Taheri, arXiv:2603.17787, 2026). The pattern is: when an agent proposes to write a memory, the system computes a structured diff between the proposed write and the current state, presents the diff to a human reviewer, and commits the write only on approval. The pattern generalizes to any state mutation; memorywire applies it to `remember`, `forget`, and `merge`, and excludes `recall` and `expire` from the default approval surface (with `recall` flagged for v0.2 reconsideration; see §6). +The governance channel in memorywire implements a human-in-the-loop diff-and-approve workflow: when an agent proposes to write a memory, the system computes a structured diff between the proposed write and the current state, presents the diff to a human reviewer, and commits the write only on approval. This contrasts with production governance for autonomous agents such as Taheri's *Governed Memory* (arXiv:2603.17787, 2026), which enforces write policy automatically rather than routing writes through a human gate. The workflow generalizes to any state mutation; memorywire applies it to `remember`, `forget`, and `merge`, and excludes `recall` and `expire` from the default approval surface (with `recall` flagged for v0.2 reconsideration; see §6). -The Co-memorize pattern is not novel to this paper. What is new is its standardization at the wire-format layer: memorywire defines a `governance` JSON schema for the diff-and-approve message and ships a reference UI that any backend adapter inherits transparently. An agent calling `remember(content="…", approval_required=true)` gets the governance flow regardless of which of the five backends actually stores the row. +Diff-and-approve as a review discipline is not itself novel — it mirrors code review and change-management gates. What is new here is its standardization at the wire-format layer: memorywire defines a `governance` JSON schema for the diff-and-approve message and ships a reference UI that any backend adapter inherits transparently. An agent calling `remember(content="…", approval_required=true)` gets the governance flow regardless of which of the five backends actually stores the row. ## 3. The memorywire Wire Format @@ -217,7 +217,7 @@ The transformer is deliberately simple. The point is not to invent a new consoli ### 4.5 The governance UI -The governance UI (`ui/src/amp_ui/`) is a Starlette server with HTMX-driven templates. It shares the sqlite-vec adapter's SQLite database, so a reviewer sees pending writes as rows with the `PENDING_APPROVAL_DELETED_AT = -1` sentinel in `deleted_at` and a `pending: yes` badge in the UI. Reviewers can approve (clear the sentinel; the row becomes live), reject (hard-delete or soft-delete depending on the policy), or apply a Co-memorize transformation — typically a `merge` against an existing canonical row or a `forget` of a similar row that the new write supersedes. +The governance UI (`ui/src/amp_ui/`) is a Starlette server with HTMX-driven templates. It shares the sqlite-vec adapter's SQLite database, so a reviewer sees pending writes as rows with the `PENDING_APPROVAL_DELETED_AT = -1` sentinel in `deleted_at` and a `pending: yes` badge in the UI. Reviewers can approve (clear the sentinel; the row becomes live), reject (hard-delete or soft-delete depending on the policy), or apply a reconciling transformation — typically a `merge` against an existing canonical row or a `forget` of a similar row that the new write supersedes. Authentication is opt-in via the `MEMORYWIRE_UI_TOKEN` environment variable. When set, every request requires either an `Authorization: Bearer ` header or an `memorywire_ui_session` cookie; comparison uses `hmac.compare_digest` for constant-time matching. CSRF is enforced through a double-submit-cookie pattern signed with HMAC-SHA256 over `nonce.ts` with a 24-hour TTL. Without `MEMORYWIRE_UI_TOKEN` the UI is unauthenticated; binding a non-loopback host without a token fires a stderr warning on boot (`ui/src/amp_ui/middleware.py:55-66`). @@ -432,7 +432,7 @@ We are honest about the bet. The algorithmic substrate (RRF, FSMs, STM/LTM, diff ## Acknowledgments -We acknowledge the open-source projects memorywire composes with: mem0, Letta (formerly MemGPT), Cognee, Zep/Graphiti, MemoryOS, MemTensor MemOS, sqlite-vec, pgvector, `pytransitions`, Starlette, HTMX, sentence-transformers, and the Model Context Protocol community. The Co-memorize diff-and-approve pattern draws on the Governed Memory line of work. RRF as a fusion primitive is from Cormack, Clarke, and Buettcher. The human-memory taxonomy mapping follows the cognitive-science literature established by Tulving and Squire. +We acknowledge the open-source projects memorywire composes with: mem0, Letta (formerly MemGPT), Cognee, Zep/Graphiti, MemoryOS, MemTensor MemOS, sqlite-vec, pgvector, `pytransitions`, Starlette, HTMX, sentence-transformers, and the Model Context Protocol community. RRF as a fusion primitive is from Cormack, Clarke, and Buettcher. The human-memory taxonomy mapping follows the cognitive-science literature established by Tulving and Squire. ## References diff --git a/docs/paper/memorywire-paper.tex b/docs/paper/memorywire-paper.tex index e0aa8df..cf3111c 100644 --- a/docs/paper/memorywire-paper.tex +++ b/docs/paper/memorywire-paper.tex @@ -69,7 +69,7 @@ \subsection{The problem: islanded memory frameworks} Agent runtimes that maintain memory across sessions are now a category. Open-source frameworks include mem0, Letta (formerly MemGPT), Cognee, Zep/Graphiti, MemoryOS, and MemTensor MemOS; closed commercial offerings include Oracle's AI Agent Memory and the memory layers shipped inside major hosted-agent platforms. What the category has not produced is a shared wire format. Each framework defines its own SDK surface, JSON shape for memory records, embedding-provider integration, taxonomy (or absence of taxonomy) for memory types, and implicit lifecycle for record creation and deletion. The heterogeneity is non-trivial to bridge: mem0 stores records under a \texttt{memories[]} list keyed by \texttt{user\_id} with a heterogeneous \texttt{created\_at} representation; Letta stores archival memory keyed by \texttt{agent\_id} and exposes a \texttt{tags} list as the only structured-metadata sink; Cognee mints internal \texttt{data\_id} UUIDs that are not surfaced through its public \texttt{add} API, making per-record deletion impossible from outside the pipeline; sqlite-vec stores tables keyed by a stable ULID-shaped string; pgvector exposes records through an application-chosen SQL schema. Re-platforming an agent from one framework to another therefore requires a bespoke migrator and field-level losses where the source framework encodes more state than the target's data model holds. -The same heterogeneity means there is no shared \emph{governance} surface. Each framework provides a write API and a read API; none mediate the write with a ``diff against current state, present to a human, commit only on approval'' workflow. The Co-memorize human-in-the-loop pattern, formalized in the \emph{Governed Memory} line of work~\cite{taheri2026governed}, has no production implementation an off-the-shelf agent can drop in. Operators who want auditability over what enters long-term memory must build it themselves and accept that the framework can bypass them. +The same heterogeneity means there is no shared \emph{governance} surface. Each framework provides a write API and a read API; none mediate the write with a ``diff against current state, present to a human, commit only on approval'' workflow. Production governance work such as Taheri's \emph{Governed Memory}~\cite{taheri2026governed} enforces write policy automatically across autonomous agents rather than through such a human-approval gate, so this human-in-the-loop workflow has no production implementation an off-the-shelf agent can drop in. Operators who want auditability over what enters long-term memory must build it themselves and accept that the framework can bypass them. This is the gap memorywire addresses. It is not ``we need a better retrieval algorithm'' --- the algorithms in the category (vector search, hybrid lexical-semantic RRF fusion, graph hop boosts, FSM-encoded procedures, STM/LTM consolidation) are well understood. It is ``we need a shared protocol so any client can talk to any backend, any agent can carry its memory across runtimes, and any write can be diffed and approved.'' Structurally it is the gap MCP closed for \emph{tool use}, applied to \emph{memory}. @@ -81,7 +81,7 @@ \subsection{Contributions} \begin{itemize}[leftmargin=*] \item \textbf{C1.} A wire format for five memory operations over four memory types, expressed as JSON Schema 2020-12~\cite{jsonschema_2020_12} (\texttt{docs/spec/v0.md}). The operations are \texttt{remember}, \texttt{recall}, \texttt{forget}, \texttt{merge}, \texttt{expire}; the types are \texttt{semantic}, \texttt{episodic}, \texttt{procedural}, \texttt{emotional}. The schemas are vendor-neutral, transport-agnostic, and explicitly versioned with a breaking-change policy through v0.5. \item \textbf{C2.} A reference implementation in Python 3.11+ with five production-backend adapters (sqlite-vec, mem0, Letta, Cognee, pgvector), all implementing a single \texttt{MemoryStore} Protocol. The reference includes a memory router that fans operations across $N$ stores in parallel and fuses recall results via Reciprocal Rank Fusion ($k=60$)~\cite{cormack2009rrf} with an optional one-hop graph boost, plus a tolerant partial-failure model where a single rogue or unavailable backend cannot crash the operation. - \item \textbf{C3.} A governance UI implementing the Co-memorize diff-and-approve pattern over \texttt{remember}, \texttt{forget}, and \texttt{merge}. Writes flagged \texttt{approval\_required} are staged behind a \texttt{PENDING\_APPROVAL\_DELETED\_AT = -1} sentinel and remain invisible to \texttt{recall} until a reviewer commits or rejects them through the UI. The same audit log is the single source of truth for all governance and mutation events. + \item \textbf{C3.} A governance UI implementing a human-in-the-loop diff-and-approve workflow over \texttt{remember}, \texttt{forget}, and \texttt{merge}. Writes flagged \texttt{approval\_required} are staged behind a \texttt{PENDING\_APPROVAL\_DELETED\_AT = -1} sentinel and remain invisible to \texttt{recall} until a reviewer commits or rejects them through the UI. The same audit log is the single source of truth for all governance and mutation events. \item \textbf{C4.} An empirical evaluation comprising (a)~a microbenchmark on 100 hand-authored facts $\times$ 50 labelled queries against a real sentence-transformer embedder, (b)~an adversarial-fusion experiment that sweeps a 1-of-$N$ rank-0 injection attack across three fusion algorithms (RRF, MAX, weighted), and (c)~a cross-adapter conformance suite of 16 protocol-invariant scenarios run against all five shipped adapters (68 PASS / 12 SKIP / 0 FAIL out of 80 cells). \item \textbf{C5.} A six-adversary threat model with line-level mitigation citations into the reference implementation, plus an open-data artifact (\texttt{docs/adversarial-results.\{rrf,max,weighted\}.json}, the labelled microbench corpus, the conformance scenario list) sufficient to reproduce every empirical claim without re-running paid evaluators. \end{itemize} @@ -94,7 +94,7 @@ \subsection{Limitations, front-loaded} \subsection{Paper roadmap} \label{sec:intro-roadmap} -\Cref{sec:related} places memorywire against prior work in agent memory frameworks, cross-vendor protocols (particularly MCP), and the Governed Memory line. \Cref{sec:spec} specifies the wire format: operations, types, the \texttt{MemoryStore} Protocol, router semantics, and the governance channel. \Cref{sec:impl} describes the reference implementation including the five backend adapters, the procedural-memory FSM backend, the STM$\leftrightarrow$LTM transformer, and the governance UI. \Cref{sec:eval} reports the empirical evaluation. \Cref{sec:threats} is the threat model. \Cref{sec:mcp} details the relationship to MCP. \Cref{sec:future} and \Cref{sec:conclusion} lay out future work and the bet we are making. +\Cref{sec:related} places memorywire against prior work in agent memory frameworks, cross-vendor protocols (particularly MCP), and prior production work on memory governance. \Cref{sec:spec} specifies the wire format: operations, types, the \texttt{MemoryStore} Protocol, router semantics, and the governance channel. \Cref{sec:impl} describes the reference implementation including the five backend adapters, the procedural-memory FSM backend, the STM$\leftrightarrow$LTM transformer, and the governance UI. \Cref{sec:eval} reports the empirical evaluation. \Cref{sec:threats} is the threat model. \Cref{sec:mcp} details the relationship to MCP. \Cref{sec:future} and \Cref{sec:conclusion} lay out future work and the bet we are making. \section{Background and Related Work} \label{sec:related} @@ -130,9 +130,9 @@ \subsection{Human-memory taxonomy} \subsection{Human-in-the-loop approval for agent actions} -The governance channel in memorywire implements the Co-memorize diff-and-approve pattern formalized in the Governed Memory line of work~\cite{taheri2026governed}. The pattern is: when an agent proposes to write a memory, the system computes a structured diff between the proposed write and the current state, presents the diff to a human reviewer, and commits the write only on approval. The pattern generalizes to any state mutation; memorywire applies it to \texttt{remember}, \texttt{forget}, and \texttt{merge}, and excludes \texttt{recall} and \texttt{expire} from the default approval surface (with \texttt{recall} flagged for v0.2 reconsideration; see \Cref{sec:threats}). +The governance channel in memorywire implements a human-in-the-loop diff-and-approve workflow: when an agent proposes to write a memory, the system computes a structured diff between the proposed write and the current state, presents the diff to a human reviewer, and commits the write only on approval. This contrasts with production governance for autonomous agents such as Taheri's \emph{Governed Memory}~\cite{taheri2026governed}, which enforces write policy automatically rather than routing writes through a human gate. The workflow generalizes to any state mutation; memorywire applies it to \texttt{remember}, \texttt{forget}, and \texttt{merge}, and excludes \texttt{recall} and \texttt{expire} from the default approval surface (with \texttt{recall} flagged for v0.2 reconsideration; see \Cref{sec:threats}). -The Co-memorize pattern is not novel to this paper. What is new is its standardization at the wire-format layer: memorywire defines a \texttt{governance} JSON schema for the diff-and-approve message and ships a reference UI that any backend adapter inherits transparently. An agent calling \texttt{remember(content="\ldots", approval\_required=true)} gets the governance flow regardless of which of the five backends actually stores the row. +Diff-and-approve as a review discipline is not itself novel---it mirrors code review and change-management gates. What is new here is its standardization at the wire-format layer: memorywire defines a \texttt{governance} JSON schema for the diff-and-approve message and ships a reference UI that any backend adapter inherits transparently. An agent calling \texttt{remember(content="\ldots", approval\_required=true)} gets the governance flow regardless of which of the five backends actually stores the row. \section{The memorywire Wire Format} \label{sec:spec} @@ -304,7 +304,7 @@ \subsection{The STM$\leftrightarrow$LTM transformer} \subsection{The governance UI} \label{sec:impl-ui} -The governance UI (\texttt{ui/src/memorywire\_ui/}) is a Starlette server with HTMX-driven templates. It shares the sqlite-vec adapter's SQLite database, so a reviewer sees pending writes as rows with the \texttt{PENDING\_APPROVAL\_DELETED\_AT = -1} sentinel in \texttt{deleted\_at} and a \texttt{pending: yes} badge in the UI. Reviewers can approve (clear the sentinel; the row becomes live), reject (hard-delete or soft-delete depending on the policy), or apply a Co-memorize transformation --- typically a \texttt{merge} against an existing canonical row or a \texttt{forget} of a similar row that the new write supersedes. +The governance UI (\texttt{ui/src/memorywire\_ui/}) is a Starlette server with HTMX-driven templates. It shares the sqlite-vec adapter's SQLite database, so a reviewer sees pending writes as rows with the \texttt{PENDING\_APPROVAL\_DELETED\_AT = -1} sentinel in \texttt{deleted\_at} and a \texttt{pending: yes} badge in the UI. Reviewers can approve (clear the sentinel; the row becomes live), reject (hard-delete or soft-delete depending on the policy), or apply a reconciling transformation --- typically a \texttt{merge} against an existing canonical row or a \texttt{forget} of a similar row that the new write supersedes. Authentication is opt-in via the \texttt{MEMORYWIRE\_UI\_TOKEN} environment variable. When set, every request requires either an \texttt{Authorization: Bearer } header or an \texttt{memorywire\_ui\_session} cookie; comparison uses \texttt{hmac.compare\_digest} for constant-time matching. CSRF is enforced through a double-submit-cookie pattern signed with HMAC-SHA256 over \texttt{nonce.ts} with a 24-hour TTL. Without \texttt{MEMORYWIRE\_UI\_TOKEN} the UI is unauthenticated; binding a non-loopback host without a token fires a stderr warning on boot (\texttt{ui/src/memorywire\_ui/middleware.py:55--66}). @@ -570,7 +570,7 @@ \section{Conclusion} \section*{Acknowledgments} -We acknowledge the open-source projects memorywire composes with: mem0, Letta (formerly MemGPT), Cognee, Zep/Graphiti, MemoryOS, MemTensor MemOS, sqlite-vec, pgvector, \texttt{pytransitions}, Starlette, HTMX, sentence-transformers, and the Model Context Protocol community. The Co-memorize diff-and-approve pattern draws on the Governed Memory line of work. RRF as a fusion primitive is from Cormack, Clarke, and B{\"u}ttcher. The human-memory taxonomy mapping follows the cognitive-science literature established by Tulving and Squire. +We acknowledge the open-source projects memorywire composes with: mem0, Letta (formerly MemGPT), Cognee, Zep/Graphiti, MemoryOS, MemTensor MemOS, sqlite-vec, pgvector, \texttt{pytransitions}, Starlette, HTMX, sentence-transformers, and the Model Context Protocol community. RRF as a fusion primitive is from Cormack, Clarke, and B{\"u}ttcher. The human-memory taxonomy mapping follows the cognitive-science literature established by Tulving and Squire. \bibliographystyle{plain} \bibliography{memorywire} diff --git a/docs/paper/memorywire.bib b/docs/paper/memorywire.bib index 589285f..68f334d 100644 --- a/docs/paper/memorywire.bib +++ b/docs/paper/memorywire.bib @@ -39,7 +39,8 @@ @misc{packer2023memgpt year = {2023}, eprint = {2310.08560}, archivePrefix = {arXiv}, - primaryClass = {cs.AI} + primaryClass = {cs.AI}, + note = {arXiv:2310.08560} } @misc{wu2024longmemeval, @@ -48,7 +49,8 @@ @misc{wu2024longmemeval year = {2024}, eprint = {2410.10813}, archivePrefix = {arXiv}, - primaryClass = {cs.CL} + primaryClass = {cs.CL}, + note = {arXiv:2410.10813} } @misc{maharana2024locomo, @@ -57,21 +59,22 @@ @misc{maharana2024locomo year = {2024}, eprint = {2402.17753}, archivePrefix = {arXiv}, - primaryClass = {cs.CL} + primaryClass = {cs.CL}, + note = {arXiv:2402.17753} } % Verified via arXiv listing 2026-05-28: arXiv:2603.17787 resolves to a real % paper by Hamed Taheri titled "Governed Memory: A Production Architecture % for Multi-Agent Workflows" (submitted 2026-03-18). The 2603 YYMM prefix % (March 2026) is unusual but valid for an arXiv submission of that month. -% TODO verify arXiv ID before submission (re-check listing close to camera-ready). @misc{taheri2026governed, author = {Taheri, Hamed}, title = {Governed Memory: A Production Architecture for Multi-Agent Workflows}, year = {2026}, eprint = {2603.17787}, archivePrefix = {arXiv}, - primaryClass = {cs.MA} + primaryClass = {cs.AI}, + note = {arXiv:2603.17787} } @misc{chhikara2025mem0, @@ -80,7 +83,8 @@ @misc{chhikara2025mem0 year = {2025}, eprint = {2504.19413}, archivePrefix = {arXiv}, - primaryClass = {cs.AI} + primaryClass = {cs.AI}, + note = {arXiv:2504.19413} } @misc{mcp_spec_2025, diff --git a/docs/spec/v0.md b/docs/spec/v0.md index 58c486b..961d627 100644 --- a/docs/spec/v0.md +++ b/docs/spec/v0.md @@ -486,7 +486,7 @@ Memory research informing the design: - LoCoMo benchmark - BEAM benchmark - "Remember Me, Refine Me" — procedural memory -- "Governed Memory" arXiv 2603.17787 — Co-memorize HITL +- "Governed Memory" arXiv 2603.17787 — automated agent-memory governance (human-gated contrast) - Mem0 architecture writeup - Letta (MemGPT) hierarchical memory paper