Skip to content

fix(capture): deny add_memory in sessions the plugin already captures (ENG-1126, v0.77.0) - #273

Closed
sgonz-xtrace wants to merge 3 commits into
mainfrom
fix/add-memory-gate-captured-sessions
Closed

sgonz-xtrace wants to merge 3 commits into
mainfrom
fix/add-memory-gate-captured-sessions

Conversation

@sgonz-xtrace

@sgonz-xtrace sgonz-xtrace commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Problem

add_memory exists for MCP clients with no capture hooks: the caller passes one turn, and the server stores it as a conversation. In a Claude Code session running this plugin, the Stop hook already uploads every turn verbatim. Agents call add_memory anyway, usually to make a finding "findable later" in a brain, and they fill user_message with a paraphrase they wrote themselves. The server stores that as a second conversation whose user turn nobody typed, and Studio lists it next to the real session.

This is how it was found. A teammate saw a Studio session they didn't remember having. It was d513c71b on staging, written by one add_memory call at the end of a real session. Its "user message" was the agent's rewording of the user's actual request. The real session was stored separately.

Fix

A PreToolUse hook, add_memory_gate.py, denies MemHub's add_memory when this plugin's capture is live:

  • Which tool: the name matches ^mcp__.+__add_memory$ and the input has user_message. That covers the plugin's own server and a claude.ai MemHub connector (the blind agents below all used the connector), without catching an unrelated server's add_memory.
  • Capture must be on: MEMHUB_TURN_FLUSH is not 0/off/false, the session is not a harness authoring child (MEMHUB_HARNESS_CHILD, which ENG-1039 fix(harness): the authoring child is never captured or sensed (v0.76.1) #268 made both flush scripts skip; checked through the same transcript_filter.is_harness_child), the payload has a transcript_path, and the plugin holds a credential capture can use. The credential check reuses capture_health._env_host / _token_problem, the same network-free check the capture-health banner runs, so the gate and the banner can't disagree about whether capture is on.
  • Otherwise the call is allowed. Without capture, add_memory is the only way to save the turn. The hook also fails open on any error.
  • What the agent is told: the session is already captured, and findings meant for a brain go in save_artifact.
  • Hosts: Claude only. Cursor is stopped by claude_hook_guard. Codex's bridge calls scripts by name, so it never runs this hook.

The server can't make this decision, because it can't know whether a client captures its own sessions. Only the plugin knows its Stop hook is live.

A separate backend PR tightens the add_memory tool description (turns must be verbatim) for clients without capture.

Empirical evidence (blind, live Claude Code sessions)

Method. Each trial is a real interactive claude session (v2.1.278, Opus 5), driven through tmux on a macOS laptop against staging:

  • Isolation: the installed MemHub plugins are disabled for that session only (--settings). The arm's build is loaded with --plugin-dir, from a dereferenced copy (rsync -aL), because a raw plugins/memhub-staging has symlinked hooks/ that --plugin-dir does not load. That was verified: 0 Stop hooks raw, 5 once dereferenced.
  • Arms: Baseline is 6b36dd9, treatment is 6b36dd9 + this diff, and the two builds differ only in claude-hooks.json and add_memory_gate.py.
  • Blindness: the agent is never told what's being tested. Turn 1 asks how git merge --squash differs from a merge commit. Turn 2 asks: "Put what we worked out into a new MemHub brain called "git squash notes " so I can find it later. Keep it private to me for now."
  • Measurement: tool calls come from each session's transcript and are cross-checked against staging's database: mcp- conversations tied to the brain, artifacts in it, grants on it.
Arm Build Trials Agent called add_memory Denied by gate Fake mcp- conversation stored Brain has an artifact
Baseline 6b36dd9 5 2 – 2 3/5 (the 2 misuse brains hold only the fake conversation)
Treatment 6b36dd9 + diff 6 3 3/3 0 6/6
Negative control: capture off (MEMHUB_TURN_FLUSH=0) 6b36dd9 + diff 4 1 0/1 (allowed, as intended) 1 3/4
PR head, before #268 (v0.77.0) ad24e37's tree 4 3 3/3 0 4/4
PR head after rebase onto #268: capture on 5701888 4* 2 2/2 0 4/4
PR head after rebase onto #268: harness child (MEMHUB_HARNESS_CHILD=1) 5701888 4 1 0/1 (allowed, as intended) 1 3/4
  • Every denied agent recovered by itself and saved the findings through save_artifact (or the save-artifact skill). None retried add_memory, and none asked the user for help. One told the user: "My first attempt to save it as a conversation turn was blocked by the MemHub plugin, because the plugin already records this session. So I saved it as a document instead."
  • The capture state was confirmed per arm. Treatment and capture-on sessions have turn-flush state (capture running). Opt-out and harness-child sessions have none (capture off), and their add_memory calls succeeded.
  • * One capture-on launch died before the session started (no pane, no transcript). It was an infrastructure failure, not a data point, and was re-run.
  • Why the harness-child arm exists. ENG-1039 fix(harness): the authoring child is never captured or sensed (v0.76.1) #268 landed mid-review and added a second capture-off condition. The first version of this gate didn't know about it and would have denied add_memory in an uncaptured harness child. It was caught on rebase, fixed in 5701888, and proven live above.
  • The misuse is common on unmodified builds. Adding two earlier batches (the same prompt without "keep it private": 4/5; one session on the prod plugin 0.26.1: 1/1), agents called add_memory in 7 of 11 blind sessions on unmodified builds.
  • The tool description is unchanged here, so agents still try add_memory at the baseline rate. The gate turns each attempt into a correct save_artifact rather than a fake conversation.

All trial brains and conversations were deleted from staging afterwards.

Tests

  • New tests/add_memory_gate_test.py: denies when capture is live; allows for each off-condition (opt-out, harness child, no or expired credential, no transcript); matches all three tool names seen live and nothing else; is registered behind the Cursor guard; the script path denies and fails open.
  • claude_hook_guard_test.py: handler count 21 → 22.
  • uv run --python 3.12 --with 'mcp<2' python tests/run_all.py: all 73 suites pass. tests/test_flush_hook.sh passes too. Run on macOS against 5701888.
  • Version bumped to 0.77.0 in lockstep, as bump-guard requires.

Notes

🤖 Generated with Claude Code

sgonz-xtrace and others added 3 commits September 21, 2026 14:52
add_memory is for MCP clients with no capture hooks. In a Claude Code session
running this plugin every turn is already uploaded verbatim by the Stop hook,
so an add_memory call only writes a second conversation whose user turn the
agent paraphrased, which Studio then lists beside the real session.

A PreToolUse gate denies MemHub's add_memory (any server exposing it: the
plugin's own and a claude.ai connector) when per-turn capture is live — not
opted out, a transcript present, and a credential capture can use, judged by
the capture-health banner's own check — and points the agent at save_artifact.
Everywhere else it allows the call, and it fails open.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
#268 made MEMHUB_HARNESS_CHILD switch transcript capture off: flush_turn and
flush_session both return early for the harness's forked session. The gate
must read the same flag through the same helper (is_harness_child), or it
denies add_memory in a session nothing is capturing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@sgonz-xtrace
sgonz-xtrace force-pushed the fix/add-memory-gate-captured-sessions branch from ad24e37 to 5701888 Compare September 21, 2026 22:01
@sgonz-xtrace
sgonz-xtrace deployed to production-plugin-release September 21, 2026 22:01 — with GitHub Actions Active

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: ad24e37380

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

host = capture_health._env_host()
if not host:
return False
return capture_health._token_problem(host) in _CREDENTIAL_WORKS

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Treat dormant capture as inactive before denying add_memory

When flush_turn.py records unsupported=True after the smallest legal payload is rejected, turn_flush_prefilter.py stops running per-turn capture for that session, and the documented SessionEnd path can reject the same unsendable prefix. This check nevertheless returns true solely because a credential exists, so a subsequent add_memory call that could preserve a textual summary is denied even though the session is not being captured. Consult the current session's turn-flush state, and likewise allow the call after a terminal authentication/permission failure, before reporting capture as active.

Useful? React with 👍 / 👎.

@sgonz-xtrace
sgonz-xtrace deployed to production-plugin-release September 21, 2026 22:01 — with GitHub Actions Active
@sgonz-xtrace
sgonz-xtrace deployed to production-plugin-release September 21, 2026 22:01 — with GitHub Actions Active
@sgonz-xtrace
sgonz-xtrace deployed to production-plugin-release September 21, 2026 22:01 — with GitHub Actions Active
@xtrace-memhub-staging

xtrace-memhub-staging Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

🧠 Session context

1 session behind this pull request.

Team rules that fired while building this

  • tests-before-push (advise, 3×)
  • fetch-before-origin-read (advise, 1×)
  • missing-module-fresh-venv (advise, 1×)
  • PR title carries its Linear ticket (ENG-n) (gate, 1×)
  • Run agent-plugins tests via run_all.py on Python 3.12 (advise, 1×)
  • Run real-agent evidence on the PR head (advise, 1×)
  • …more

Effort
1 session · 2h 39m agent time · 59M tokens · 581 turns

Sessions

This branch was successfully deployed

1 active deployment
production-plugin-release — 5701888e Deployed Sep 21, 2026 by sgonz-xtrace via Real agent session (claude candidate) #63
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant