Skip to content

ENG-1039 fix(harness): the authoring child is never captured or sensed (v0.76.1) - #268

Merged
sgonz-xtrace merged 2 commits into
mainfrom
fix/harness-child-not-captured
Sep 21, 2026
Merged

sgonz-xtrace merged 2 commits into
mainfrom
fix/harness-child-not-captured

Conversation

@sgonz-xtrace

@sgonz-xtrace sgonz-xtrace commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Problem

On staging, one teammate's Claude Code sessions showed up as dozens of conversations with the same title. "Dependency updates" has 80 rows and "Remote box setup for Claude code" has 39. Each row has a different source_id. This is not the known cwd-routing fork, where the same sid is stored twice.

The rows come from the harness author lane (v0.66.0+; on by default in v0.69.0–v0.75.x, opt-in again since v0.76.0 / #270). harness_stop.run_author runs one claude -p … --resume <owner> --fork-session per flagged moment. --fork-session gives the child a new session id and copies the parent's whole history into its transcript: same record uuids and timestamps, sessionId rewritten, title records copied.

The child is launched with MEMHUB_HARNESS_CHILD=1 # its hooks stay silent, but no hook read that variable (git grep MEMHUB_HARNESS_CHILD on main finds only the line that sets it). So:

  1. Every fork is captured as a new conversation. The child's Stop and SessionEnd flushes (flush_turn / flush_session) ship its transcript under its own session id, which the backend keys as a new row. The row carries the parent's title.
  2. Forks spawn forks. The child also gets MEMHUB_HARNESS_EXTRACT=0, but Claude Code applies a settings.json env value over what a process inherits. Installs that opted in before v0.69 with MEMHUB_HARNESS_EXTRACT=1 in settings re-armed the harness inside every child. The fork classified its own turns and drained the repo's pending moments, which spawned more forks.
  3. Failed passes multiply this. A pass that fails, for example "staging MCP disconnected" or "expired token", leaves its moment pending by design. Every later Stop in the repo forks it again.

After #270 (opt-in) this still matters, and arguably more. The harness now runs only for installs that set MEMHUB_HARNESS_EXTRACT=1. The settings.json env is the usual place to set it, and that value overrides the child's EXTRACT=0. So for every opted-in install, the fork's own harness is re-armed unless the child flag is honored (point 2). Capture of the fork (point 1) applies to every opted-in install regardless.

Why Claude Code only: only hooks/claude-hooks.json wires harness_stop.py. The Codex and Cursor hooks don't.

Evidence

1. Staging DB forensics (read-only)

Dependency updates Remote box setup
rows / distinct source_id 80 / 80 39 / 39
parent messages present in every row 215 (same source_message_ids) 31
row size 215–334 msgs, not increasing (a tree of branches) 38–239
updated_at − created_at ~1–2 s (flushed once, when the fork exited) same
first new user message in every tail MemHub harness: … turn N was flagged … + create-rule SKILL body same (38/39 rows)
last assistant message HARNESS-RESULT: none | failed | filed … same
outcomes 37 failed / 36 none / 4 filed 33 none / 5 filed
  • Human prompts in the tails: none. Every tail is harness prompts, <task-notification> records and "Continue from where you left off."
  • Forks of forks: many source_refs name earlier fork ids, not the root (56a28e64#2, e58bb72b#2, 2150b7a0#3, …).
  • Retries: the same moment was re-forked up to 7×. The failures were "staging MCP disconnected" and "expired staging token".
  • Cadence: bursts, e.g. 44 rows in one hour, with median gaps of 45 s and 112 s.
  • Scope: across staging (last 10 days), rows containing a harness prompt belong almost entirely to this one user: 117 rows, plus 2 from another teammate's first run on 09-19.

2. Live, blind reproduction (before the fix)

Conditions for every run:

  • Session: a real interactive Claude Code 2.1.278 session, driven through tmux on macOS.
  • Repo: a scratch git repo with a unique fake origin, so its moments are isolated from real repos (the drain selects by repo).
  • Plugin: installed at local scope for that folder only. The user-scope staging plugin and the synced prod plugin were disabled there, so only the build under test loaded.
  • Blind prompts: the session's agent got ordinary requests only ("the tests are failing, can you fix it?", "commit this"). Nothing told it what was under test. Each repo contained a real failing test, so a natural fail→fix→pass arc occurred.
  • Measurement: state checked in ~/.config/memhub-plugin/{harness,turnflush}, ~/.claude/projects/<repo>/*.jsonl, and staging team_conversations / team_memory_messages (read-only).
Run Build settings env Result
A v0.71.0 none Moment flagged on turn 1 (error_arc). On turn 2's Stop the drain spawned fork 1ef71624 (entrypoint: sdk-cli). Staging got a second row: same name "Fix failing tests", created 14 s after the drain, updated 0.5 s later, all 12 parent messages + 7 harness messages. The fork's own harness stayed silent (no extract 1ef71624 line).
B v0.71.0 MEMHUB_HARNESS_EXTRACT=1 Fork dee486cf was again shipped as a second row, same name "Due date test", 16 msgs ⊇ parent's 9. stop.log shows extract dee486cf t3, meaning the child's harness ran despite run_author setting EXTRACT=0. The settings value overrode it. The classifier found no moment in that fork's turn, so no grandchild this time. When a fork's turn is flagged (the failing passes above), the next drain forks the fork.

3. Live re-test with this branch (same conditions as run B, the harder case)

before (run B) after (this branch, 0.74.1; now 0.76.1)
fork created by the author pass dee486cf 19abfc8b (sdk-cli, carries the parent history)
fork shipped to staging yes, new row no row, no turnflush state
fork's own harness ran extract dee486cf t3 0 lines in stop.log / extract.log, no harness files
author pass still works outcome none outcome none (the child ran its create-rule flow and returned a result)
parent captured normally yes yes: 41b05c92, per-turn + SessionEnd (seq 19 → 20)

4. Live retry-loop reproduction: before vs after, identical conditions

The staging passes failed with "staging MCP disconnected". To recreate that, a claude wrapper was put first on the test session's PATH only. harness_stop.claude_bin() resolves claude from PATH. The wrapper rewrites only the author child's --mcp-config to a dead port (127.0.0.1:9) and then execs the real claude with every other argument unchanged. So the fork, its transcript and its hooks are all real, and capture (the plugin's own credential) keeps working.

Both runs used:

  • MEMHUB_HARNESS_EXTRACT=1 in the repo's local settings env.
  • The same starting repo (3 failing tests).
  • The same blind turn sequence: fix the tests / commit / a standing instruction for this repo / an ordinary question / add a docstring / commit.
v0.74.0 (main) this branch (0.74.1; now 0.76.1)
author passes (claude -p --fork-session) 5 3
outcome of the moments that needed the server failed: memhub rulebook server unreachable (ConnectionRefused) same, failed ×3 on the same moment
a failed moment re-forked at later Stops yes yes: the same #3 re-forked on 3 consecutive Stops
fork's own harness ran yes: extract <fork> t4/t5/t6/t7, each moment=1 no: 0 log lines for any fork
forks of forks yes: --resume 6238ac78 → 8207d302 → 29618a76, each a previous fork no: every pass resumed the parent
pending backlog per drain 1 → 1 → 3 → 4 → 6 waiting 1 → 1 → 1
staging rows for the session 6, all named "Shipping calculation tests", one per Stop, each shipped once at birth (19 → 22 → 25 → 27 → 29 msgs) 1, the parent (seq 35)

So the whole staging pattern reproduces from one blind session on current main: distinct sids, identical name, rows born about 2 min apart, source_refs naming forks.

Two observations this branch does not change:

  • Retries still cost quota. A pass that fails on an unreachable server leaves its moment pending, and every later Stop in the repo spends one claude -p on it until the 14-day TTL. With this fix those retries produce no rows and no new moments, but they are not free. A retry cap or backoff is worth a follow-up.
  • The detail NameError is currently throttling the cascade. harness_stop.py:751 raises after the first pass in every drain. The except then ends the loop, so only one of the "6 of 6 waiting" moments is authored per Stop. Fixing that NameError without this PR would let each drain fork every waiting moment. It should land after this fix, not before.

Fix

MEMHUB_HARNESS_CHILD is now honored by the lanes the comment promised would be silent:

  • Capture. transcript_filter.is_harness_child() is a new helper. It is shared, like the filter, by both upload paths. flush_turn.main and flush_session.main return 0 immediately when it is set. That covers Stop, SessionEnd and the commit/PR flush, whatever hook wiring calls them.
  • Harness. harness_extract.extract_enabled and rulebook_hook.harness_extract_on (the one switch, in two copies) read off inside a child whatever MEMHUB_HARNESS_EXTRACT says. A settings env can no longer re-arm it.
  • Rulebook. The child's rulebook_hook is deliberately untouched. The create-rule forward test runs in it against MEMHUB_RULEBOOK_BASE.
  • Version. 0.76.1 in all nine version files (the cache is keyed by version). The branch was merged with main after fix(harness): back to opt-in — off unless MEMHUB_HARNESS_EXTRACT is on (v0.76.0) #270: the child gate sits in front of fix(harness): back to opt-in — off unless MEMHUB_HARNESS_EXTRACT is on (v0.76.0) #270's opt-in parsing and uses the same explicit on-values (1/on/true/yes).

A backend dedup-by-uuid would only have hidden the symptom. The copy is the plugin's own work, never the person's, so the plugin must not ship it.

Tests

  • Unit tests: four new tests. Each fails on main and passes here (defeat-tested by restoring the six origin/main scripts under the new tests):

    • flush_turn_test::test_a_harness_child_is_never_captured
    • flush_session_test: the harness-child block
    • harness_extract_test::test_an_authoring_child_is_never_sensed_whatever_the_flag_says
    • harness_stop_test::test_an_authoring_child_neither_senses_nor_drains_when_settings_re_arm_the_flag

    Each includes the no-flag control, so it can't pass vacuously. test_the_author_child_is_launched_as_a_harness_child pins that run_author sets the flag.

  • Full check: bash scripts/check-plugin.sh passes: all 72 suites under bare Python and under mcp<2, plus test_flush_hook.sh. Run on macOS 26.4.1 from this branch's source, before and after the merge with main (head c906219).

  • Real-agent evidence: real-agent-evidence.yml dispatched on head c906219 (run 35653938069).

Not in this PR

  • Separate bug seen in every live run. harness_stop.py:751 references detail, which is undefined, after every pass. The outcome row is written first, then the except appends a spurious failed row, logs "could not start", and abandons the rest of the drain. See §4 for why its fix should land after this one.
  • Existing junk rows. Staging needs a cleanup (rows whose last assistant message starts with HARNESS-RESULT:). That is backend data work.

🤖 Generated with Claude Code

`run_author` runs `claude -p --resume <owner> --fork-session`. That is a NEW
session id whose transcript copies the person's whole history. The child was
launched with MEMHUB_HARNESS_CHILD=1 ("its hooks stay silent"), but nothing
read that flag:

- flush_turn / flush_session shipped every fork as another conversation under
  the parent's title. On staging one session reached 80 rows.
- The child's EXTRACT=0 is overridden by a settings.json `env` value. Installs
  that opted in with MEMHUB_HARNESS_EXTRACT=1 re-armed the sensor in every
  child, so forks classified their own turns, drained the repo's moments, and
  spawned more forks.

Now MEMHUB_HARNESS_CHILD gates both transcript-capture scripts
(transcript_filter.is_harness_child) and both harness switches
(harness_extract.extract_enabled, rulebook_hook.harness_extract_on). The
child's rulebook hook still runs, for the forward test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@sgonz-xtrace
sgonz-xtrace deployed to production-plugin-release September 21, 2026 19:27 — with GitHub Actions Active
@xtrace-memhub-staging

xtrace-memhub-staging Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

🧠 Session context

1 session behind this pull request.

Team rules that fired while building this

  • fetch-before-origin-read (advise, 1×)

Effort
1 session · 2h 36m agent time · 62.1M tokens · 557 turns

Sessions

felix-xtrace added a commit that referenced this pull request Sep 21, 2026
…n (v0.76.0) (#270)

#260 (v0.69.0) made the harness lane default-on for every install. On costs
the person a classifier call per flagged turn and a headless authoring run per
moment on THEIR OWN model quota, and files proposals into a shared team
rulebook. Felix's call, 2026-09-21: an install should not start that without
being asked. Opt-in again.

Three gates, flipped together because they are one switch and must never
disagree: `extract_enabled` in harness_extract, `harness_extract_on` in
rulebook_hook (the error-arc pairing), and the shell `case` in
claude-hooks.json that runs before either. The shell gate proceeds only on an
explicit on value (1/on/true/yes, any case), so unset, blank and unrecognised
all exit before python starts.

The empty string moved back with the default, and unrecognised values moved
with it: under default-on a typo ran the lane, under opt-in a typo must not
start the spend. Tests assert both sides, including the spellings #260 added.

Docstrings, the module headers, the rulebook_hook comment and the README moved
in the same change. The README section was still titled "(flagged off)" and
still ended "With the variable unset, the default, none of this runs" — both
false since #260, true again now. `_OFF` is removed; nothing reads it.

Anyone who was relying on default-on since v0.69.0 must now set
MEMHUB_HARNESS_EXTRACT=1.

check-plugin.sh 69 of 72, the same three pre-existing failures as origin/main
(codex_history, readers_cli, readers_validation). Nine version files to 0.76.0
(0.74.1 and 0.75.0 are taken by open #268 and #267).

Refs ENG-1107, reverts the default from #260.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…-captured

Conflicts resolved:
- extract_enabled / harness_extract_on: keep #270's opt-in parsing
  (explicit on-values only) and put the child gate in front of it.
- The child flag uses the same on-values, as does
  transcript_filter.is_harness_child, so the switch reads one way.
- Version 0.76.0 -> 0.76.1 in all nine version files.

Opt-in makes the child gate matter more, not less: settings.json `env` is
where an install opts in with MEMHUB_HARNESS_EXTRACT=1, and that value
overrides the child's EXTRACT=0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@sgonz-xtrace
sgonz-xtrace deployed to production-plugin-release September 21, 2026 20:53 — with GitHub Actions Active
@sgonz-xtrace
sgonz-xtrace deployed to production-plugin-release September 21, 2026 20:54 — with GitHub Actions Active
@sgonz-xtrace
sgonz-xtrace deployed to production-plugin-release September 21, 2026 20:54 — with GitHub Actions Active
@sgonz-xtrace
sgonz-xtrace deployed to production-plugin-release September 21, 2026 20:54 — with GitHub Actions Active
@sgonz-xtrace sgonz-xtrace changed the title fix(harness): the authoring child is never captured or sensed (v0.74.1) fix(harness): the authoring child is never captured or sensed (v0.76.1) Sep 21, 2026
@sgonz-xtrace
sgonz-xtrace merged commit 31d7dbd into main Sep 21, 2026
12 checks passed
@sgonz-xtrace
sgonz-xtrace deleted the fix/harness-child-not-captured branch September 21, 2026 21:47
@sgonz-xtrace sgonz-xtrace changed the title fix(harness): the authoring child is never captured or sensed (v0.76.1) ENG-1039 fix(harness): the authoring child is never captured or sensed (v0.76.1) Sep 21, 2026
sgonz-xtrace added a commit that referenced this pull request Sep 21, 2026
#268 made MEMHUB_HARNESS_CHILD switch transcript capture off: flush_turn and
flush_session both return early for the harness's forked session. The gate
must read the same flag through the same helper (is_harness_child), or it
denies add_memory in a session nothing is capturing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
production-plugin-release — c9062198 Deployed Sep 21, 2026 by sgonz-xtrace via Real agent session (codex) #62
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant