ENG-1039 fix(harness): the authoring child is never captured or sensed (v0.76.1) - #268
Merged
Merged
Conversation
`run_author` runs `claude -p --resume <owner> --fork-session`. That is a NEW
session id whose transcript copies the person's whole history. The child was
launched with MEMHUB_HARNESS_CHILD=1 ("its hooks stay silent"), but nothing
read that flag:
- flush_turn / flush_session shipped every fork as another conversation under
the parent's title. On staging one session reached 80 rows.
- The child's EXTRACT=0 is overridden by a settings.json `env` value. Installs
that opted in with MEMHUB_HARNESS_EXTRACT=1 re-armed the sensor in every
child, so forks classified their own turns, drained the repo's moments, and
spawned more forks.
Now MEMHUB_HARNESS_CHILD gates both transcript-capture scripts
(transcript_filter.is_harness_child) and both harness switches
(harness_extract.extract_enabled, rulebook_hook.harness_extract_on). The
child's rulebook hook still runs, for the forward test.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
sgonz-xtrace
deployed
to
production-plugin-release
September 21, 2026 19:27 — with
GitHub Actions
Active
🧠 Session context1 session behind this pull request. Team rules that fired while building this
Effort Sessions
|
felix-xtrace
added a commit
that referenced
this pull request
Sep 21, 2026
…n (v0.76.0) (#270) #260 (v0.69.0) made the harness lane default-on for every install. On costs the person a classifier call per flagged turn and a headless authoring run per moment on THEIR OWN model quota, and files proposals into a shared team rulebook. Felix's call, 2026-09-21: an install should not start that without being asked. Opt-in again. Three gates, flipped together because they are one switch and must never disagree: `extract_enabled` in harness_extract, `harness_extract_on` in rulebook_hook (the error-arc pairing), and the shell `case` in claude-hooks.json that runs before either. The shell gate proceeds only on an explicit on value (1/on/true/yes, any case), so unset, blank and unrecognised all exit before python starts. The empty string moved back with the default, and unrecognised values moved with it: under default-on a typo ran the lane, under opt-in a typo must not start the spend. Tests assert both sides, including the spellings #260 added. Docstrings, the module headers, the rulebook_hook comment and the README moved in the same change. The README section was still titled "(flagged off)" and still ended "With the variable unset, the default, none of this runs" — both false since #260, true again now. `_OFF` is removed; nothing reads it. Anyone who was relying on default-on since v0.69.0 must now set MEMHUB_HARNESS_EXTRACT=1. check-plugin.sh 69 of 72, the same three pre-existing failures as origin/main (codex_history, readers_cli, readers_validation). Nine version files to 0.76.0 (0.74.1 and 0.75.0 are taken by open #268 and #267). Refs ENG-1107, reverts the default from #260. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…-captured Conflicts resolved: - extract_enabled / harness_extract_on: keep #270's opt-in parsing (explicit on-values only) and put the child gate in front of it. - The child flag uses the same on-values, as does transcript_filter.is_harness_child, so the switch reads one way. - Version 0.76.0 -> 0.76.1 in all nine version files. Opt-in makes the child gate matter more, not less: settings.json `env` is where an install opts in with MEMHUB_HARNESS_EXTRACT=1, and that value overrides the child's EXTRACT=0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
sgonz-xtrace
deployed
to
production-plugin-release
September 21, 2026 20:53 — with
GitHub Actions
Active
sgonz-xtrace
deployed
to
production-plugin-release
September 21, 2026 20:54 — with
GitHub Actions
Active
sgonz-xtrace
deployed
to
production-plugin-release
September 21, 2026 20:54 — with
GitHub Actions
Active
sgonz-xtrace
deployed
to
production-plugin-release
September 21, 2026 20:54 — with
GitHub Actions
Active
sgonz-xtrace
added a commit
that referenced
this pull request
Sep 21, 2026
#268 made MEMHUB_HARNESS_CHILD switch transcript capture off: flush_turn and flush_session both return early for the harness's forked session. The gate must read the same flag through the same helper (is_harness_child), or it denies add_memory in a session nothing is capturing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
On staging, one teammate's Claude Code sessions showed up as dozens of conversations with the same title. "Dependency updates" has 80 rows and "Remote box setup for Claude code" has 39. Each row has a different
source_id. This is not the known cwd-routing fork, where the same sid is stored twice.The rows come from the harness author lane (v0.66.0+; on by default in v0.69.0–v0.75.x, opt-in again since v0.76.0 / #270).
harness_stop.run_authorruns oneclaude -p … --resume <owner> --fork-sessionper flagged moment.--fork-sessiongives the child a new session id and copies the parent's whole history into its transcript: same recorduuids and timestamps,sessionIdrewritten, title records copied.The child is launched with
MEMHUB_HARNESS_CHILD=1 # its hooks stay silent, but no hook read that variable (git grep MEMHUB_HARNESS_CHILDonmainfinds only the line that sets it). So:flush_turn/flush_session) ship its transcript under its own session id, which the backend keys as a new row. The row carries the parent's title.MEMHUB_HARNESS_EXTRACT=0, but Claude Code applies a settings.jsonenvvalue over what a process inherits. Installs that opted in before v0.69 withMEMHUB_HARNESS_EXTRACT=1in settings re-armed the harness inside every child. The fork classified its own turns and drained the repo's pending moments, which spawned more forks.After #270 (opt-in) this still matters, and arguably more. The harness now runs only for installs that set
MEMHUB_HARNESS_EXTRACT=1. The settings.jsonenvis the usual place to set it, and that value overrides the child'sEXTRACT=0. So for every opted-in install, the fork's own harness is re-armed unless the child flag is honored (point 2). Capture of the fork (point 1) applies to every opted-in install regardless.Why Claude Code only: only
hooks/claude-hooks.jsonwiresharness_stop.py. The Codex and Cursor hooks don't.Evidence
1. Staging DB forensics (read-only)
source_idsource_message_ids)updated_at − created_atMemHub harness: … turn N was flagged …+ create-rule SKILL bodyHARNESS-RESULT: none | failed | filed …<task-notification>records and "Continue from where you left off."source_refs name earlier fork ids, not the root (56a28e64#2,e58bb72b#2,2150b7a0#3, …).2. Live, blind reproduction (before the fix)
Conditions for every run:
origin, so its moments are isolated from real repos (the drain selects by repo).~/.config/memhub-plugin/{harness,turnflush},~/.claude/projects/<repo>/*.jsonl, and stagingteam_conversations/team_memory_messages(read-only).enverror_arc). On turn 2's Stop the drain spawned fork1ef71624(entrypoint: sdk-cli). Staging got a second row: same name "Fix failing tests", created 14 s after the drain, updated 0.5 s later, all 12 parent messages + 7 harness messages. The fork's own harness stayed silent (noextract 1ef71624line).MEMHUB_HARNESS_EXTRACT=1dee486cfwas again shipped as a second row, same name "Due date test", 16 msgs ⊇ parent's 9.stop.logshowsextract dee486cf t3, meaning the child's harness ran despiterun_authorsettingEXTRACT=0. The settings value overrode it. The classifier found no moment in that fork's turn, so no grandchild this time. When a fork's turn is flagged (the failing passes above), the next drain forks the fork.3. Live re-test with this branch (same conditions as run B, the harder case)
dee486cf19abfc8b(sdk-cli, carries the parent history)extract dee486cf t3nonenone(the child ran its create-rule flow and returned a result)41b05c92, per-turn + SessionEnd (seq 19 → 20)4. Live retry-loop reproduction: before vs after, identical conditions
The staging passes failed with "staging MCP disconnected". To recreate that, a
claudewrapper was put first on the test session'sPATHonly.harness_stop.claude_bin()resolvesclaudefromPATH. The wrapper rewrites only the author child's--mcp-configto a dead port (127.0.0.1:9) and thenexecs the realclaudewith every other argument unchanged. So the fork, its transcript and its hooks are all real, and capture (the plugin's own credential) keeps working.Both runs used:
MEMHUB_HARNESS_EXTRACT=1in the repo's local settingsenv.main)claude -p --fork-session)failed: memhub rulebook server unreachable (ConnectionRefused)failed×3 on the same moment#3re-forked on 3 consecutive Stopsextract <fork> t4/t5/t6/t7, eachmoment=1--resume 6238ac78→8207d302→29618a76, each a previous fork1 → 1 → 3 → 4 → 6 waiting1 → 1 → 1So the whole staging pattern reproduces from one blind session on current
main: distinct sids, identical name, rows born about 2 min apart,source_refs naming forks.Two observations this branch does not change:
claude -pon it until the 14-day TTL. With this fix those retries produce no rows and no new moments, but they are not free. A retry cap or backoff is worth a follow-up.detailNameError is currently throttling the cascade.harness_stop.py:751raises after the first pass in every drain. Theexceptthen ends the loop, so only one of the "6 of 6 waiting" moments is authored per Stop. Fixing that NameError without this PR would let each drain fork every waiting moment. It should land after this fix, not before.Fix
MEMHUB_HARNESS_CHILDis now honored by the lanes the comment promised would be silent:transcript_filter.is_harness_child()is a new helper. It is shared, like the filter, by both upload paths.flush_turn.mainandflush_session.mainreturn 0 immediately when it is set. That covers Stop, SessionEnd and the commit/PR flush, whatever hook wiring calls them.harness_extract.extract_enabledandrulebook_hook.harness_extract_on(the one switch, in two copies) read off inside a child whateverMEMHUB_HARNESS_EXTRACTsays. A settingsenvcan no longer re-arm it.rulebook_hookis deliberately untouched. The create-rule forward test runs in it againstMEMHUB_RULEBOOK_BASE.mainafter fix(harness): back to opt-in — off unless MEMHUB_HARNESS_EXTRACT is on (v0.76.0) #270: the child gate sits in front of fix(harness): back to opt-in — off unless MEMHUB_HARNESS_EXTRACT is on (v0.76.0) #270's opt-in parsing and uses the same explicit on-values (1/on/true/yes).A backend dedup-by-uuid would only have hidden the symptom. The copy is the plugin's own work, never the person's, so the plugin must not ship it.
Tests
Unit tests: four new tests. Each fails on
mainand passes here (defeat-tested by restoring the sixorigin/mainscripts under the new tests):flush_turn_test::test_a_harness_child_is_never_capturedflush_session_test: the harness-child blockharness_extract_test::test_an_authoring_child_is_never_sensed_whatever_the_flag_saysharness_stop_test::test_an_authoring_child_neither_senses_nor_drains_when_settings_re_arm_the_flagEach includes the no-flag control, so it can't pass vacuously.
test_the_author_child_is_launched_as_a_harness_childpins thatrun_authorsets the flag.Full check:
bash scripts/check-plugin.shpasses: all 72 suites under bare Python and undermcp<2, plustest_flush_hook.sh. Run on macOS 26.4.1 from this branch's source, before and after the merge withmain(headc906219).Real-agent evidence:
real-agent-evidence.ymldispatched on headc906219(run 35653938069).Not in this PR
harness_stop.py:751referencesdetail, which is undefined, after every pass. The outcome row is written first, then theexceptappends a spuriousfailedrow, logs "could not start", and abandons the rest of the drain. See §4 for why its fix should land after this one.HARNESS-RESULT:). That is backend data work.🤖 Generated with Claude Code