Repository navigation
fix(runtime-host): Host recovery decodes every Session's full history before readiness, breaching the 45s election deadline #4032
Description
Activity
Working on a first, semantics-preserving step:
prepareHostedExecutionRecoverycurrently triggers two full message decodes per Session — one discarded pre-read insidelistForRecovery(), one explicitreadMessagesForRecoveryin the per-session loop. The facade already exposeslistHeaders(), and all production callers oflistForRecovery()only consume the headers, so the recovery coordinator can list headers directly and keep exactly one decode per Session (the fail-closed message validation still happens in the per-session read).Skipping per-run validation reads for terminal runs (the dominant cost here) changes the startup-validation contract, so I am leaving that for maintainer discussion in this issue before sending a PR.
Measured end-to-end on the same long-lived workspace (~80 recoverable Sessions, ~500 runs — all terminal, ~100k event/message rows), cold Host start to interactive TUI:
Stage Cold start Before #4031 (artifact recovery storm + full validation decode) 10+ min With #4031 only ~157 s With #4031 + #4033 ~135 s Profiling during the remaining 135 s confirms the residual cost is the discarded per-run validation decode (
readEventsForRecovery/readRuntimeEventsover every run), exactly the item flagged for a design decision above.Two additional findings from the measurements:
recoverabledoes not exclude archived Sessions (sqliteRecoverableSessionRolePredicateis ordinary + coordination only), so archiving old Sessions does not reduce recovery work today.MAKA_RUNTIME_HOST_ELECTION_DEADLINE_MSis capped at 120000 ms, below the measured 135 s recovery time — even the escape hatch cannot cover a cold start on this workspace until recovery gets cheaper.
- added a commit that references this issue
on Aug 27, 2026 - addedbugSomething isn't workingSomething isn't workinghelp wantedExtra attention is neededExtra attention is needed
on Aug 29, 2026 @me2seeks @Astro-Han — I think the remaining issue may benefit from separating Host readiness from full-history validation, rather than continuing to reduce the constant factor of an
O(total history)startup scan.A possible recovery contract:
-
Startup-critical recovery before
ready- non-terminal Runs;
- unresolved admissions, pending closures, and interrupted repair state;
- scheduled-task / bot work whose ownership or replay semantics must be settled before the Host accepts work.
-
Demand-driven Session recovery
- the Session explicitly requested by
maka --resume <id>; - Sessions currently opened/visible in the client;
- optionally a small recent working set.
An
ensureSessionRecovered(sessionId)barrier could deduplicate concurrent recovery and block only that Session, instead of blocking the whole Host. - the Session explicitly requested by
-
Cold terminal history
- settled, inactive Sessions should not block Host readiness;
- archived Sessions should normally stay cold;
- validation can happen when a Session is opened, or in bounded idle maintenance. A full background scan on every startup is probably still unnecessary work.
This likely needs a lightweight durable recovery index containing only actionable state (non-terminal Run, unresolved admission/closure/repair, archive state, and perhaps last access), plus per-Session states such as
cold | recovering | ready | corrupt. Transcript bodies should remain paged/lazy even for the visible working set.The important safety boundaries would be:
- delayed recovery must never replay already-settled work;
- non-terminal and scheduled work cannot be deferred as ordinary cold history;
- corruption in a cold terminal Session should quarantine that Session rather than poison global Host readiness;
- opening or sending to a cold Session must cross the per-Session recovery barrier first.
This would change startup cost from
O(all Session history)towardO(actionable state + demanded Sessions). For the reported workspace, that could mean recovering the handful of actively used Sessions rather than decoding all ~80 Sessions / ~100k rows before readiness.Would you be comfortable with the contract that corruption in terminal/archived history becomes a Session-local failure discovered on access, while only actionable recovery state remains globally fail-closed? If so, this could be split into incremental PRs: remove semantics-preserving duplicate reads first, then add the actionable index / scoped recovery, and finally the per-Session on-demand barrier.
中文翻译
@me2seeks @Astro-Han——我认为,#4032 剩余问题更适合把 Host 就绪 与 全历史校验 分开,而不是继续优化启动时
O(全部历史)扫描的常数。可以考虑如下恢复契约:
-
进入
ready前必须完成的关键恢复- 非终态 Run;
- 未解决的 admission、待关闭记录和中断的修复状态;
- scheduled task / bot 等必须在 Host 接受新工作前确定所有权或重放语义的任务。
-
按需恢复 Session
maka --resume <id>明确请求的 Session;- 客户端当前打开或可见的 Session;
- 可选的一小组最近使用 Session。
可以增加
ensureSessionRecovered(sessionId)屏障,对同一 Session 的并发恢复进行去重,并且只阻塞该 Session,而不是阻塞整个 Host。 -
冷的终态历史
- 已结束且不活跃的 Session 不应阻塞 Host 就绪;
- 归档 Session 通常应保持冷状态;
- 可以在用户打开 Session 时校验,或在系统空闲时进行有界维护。每次启动都在后台完整扫描一次,可能仍然属于不必要的工作。
这可能需要一个轻量的持久化 recovery index,只记录需要采取动作的状态,例如非终态 Run、未解决的 admission/closure/repair、归档状态,以及可选的最近访问时间;同时为每个 Session 维护类似
cold | recovering | ready | corrupt的状态。即使是当前可见的工作集,transcript 正文也应该继续分页、延迟加载。需要守住的安全边界包括:
- 延迟恢复绝不能重新执行已经结束的工作;
- 非终态任务和 scheduled task 不能被当作普通冷历史延迟处理;
- 冷的终态 Session 如果损坏,应隔离该 Session,而不是阻止整个 Host 就绪;
- 打开冷 Session 或向其发送消息前,必须先通过该 Session 的恢复屏障。
这样可以把启动成本从
O(所有 Session 历史)降到接近O(待处理状态 + 当前请求的 Session)。对于 issue 中的工作区,这意味着先恢复实际使用的少数 Session,而不是在就绪前解码约 80 个 Session、10 万行历史。你们是否接受这样的契约:终态或归档历史的损坏变成在访问时发现的 Session 级错误,而只有需要采取动作的恢复状态保持全局 fail-closed?如果可以,可以拆成几个增量 PR:先移除不改变语义的重复读取,再加入 actionable index / 范围化恢复,最后加入按 Session 的按需恢复屏障。
-
take
The duplicate-read slice described in the existing discussion is already present on current main (
2df11b25b2): recovery now useslistHeaders()and performs onereadMessagesForRecoveryper Session. I verified the current source and will not open a duplicate PR; the remaining O(total history) validation needs the maintainer contract decision described above.untake
take
@me2seeks @Astro-Han Following my
take, I audited the remaining recovery path and built a reusable startup benchmark. Before changing production recovery, I'd like agreement on two contracts below, building on @sunheyi6's earlier readiness/validation split.What remains on main
The audit and benchmark use
19971f3f2ce4a865d1479e2b12f1eabc5f03980c. The subsequent main commit914242b34does not change the recovery paths discussed here.- perf(runtime-host): list recovery Session headers without the duplicate message decode (#4032) #4033 already removed the old duplicate message pre-read. Since refactor(runtime): converge AgentRun metadata into the RuntimeInvocation event spine #4631/refactor(runtime): derive Session transcripts from RuntimeEvents #4879, invocation boundaries and Session transcripts are derived from RuntimeEvents, so this work should use those current authorities.
- Full-history work remains in storage initialization, steering-proof validation,
prepareHostedExecutionRecovery, strict SessionManager recovery, and root-coordinator recovery. In particular,refreshToolLedgerHealth()reads and decodes the whole event table on store construction. Removing only the discarded reads in prepare would leave this cost. - The default election deadline is now 75 s after fix(runtime-host): tolerate slow Windows ACL startup #5323, which accommodated Windows ACL startup. That change does not remove the history scan.
Two behavior changes to decide
- Readiness establishes safe ownership and recovery of unsettled work; it does not certify every cold historical payload. Validate settled history when it is displayed or used as execution/recovery evidence. A corrupt, unrelated historical record would then fail the operation that uses it rather than global Host readiness. Shared ownership/control-state corruption and corruption in facts required for startup recovery remain fail-closed.
- Tool-ledger health checks cover the affected execution and its explicit dependencies, rather than every Session in the workspace. This changes the existing contract: a current test intentionally rejects a healthy Session's tool write because another Session has an orphan response. Keeping that global gate but deferring it until the first tool call would move the stall rather than remove it. Existing corruption must remain distinct from an invalid new candidate; cross-invocation identity constraints and referenced evidence must still be checked.
I favor both changes. They preserve RuntimeEvents as the authority while changing when unrelated history is read and how far its failure propagates. They are contract changes, not a semantics-preserving read deduplication.
The intended failure boundary would be:
Condition Proposed behavior State Root ownership, schema/database access, or shared control state cannot be trusted Refuse readiness or the affected authority operation; do not interpret an unreadable authority as an empty one. An unsettled execution requires a corrupt prompt, receipt, tool fact or continuation source Fail closed on that recovery path. Deferring cold history must not bypass a required dependency. An unrelated, settled historical payload is corrupt Detect and report it when that range is consumed; keep independent Sessions usable. A new tool event is invalid but its existing ledger is healthy Preserve candidate rejection, rather than poisoning the Run as if the store were already corrupt. A read exceeds its budget, is cancelled, or encounters a transient failure Report that condition separately; do not permanently label the Session corrupt or silently truncate execution evidence. "Ready" would mean that old execution obligations have a verified owner and a safe recovery/parking path, not that every long-running task has finished. Opening a Session would not itself certify the entire Session: display reads validate their requested range, while execution validates the complete evidence it actually depends on. A UI page is not proof that the rest of the history is absent or safe.
This necessarily gives up eager discovery of arbitrary corruption in unrelated cold history. It does not authorize rewriting historical facts, replaying an already settled effect, or trusting a projection over contradictory canonical evidence. Physical database failures and cross-execution identity conflicts are not reclassified as harmless historical errors.
Proposed implementation boundary
Start with existing opening/terminal indexes and canonical admission, claim, message and domain records. Select recovery work at Run/Turn granularity, then load its required evidence. Filtering only by Session would still load a long active Session's entire settled history.
The startup set must include admission-without-Run, non-terminal execution, incomplete continuation/handoff, pending message handoff, interactions/tool outcomes/resource closures, and outstanding scheduled-task/Goal/Graph/WorkHub obligations. A physical terminal is not sufficient: a handoff successor or domain settlement may still be pending. Archive status is not an exclusion rule.
Concretely, the selection should be a conservative superset:
Evidence of unfinished work Why it remains startup-critical Durable admission with no invocation opening The request was accepted before the process died; it must not disappear. Opening without a valid ending The model/tool execution may have stopped between durable boundaries. Handoff seal, pending continuation claim or incomplete repair start A physical Run may be terminal while its logical Turn still needs a successor; generic failure repair must not take ownership from the claim saga. Pending Message admission or unmaterialized root/steering handoff An execution can be terminal before message-consumption bookkeeping is settled. Pending interaction, unknown client-capability/tool outcome, orphan resource The previous process's ownership is gone; an effect may have happened without a recorded result. Pending scheduled fire, Goal continuation, Graph/WorkHub claim/wake/assignment The domain's obligation may outlive the underlying Run. Its owner must settle or take over that obligation before new scheduling races it. All of these need their existing domain validators. Selection is not a second RecoveryResolver. In particular, dispatch-without-outcome must retain its current conservative disposition rather than become automatic retry.
I would retain admission-chain/tip initialization, existing admission gates, HostEpoch ownership, recovery phase ordering, and the current side-effect/continuation proof rules. The first version would still pay for Session headers, invocation boundaries and admission control records; it would eliminate cold payload traversal, not claim constant startup cost or strict O(active work).
The code changes would narrow storage health validation, replace Session-wide prompt/steering scans with keyed evidence reads, and make prepare, strict recovery and root recovery consume the same selected scope. Necessary decoded evidence should be reused within a recovery step. New reads would go through the authenticated storage interfaces with equivalent SQLite/Memory behavior. As discussed in #4876, row/byte budgets must apply before materialization, and truncation must never be interpreted as absence.
Several details seem important to keep explicit:
- Both native opening events and legacy invocation openings must participate. Admission/claim/resource references to a missing opening are gaps to investigate, not evidence of no work.
- Preserve the existing first-terminal interpretation for legacy history; do not substitute "the last row has terminal status." Actual handoff/continuation still requires the original seal, lineage, high-water and prefix-digest checks.
- Match a recovered root prompt by Turn and content, preserving older event-ID conventions and folded-message behavior. Looking only for the newly derived prompt ID could append a second prompt.
- For tool health, the current transition path already reads the candidate invocation after the global gate. Reuse the common scanner for existing/prospective facts, preserving
ToolLedgerCorruptionErrorversusToolLedgerRejectionError. Event/operation identity and explicit cross-invocation references still require canonical checks; a missingtool_operationsprojection cannot establish that an operation never existed. - Control-record inventory is still work. In particular, admission inputs can be large. Measure their bytes as well as their count, and make any necessary full traversal explicit about its boundary and restart behavior.
- Keep mutations serialized by the existing owners/gates, and revalidate evidence after recovery writes. A startup inventory cannot be treated as permanently current once repairs or resumed work append new facts.
Initially I would avoid a new durable recovery-index table, a persistent
cold/recovering/readystate machine, a Worker pool, and an automatic full-history scan after ready. Add targeted indexes only where query plans justify them, and measure their one-time migration cost separately. Existing facts appear sufficient for the first step; a durable projection can be reconsidered if control-record enumeration itself becomes the measured bottleneck.The reason to defer those alternatives is specific:
Alternative Trade-off A larger deadline or concurrent Promises Leaves the amount of synchronous SQLite/JSON work unchanged. Move the whole scan to a Worker or immediately after ready Can improve responsiveness or the ready timestamp, but still pays the scan and may compete with the first user operation. Recover only recently opened Sessions Misses unattended domain work and does not fix a large active Session's old history. A durable pending-work index Could later reduce control-record enumeration, but requires atomic maintenance across admission, events, claims and domain settlement, plus complete backfill/rebuild rules. A terminal alone cannot clear every obligation. A per-Session recovery state machine Adds another lifecycle before it is needed: this first version still recovers all unsettled obligations and initializes admission ownership before ready; only historical reads are demand-driven. If admission ownership itself becomes lazy later, that would need a deduplicated authority-loading barrier across every execution entry point. It should not be smuggled into this change as a full-transcript load on first click. I would also leave the recent Desktop transcript/virtualization work in #5366 out of this issue's implementation scope.
Baseline already measured; no recovery optimization implemented
The local benchmark starts the real Host candidate/kernel/composition against isolated SQLite data: 80 unarchived Sessions, 500 completed Runs/admissions, 100,000 RuntimeEvents and 25,000 operational diagnostic records. Test text/diagnostic bodies are 512 ASCII bytes. It is a text-dominated settled-history case, with no tools, pending domain work or model calls.
On Apple M2 Pro / macOS arm64 / Node 24.19.0 / SQLite 3.53.3, after one warmup, three uninstrumented fresh-process trials took 5.163 / 5.735 / 6.060 s from process launch to Host ready. Three alternating instrumented trials had a 5.733 s median. For the middle instrumented sample:
Stage Time Share of launch → ready Prepare recovery 1.702 s 29.7% Strict SessionManager recovery 1.665 s 29.0% Root-coordinator recovery 0.613 s 10.7% Store-wide tool-ledger health scan 0.431 s 7.5% Steering-proof validation 0.285 s 5.0% Union of these stages 4.696 s 81.9% Other startup work 1.037 s 18.1% Holding the other counts fixed and reducing RuntimeEvents to 10,000 lowered the uninstrumented median to 2.955 s. The 100k fixture's median peak RSS was approximately 781 MiB, including module initialization.
These are fresh processes with warm OS filesystem caches, not disk-cold measurements or a reproduction of the reporter's 135 s workload. They exclude CLI/TUI rendering and the full client election loop. Stage attribution uses interval unions to avoid nested double-counting; 81.9% is time in scope, not a promised saving. Three samples do not establish p95. Each trial asserts zero model activation and unchanged canonical event/admission/operational row counts and SHA-256 digests.
For reproducibility: the measured modules were transpiled from the pinned source with esbuild, without bundling, using the installed workspace dependencies and offline generated metadata. This is not a packaged-release measurement or a substitute for the repository build/typecheck/test gates. Fixture creation is outside the timer: admissions use the real store; RuntimeEvents pass the canonical encoder and are bulk-inserted into the real schema with their ordinal/Turn indexes. The workload fixes counts, event timestamps and body sizes; temporary paths and Session UUIDs vary between benchmark invocations, while all trials within an invocation use the same canonical history. Runtime policy bootstrapping is disabled, and the benchmark fails if a model backend is activated.
The runner and raw samples are currently local; I would include them in the first PR and retain this fixture as the before/after baseline for each optimization. Comparisons would keep fixture version, runtime/dependencies, hardware and cache conditions comparable, report readiness time and RSS, and repeat the 10k/100k growth control. First Session access/tool execution and index migration should also be measured so work is not merely shifted past ready. Tool-rich and active-recovery fixtures need separate coverage; this settled-history benchmark does not replace crash/restart, unknown-outcome, legacy receipt, handoff, archive or concurrency tests.
The fixture currently demonstrates where the time goes, not the correctness or achieved speedup of the proposed design. I would use the following acceptance checks together:
- Compare the unchanged 80-Session/500-Run fixture before/after each optimization, retaining individual samples and both instrumented and uninstrumented results.
- Keep the 10k/100k cold-history comparison with control-state counts fixed. Separately increase Session/Run/admission counts so an expensive control scan is not hidden by only testing event density.
- Verify that a long Session with one unsettled Run does not cause every settled Run in that Session to be decoded.
- Inspect actual query plans and rows/bytes read. A small result set or a cursor loop that ultimately loads everything is not evidence of bounded work.
- Preserve zero unnecessary model activation and unchanged canonical settled history in the baseline. Add crash-window and repeated-recovery tests for the paths that do legitimately write repairs.
- Exercise cold corruption versus required-evidence corruption, duplicate identities, legacy receipts, terminal-but-pending handoff/domain work, and concurrent opening/sending/scheduling. Keep SQLite/Memory behavior aligned.
- Report any budget-exceeded behavior and residual costs explicitly. Do not count an early failure or incomplete recovery as a faster successful startup.
Suggested delivery sequence
- Publish the reusable benchmark, raw baseline samples and measurement recipe; record the agreed contract and add the relevant behavioral coverage. The current benchmark can be reviewed without first accepting an optimization.
- Narrow storage health checking and failure propagation, preserving the common tool-ledger rules and identity constraints. This alone does not resolve the Host-layer scans, so it would not close fix(runtime-host): Host recovery decodes every Session's full history before readiness, breaching the 45s election deadline #4032.
- Select recovery work from existing authorities and apply that scope across prepare, strict recovery and root recovery; replace the prompt/steering full-history reads with targeted evidence queries. Keep legacy and domain ownership behavior explicit.
- Verify the complete ready/first-access path with the same baseline, query-plan checks and crash/concurrency coverage. Add batching, event-loop yields or further indexing only where the resulting profile justifies them.
Each behavioral PR would report its before/after measurements and relevant correctness checks. The main acceptance condition would be no cold Run body/operational full-read before ready unless an unsettled obligation actually depends on it. This fits the log-as-authority design without turning routine restart into a mandatory historical audit.
Would you accept these two contract changes and that implementation boundary? In particular, is there an intended guarantee behind the workspace-wide tool-ledger failure gate that requires retaining it? If so, I would revise the scope around that guarantee before implementation.
I would also appreciate corrections to the startup-obligation list or a preference for a smaller first behavioral slice. The ownership/failure contract should determine that boundary; the measured 82% is a way to prioritize work, not a reason to weaken it.
中文要点
认领后重新核对了 main,并完成真实 Host 的独立启动基准;目前没有修改产品恢复逻辑。希望先确认两个契约,再实施:
- ready 保证未结算工作有可信的所有者和恢复路径,不保证所有冷历史正文都已完整审计。 已结束历史在展示或作为执行/恢复证据使用时校验。共享控制状态损坏、启动恢复真实依赖的证据损坏仍 fail-closed。
- 工具账本健康检查限定到受影响执行及其明确依赖。 当前测试刻意让其他 Session 的 orphan response 阻止健康 Session 的工具写入,建议改变这一全局影响范围;已有损坏与新候选非法必须继续区分,跨 invocation 身份和证据约束不能丢。
这两项都会改变当前契约。只将全局检查延迟到第一次工具调用,会把卡顿移到首次执行,不能达成本次目标。
旧的重复消息预读已由 #4033 修复。当前范围包括 store 构造时的全表工具账本检查、steering 历史校验、prepare、SessionManager strict recovery 和 root recovery。默认等待上限从 45 秒改成 75 秒属于 #5323 对 Windows 启动的兼容,不是历史扫描的修复。
实施上优先复用现有 opening/terminal 索引和 admission/claim/message/domain 权威记录,按 Run/Turn 及其依赖选择恢复范围。不能只看 terminal 或 archived:admission 无 Run、handoff successor、continuation claim、消息消费和定时任务等领域收尾都可能尚未完成。保留 admission 链/tip、接纳门、HostEpoch、恢复阶段顺序和原有副作用/前缀证明。首版仍有控制记录枚举成本,不宣称严格 O(active work)。
初期不引入通用持久恢复索引、Session 加载状态机、Worker pool 或 ready 后的自动全历史扫描。它们需要额外的事务维护和生命周期,应等测量证明必要。新的读取经现有存储接口提供,SQLite/Memory 语义一致;字节和条数预算在物化前生效,截断不代表不存在。#5366 的 Desktop 历史策略不纳入本次重做。
基准固定在
19971f3f2:80 个未归档 Session、500 个已结束 Run/admission、100k RuntimeEvents、25k operational diagnostics,测试正文 512 字节,没有工具调用或待处理领域工作。M2 Pro / Node 24.19.0 / SQLite 3.53.3 上,三次无插桩进程启动到 ready 为 5.163 / 5.735 / 6.060 秒,中位数 5.735 秒。居中的插桩样本为 5.733 秒,其中本次范围的区间并集为 4.696 秒,占 81.9%。保持其他数量不变,将 RuntimeEvents 降为 10k,启动中位数为 2.955 秒。这是新进程、暖 OS 文件缓存的合成测试,不是原作者 135 秒工作区的复现,也不包括完整选举和 CLI/TUI。82% 不是承诺可以全部节省,三个样本也不能推导可靠 p95。每次都验证无模型激活、canonical 三类历史的行数与 SHA-256 不变。构建方式、数据生成方式和测量限制已在英文正文说明。
脚本及原始样本目前在本地,将纳入第一个 PR,作为之后每次优化的统一 before/after 基准。除了 ready 时间、RSS 和 10k/100k 增长对照,还要测首次访问/工具执行、索引迁移、控制记录增多,以及工具密集/活跃恢复场景,防止把开销移到别处。基准不能代替 crash window、legacy receipt、handoff、未知副作用和并发正确性测试。
建议先提交基准和契约覆盖,再分别处理存储健康检查范围、Host 恢复范围与定向查询,最后验证完整路径。主要请确认两项契约,尤其全局工具账本失败门是否承载必须保留的保证;也欢迎补充遗漏的恢复义务或建议更小的行为切片。
AI disclosure: Codex assisted with the source audit, benchmark and drafting; this comment is posted by Codex on behalf of @testikun.
Decision — contract changes in #4032
To be precise about the guarantee question, since it changes the framing: the workspace-wide tool-ledger fail-stop is not an accident. It is a deliberate first-version choice recorded in
docs/architecture/runtime-resume-extraction-ledger.zh-CN.md§3.1, anddocs/architecture/runtime-resume-phase3-phase4-workspace-checkpoint-design.zh-CN.md§85-89 already records the intended narrowing path — an incremental reducer rebuilt from immutable events, scoped to the candidate execution spine, not a mutable cache replacing the fact authority.So the answer is: no rule requires the workspace scope. What the design requires is that tool-fact validation stays derived from immutable events, runs the same shared prospective transition validator inside the write transaction, and stays fail-closed for the affected scope. Non-tool events remaining writable was already part of the PR A trade-off.
I accept both contract changes and the implementation boundary, with these constraints:
- Readiness establishes a verified owner and a safe recovery/parking path for unsettled work; cold payloads are validated when consumed. Shared ownership/control state and startup-required evidence remain fail-closed.
- Scoped ledger health: scope the check to the candidate execution spine plus explicit dependencies (
parentOperationId/parentToolCallId), computed by an incremental reducer over immutable RuntimeEvents. KeepToolLedgerCorruptionErrorvsToolLedgerRejectionError, the bug(runtime): a corrupt ledger costs a run its terminal fact via the latch, not via the corruption #2313 terminal-write behavior, and cross-invocation identity/reference checks. Corrupt canonical history must never be read as absent. - Boundary: Run/Turn-granularity selection, the conservative superset of unfinished work, and deferring a durable recovery index, Session state machine, Worker pool and post-ready full scan are approved as proposed; revisit only with measurements.
- Sequence: benchmark + contract coverage → scoped storage health (first) → workset selection (this closes fix(runtime-host): Host recovery decodes every Session's full history before readiness, breaching the 45s election deadline #4032) → end-to-end verification. Require before/after on the retained fixture, add a tool-rich / active-recovery fixture, and measure first access and first tool execution so cost is not merely moved past ready.
- Documentation and tests ride with the change: update the two design docs to record the narrowed contract, reconcile the comments that currently disagree (
tool-ledger-scanner.tsclass doc vssqlite-runtime-store.ts:3771-3776), and update the two tests that pin workspace-wide behavior (sqlite-runtime-store.test.ts:488-534,recovery-persistence-authority.test.ts:869-891) with explicit intent.
Post-merge Host startup check on the reported machine
I ran the merged #5556 Host startup path on the same WSL machine used for the original #4032 report (Intel N100, 4 vCPUs, 7 GiB RAM, Linux x86_64, Node 26.3.0).
To avoid stopping the Host currently in use or modifying its data, I took a SQLite-consistent snapshot of the current workspace database and started fresh Runtime Host candidate processes against an isolated copy. The original Host remained running. The snapshot was schema 20, 1.83 GiB, with 145 Session metadata rows, 809 core Agent Run rows, 766 root-turn admissions, 74,709 RuntimeEvents, 47,499 Session messages, and 17,566 tool operations. All 766 admissions had terminal Runtime evidence; there were no queued message admissions or continuation claims. I paused the two active scheduled tasks in the test copy and removed its pending supervisor wake rows to prevent duplicate work. Credential-vault and MCP configuration were not copied; no model/tool execution was requested.
Three fresh Node processes reached the Runtime Host registration state
readyin:Trial Host process start → ready1 2,327.2 ms 2 2,378.2 ms 3 2,357.4 ms Median 2,357.4 ms The original database and its WAL were left untouched. This was a fresh-process test with a warm OS file cache, and measured Host readiness only—not TUI rendering or time to selected history. Because this is a later snapshot (145 Sessions / 809 Runs versus the issue-time ~80 / ~500), it is not a paired before/after benchmark on identical rows, so I am not presenting a direct percentage reduction against the issue's ~135 s observation. It does confirm that the merged Host reached
readyin about 2.36 s on this machine using the actual current workspace history, comfortably within the 45 s election deadline.The isolated run omitted provider credentials and logged a models.dev catalog fallback warning (missing
kimi-for-coding); it still reachedreadyon all three trials using the bundled catalog snapshot.
What happened
Runtime Host startup recovery decodes the entire durable execution history of every recoverable Session before reporting readiness. On a long-lived workspace (~80 recoverable Sessions, ~500 runs, ~100k event/message rows in
runtime.sqlite) recovery takes ~150s of mostly main-thread CPU, which breaches the client's default 45000ms election deadline:maka --resume <id>fails with "did not become ready before the startup deadline elapsed" (lastRegistration.state: "recovering",readyWaitFailed: 1) even though the Host is making progress. Retrying after the Host settles works.This is the residual bottleneck after the artifact-store costs in #4027 are removed; CPU profile of a recovering Host (CDP Profiler on the main thread) shows:
readSqliteAgentRunEvents(packages/storage/dist/agent-run-store.js)all()callsdecodeStoredRuntimeEvent, GC pressureRoot causes in
packages/runtime-host/src/server/hosted-execution-recovery.ts(prepareHostedExecutionRecovery):readEventsForRecovery+readRuntimeEventsand discards the results — a full decode of every event of every run, purely as a fail-closed consistency check.listForRecovery()(packages/storage/src/session-store.ts) already callsreadMessagesForRecoveryper header and discards the result;prepareHostedExecutionRecoverythen re-reads the same messages per Session.all()and JSON decoding are synchronous main-thread work, so concurrency alone would not help — the fix has to reduce the amount of work, not interleave it.Because admission/message repair decisions only involve Sessions with admissions, non-terminal runs, or pending closures, decoding the history of long-settled Sessions appears unnecessary for correctness — it is startup-time insurance with O(total history) cost paid on every Host start (upgrade, epoch cutover, crash recovery).
How to reproduce
maka --resume <session-id>within the first ~2 minutes.Environment
d27c02af1(main)Logs, screenshots, or additional context
Election diagnostic from the client:
{"deadlineMs":45000,"elapsedMs":45008,"candidateLaunches":1,"sawEndpointConnected":true, "observations":{"readyWaitFailed":1,"connected":1}, "lastRegistration":{"state":"recovering","lifecycleMode":"ephemeral"}}Related: #4027 (artifact-store cold-start costs; fix in #4031).
Suggested directions (needs a design decision on the startup-validation contract):