Skip to content

perf(runtime): batch Recall session reads to avoid N+1 queries and repeated decoding #5877

Description

@liuxiaocs7

What happened

Recall narrows candidates to Sessions, then reconstructs each candidate's complete message view before ranking passages. A long Session therefore pays for repeated event reads and per-run SQL even when only ten passages match.

At ed38ccbb6bac474b93e4517f2a12a30209c108d5, a 20,000-Turn fixture returning ten passages takes about 1.60 seconds. One complete RuntimeReadModel read executes 80,002 SQL reads and 160,000 JSON.parse calls:

  • The invocation inventory reads openings and separately fetches each invocation's terminal event.
  • The ordinal reader loads and decodes all event payloads just to build an event-ID-to-ordinal map.
  • The read model then rereads each run's events, partial snapshots and partial segments, including empty partial queries for completed runs.

There is also a fallback duplication: if the candidate path's corpus count returns null, collectHits restarts a full scan after already reading its candidates. A two-Session fixture reads first, first, second.

Expected: batch the complete-view reads within a consistent snapshot, decode each durable event once, and resolve count fallback without rereading candidate Sessions. Preserve ranking, redaction, passage contents/navigation, event order, legacy openings, partial presentation and corruption checks.

How to reproduce

  1. Seed one Session with 20,000 completed Turns. Each Turn contains an invocation-opening event, one user text event and a terminal event, with Session ordinals in commit order.
  2. Put deploy target in the first ten user events and unrelated text in the rest.
  3. Run Recall with { terms: ['deploy'], limit: 10 } through SQLite and RuntimeReadModel.
  4. Count executed all/get/iterate calls and JSON.parse calls during a separate getSessionMessages read. Counts are deterministic: 4 * turns + 2 SQL reads and 8 * turns parses for this fixture. Even two Turns require ten SQL reads and sixteen parses.
  5. For the fallback case, offer only the first of two eligible Sessions as a candidate and return null from countSearchableMessages; record readMessages calls.

Implementation PR #5880 includes regression tests; benchmark results and measurement methodology are recorded in its description. SQL counts exclude transaction-control statements, candidate selection and corpus counting; timings cover Core Recall, SQLite and projection, excluding IPC, UI and fixture setup.

Environment

  • Commit: ed38ccbb6bac474b93e4517f2a12a30209c108d5
  • macOS 15.7.3, Apple M4 Pro, arm64
  • Node.js v24.14.0; in-memory SQLite synthetic fixture
  • Surface: Runtime Host Recall, shared by model tools and Desktop history search

Logs, screenshots, or additional context

Relevant sources: complete read model, invocation inventory, and Recall collection/fallback.

Related background: #5368, #5531 and #4677. This is separate from the long-term-memory item batching fixed by #5263. Search indexing and Host request cancellation are outside this issue; stopping after the first ten matches would change the existing ranking contract.

Prepared and submitted with Codex assistance on behalf of @liuxiaocs7.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions