Skip to content

fix: start the answer right after research, read single videos whole, and name finalization phases - #169

Open
devhims wants to merge 6 commits into
mainfrom
fix/finalization-latency
Open

devhims wants to merge 6 commits into
mainfrom
fix/finalization-latency

Conversation

@devhims

@devhims devhims commented Oct 7, 2026 •

Copy link
Copy Markdown
Owner

Problem / Motivation

Users wait about 20 seconds between research finishing and the first answer text, under a label that reads "Writing and checking the answer." the whole time. Measured on session 028dcd5d (run 43c01730, 66.8 seconds total):

Seconds into the run What happened Duration
39.2 Research hands off to the finalizer
39.2 to 52.4 Context gathering: three sequential model calls 13.2 s
52.4 to about 59 Answer call reasoning before any visible text about 6 s (estimate)
about 59 to 65.2 Answer text streams about 6 s
65.2 to 66.8 Checks and save 1.5 s

The context-gathering step found nothing. It read history from a one-message session, ran three evidence searches that returned nothing, and re-read the four transcripts research had analyzed seconds earlier; all four reads returned zero excerpts. The third call spent 6.7 seconds and about 1,000 tokens of reasoning to decide to stop. The reasoning estimate assumes an even output rate, because logs record total reasoning and text characters, not when text began.

What changed

1. Answer right after research, when nothing would be lost. Research and the answer are separate model calls: the answer receives the conversation window and the evidence projection, never research's own tool results. When research ran in the same process immediately before finalization, the finalizer skips the model-driven context loop only if both hold:

  • every stored history message is the current request or part of a turn already in the answer prompt, compared by stored message ID, so research could not have read older conversation the answer would miss. Counting messages is not enough: evidence deletion removes stored answers while the prompt keeps their turns. Without a request ID, the finalizer assumes older conversation and gathers;
  • the evidence projection fits the 40,000-character budget without sampling, shortening or dropping a packet, so every passage research saw reaches the answer.

Otherwise it gathers as before. Each decision is logged as agent_finalizer_context_plan (skip, gather, gather_older_conversation, gather_evidence_cut).

Unchanged:

  • Deterministic comparison preloads still load every compared video's saved transcript.
  • finalize routes (history and context questions) still gather.
  • Resumed runs still gather, because their research ran before a restart and stored context is how they recover it.

A fuller fix would have research forward what it read (for example an older budget constraint, or the passages it relied on) so the shortcut can apply to long conversations too. The plan log will show how often the conservative gate gathers before deciding whether that is worth building.

Single-video questions are answered from the whole transcript. Inspection gave the answer model the transcript trimmed to 40,000 characters. Per-caption JSON overhead fills that by about ten minutes of video (a 13-minute transcript projects to 47,000 characters, of which 11,000 are words), so the finalizer kept guessing keyword searches to find passages. Inspection now uses a 250,000-character budget, about an hour of transcript, so the whole transcript reaches the answer model and the stored-context pass is skipped. A transcript returned in pages counts as incomplete and keeps the pass.

Live runs with real Fireworks models (production settings), real research and finalizer code, and two production transcripts from slow runs:

Case Condition Research → first text Total Correct
Screening times near the end of a 13-min video (run 8e11bee1) Before 41.4 s / 13.0 s 64.8 s / 26.0 s 4/4, 4/4
After 2.8 s / 2.3 s 13.2 s / 11.0 s 4/4, 4/4
Key takeaways of a 22-min video (run df7b32ab) Before 34.5 s / 40.8 s 51.4 s / 56.8 s n/a
After 9.9 s / 6.3 s 30.8 s / 22.3 s n/a
Code word at 90% of a synthetic 1-hour transcript After 8.0 s / 11.3 s 22.1 s / 32.7 s 1/1, 1/1

Limits found while testing:

  • A 2-hour transcript is about 237,000 tokens. First content took 9.7 s on DeepSeek and exceeded the 10-second failover limit on GLM, so one run failed. That is why the budget stops at about an hour; longer videos keep the existing path.
  • Research also puts the full transcript in its own prompt (about 100,000 tokens for an hour), taking 7 to 10 s and causing one GLM fallback. Not changed here: research sometimes needs the transcript, for example to pick frame timestamps.

2. Honest progress labels. Streaming drafts carry an optional activity, and the dashboard names it:

Activity Label
gathering (stored-context loop running) Gathering context.
thinking (answer call started, no answer text yet) Thinking.
writing (answer text arriving) Writing the answer.
repair attempt Revising the answer.
no draft yet Preparing the answer.

activity is optional in both the platform and web schemas, so either side can deploy first. The OpenAPI description and generated spec include it.

Also corrected three statements in SESSION_EVIDENCE.md made stale by #167: first-message handling, ignored inspection requests, and acceptance having no await.

Expected effect

Up to about 13 seconds less before the answer starts on runs like the measured one, a new session whose evidence fits the budget, and the remaining reasoning wait is labelled "Thinking." instead of looking stalled. Not yet verified in production: the measured run's evidence size was not captured, so whether it would skip is unconfirmed. After deploy, agent_finalizer_context_plan shows the share of runs that skip and why the others gather.

Tests

All 1,617 platform unit tests, 350 runtime/session integration tests and 106 web tests pass, after merging main (#171 to #177). Platform and web type checks and the docs check pass. New tests cover:

  • a single video answered from its whole transcript (110,000 characters) in one call, a transcript past the single-video limit still gathering, and a paged transcript still gathering
  • fresh research going straight to one structured answer call, with no context model call, searches or reads
  • review regressions: fresh research still gathers when the conversation is older than the prompt window (the ₹13,750 budget case) and when the evidence budget would cut a passage
  • the projection's complete flag, including comparison evidence outside every subject
  • review regression on the real session store: 10 turns with 4 deleted answers, where a count matches (17) but older messages remain, is detected by ID
  • the IDs the finalizer treats as covered (current request plus each prompt turn), and a run without a request ID keeps gathering
  • a resumed run that still gathers
  • a history/context question that still gathers and emits gathering before thinking
  • the draft activity sequence, including repair
  • every progress label, including drafts from an older platform without activity

The session-evidence integration test now expects one finalizer call after fresh research (previously three) and keeps three for the resumed variant.

Scope

Not included: lowering finalizer reasoning effort (a config change that needs a quality check), and the research-phase delay in the same run, where a GLM transcript analysis hit its 14.8-second budget and the run fell back to DeepSeek.

Compatibility and deployment

No migration. Additive optional field in run progress.

After research, the finalizer ran a model-driven pass over stored context
before writing the answer. In run 43c01730 that pass made three model
calls over 13 seconds, read history from a one-message session, ran
three empty searches and re-read the four transcripts research had just
analyzed, adding no excerpts. Research routes never ask about earlier
conversation and research has already loaded this turn's evidence, so
when research ran in the same process the finalizer now answers
directly. Deterministic comparison preloads still run. Finalize routes
and resumed runs, whose research ran before a restart, still gather.

The dashboard showed "Writing and checking the answer." for the whole
finalization, including about 6 seconds of invisible reasoning before
any text. Streaming drafts now carry an activity (gathering, thinking
or writing) and the progress label names it.
@vercel

vercel Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
video2ctx-web Ready Ready Preview Oct 9, 2026 4:23pm UTC

…arch saw

Research and the answer are separate model calls. The answer receives the
conversation window and the evidence projection, never research's own
tool results. Skipping the stored-context pass after any fresh research
could therefore drop:

- a constraint from conversation older than the prompt window, which
  research read but the answer model never receives;
- a passage research saw that the 40,000-character finalization budget
  samples, shortens or drops.

The shortcut now also requires that the session holds no messages beyond
the prompt's turns and that the evidence projection fits its budget
without cutting anything. Otherwise the finalizer gathers as before.
Each decision is logged as agent_finalizer_context_plan so the share of
runs that skip can be measured.
The history guard compared the stored message count with two messages
per prompt turn plus the request. Evidence deletion removes affected
assistant answers from stored history while the prompt keeps those
turns, so the count can match while older messages remain. With ten
turns and four deleted answers, 21 - 4 = 17 = 8 x 2 + 1, the guard
skipped gathering and an older budget never reached the answer.

The run context now carries the stored ID of the current request, and
the session store reports whether any stored message lies outside that
request and the prompt turns' messages. Without a request ID or that
check, the finalizer assumes older conversation and keeps gathering.
Since #173, research with a configured finalizer ends by calling
complete_research; finalize_answer is no longer offered. The context
gathering tests still simulated finalize_answer and passed through an
unknown-tool path. They now use complete_research, the production path.
Single-video inspection gave the answer model the transcript trimmed to
40,000 characters, which per-caption JSON overhead fills by about ten
minutes of video. The answer model then ran a stored-context pass to
find passages by keyword. In live runs with real models on two
production questions, that pass took 7 to 29 seconds.

Inspection now uses a 250,000-character evidence budget, about an hour
of transcript. The whole transcript reaches the answer model, nothing is
cut, and the existing rule skips the stored-context pass. Measured from
research completion to first answer text: 41.4/13.0 s to 2.8/2.3 s and
34.5/40.8 s to 9.9/6.3 s, with the needle question answered correctly
in every run. One-hour transcripts reached first content in about 7
seconds, inside the 10-second failover limit. Longer transcripts exceed
the budget and keep the stored-context pass.

A transcript returned in pages now counts as incomplete, so it also
keeps the stored-context pass.
@devhims devhims changed the title fix: start the answer right after research and name finalization phases fix: start the answer right after research, read single videos whole, and name finalization phases Oct 9, 2026

This branch was successfully deployed

1 active deployment
Preview — b472f4b3 Deployed Oct 9, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant