Repository navigation
fix: start the answer right after research, read single videos whole, and name finalization phases - #169
Open
devhims wants to merge 6 commits into
Open
fix: start the answer right after research, read single videos whole, and name finalization phases#169devhims wants to merge 6 commits into
devhims wants to merge 6 commits into
Conversation
After research, the finalizer ran a model-driven pass over stored context before writing the answer. In run 43c01730 that pass made three model calls over 13 seconds, read history from a one-message session, ran three empty searches and re-read the four transcripts research had just analyzed, adding no excerpts. Research routes never ask about earlier conversation and research has already loaded this turn's evidence, so when research ran in the same process the finalizer now answers directly. Deterministic comparison preloads still run. Finalize routes and resumed runs, whose research ran before a restart, still gather. The dashboard showed "Writing and checking the answer." for the whole finalization, including about 6 seconds of invisible reasoning before any text. Streaming drafts now carry an activity (gathering, thinking or writing) and the progress label names it.
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
…arch saw Research and the answer are separate model calls. The answer receives the conversation window and the evidence projection, never research's own tool results. Skipping the stored-context pass after any fresh research could therefore drop: - a constraint from conversation older than the prompt window, which research read but the answer model never receives; - a passage research saw that the 40,000-character finalization budget samples, shortens or drops. The shortcut now also requires that the session holds no messages beyond the prompt's turns and that the evidence projection fits its budget without cutting anything. Otherwise the finalizer gathers as before. Each decision is logged as agent_finalizer_context_plan so the share of runs that skip can be measured.
The history guard compared the stored message count with two messages per prompt turn plus the request. Evidence deletion removes affected assistant answers from stored history while the prompt keeps those turns, so the count can match while older messages remain. With ten turns and four deleted answers, 21 - 4 = 17 = 8 x 2 + 1, the guard skipped gathering and an older budget never reached the answer. The run context now carries the stored ID of the current request, and the session store reports whether any stored message lies outside that request and the prompt turns' messages. Without a request ID or that check, the finalizer assumes older conversation and keeps gathering.
Since #173, research with a configured finalizer ends by calling complete_research; finalize_answer is no longer offered. The context gathering tests still simulated finalize_answer and passed through an unknown-tool path. They now use complete_research, the production path.
Single-video inspection gave the answer model the transcript trimmed to 40,000 characters, which per-caption JSON overhead fills by about ten minutes of video. The answer model then ran a stored-context pass to find passages by keyword. In live runs with real models on two production questions, that pass took 7 to 29 seconds. Inspection now uses a 250,000-character evidence budget, about an hour of transcript. The whole transcript reaches the answer model, nothing is cut, and the existing rule skips the stored-context pass. Measured from research completion to first answer text: 41.4/13.0 s to 2.8/2.3 s and 34.5/40.8 s to 9.9/6.3 s, with the needle question answered correctly in every run. One-hour transcripts reached first content in about 7 seconds, inside the 10-second failover limit. Longer transcripts exceed the budget and keep the stored-context pass. A transcript returned in pages now counts as incomplete, so it also keeps the stored-context pass.
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem / Motivation
Users wait about 20 seconds between research finishing and the first answer text, under a label that reads "Writing and checking the answer." the whole time. Measured on session
028dcd5d(run43c01730, 66.8 seconds total):The context-gathering step found nothing. It read history from a one-message session, ran three evidence searches that returned nothing, and re-read the four transcripts research had analyzed seconds earlier; all four reads returned zero excerpts. The third call spent 6.7 seconds and about 1,000 tokens of reasoning to decide to stop. The reasoning estimate assumes an even output rate, because logs record total reasoning and text characters, not when text began.
What changed
1. Answer right after research, when nothing would be lost. Research and the answer are separate model calls: the answer receives the conversation window and the evidence projection, never research's own tool results. When research ran in the same process immediately before finalization, the finalizer skips the model-driven context loop only if both hold:
Otherwise it gathers as before. Each decision is logged as
agent_finalizer_context_plan(skip,gather,gather_older_conversation,gather_evidence_cut).Unchanged:
finalizeroutes (history and context questions) still gather.A fuller fix would have research forward what it read (for example an older budget constraint, or the passages it relied on) so the shortcut can apply to long conversations too. The plan log will show how often the conservative gate gathers before deciding whether that is worth building.
Single-video questions are answered from the whole transcript. Inspection gave the answer model the transcript trimmed to 40,000 characters. Per-caption JSON overhead fills that by about ten minutes of video (a 13-minute transcript projects to 47,000 characters, of which 11,000 are words), so the finalizer kept guessing keyword searches to find passages. Inspection now uses a 250,000-character budget, about an hour of transcript, so the whole transcript reaches the answer model and the stored-context pass is skipped. A transcript returned in pages counts as incomplete and keeps the pass.
Live runs with real Fireworks models (production settings), real research and finalizer code, and two production transcripts from slow runs:
8e11bee1)df7b32ab)Limits found while testing:
2. Honest progress labels. Streaming drafts carry an optional
activity, and the dashboard names it:gathering(stored-context loop running)thinking(answer call started, no answer text yet)writing(answer text arriving)activityis optional in both the platform and web schemas, so either side can deploy first. The OpenAPI description and generated spec include it.Also corrected three statements in
SESSION_EVIDENCE.mdmade stale by #167: first-message handling, ignored inspection requests, and acceptance having noawait.Expected effect
Up to about 13 seconds less before the answer starts on runs like the measured one, a new session whose evidence fits the budget, and the remaining reasoning wait is labelled "Thinking." instead of looking stalled. Not yet verified in production: the measured run's evidence size was not captured, so whether it would skip is unconfirmed. After deploy,
agent_finalizer_context_planshows the share of runs that skip and why the others gather.Tests
All 1,617 platform unit tests, 350 runtime/session integration tests and 106 web tests pass, after merging
main(#171 to #177). Platform and web type checks and the docs check pass. New tests cover:completeflag, including comparison evidence outside every subjectgatheringbeforethinkingactivityThe session-evidence integration test now expects one finalizer call after fresh research (previously three) and keeps three for the resumed variant.
Scope
Not included: lowering finalizer reasoning effort (a config change that needs a quality check), and the research-phase delay in the same run, where a GLM transcript analysis hit its 14.8-second budget and the run fell back to DeepSeek.
Compatibility and deployment
No migration. Additive optional field in run progress.