Repository navigation
feat: stream model attempts and index their timing in D1 - #171
Merged
Merged
Conversation
Every failover-wrapped model call now streams, including generateText callers. The wrapper assembles the generate result from the provider stream, so a silent provider fails at its first-content limit while a call that keeps producing reasoning, text or tool arguments is no longer cut off at a fixed whole-response limit. Research roles allow 10 seconds to first content and 30 seconds overall. The classifier, visual analyst and memory updater keep their earlier whole-response limits, because Fireworks often takes 2 to 3 seconds to first content and their phases have fixed budgets. Each finished attempt records time to first content, time to completion, token counts and the provider request ID. The tool-trace publisher moves these rows to a new D1 table, agent_model_attempts, and the admin run trace returns them as modelAttempts. Run e3889923 failed after GLM and then DeepSeek each hit the earlier 10-second limit while reading a full transcript; its retry finished the same step in 7.7 seconds.
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem / Motivation
Session
9d0a2e06(rune3889923, 2026-10-09 06:35 UTC) failed withMODEL_FALLBACK_EXHAUSTEDafter all four tool calls had succeeded. Afterget_video_transcriptreturned, the next research step hitresponse_timeoutat 10,000 ms on GLM 5.3 Flash. The fallback fired, and DeepSeek V4.1 Flash then hit the same 10-second limit. In the retry run, the same step took 7.7 s (8,214 input and 1,332 output tokens).Research calls were not streamed, so the 10-second limit covered the whole response: queueing, prompt processing, reasoning and output. That made it impossible to tell a stalled provider from a model that was still writing. It also failed calls that were making steady progress.
What changed
Every failover-wrapped call streams.
withModelFailover'sdoGeneratenow reads the provider'sdoStreamand assembles the generate result itself (collectStream). Callers are unchanged, and switching to the backup is still invisible to them because nothing has been returned yet.agent_core,finalizer(generate)transcript_analystclassifiervisual_analyst,memory_updaterAll limits are still clamped to the phase deadline, less the 10 s reserve for the backup. "First content" means the first text, reasoning or tool-argument delta. Fireworks streams GLM reasoning progressively (probed at
higheffort: 180 reasoning chunks, the first one arriving as soon as response headers did), so this measures queueing plus prompt processing. A provider that stays silent until the overall limit is labelledfirst_content_timeout.I first tried a 3 s first-content limit for the classifier. A live run tripped it: Fireworks took 2 to 4 s to first content even for 58-token prompts, and a classifier fallback moves every later role in the run to DeepSeek. Roles with fixed phase budgets therefore keep their earlier whole-response limits.
Attempt timing is recorded. Each
attempt_finisheddiagnostic now carriesstartedAt,firstContentMs,elapsedMs, input, cached-input, output and reasoning tokens, the provider request ID, and the limits the attempt actually ran under. Finished attempts go into a Durable Object outbox,agent_model_attempt_outbox. The existing tool-trace publisher moves them to the new D1 tableagent_model_attempts(migration0023), with the same 15-second alarm retry. Rows hold IDs, timings and token counts only, never prompts or output.GET /v1/admin/agent-traces/{runId}returns them asmodelAttempts.Example query, time to first content by model and role:
Behavior change to review
An attempt canceled before its stream's
finishevent no longer reports usage. Before, a non-streamed primary that answered after its timeout still reported token usage, which went into the memory-update cost ledger. Now the canceled primary keeps its existingMEMORY_UPDATE_COST_RESERVE_MICROSreservation against the run budget, but its actual token cost is not recorded. This affects internal provider-cost tracking only. Memory updates charge users 0 credits.Deployment
Apply migration
0023before deploying the Worker (deploy:productionalready runsdb:migrate:productionfirst). If the Worker runs before the migration, D1 writes fail and rows stay in the outbox, retried every 15 s.Testing
npm run build(includes test type checking)npx vitest run: 1597 passedvitest.user-account.config.ts: 349 passed (adds a DO test for D1 publication with a failed D1 write followed by a retry)vitest.auth.config.ts: 39 passed (admin trace detail returnsmodelAttempts)vitest.video-catalog.config.ts: 11 passednpm run docs:checkNew unit tests: a call that keeps streaming finishes at 15 s without fallback and records its timing and usage; a mid-answer stall fails over without leaking partial text; a silent classifier and a still-writing classifier get different reasons; streamed tool calls are assembled.
Live, opt-in
AGENT_STREAMING_LIVE=1 npx vitest run test/model-streaming.live.test.tsagainst Fireworks: 4 passed. It covers a GLM tool loop with reasoning carried across steps, structured output from GLM and from DeepSeek, and the classifier:agent_corestep 1agent_corestep 2transcript_analysttranscript_analystclassifierFollow-ups (not in this PR)
modelAttemptsin the admin trace inspector (web/app/dashboard/admin/AdminTraceInspector.tsx).