Skip to content

feat: stream model attempts and index their timing in D1 - #171

Merged
devhims merged 1 commit into
mainfrom
feat/streamed-model-attempts
Oct 9, 2026
Merged

devhims merged 1 commit into
mainfrom
feat/streamed-model-attempts

Conversation

@devhims

@devhims devhims commented Oct 9, 2026

Copy link
Copy Markdown
Owner

Problem / Motivation

Session 9d0a2e06 (run e3889923, 2026-10-09 06:35 UTC) failed with MODEL_FALLBACK_EXHAUSTED after all four tool calls had succeeded. After get_video_transcript returned, the next research step hit response_timeout at 10,000 ms on GLM 5.3 Flash. The fallback fired, and DeepSeek V4.1 Flash then hit the same 10-second limit. In the retry run, the same step took 7.7 s (8,214 input and 1,332 output tokens).

Research calls were not streamed, so the 10-second limit covered the whole response: queueing, prompt processing, reasoning and output. That made it impossible to tell a stalled provider from a model that was still writing. It also failed calls that were making steady progress.

What changed

Every failover-wrapped call streams. withModelFailover's doGenerate now reads the provider's doStream and assembles the generate result itself (collectStream). Callers are unchanged, and switching to the backup is still invisible to them because nothing has been returned yet.

Role Before Now
agent_core, finalizer (generate) 10 s whole response 10 s to first content, 5 s stall limit after that, 30 s overall
transcript_analyst 15 s whole response 10 s to first content, 5 s stall limit, 30 s overall
classifier 5 s whole response 5 s, unchanged
visual_analyst, memory_updater 8 s whole response 8 s, unchanged

All limits are still clamped to the phase deadline, less the 10 s reserve for the backup. "First content" means the first text, reasoning or tool-argument delta. Fireworks streams GLM reasoning progressively (probed at high effort: 180 reasoning chunks, the first one arriving as soon as response headers did), so this measures queueing plus prompt processing. A provider that stays silent until the overall limit is labelled first_content_timeout.

I first tried a 3 s first-content limit for the classifier. A live run tripped it: Fireworks took 2 to 4 s to first content even for 58-token prompts, and a classifier fallback moves every later role in the run to DeepSeek. Roles with fixed phase budgets therefore keep their earlier whole-response limits.

Attempt timing is recorded. Each attempt_finished diagnostic now carries startedAt, firstContentMs, elapsedMs, input, cached-input, output and reasoning tokens, the provider request ID, and the limits the attempt actually ran under. Finished attempts go into a Durable Object outbox, agent_model_attempt_outbox. The existing tool-trace publisher moves them to the new D1 table agent_model_attempts (migration 0023), with the same 15-second alarm retry. Rows hold IDs, timings and token counts only, never prompts or output. GET /v1/admin/agent-traces/{runId} returns them as modelAttempts.

Example query, time to first content by model and role:

SELECT model_id, role, COUNT(*) AS attempts,
  AVG(first_content_ms) AS avg_first_content_ms,
  SUM(reason = 'first_content_timeout') AS silent_timeouts
FROM agent_model_attempts WHERE started_at > ?
GROUP BY model_id, role;

Behavior change to review

An attempt canceled before its stream's finish event no longer reports usage. Before, a non-streamed primary that answered after its timeout still reported token usage, which went into the memory-update cost ledger. Now the canceled primary keeps its existing MEMORY_UPDATE_COST_RESERVE_MICROS reservation against the run budget, but its actual token cost is not recorded. This affects internal provider-cost tracking only. Memory updates charge users 0 credits.

Deployment

Apply migration 0023 before deploying the Worker (deploy:production already runs db:migrate:production first). If the Worker runs before the migration, D1 writes fail and rows stay in the outbox, retried every 15 s.

Testing

  • npm run build (includes test type checking)

  • npx vitest run: 1597 passed

  • vitest.user-account.config.ts: 349 passed (adds a DO test for D1 publication with a failed D1 write followed by a retry)

  • vitest.auth.config.ts: 39 passed (admin trace detail returns modelAttempts)

  • vitest.video-catalog.config.ts: 11 passed

  • npm run docs:check

  • New unit tests: a call that keeps streaming finishes at 15 s without fallback and records its timing and usage; a mid-answer stall fails over without leaking partial text; a silent classifier and a still-writing classifier get different reasons; streamed tool calls are assembled.

  • Live, opt-in AGENT_STREAMING_LIVE=1 npx vitest run test/model-streaming.live.test.ts against Fireworks: 4 passed. It covers a GLM tool loop with reasoning carried across steps, structured output from GLM and from DeepSeek, and the classifier:

    Role Model First content Total
    agent_core step 1 GLM 3.5 s 3.5 s
    agent_core step 2 GLM 2.4 s 2.4 s
    transcript_analyst GLM 2.8 s 3.1 s
    transcript_analyst DeepSeek 1.3 s 1.4 s
    classifier GLM 0.7 s 0.7 s

Follow-ups (not in this PR)

  • Show modelAttempts in the admin trace inspector (web/app/dashboard/admin/AdminTraceInspector.tsx).
  • When both models fail after tools have collected evidence, hand off to the finalizer instead of failing the run.

Every failover-wrapped model call now streams, including generateText
callers. The wrapper assembles the generate result from the provider
stream, so a silent provider fails at its first-content limit while a
call that keeps producing reasoning, text or tool arguments is no longer
cut off at a fixed whole-response limit.

Research roles allow 10 seconds to first content and 30 seconds overall.
The classifier, visual analyst and memory updater keep their earlier
whole-response limits, because Fireworks often takes 2 to 3 seconds to
first content and their phases have fixed budgets.

Each finished attempt records time to first content, time to completion,
token counts and the provider request ID. The tool-trace publisher moves
these rows to a new D1 table, agent_model_attempts, and the admin run
trace returns them as modelAttempts.

Run e3889923 failed after GLM and then DeepSeek each hit the earlier
10-second limit while reading a full transcript; its retry finished the
same step in 7.7 seconds.
@vercel

vercel Bot commented Oct 9, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
video2ctx-web Ready Ready Preview Oct 9, 2026 7:37am UTC

@devhims
devhims merged commit 553b5b7 into main Oct 9, 2026
10 checks passed
@devhims
devhims deleted the feat/streamed-model-attempts branch October 9, 2026 08:16

This branch was successfully deployed

1 active deployment
Preview — af6d91ce Deployed Oct 9, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant