Skip to content

feat: fit runs to the model's real context window; local vision; provider capabilities (Oct 2026); release 0.0.22 - #20

Merged
siarheidudko merged 2 commits into
mainfrom
claude/agents-review-improvements-2m2406
Oct 5, 2026
Merged

siarheidudko merged 2 commits into
mainfrom
claude/agents-review-improvements-2m2406

Conversation

@siarheidudko

@siarheidudko siarheidudko commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Problem

In the demo, a local 13B model on the MCP tab failed with:

WebLLM generation failed: Prompt tokens exceed context window size: number of prompt tokens: 4124; context window size: 4096

WebLLM loads every prebuilt model with a 4096-token window (its default, to save VRAM; "-1k" builds get 1024), whatever the model was trained on. That holds from Qwen3.5 0.8B to Llama 2 13B. The agent, meanwhile, sized compaction from the configured window (8k in the demo; 128k by default). So the tool definitions plus a short conversation overflowed.

Change: the window

  • The window in use is the model's real one.
    • webLLMContextWindow(model) reads it from the loaded engine (loadedModelIdToPipeline), else from the app config the model was built with, else from WebLLM's defaults.
    • The agent uses the configured window, but never more than the model's own. agent.contextWindowTokens exposes it.
    • A known window means compacting by tokens even without a compaction config.
  • Sizes follow the window. These shrink with it, also when the host set them for a larger window:
    • the compaction threshold (½ of the window);
    • tool-result clearing (¼);
    • the cap on one tool result;
    • the results kept verbatim (1 below 16k).
  • Tool search by size. 'auto' tool selection also switches to search once the tool definitions would take ¼ of the window.
  • A clear overflow error. When a provider still refuses a prompt for length, the run reports ContextWindowExceededError with the sizes and what to do. This covers WebLLM, OpenAI, Claude, Gemini and llama.cpp.
  • A larger window at load.
    • createWebLLMModel(id, { contextWindowTokens }).
    • withWebLLMContextWindow(appConfig, id, tokens) for a host's own factory, through engineConfig.appConfig. @browser-ai/web-llm silently ignores its top-level appConfig.
  • Local vision models take images. That is WebLLM's Phi-3.5-vision-instruct and ids naming vision / VL / LLaVA / SmolVLM / Gemma 3n, on the prompted path too. Local models still take no PDFs.

Change: providers and synthesis

  • CORS. directBrowserOk is now true for Google, OpenAI, xAI and DeepSeek, which all send CORS headers. This was checked with a preflight from a page origin in October 2026. The docs drop "OpenAI/xAI/DeepSeek need a proxy" and list the other providers a page can call directly: Kimi, Groq, Mistral, OpenRouter and Cerebras.
  • Images. Moonshot (Kimi) and Mistral read images, and so does DeepSeek's Flash line. Kimi takes no PDF parts.
  • The built-in model's window. contextWindowOf also reads the window of the browser's built-in model (@browser-ai/core).
  • No more stock answer. synthesize: false, or a failed synthesizer, used to answer "Done — the changes have been applied.". The answer is now what the steps themselves said; the stock sentence remains only when they said nothing.

New exports: ContextWindowExceededError, contextOverflowOf, contextWindowOf, webLLMContextWindow, withWebLLMContextWindow, fitCompactionToWindow, toolDefinitionTokens, TOOL_SEARCH_WINDOW_SHARE, WEBLLM_DEFAULT_CONTEXT_WINDOW. Docs: docs/capabilities.md → "The model's window is a hard limit", docs/providers.md, the README. Version 0.0.22.

Checks

  • typecheck, format:check, build.
  • npm test 155/155. That includes 11 new tests in tests/context-window.test.ts:
    • the app-config override;
    • window detection;
    • fitting and clamping;
    • overflow translation;
    • the window for local and cloud agents;
    • tool search by definition size;
    • a run hitting WebLLM's overflow error;
    • local vision;
    • the built-in model's window;
    • provider capabilities;
    • the answer without a synthesizer.
  • Real WebLLM. In Chromium (SwiftShader WebGPU, SmolLM2-360M-Instruct-q4f32_1-MLC), a model loaded with defaults reports 4096 from the live engine, and an agent configured for 128k uses 4096.

🤖 Generated with Claude Code

https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5

claude added 2 commits October 5, 2026 20:51
…s; release 0.0.22

A local WebLLM model is loaded with a 4096-token window (WebLLM's default,
to save VRAM), while the agent sized compaction from the configured window
(128k by default, or whatever the host said) — so a few MCP tools plus a
short conversation overflowed with "Prompt tokens exceed context window
size: 4124; 4096".

- The window in use is the configured one, never more than the model's own:
  webLLMContextWindow(model) reads it from the loaded engine, else from the
  app config it was built with, else WebLLM's defaults (4096; 1024 for -1k
  builds). agent.contextWindowTokens exposes it. Compaction threshold, tool
  result clearing, the cap on one tool result and the results kept verbatim
  shrink with it (also when the host set them for a larger window), and a
  known window compacts by tokens even without a compaction config.
- 'auto' tool selection also switches to search once the tool definitions
  would take a quarter of the window.
- A prompt the provider refuses for length (WebLLM, OpenAI, Claude, Gemini,
  llama.cpp) is reported as ContextWindowExceededError with the sizes and what
  to do, instead of the provider's text.
- createWebLLMModel({ contextWindowTokens }) loads with a larger window, and
  withWebLLMContextWindow(appConfig, id, tokens) does it for a host's own
  factory (through engineConfig.appConfig — @browser-ai/web-llm ignores its
  top-level appConfig).
- Local vision models (WebLLM's Phi-3.5-vision; ids naming vision / VL /
  LLaVA / SmolVLM / Gemma 3n) take images, also on the prompted path; local
  models still take no PDFs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
…' own words without a synthesizer

- directBrowserOk: Google, OpenAI, xAI and DeepSeek send CORS headers now
  (checked with a preflight from a page origin, October 2026); the docs drop
  "OpenAI/xAI/DeepSeek need a proxy" and list the other providers a page can
  call directly (Kimi, Groq, Mistral, OpenRouter, Cerebras).
- supportsImages: Moonshot (Kimi) and Mistral read images, DeepSeek's Flash
  line too; Kimi takes no PDF parts.
- contextWindowOf also reads the browser's built-in model's window
  (@browser-ai/core reports it once its session exists).
- synthesize: false (or a failed synthesizer) answered with a stock sentence,
  "Done — the changes have been applied."; the answer is now what the steps
  themselves said, the stock sentence only when they said nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
@siarheidudko siarheidudko changed the title feat: fit runs to the model's real context window; local vision models; release 0.0.22 feat: fit runs to the model's real context window; local vision; provider capabilities (Oct 2026); release 0.0.22 Oct 5, 2026
@siarheidudko
siarheidudko marked this pull request as ready for review October 5, 2026 21:48
@siarheidudko
siarheidudko merged commit 4451849 into main Oct 5, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants