Repository navigation
feat: fit runs to the model's real context window; local vision; provider capabilities (Oct 2026); release 0.0.22 - #20
Merged
Conversation
…s; release 0.0.22
A local WebLLM model is loaded with a 4096-token window (WebLLM's default,
to save VRAM), while the agent sized compaction from the configured window
(128k by default, or whatever the host said) — so a few MCP tools plus a
short conversation overflowed with "Prompt tokens exceed context window
size: 4124; 4096".
- The window in use is the configured one, never more than the model's own:
webLLMContextWindow(model) reads it from the loaded engine, else from the
app config it was built with, else WebLLM's defaults (4096; 1024 for -1k
builds). agent.contextWindowTokens exposes it. Compaction threshold, tool
result clearing, the cap on one tool result and the results kept verbatim
shrink with it (also when the host set them for a larger window), and a
known window compacts by tokens even without a compaction config.
- 'auto' tool selection also switches to search once the tool definitions
would take a quarter of the window.
- A prompt the provider refuses for length (WebLLM, OpenAI, Claude, Gemini,
llama.cpp) is reported as ContextWindowExceededError with the sizes and what
to do, instead of the provider's text.
- createWebLLMModel({ contextWindowTokens }) loads with a larger window, and
withWebLLMContextWindow(appConfig, id, tokens) does it for a host's own
factory (through engineConfig.appConfig — @browser-ai/web-llm ignores its
top-level appConfig).
- Local vision models (WebLLM's Phi-3.5-vision; ids naming vision / VL /
LLaVA / SmolVLM / Gemma 3n) take images, also on the prompted path; local
models still take no PDFs.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
…' own words without a synthesizer - directBrowserOk: Google, OpenAI, xAI and DeepSeek send CORS headers now (checked with a preflight from a page origin, October 2026); the docs drop "OpenAI/xAI/DeepSeek need a proxy" and list the other providers a page can call directly (Kimi, Groq, Mistral, OpenRouter, Cerebras). - supportsImages: Moonshot (Kimi) and Mistral read images, DeepSeek's Flash line too; Kimi takes no PDF parts. - contextWindowOf also reads the browser's built-in model's window (@browser-ai/core reports it once its session exists). - synthesize: false (or a failed synthesizer) answered with a stock sentence, "Done — the changes have been applied."; the answer is now what the steps themselves said, the stock sentence only when they said nothing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
siarheidudko
marked this pull request as ready for review
October 5, 2026 21:48
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
In the demo, a local 13B model on the MCP tab failed with:
WebLLM loads every prebuilt model with a 4096-token window (its default, to save VRAM; "-1k" builds get 1024), whatever the model was trained on. That holds from Qwen3.5 0.8B to Llama 2 13B. The agent, meanwhile, sized compaction from the configured window (8k in the demo; 128k by default). So the tool definitions plus a short conversation overflowed.
Change: the window
webLLMContextWindow(model)reads it from the loaded engine (loadedModelIdToPipeline), else from the app config the model was built with, else from WebLLM's defaults.agent.contextWindowTokensexposes it.compactionconfig.ContextWindowExceededErrorwith the sizes and what to do. This covers WebLLM, OpenAI, Claude, Gemini and llama.cpp.createWebLLMModel(id, { contextWindowTokens }).withWebLLMContextWindow(appConfig, id, tokens)for a host's own factory, throughengineConfig.appConfig.@browser-ai/web-llmsilently ignores its top-levelappConfig.Phi-3.5-vision-instructand ids naming vision / VL / LLaVA / SmolVLM / Gemma 3n, on the prompted path too. Local models still take no PDFs.Change: providers and synthesis
directBrowserOkis now true for Google, OpenAI, xAI and DeepSeek, which all send CORS headers. This was checked with a preflight from a page origin in October 2026. The docs drop "OpenAI/xAI/DeepSeek need a proxy" and list the other providers a page can call directly: Kimi, Groq, Mistral, OpenRouter and Cerebras.contextWindowOfalso reads the window of the browser's built-in model (@browser-ai/core).synthesize: false, or a failed synthesizer, used to answer "Done — the changes have been applied.". The answer is now what the steps themselves said; the stock sentence remains only when they said nothing.New exports:
ContextWindowExceededError,contextOverflowOf,contextWindowOf,webLLMContextWindow,withWebLLMContextWindow,fitCompactionToWindow,toolDefinitionTokens,TOOL_SEARCH_WINDOW_SHARE,WEBLLM_DEFAULT_CONTEXT_WINDOW. Docs:docs/capabilities.md→ "The model's window is a hard limit",docs/providers.md, the README. Version 0.0.22.Checks
npm test155/155. That includes 11 new tests intests/context-window.test.ts:SmolLM2-360M-Instruct-q4f32_1-MLC), a model loaded with defaults reports 4096 from the live engine, and an agent configured for 128k uses 4096.🤖 Generated with Claude Code
https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5