Repository navigation
fix(server): 馃悰 get AI insights working again on Ollama's cloud models - #137
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The story
Rolling out the status page moved the production server from an image built in May straight to today's, and AI insights stopped working. Every request came back 502 with "The model gpt-oss:120b-cloud may be too small for structured JSON output".
The cause is #114, from 22 September. It moved the insight schema out of the prompt and into
response_format: json_schema, trusting that a provider without schema support would refuse with a 4xx and trigger the fallback. Ollama's cloud models answer 200 and ignore the schema, and on that server every chat model is a cloud model. With the schema gone from the prompt, the model invents a shape of its own:authoras an object, comparison items with their own keys. The retry fails the same way, twice, every time.A probe on the cluster, same model, asking for a made-up shape it cannot guess:
response_format(what the server sends today)formatparameterjson_object(before #114)What changes in practice
json_schemaanswer leaves out a required key or adds one the schema forbids, the provider ignored the schema, because any runtime that enforces it rules both out. On that proof the retry spells the schema out in the prompt, still injson_schemamode, and so does every later call in that process. The switch is logged once, asai.client.schema_in_prompt model=... reason=provider_ignored_json_schema.What does not change
Providers that do enforce the schema (local Ollama, llama.cpp, OpenAI) never trip the switch and keep #114's short prompt. The 4xx fallback is unchanged. It now also sets the new flag, since it always sent the schema in the prompt.
How it was checked
In short: when a provider says yes to a schema and then ignores it, the server notices from the answer and spells the schema out, instead of failing every insight.