Skip to content

fix(server): 馃悰 get AI insights working again on Ollama's cloud models - #137

Merged
vitofico merged 2 commits into
mainfrom
fix/ai-schema-in-prompt
Sep 28, 2026
Merged

vitofico merged 2 commits into
mainfrom
fix/ai-schema-in-prompt

Conversation

@vitofico

Copy link
Copy Markdown
Owner

The story

Rolling out the status page moved the production server from an image built in May straight to today's, and AI insights stopped working. Every request came back 502 with "The model gpt-oss:120b-cloud may be too small for structured JSON output".

The cause is #114, from 22 September. It moved the insight schema out of the prompt and into response_format: json_schema, trusting that a provider without schema support would refuse with a 4xx and trigger the fallback. Ollama's cloud models answer 200 and ignore the schema, and on that server every chat model is a cloud model. With the schema gone from the prompt, the model invents a shape of its own: author as an object, comparison items with their own keys. The retry fails the same way, twice, every time.

A probe on the cluster, same model, asking for a made-up shape it cannot guess:

Request Result
Schema only in response_format (what the server sends today) wrong shape, 2 of 2
Ollama's own format parameter wrong shape
Schema in the prompt, json_object (before #114) right shape
Schema in both right shape, 3 of 3

What changes in practice

  • A shape the schema forbids now counts as proof. When a json_schema answer leaves out a required key or adds one the schema forbids, the provider ignored the schema, because any runtime that enforces it rules both out. On that proof the retry spells the schema out in the prompt, still in json_schema mode, and so does every later call in that process. The switch is logged once, as ai.client.schema_in_prompt model=... reason=provider_ignored_json_schema.
  • Wrong values and cut-off answers do not count. They can slip past enforcement, the retry's error message handles them, and the schema written out is about 7,000 characters: sending it every time would more than double each prompt, which a CPU-only model like the one in Web UI for Configuration聽#102 pays for in seconds.
  • On a cloud model the cost is one failed attempt per server start, spent inside the retry the client already makes, so no insight fails because of it.
  • The parse-error hint stops going in a circle. It no longer recommends gpt-oss:120b-cloud to someone running gpt-oss:120b-cloud. For that model it points at the server log instead.

What does not change

Providers that do enforce the schema (local Ollama, llama.cpp, OpenAI) never trip the switch and keep #114's short prompt. The 4xx fallback is unchanged. It now also sets the new flag, since it always sent the schema in the prompt.

How it was checked

  • 5 new tests: an ignored schema is spelled out on the retry, later calls keep it, a wrong value and a cut-off answer keep the prompt short, and the hint. The two behaviour tests failed before the fix.
  • I broke the fix three ways (wrong values counted as proof, cut-off answers counted, the switch forgotten between calls), and a test failed each time.
  • The full server suite passes, 886 tests, and ruff 0.16.8 is clean.
  • I ran the patched client inside the production pod, against the real model with the real system prompt and payload, on three books from the failing logs. The first failed exactly as in the logs, switched, and succeeded on the retry. The next two succeeded on the first request.

In short: when a provider says yes to a schema and then ignores it, the server notices from the answer and spells the schema out, instead of failing every insight.

@vitofico
vitofico merged commit 1219282 into main Sep 28, 2026
8 checks passed
@vitofico
vitofico deleted the fix/ai-schema-in-prompt branch October 9, 2026 09:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant