Skip to content

fix: make documentation reliable on small local reasoning models - #36

Closed
arrase wants to merge 1 commit into
mainfrom
feat/llm-quality-and-token-safety
Closed

arrase wants to merge 1 commit into
mainfrom
feat/llm-quality-and-token-safety

Conversation

@arrase

@arrase arrase commented Sep 29, 2026

Copy link
Copy Markdown
Owner

The pipeline assumed a non-reasoning model and a 4 chars/token ratio, which silently corrupted output on the models this tool actually targets.

Reasoning models: ornith and gemma4 emit a thinking block that the client never parsed and never bounded. Measured on an RTX 5060 Ti, a single-line slot cost 400 tokens with thinking enabled but 34 with think: false, and a paragraph slot 795 against 63. The reasoning was invisible to the tool, ate the generation budget and could return empty or confidently false content. The client now parses the envelope, sends think and num_predict, and fails loudly when a slot yields no visible output.

Context budget: budgeting in characters with a 4 chars/token assumption overshot the real 3.05 ratio, so a maximum-size payload consumed 19,997 of 20,000 tokens and the document was written mid-sentence. Three of the nine existing wiki files were already truncated, and the truncated text was stored in .metadata.json so update never repaired it. Budgeting now uses a configurable ratio and reserves tokens for generation, and truncated or empty output is never written or cached.

Document structure: the model chose the page layout, which failed roughly half the time, and tightening the instruction made compliance worse (17%). Factual retention was never the problem, so structure moved into Go: the code renders the headings, component names and ordering, and the LLM fills one prose slot per call. Module pages now share an identical structure by construction.

Other changes: test files are excluded by default (22% of calls for restated assertions), the system message is constant across a run so Ollama's prefix cache can hit, quickstart.md is derived from the architecture instead of the root summary, and parent pages describe subsystem interaction instead of restating their children.

Verified on a copy of this repository with ornith:9b at the shipped defaults: 134 calls, no truncation, no empty files, and identical headings across all seven module pages.

The pipeline assumed a non-reasoning model and a 4 chars/token ratio, which
silently corrupted output on the models this tool actually targets.

Reasoning models: ornith and gemma4 emit a `thinking` block that the client
never parsed and never bounded. Measured on an RTX 5060 Ti, a single-line
slot cost 400 tokens with thinking enabled but 34 with `think: false`, and a
paragraph slot 795 against 63. The reasoning was invisible to the tool, ate
the generation budget and could return empty or confidently false content.
The client now parses the envelope, sends `think` and `num_predict`, and
fails loudly when a slot yields no visible output.

Context budget: budgeting in characters with a 4 chars/token assumption
overshot the real 3.05 ratio, so a maximum-size payload consumed 19,997 of
20,000 tokens and the document was written mid-sentence. Three of the nine
existing wiki files were already truncated, and the truncated text was stored
in .metadata.json so `update` never repaired it. Budgeting now uses a
configurable ratio and reserves tokens for generation, and truncated or empty
output is never written or cached.

Document structure: the model chose the page layout, which failed roughly half
the time, and tightening the instruction made compliance worse (17%). Factual
retention was never the problem, so structure moved into Go: the code renders
the headings, component names and ordering, and the LLM fills one prose slot
per call. Module pages now share an identical structure by construction.

Other changes: test files are excluded by default (22% of calls for
restated assertions), the system message is constant across a run so Ollama's
prefix cache can hit, quickstart.md is derived from the architecture instead of
the root summary, and parent pages describe subsystem interaction instead of
restating their children.

Verified on a copy of this repository with ornith:9b at the shipped defaults:
134 calls, no truncation, no empty files, and identical headings across all
seven module pages.
@arrase arrase closed this Sep 29, 2026
@sonarqubecloud

Copy link
Copy Markdown

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant