Conversation
The pipeline assumed a non-reasoning model and a 4 chars/token ratio, which silently corrupted output on the models this tool actually targets. Reasoning models: ornith and gemma4 emit a `thinking` block that the client never parsed and never bounded. Measured on an RTX 5060 Ti, a single-line slot cost 400 tokens with thinking enabled but 34 with `think: false`, and a paragraph slot 795 against 63. The reasoning was invisible to the tool, ate the generation budget and could return empty or confidently false content. The client now parses the envelope, sends `think` and `num_predict`, and fails loudly when a slot yields no visible output. Context budget: budgeting in characters with a 4 chars/token assumption overshot the real 3.05 ratio, so a maximum-size payload consumed 19,997 of 20,000 tokens and the document was written mid-sentence. Three of the nine existing wiki files were already truncated, and the truncated text was stored in .metadata.json so `update` never repaired it. Budgeting now uses a configurable ratio and reserves tokens for generation, and truncated or empty output is never written or cached. Document structure: the model chose the page layout, which failed roughly half the time, and tightening the instruction made compliance worse (17%). Factual retention was never the problem, so structure moved into Go: the code renders the headings, component names and ordering, and the LLM fills one prose slot per call. Module pages now share an identical structure by construction. Other changes: test files are excluded by default (22% of calls for restated assertions), the system message is constant across a run so Ollama's prefix cache can hit, quickstart.md is derived from the architecture instead of the root summary, and parent pages describe subsystem interaction instead of restating their children. Verified on a copy of this repository with ornith:9b at the shipped defaults: 134 calls, no truncation, no empty files, and identical headings across all seven module pages.
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



The pipeline assumed a non-reasoning model and a 4 chars/token ratio, which silently corrupted output on the models this tool actually targets.
Reasoning models: ornith and gemma4 emit a
thinkingblock that the client never parsed and never bounded. Measured on an RTX 5060 Ti, a single-line slot cost 400 tokens with thinking enabled but 34 withthink: false, and a paragraph slot 795 against 63. The reasoning was invisible to the tool, ate the generation budget and could return empty or confidently false content. The client now parses the envelope, sendsthinkandnum_predict, and fails loudly when a slot yields no visible output.Context budget: budgeting in characters with a 4 chars/token assumption overshot the real 3.05 ratio, so a maximum-size payload consumed 19,997 of 20,000 tokens and the document was written mid-sentence. Three of the nine existing wiki files were already truncated, and the truncated text was stored in .metadata.json so
updatenever repaired it. Budgeting now uses a configurable ratio and reserves tokens for generation, and truncated or empty output is never written or cached.Document structure: the model chose the page layout, which failed roughly half the time, and tightening the instruction made compliance worse (17%). Factual retention was never the problem, so structure moved into Go: the code renders the headings, component names and ordering, and the LLM fills one prose slot per call. Module pages now share an identical structure by construction.
Other changes: test files are excluded by default (22% of calls for restated assertions), the system message is constant across a run so Ollama's prefix cache can hit, quickstart.md is derived from the architecture instead of the root summary, and parent pages describe subsystem interaction instead of restating their children.
Verified on a copy of this repository with ornith:9b at the shipped defaults: 134 calls, no truncation, no empty files, and identical headings across all seven module pages.