Skip to content

Fix empty-output on hard batches: raise max_completion_tokens to 16384 (tunable) - #4

Merged
msitarzewski merged 2 commits into
mainfrom
fix/reasoning-token-ceiling
Aug 26, 2026
Merged

msitarzewski merged 2 commits into
mainfrom
fix/reasoning-token-ceiling

Conversation

@msitarzewski

Copy link
Copy Markdown
Owner

Reasoning models (GPT-5.x) spend max_completion_tokens on hidden reasoning before output. At 4096 a hard 50-post batch burns the whole budget reasoning (finish_reason=length, empty content) → the client raises and the batch errors; easy batches pass, so it looks intermittent. Verified against the live failing batch: 4096→empty, 16384→full output (~5-7k used). Raises default to 16384 and adds FINTICK_LLM_MAX_TOKENS to tune per model without code changes.

msitarzewski and others added 2 commits August 25, 2026 19:10
… tunable)

GPT-5.x reasoning models spend max_completion_tokens on hidden reasoning BEFORE
emitting output. At 4096 a hard 50-post batch consumes the entire budget reasoning
(finish_reason=length, content_len=0) -> the client raises 'empty content' -> the
whole batch errors. Easy batches fit and pass, so the failure looks intermittent.
Raise the default to 16384 (observed usage ~5-7k, ample headroom) and make it tunable
via FINTICK_LLM_MAX_TOKENS so no model-specific ceiling needs another code change.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@msitarzewski
msitarzewski merged commit 8c390d6 into main Aug 26, 2026
12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant