Skip to content

fix: GLM reasoning effort, [1m] request strip, and catalog windows - #30

Closed
xz-dev wants to merge 4 commits into
TriDefender:masterfrom
xz-dev:fix/reasoning-effort-forwarding
Closed

xz-dev wants to merge 4 commits into
TriDefender:masterfrom
xz-dev:fix/reasoning-effort-forwarding

Conversation

@xz-dev

@xz-dev xz-dev commented Aug 23, 2026

Copy link
Copy Markdown

What

Aligns the proxy with current GLM Coding Plan contracts for reasoning effort and the official [1m] listing aliases.

  1. Reasoning effort — model-specific maps (GLM-5.3 none/minimal/low → low, medium/high → high, xhigh/max → max; GLM-5.2 can disable). Shared translator covers Chat, Responses, Claude Messages, and async. Catalog is the single source of truth: Codex ?client_version and Anthropic anthropic-version advertise the official efforts.
  2. [1m] request path — catalogs keep glm-5.2[1m] / glm-5.3[1m]. Shared transformRequestBody strips trailing [1m] before upstream so Chat/Responses/Claude/async stop hitting 1214 / modelCode不存在.
  3. Catalog windows — bare glm-5.2 / glm-5.3 advertise 200k. Only [1m] aliases advertise 1M. Official Coding Plan 1M is opt-in via the suffix (Enable 1M Context, ZCode Connect Models).

Why

  • Non-none effort previously collapsed to Anthropic thinking:{type:"enabled"}, so low and max looked identical upstream.
  • Listing [1m] aliases were forwarded as modelCode, which the provider rejects (1214).
  • Catalog advertised the 1M ceiling on the bare ids, so clients compacted/budgeted as if every glm-5.3 call were 1M.

Tests

Focused suites for the translator, body transformer, catalogs, and providers pass. Full bun test: 550 pass / 1 fail. The failure is pre-existing (src/integration.test.ts still loads config.test.yaml, removed in f6aa147) and is unrelated to this branch.

Notes

  • No plan / async behavior change.
  • README model table still lists glm-5.2/5.3 as 1M; I can follow with a docs-only commit if wanted (GPG pinentry timed out on that commit).

xz-dev and others added 4 commits August 23, 2026 14:08
Keep glm-*[1m] in the listing catalog, but send the base modelCode
upstream so Chat/Responses/Claude/async no longer hit 1214.
Coding Plan 1M context is opt-in via glm-*[1m]. Bare ids inherit the
200k catalog window; listing still keeps the [1m] aliases.
@TriDefender

TriDefender commented Aug 23, 2026 •

Copy link
Copy Markdown
Owner

Hey I checked this commit and found a few issues:

Blocking issues:

  • Regression: reasoning_effort: "none" now forces thinking ON for every model without an effort map (glm-4.5-air, 4.6, 4.6v, 4.7, 5, 5-turbo, 5v-turbo, 5.1, unknown ids). I re-confirmed in source: normalizeReasoningEffort returns undefined for mapless models → translateReasoning falls through to {thinking: enabled}. Master sent disabled. Live probe proved upstream honors disabled for these models — genuine regression. Fix: global fallback none/minimal + no map → {type:"disabled"}.

@TriDefender

TriDefender commented Aug 23, 2026 •

Copy link
Copy Markdown
Owner

After careful review this PR will be closed without merging as it's considered NOT A BUG

Breaking down the PR it's obvious that 3 issues are adressed, respectivly the 1M marker, disable thinking and proxy side reasoning effort map.

Let's start with the effort map.

  • Now, if you've read the original docs, it would be hard to miss that the server side handles reasoning effort already, thus it would be redundant to implement proxy side adapter.
  • Then the reasoning part, as mentioned above and per test results shows, PR 30 has a serious regression. From my one-off test all models that previously behaved fine on the current branch leaked thinking, therefore this is not a "fix", it's rather the bug itself.
  • Last but not the least, the 1M marker is no longer necessary in Claude Code or Codex. You may use CLAUDE_CODE_MAX_CONTEXT_TOKENS to enable 1M long context in claude code harness, for codex set it explicitly in config.toml.
  • Also the [1m] marker won't be sent out anyways on CC since it's just local handling so I don't understand why you need a fix to that.

xz-dev added a commit to xz-dev/zcode-api that referenced this pull request Aug 24, 2026
Fork-only overlay. Upstream closed PR TriDefender#30. Includes f76599e max_tokens
catalog + request clamp so Z.AI 1210 does not fire on 200k/1M context.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants