A reverse proxy for Z.AI / Bigmodel.cn coding-plan APIs that exposes both OpenAI-compatible and Anthropic-format endpoints.
# Install dependencies
bun install
# Copy and edit config
cp config.example.yaml config.yaml
# Log in (OAuth, browser-based)
bun run src/index.ts auth login zai # or: bigmodel
# Start the proxy
bun run src/index.ts
# Or specify a config path
bun run src/index.ts /path/to/config.yamlUpstream credentials come exclusively from the OAuth login flow — run
auth login before starting the proxy. There is no apikey mode anymore.
# Z.AI server-mediated CLI login (3.10 parity: init/poll at zcode.z.ai, no local callback)
bun run src/index.ts auth login zai
# Bigmodel auth-code flow (bigmodel.cn authorize → localhost callback → zcode.z.ai token exchange)
bun run src/index.ts auth login bigmodel
# This will:
# 1. Print an authorize URL and open your browser
# 2. Z.AI: poll the server until authorization completes;
# Bigmodel: receive the browser callback and exchange the auth code
# 3. Resolve your coding-plan API key automatically
# 4. Save encrypted credentials to ~/.zcode-proxy/credentials.jsonThe encrypted credential store is keyed to the machine
(homedir-platform-arch seed, or set ZCODE_PROXY_CREDENTIAL_SECRET for a
portable seed). Then set provider in config.yaml and start the proxy.
If you already use the ZCode desktop app, import the API key directly:
bun run src/index.ts auth login bigmodel --import| Method | Path | Description |
|---|---|---|
POST |
/v1/chat/completions |
OpenAI-compatible chat completions (streaming + non-streaming) |
POST |
/v1/messages |
Anthropic-format messages (streaming + non-streaming) |
POST |
/v1/responses |
OpenAI Responses API (Codex CLI / Agents SDK; translates Responses → Chat → Anthropic upstream) |
POST |
/async/v1/messages |
Async (off-peak) Anthropic-format — routes to free idle-compute pool |
POST |
/async/v1/chat/completions |
Async (off-peak) OpenAI-format — same backend, translates request/response |
GET |
/async/v1/health |
Probe off-peak queue availability |
GET |
/v1/models |
List available models |
GET |
/webui |
Built-in chat web UI (served without the proxy key; see below) |
GET |
/health |
Health check |
/async/* routes are gated by async.enabled: true in config (default false)
and are a coding-plan feature: when plan: start-plan, the routes return
400 async_plan_unsupported even when enabled.
They require a logged-in credential that carries a JWT (the off-peak backend
needs both the JWT from login and the coding-plan API key — a JWT-less
credential, e.g. from auth login --import, returns 400
async_credentials_unavailable).
When enabled, requests are routed through ZCode's off-peak ticket-queue backend:
the proxy takes a ticket, holds the connection open with SSE keepalive comments
while waiting for a free slot, then streams the upstream response through. If
the ticket expires mid-run (server reclaims the slot), the proxy automatically
takes a new ticket and resends the original request (up to async.maxRetries,
default 3). Client disconnect triggers a fire-and-forget /ticket/{id}/settle
call as the universal close-out signal.
Streaming (stream: true) is the expected mode for coding harnesses; non-stream
is supported as a fallback (the proxy internally still consumes upstream as a
stream, then emits one aggregated JSON body).
Off-peak is one-shot, not conversational. Each /async/* request is an
independent task with its own ticket; the proxy does NOT preserve conversation
history across requests. To do multi-turn, send the full conversation in each
request (typical for stateless chat completions clients), or use the synchronous
/v1/* endpoints which can leverage server-side session affinity.
Phase 1 limitations (planned for Phase 2):
- No native async task API (
POST /async/v1/taskswith persistent store) — bridge mode only - No concurrency cap on
/async/*routes — body size is capped at 4 MiB but unlimited simultaneous connections are allowed - No persistent state across proxy restarts
curl http://localhost:8080/async/v1/messages \
-H "Authorization: Bearer your-proxy-secret" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-4.6",
"max_tokens": 1024,
"stream": true,
"messages": [{"role": "user", "content": "Hello!"}]
}'ZCode 3.10 exposes limited-quota trial plans (e.g. weekend packages) that are
claimed first-come-first-served and activate at a future time. The proxy can
grab them for you automatically — it polls the preview endpoint every 5 minutes
and claims the highest-priority plan the moment the campaign endpoint goes live
(a 404 before launch is the expected pre-campaign state). Requires a logged-in
credential and identity.appVersion >= 3.10.0.
claim:
enabled: true # start the auto-claim scheduler while serving
# auto: true # set false to only use the CLI command
# planId: "" # claim a specific plan; empty = highest priorityOne-shot usage:
zcode-proxy claim list # show currently claimable plans
zcode-proxy claim # claim now (highest priority, or claim.planId)Backoff: already_claimed/quota_exhausted wait for the server-provided next
window; other failures use claim.cooldownMs (default 10 min).
curl http://localhost:8080/v1/chat/completions \
-H "Authorization: Bearer your-proxy-secret" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-4.6",
"messages": [{"role": "user", "content": "Hello!"}],
"stream": false
}'curl http://localhost:8080/v1/messages \
-H "x-api-key: your-proxy-secret" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-4.6",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Hello!"}]
}'curl http://localhost:8080/v1/chat/completions \
-H "Authorization: Bearer your-proxy-secret" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-4.6",
"messages": [{"role": "user", "content": "Write a poem"}],
"stream": true
}'curl http://localhost:8080/v1/models \
-H "Authorization: Bearer your-proxy-secret"Open http://localhost:8080/webui in a browser for a built-in, ChatGPT-style
chat client. The page is served without the proxy API key (so it can load
and present the key input); it then sends the key on its own /v1/* calls.
Features: streaming responses (SSE), model picker (auto-populated from
/v1/models), editable system prompt, temperature / top-p / max-tokens /
do_sample, deep-thinking toggle with reasoning_effort (GLM-5.2+), image
upload (auto-enabled for models whose id contains v), MCP HTTP servers,
markdown + code-highlight rendering, light/dark theme, and per-browser
multi-session autosave (localStorage). Open Settings (⚙) to configure.
The TUI is the default mode — launching without arguments opens it:
bun run src/index.ts # interactive panel (default)
bun run src/index.ts debug # with per-request debug diagnostics
bun run src/index.ts serve # classic CLI mode: headless serverzcode-proxy --cli opts out of the TUI entirely and restores the classic CLI
dispatch (bare --cli = serve).
A PC terminal control panel mirroring the Android app's layout — three cards (Settings & Login / Proxy Server / Logs) rendered in the alternate screen with zero extra dependencies. The proxy auto-starts on launch when credentials are available; the log card shows live per-request rows (console output is captured in-process) with scrollback and tail-following.
| Key | Action |
|---|---|
s |
Start / stop the proxy server |
l |
OAuth login for the current provider (opens the browser) |
o |
Logout (cancels an in-flight login and releases its callback port) |
p / t |
Switch provider (zai ↔ bigmodel) / plan (coding-plan ↔ start-plan) — requires the proxy stopped, persisted to config.yaml |
↑↓ / PgUp/PgDn / Home/End |
Scroll the log pane |
g |
Jump back to the tail (re-enable following) |
c |
Clear the log pane |
q / Ctrl+C |
Quit |
Set ZCODE_TUI_LOGFILE=/path/to/file.log to also tee every captured line to a
file (best-effort; never affects handling). Requires a TTY with at least
40x14 cells; Windows Terminal, mintty and common Unix terminals are supported.
| Field | Env Var | Default | Description |
|---|---|---|---|
server.port |
ZCODE_PROXY_PORT |
8080 |
Listen port |
auth.proxyApiKey |
ZCODE_PROXY_API_KEY |
— | Client auth key |
auth.oauthCredentialsPath |
— | — | Parsed but currently not honored — the credential store path is fixed at ~/.zcode-proxy/credentials.json |
provider |
ZCODE_PROVIDER |
zai |
Upstream provider |
plan |
— | coding-plan |
Plan tier: coding-plan (direct upstream) or start-plan (zcode.z.ai gateway + JWT + captcha) |
identity.appVersion |
ZCODE_APP_VERSION |
3.11.2 |
User-Agent: ZCode/{version} |
identity.deviceMid |
ZCODE_IDENTITY_DEVICE_MID |
auto-generated | Device identity (X-Device-Mid); UUIDv4 generated once at first auth login / config creation and reused forever |
identity.sourceTitle |
ZCODE_SOURCE_TITLE |
cli |
X-Title: Z Code@{title} |
identity.refererOrigin |
ZCODE_REFERER_ORIGIN |
https://zcode.z.ai |
HTTP-Referer URL |
endpointRouting.enabled |
ZCODE_ENDPOINT_ROUTING |
true |
Server-controlled upstream URL remapping via zcode.z.ai/api/v1/agent/configs (mirrors ZCode's ProviderEndpointRoutingService; fail-open) |
clientSigning.enabled |
ZCODE_CLIENT_SIGNING |
true |
Client request signing V4 (Ed25519 + proof-of-work, gate-driven; only activates when the server sets codingPlanSignature.enable=true; fail-open) |
claim.enabled |
ZCODE_CLAIM_ENABLED |
true |
Weekend-plan auto-claim: poll zcode.z.ai/api/v1/zcode-plan/billing/preview and claim trial packages (see below) |
async.enabled / origin / maxRetries / maxWaitMs |
ZCODE_ASYNC_ENABLED / ZCODE_ASYNC_ORIGIN / ZCODE_ASYNC_MAX_RETRIES / ZCODE_ASYNC_MAX_WAIT_MS |
false / https://zcode.z.ai / 3 / 0 |
Async off-peak bridge gating + tuning (see Async section) |
| config file path | ZCODE_PROXY_CONFIG |
config.yaml |
Config file to load on serve |
Start-plan captcha tunables (env only): ZCODE_CAPTCHA_RETRIES (per-token solve retries), CAPTCHA_POOL_MIN / CAPTCHA_POOL_MAX (pre-solved token pool sizing).
Client Request
│
▼
Proxy API Key Auth (shared secret)
│
▼
Route Detection + Plan-aware Routing (both plans post Anthropic upstream, v4.5.0+)
/v1/chat/completions (OpenAI client format)
├─ coding-plan → TRANSLATE OpenAI→Anthropic → provider's anthropic endpoint
│ (remapped to zcode.z.ai ultra endpoints via server-controlled mapping)
└─ start-plan → TRANSLATE OpenAI→Anthropic → zcode.z.ai
/api/v1/zcode-plan/anthropic/v1/messages (JWT + captcha)
/v1/messages (Anthropic client format)
├─ coding-plan → NATIVE PASSTHROUGH to the provider's anthropic endpoint (same format)
└─ start-plan → NATIVE PASSTHROUGH → zcode.z.ai
/api/v1/zcode-plan/anthropic/v1/messages (JWT + captcha)
/v1/responses (Responses client format)
├─ both plans → TRANSLATE Responses→Chat→Anthropic → plan's anthropic endpoint
│
▼
Body Transformation (ZCode-equivalent mutations)
Anthropic upstream → cache_control on last message + metadata.user_id
start-plan → prepend ZCode system messages
│
▼
Auth + Identity Header Injection
Anthropic upstream: x-api-key: {credential} + anthropic-version
start-plan: Authorization: Bearer {jwt}
Both: User-Agent: ZCode/{version} + X-ZCode-* + trace headers
│
▼
Endpoint Routing (server-controlled, fail-open)
GET zcode.z.ai/api/v1/agent/configs → proxyEndpoint.mapping rewrites the upstream URL
│
▼
Client Signing V4 (gate-driven, fail-open)
gate says codingPlanSignature.enable → handshake + Ed25519 + PoW headers per request
│
▼
Upstream Forward (fetch, or ordered raw-TCP transport for session affinity)
Translation mode: decompress enabled (proxy reads + translates body)
Passthrough: decompress disabled (raw gzip bytes stream through)
│
▼
Response Handling
Passthrough: raw bytes → client (content-encoding preserved)
Translation batch: Anthropic JSON ↔ OpenAI JSON (gzip if client accepts)
Translation SSE stream: translated chunk-for-chunk in the client's format
# Run tests
bun test
# Type check
bun x tsc --noEmit
# Run in dev mode (opens the interactive TUI; add `--cli serve` for headless)
bun run src/index.ts config.yaml
# Compile a single-file binary (→ zcode-proxy.exe, gitignored)
bun run buildPull the multi-arch image from GitHub Packages (ghcr.io):
docker pull ghcr.io/tridefender/zcode-proxy:latestUpstream credentials come only from auth login, so the container needs the
encrypted credential store. Log in on the host with a fixed encryption seed
(the container cannot derive the default machine seed), then mount the store:
# 1. Log in on the host with a portable seed
ZCODE_PROXY_CREDENTIAL_SECRET="a-long-random-secret" \
bun run src/index.ts auth login zai
# 2. Mount the store at the path the proxy reads. The image runs as user `bun`
# (home /home/bun) and the store path is fixed, not configurable:
docker run --rm -p 8080:8080 \
-v "$(pwd)/config.yaml:/data/config.yaml:ro" \
-v "$(HOME)/.zcode-proxy/credentials.json:/home/bun/.zcode-proxy/credentials.json:ro" \
-e ZCODE_PROXY_CREDENTIAL_SECRET="a-long-random-secret" \
ghcr.io/tridefender/zcode-proxy:latestNote:
/healthand all routes sit behind the proxy-API-key check, so health probes must sendx-api-key: <ZCODE_PROXY_API_KEY>.
Common environment variables (see the Configuration table above for the full list):
| Env Var | Description |
|---|---|
ZCODE_PROVIDER |
zai or bigmodel |
ZCODE_PROXY_API_KEY |
Client auth shared secret |
ZCODE_PROXY_CREDENTIAL_SECRET |
Encryption seed for the credential store (must match the seed used at auth login) |
ZCODE_PROXY_PORT |
Listen port (default 8080) |
docker-compose:
services:
zcode-proxy:
image: ghcr.io/tridefender/zcode-proxy:latest
ports:
- "8080:8080"
volumes:
- ./config.yaml:/data/config.yaml:ro
- ./credentials.json:/home/bun/.zcode-proxy/credentials.json:ro
environment:
ZCODE_PROVIDER: zai
ZCODE_PROXY_API_KEY: "your-proxy-secret"
ZCODE_PROXY_CREDENTIAL_SECRET: "a-long-random-secret"
restart: unless-stoppedThe proxy lists these models on GET /v1/models (pinned to the GLM coding-plan tier):
| Model | Context | Max Output |
|---|---|---|
glm-4.5-air |
200K | 128K |
glm-4.6 |
200K | 128K |
glm-4.6v |
200K | 128K |
glm-4.7 |
200K | 128K |
glm-5 |
200K | 128K |
glm-5-turbo |
200K | 128K |
glm-5v-turbo |
200K | 128K |
glm-5.1 |
200K | 128K |
glm-5.2 |
1M | 128K |
glm-5.3 |
1M | 128K |
glm-5.3-flash |
1M | 128K |
Requests for models not in this list are still forwarded upstream — the listing is informational, not a gate.
MIT