Skip to content

[Bug] start-plan: captcha pool refill freezes the main thread after idle; every connection hangs until restart #54

Description

@Jasmine-Lee-2026

前置确认

  • 我已在最新版本上复现(master 或最新 Release)
  • 我已搜索过现有 issue(含已关闭的)
  • 我已用 serve debug 模式复现过一次并保留了日志
  • 我已确认粘贴内容里的 API key / JWT / proxy key 已脱敏

问题症状

其他(请在"补充说明"里描述)

一句话描述问题

start-plan 下闲置 >= 15 分钟后,第一条请求触发验证码补池,8 路并发解码全部跑在代理主线程(happy-dom + Atomics.wait 同步 XHR),事件循环被反复冻结数十秒,所有连接卡"思考中"直到重启 proxy。

zcode-proxy 版本

v4.7.0 (master 6338d13)

运行方式

Bun 编译二进制(zcode-proxy.exe / linux-x64 / darwin-arm64 等)

操作系统

Windows

Provider(上游服务商)

bigmodel(智谱)

Plan(计费层级)

start-plan(zcode.z.ai 网关,JWT + captcha)

客户端请求格式

OpenAI(POST /v1/chat/completions)

客户端类型 + 版本

Kilo via OpenAI-compatible endpoint(curl 复现脚本同样触发)

使用的模型

glm-5.3-flash

复现步骤

1. config.yaml: provider: bigmodel, plan: start-plan, auth: oauth(或等价的 start-plan 配置)
2. 启动:zcode-proxy serve(普通模式即可,bug 与 debug 模式无关)
3. 发任意一条请求确认 200 正常
4. 完全闲置 >= 15 分钟(验证码池 scale-down 120s 降池、900s 深度停摆后完全排空)
5. 回来发任意请求(新对话或续接均可)
6. 现象:所有请求 40-100s+ 无响应;期间新开的对话同样卡死;重启 proxy 立即恢复

期望行为

闲置唤醒后第一条请求最多等一次解码时长(秒级),其余请求不受影响;事件循环永不因验证码解码而冻结。

实际行为

唤醒瞬间池里 0 token,第一条请求触发并发补池(up to 8 并发 x 6 重试),全部在主线程上串行执行 happy-dom 解码 + Atomics.wait 同步 XHR(单次最长 30s)。事件循环被反复冻结,所有在途请求(包括不依赖验证码的)全部超时无响应。

serve debug 模式日志

[captcha-pool] parallel solve failed: warn: captcha failed after 6 attempts: captcha solve stall pe=? lastXhr=10175ms reqs=["8661ms POST upload.captcha-open.aliyuncs.com/","8789ms POST cloudauth-device-dualstack.cn-shanghai.aliyuncs.com/","8824ms POST cloudauth-device-dualstack.cn-shanghai.aliyuncs.com/","9015ms POST no8xfe-verify.captcha-open.aliyuncs.com/","9128ms POST cloudauth-device-dualstack.cn-shanghai.aliyuncs.com/","9278ms POST no8xfe-verify.captcha-open-b.aliyuncs.com/","9303ms POST upload.captcha-open-b.aliyuncs.com/","9409ms POST cloudauth-device-dualstack.cn-shanghai.aliyuncs.com/","9552ms POST no8xfe-verify.captcha-open-b.aliyuncs.com/","10115ms POST upload.captcha-open-b.aliyuncs.com/","10143ms POST upload.captcha-open.aliyuncs.com/","10175ms POST upload.captcha-open.aliyuncs.com/"]

[captcha-pool] parallel solve failed: warn: captcha failed after 6 attempts: captcha solve stall pe=? lastXhr=-6441ms reqs=[]

[captcha-pool] parallel solve failed: warn: captcha failed after 6 attempts: captcha solve stall pe=? lastXhr=-6346ms reqs=[] | guestErrors(1): WINDOW-ERROR: B[$] is not a function. (In 'B[$](F.bind(1,30,18)())', 'B[$]' is undefined)

[captcha-guest-uncaught] undefined is not an object (evaluating 'window.AliyunCaptcha.prototype')
(上条在日志中重复出现 5 次)

注:lastXhr=10175ms 显示单条同步 XHR 阻塞 10+ 秒;负值 lastXhr 与空 reqs 显示事件循环被冻结到连请求记录都来不及写入。以上为 serve stderr 现场摘录(非 debug 模式),每次闲置唤醒后稳定复现。

ZCODE_DUMP_UPSTREAM trace(强烈建议)

未开启

config.yaml 关键片段(脱敏)

auth:
  mode: oauth
provider: bigmodel
plan: start-plan

是否使用了 clientIdentity(session affinity)?

None

这是首次出现还是偶发?

首次出现

如果是回归,最后能正常工作的版本

首次报告,但非回归——v4.7.0 上 100% 复现(每次闲置 ≥15 分钟后唤醒必卡)

其他补充信息

注:GitHub 重复提示中的 #50(内存增长)日志里同样出现了 captcha solve stall 现象,但该 issue 只解决了内存回收,主线程冻结问题在 v4.7.0 仍然存在。

根因定位(源码级):src/proxy/captcha-happy.ts 的 solveTraceless 在主线程跑 happy-dom,同步 XHR 用 Atomics.wait 阻塞事件循环(SYNC_FETCH_TIMEOUT_MS 默认 30s)。已用主线程心跳探针实测:闲置唤醒并发补池期间主线程心跳间隔达数十秒。

本人已在本地分支验证了修复方向(每次 solve 挪进独立 worker + 硬超时,主线程零阻塞,6 路并发补池期间心跳间隔仅 205ms),生产运行多日零复现。如果方向可以接受,我可以整理提交 PR。

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions