Repository navigation
Conversation
Publish the existing real-hardware-tested Workspace change for FlashML-org#600. Keep source residency within the existing automatic cache budget.
Publish the existing tested adaptive loader from the matching Coder Workspace for FlashML-org#600. Preserve checkpoint bytes and reuse the current copy paths.
Add tests for FP8 banks fitting in host RAM and available bytes calculation.
yunwei37
marked this pull request as ready for review
October 4, 2026 09:53
KarrAcaRn
pushed a commit
to KarrAcaRn/FreeToken-ByAI
that referenced
this pull request
Oct 4, 2026
FlashML-org#601 adopted with fixups, FlashML-org#599 adopted, FlashML-org#602 deferred, FlashML-org#596 own. Assisted-by: Claude Opus 5.5
4 tasks done
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes #600: native Qwen3.8-Flash-Next-FP8 expert loading can exhaust host RAM even after the parallel reader falls back to serial loading. On one RTX 5090 (32 GB) with 128 GB RAM, this patch lets the unchanged serve command finish loading by keeping the FP8 source layers that do not fit in usable host RAM on the GPU.
This is a loading/capacity fix, not a quantization conversion or a measured speedup over working upstream inference. Upstream main did not complete loading on this machine.
What changes
No new CLI option, lower-precision weights, host setting, scheduler limit or serving override. Dummy loads, converter sinks and other quantization formats keep their existing placement path. Three files: 92 additions / 10 deletions.
Results on the actual checkpoint
Complete measured performance, capacity, concurrency and memory results
RTX 5090 FP8 results
The candidate preserves the official checkpoint precision and uses the same command:
ft serve --model /tmp/freetoken-eval/checkpoint. Native automatic choices: TP1, disk PLE, offload MoE, hybrid radix cache, and four request slots. No server tuning flag was added for the two fresh-load runs.Short/concurrent throughput uses the upstream streaming benchmark helper with 128 output tokens and ignore_eos=true, so it measures throughput rather than natural answer completion. Long-context tests use synthetic repeated-word marker prompts with natural stop, not general agent-quality evaluation. Their cache pool was temporarily rebuilt to 1,024 expert slots and 16 usable mamba slots, then restored to the exact default geometry. These results describe two successful candidate runs, not a speedup over main (main could not serve here).
Loader/copy/engine tests: 154 passed, two skipped. A changing-RAM placement regression fails on the earlier candidate and passes on the adaptive candidate. A real checkpoint layer test compared gate/up/down weights and scales across eight experts before/after existing copy paths, with exact byte equality. Tests and measurements were AI-assisted and run on real hardware in a dedicated Coder Workspace. Full repository tests were not run. No quality, vision, tool-use, 262K or speculative-decoding result is asserted here.
Reproduction environment
Qwen/Qwen3.8-Flash-Next-FP8, revision236dfdf285828023ca3bcd3f37366c58a3469b13.d3512b43affe981465e03ee28cbd88f49c39b9aa.The two complete loading/generation runs used the loader code in this PR, with the original serve defaults. Larger-context tests used a temporary cache rebuild and restored the exact default geometry afterwards; they are not default 91K-capacity claims.
Validation
pytest -q tests/moe/test_offload.py tests/moe/test_fused_copy.py tests/kernels/test_pinned_tensor.py tests/engine # 154 passed, 2 skipped4; natural-stop long-context requests retrieved the expected marker.Full-repository CI has not run. These results establish text loading/generation and the listed tests, not untested vision, tool-use, 262K or speculative-decode behavior. The placement currently does not pre-reserve future cache/graph minima before allocating source banks; the existing accounting sizes caches afterwards, so this is not a general guarantee for arbitrary smaller GPUs or changing shared-host loads.
Related work
#214 already covers the model architecture. #334 concerns finite cgroup limits; this Workspace has no finite memory cap. #563 concerns hot expert caching on a 48 GB GPU, not source-bank placement. The separate stats-page-size repair is #599 and is not included here.
AI-assisted patch and test/report preparation, validated on the hardware and official checkpoint above. Published as an operator-directed draft for review.