Skip to content

fix(moe): fit Qwen3.8 Flash Next FP8 banks across host and GPU memory - #602

Open
yunwei37 wants to merge 3 commits into
FlashML-org:mainfrom
yunwei37:codex/fp8-host-gpu-banks
Open

yunwei37 wants to merge 3 commits into
FlashML-org:mainfrom
yunwei37:codex/fp8-host-gpu-banks

Conversation

@yunwei37

@yunwei37 yunwei37 commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Summary

Fixes #600: native Qwen3.8-Flash-Next-FP8 expert loading can exhaust host RAM even after the parallel reader falls back to serial loading. On one RTX 5090 (32 GB) with 128 GB RAM, this patch lets the unchanged serve command finish loading by keeping the FP8 source layers that do not fit in usable host RAM on the GPU.

This is a loading/capacity fix, not a quantization conversion or a measured speedup over working upstream inference. Upstream main did not complete loading on this machine.

What changes

  • Measure usable host memory, excluding free CMA, and leave one layer of packing space.
  • Place only the necessary FP8 source layers on GPU; recheck capacity before filling each untouched layer because availability changes during loading.
  • Use the existing FP8 packer and prefill/decode copy paths, without registering the unused host banks.
  • Charge GPU source-bank bytes to the existing expert/KV memory accounting.

No new CLI option, lower-precision weights, host setting, scheduler limit or serving override. Dummy loads, converter sinks and other quantization formats keep their existing placement path. Three files: 92 additions / 10 deletions.

Results on the actual checkpoint

Measurement Observed result
Complete fresh-process loads 2/2 completed all 48 expert layers
Short-request decode 29.7–32.8 tokens/s
Long-context input 91,610 input tokens, correct marker retrieval, about 28 tokens/s
Long-prefix TTFT 29.65 s first request → 3.04 s with prefix reuse
Concurrent requests 2 and 4 both completed; more concurrency did not improve total throughput in these samples
OOM-kill counters during successful runs Zero
Complete measured performance, capacity, concurrency and memory results

RTX 5090 FP8 results

The candidate preserves the official checkpoint precision and uses the same command: ft serve --model /tmp/freetoken-eval/checkpoint. Native automatic choices: TP1, disk PLE, offload MoE, hybrid radix cache, and four request slots. No server tuning flag was added for the two fresh-load runs.

Test Result
Unmodified main Serial expert loading reached 45/48 layers; host allocation/reclaim and I/O stalled; no generation
Candidate fresh loads Both completed all 48 layers and genuine generation; serial loading phases approximately 411 s and 469 s
Arithmetic, natural stop Returned 4 for 2+2; second load 2.5900 s, two output tokens
Short input, first load 29.86 tok/s, TTFT 3.81 s; repeated prompt 29.71 tok/s, TTFT 2.91 s
Short input, second load 32.84 tok/s, TTFT 2.43 s; repeated prompt 32.63 tok/s, TTFT 2.39 s
Two concurrent requests, second load Both completed: 18.43 / 18.58 tok/s per request; 256 output tokens in 11.60 s including prefill
Four concurrent requests, first load All completed: approximately 5.92 tok/s per request; 512 output tokens in 26.52 s including prefill
Default KV capacity 130 pages x 64 = 8,320 tokens; oversized 8,588-token input returned HTTP 400 context_length_exceeded
Rebuilt 65,536-token KV pool 65,178 actual input; TTFT first/repeat 20.62 / 2.95 s; decode 28.60 / 28.68 tok/s; marker retrieval succeeded
Rebuilt 91,968-token KV pool 91,610 actual input; TTFT first/repeat 29.65 / 3.04 s; decode 27.86 / 28.59 tok/s; marker retrieval succeeded
Prefix reuse Repeated 72-token prompt logged 64 cached tokens and eight new tokens; long-prefix reuse reduced TTFT substantially
GPU temperature Per-run maxima 56 C and 51 C
Memory OOM/kill counters stayed zero in both successful runs; second-run max sampled cgroup current 123,064,684,544 bytes including reclaimable file cache; GPU allocation 31,184,781,312 bytes

Short/concurrent throughput uses the upstream streaming benchmark helper with 128 output tokens and ignore_eos=true, so it measures throughput rather than natural answer completion. Long-context tests use synthetic repeated-word marker prompts with natural stop, not general agent-quality evaluation. Their cache pool was temporarily rebuilt to 1,024 expert slots and 16 usable mamba slots, then restored to the exact default geometry. These results describe two successful candidate runs, not a speedup over main (main could not serve here).

Loader/copy/engine tests: 154 passed, two skipped. A changing-RAM placement regression fails on the earlier candidate and passes on the adaptive candidate. A real checkpoint layer test compared gate/up/down weights and scales across eight experts before/after existing copy paths, with exact byte equality. Tests and measurements were AI-assisted and run on real hardware in a dedicated Coder Workspace. Full repository tests were not run. No quality, vision, tool-use, 262K or speculative-decoding result is asserted here.

Reproduction environment

  • Model: Qwen/Qwen3.8-Flash-Next-FP8, revision 236dfdf285828023ca3bcd3f37366c58a3469b13.
  • Baseline: FreeToken 0.1.3, main d3512b43affe981465e03ee28cbd88f49c39b9aa.
  • RTX 5090 32 GB; Intel Core Ultra 9 285K; 128 GB physical RAM; local NVMe checkpoint/PLE storage.
  • Ubuntu 24.04.3, Linux 7.3.0-070300rc3-generic, NVIDIA 610.57.04, CUDA 13.0.88, Python 3.12.3, torch 2.11.0+cu130; dedicated Coder Kubernetes Workspace.
ft serve --model /tmp/freetoken-eval/checkpoint

The two complete loading/generation runs used the loader code in this PR, with the original serve defaults. Larger-context tests used a temporary cache rebuild and restored the exact default geometry afterwards; they are not default 91K-capacity claims.

Validation

pytest -q tests/moe/test_offload.py tests/moe/test_fused_copy.py tests/kernels/test_pinned_tensor.py tests/engine
# 154 passed, 2 skipped
  • Original main fails the fixed-capacity placement regression; the first candidate also fails changing-capacity placement. The adaptive candidate passes both.
  • CUDA regressions compare exact FP8 matrix/scale bytes after prefill and decode copies for CPU-only, mixed and GPU-only bank placement.
  • A separate real-checkpoint layer-0 test loaded all 512 experts and compared matrices/scales for eight selected experts against the source tensors; all comparisons passed.
  • Full-model arithmetic returned 4; natural-stop long-context requests retrieved the expected marker.

Full-repository CI has not run. These results establish text loading/generation and the listed tests, not untested vision, tool-use, 262K or speculative-decode behavior. The placement currently does not pre-reserve future cache/graph minima before allocating source banks; the existing accounting sizes caches afterwards, so this is not a general guarantee for arbitrary smaller GPUs or changing shared-host loads.

Related work

#214 already covers the model architecture. #334 concerns finite cgroup limits; this Workspace has no finite memory cap. #563 concerns hot expert caching on a 48 GB GPU, not source-bank placement. The separate stats-page-size repair is #599 and is not included here.

AI-assisted patch and test/report preparation, validated on the hardware and official checkpoint above. Published as an operator-directed draft for review.

Publish the existing real-hardware-tested Workspace change for FlashML-org#600. Keep source residency within the existing automatic cache budget.
Publish the existing tested adaptive loader from the matching Coder Workspace for FlashML-org#600. Preserve checkpoint bytes and reuse the current copy paths.
Add tests for FP8 banks fitting in host RAM and available bytes calculation.
@yunwei37
yunwei37 marked this pull request as ready for review October 4, 2026 09:53
Copilot AI balanced review requested due to automatic review settings October 4, 2026 09:53

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

KarrAcaRn pushed a commit to KarrAcaRn/FreeToken-ByAI that referenced this pull request Oct 4, 2026
FlashML-org#601 adopted with fixups, FlashML-org#599 adopted, FlashML-org#602 deferred, FlashML-org#596 own.

Assisted-by: Claude Opus 5.5

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Engine] Qwen3.8-Flash-Next-FP8 exhausts host RAM on RTX 5090 + 128 GB after serial loading fallback

2 participants