Skip to content

open_jev: add native Rust/CUDA text worker - #55

Open
hsliuustc0106 wants to merge 11 commits into
mainfrom
codex-open-jev-native
Open

hsliuustc0106 wants to merge 11 commits into
mainfrom
codex-open-jev-native

Conversation

@hsliuustc0106

@hsliuustc0106 hsliuustc0106 commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Closes #54

7.47× faster than raw HF Transformers on one H200: mean warm HTTP latency drops 362.21→48.50 ms (86.61% lower). The matched comparison covers 74 single-candidate JevBench noul requests per pass, with two measured passes per backend (BF16, concurrency 1).

Serve Open-Jev-27B-v1.1 through the existing Rust frontend with a native Rust/CUDA text worker. It compiles and tokenizes independent candidate prompts, applies the trained scalar decision head and saved temperature, and returns typed choice, score and noul answers.

  • Extract the Qwen prefill implementation shared by Cua-S1 and Open-Jev, preserving Cua-S1's public imports.
  • Add a CPU checkpoint export recipe and real inference before worker readiness.
  • Fuse attention gating, cache residual RMSNorm values and pack BF16 MLP SiLU loads/stores while preserving rounding and reduction order. Other SiLU layouts use the scalar fallback.
  • Retain up to 64 exact-length CUDA Graphs for opt-in replay across repeated request lengths.
  • Add reference contract, tokenization and kernel coverage, plus setup and H200 results.

The CUDA ABI is now 4; rebuild the shared library and both native workers together.

Test Plan

System1-Omni Version / Commit: target main at 58b8cbe9d4738f6c217f6fe584852dc9361b2c24; merge base 1be7d41eb74d5b49fd25194042041319950e5dce; head 3685d2037c6910a10ffa033db2262f734ba3fc11.

Open-Jev tests and fixtures live under the repository-level tests/open_jev/; shared Qwen JSON/configuration and CUDA tests live under tests/qwen3_5/. Cargo explicitly registers the integration tests, and private unit tests load their files from that same top-level tree. Frontend CPU mock-worker coverage is tracked separately in issue #46 and PR #58.

Check prompts, typed answers and token IDs against pinned Open-Jev fixtures. Run the workspace checks and reserved-GPU kernel/worker validation. The October 2 comparison isolates scalar versus packed SiLU with cached RMSNorm fixed. The October 3 comparison adds the raw HF Transformers / unmerged-LoRA / PyTorch-fallback baseline and newly times native and original Fast. Both use the same 74 single-candidate JevBench requests, one excluded feasibility pass and two measured passes per configuration; their timing instrumentation and results are reported separately below.

Test Result

Passed locally after cleanup and the CUDA Graph cache update:

cargo fmt --all --check
cargo clippy --workspace --locked --offline --all-targets -- -D warnings
cargo test --workspace --locked --offline
cargo build --workspace --release --locked --offline
  • Documentation update: README news, the validation page and this summary highlight the raw HF comparison. All three latency-table rows and rounded speedup figures match the archived raw results; 22 local links were checked and mkdocs build --strict passed. The October 2 evidence is preserved separately. Runtime code is unchanged by these documentation updates.
  • After moving the tests: all workspace Rust checks passed, including 29 CPU tests; the pinned-checkpoint tokenizer test and strict MkDocs build passed. All 38 registered test names are preserved, including ignored GPU/checkpoint cases. The relocated six-test GPU suite compiled in release mode without device execution.
  • Six CUDA ABI 4 tests and live worker/frontend smoke checks previously passed on reserved H200 GPU 2 at 61b83b3. The CUDA source and GPU test contents are unchanged; their location and registration changed, and device checks were not repeated. SM89 compilation passed previously; SM89 device execution and full Cua-S1 checkpoint inference were not checked.

Latest warm HTTP latency — H200, 2026-10-03

74 real JevBench noul requests, one candidate each, BF16, max length 16384 and concurrency 1 on exact H200 GPU 2 (UUID GPU-cbf66259-f4ab-0ede-1811-82037dde5924; NUMA 0, CPUs 0–15). All three backends were newly timed, with one reused server, one excluded feasibility pass and two measured passes each: 148 measured requests per configuration.

Configuration Overall mean (ms) Mean, pass 1 / pass 2 (ms) P50, pass 1 / pass 2 (ms) P95, pass 1 / pass 2 (ms) Correct / 74
Raw HF Transformers (unmerged LoRA; PyTorch fallback) 362.209 362.238 / 362.180 324.822 / 321.321 651.999 / 652.911 64
Native Rust/CUDA (cached RMSNorm + packed SiLU) 48.503 48.471 / 48.535 25.797 / 24.895 215.696 / 218.882 64
Original OpenJev-Fast 50.936 51.097 / 50.775 24.651 / 24.622 248.121 / 251.193 63

Native's mean latency gives a 7.47× speedup over raw HF (86.61% lower latency) on this subset. Native and HF agree on all 74 thresholded decisions and both score 64/74, with maximum probability difference 0.020423. Fast scores 63/74, with one different decision (hard-opus-a-temporal_numeric-09); native/Fast maximum probability difference is 0.034353. Each backend's probabilities are exactly unchanged across its feasibility and measured passes. These counts do not establish statistical accuracy superiority or full numerical parity.

Raw HF baseline: original Open-Jev server and DecisionModel with the unmerged PEFT LoRA adapter, trained scalar head and saved temperature. Full attention uses stock SDPA; runtime checks verify all 48 linear-attention layers use torch_chunk_gated_delta_rule, stock convolution and Qwen3_5RMSNormGated, with 160 unmerged LoRA modules. The baseline disables the optional FLA/causal-conv1d availability checks before importing model classes, while reusing the same prepared environment. It uses no custom Fast model, torch.compile, CUDA Graph replay or prefix cache. Native retains its accepted merged-LoRA CUDA path; original Fast retains its custom kernels and CUDA Graph stack, reporting 30 retained graphs.

Timing includes localhost HTTP through the same frozen Rust frontend, tokenization, worker execution and UTF8 response decoding. Client body serialization and response JSON parsing are excluded. Downloads, preparation, process-to-readiness, warmup, first inference after readiness and the complete feasibility pass are excluded. No Nsight launcher or trace collection was used. P50 is the median; P95 uses JevBench's sorted[int(0.95*N)-1] rank. Pass ranges are observed variability, not confidence intervals.

Native's observed mean is 4.78% below Fast in this run; Fast has a lower median. The mean pass ranges are disjoint, but the fixed two-pass budget and single-candidate workload do not establish a general winner. Multi-candidate prefix sharing and broader JevBench coverage remain unvalidated. Native graph replay is validated separately below. The author's 17.3 ms B300 result uses different hardware/workload.

Frozen native worker/frontend 202c0e1 and the accepted packed-SiLU library were reused; PR source at measurement was ad1cb81. Open-Jev 3308a15, Fast c52b8bb, JevBench f8ce713; pinned base 1d4bf0f, adapter 28cf730, temperature 2.5343690298472983. Torch 2.13.0+cu130, Transformers 5.10.2 and PEFT 0.19.1. Raw commands, rows, token IDs, source snapshots and 48 input hashes are archived locally in profile/jev-hf-transformers-comparison-20261003/.

A collector cleanup assertion treated native's intentional SIGTERM exit as failure after all its measurements were saved. Only the remaining Fast configuration continued on a second reservation of the same GPU/affinity; no extra measured passes were added. All task-owned processes exited and GPU 2 returned to 0 MB used.

Native CUDA Graph replay — H200, 2026-10-03

Warm mixed-length HTTP: 48.086→47.112 ms (2.03% lower) with the new bounded 64-entry exact-length cache; fixed 107-token requests improve 3.77%. The former eight-entry cache regresses the mixed workload to 95.289 ms through repeated capture. Graph mode stays opt-in with CUA_S1_GRAPH=1; new lengths and scratch growth still incur preparation costs.

Workload Configuration Mean (ms) Mean, pass 1 / pass 2 (ms)
107 tokens Eager 19.829 19.835 / 19.823
107 tokens Graph, eight entries 18.882 18.902 / 18.863
107 tokens Graph, 64 entries 19.081 18.968 / 19.194
74 mixed-length cases Eager 48.086 48.062 / 48.110
74 mixed-length cases Graph, eight entries 95.289 95.425 / 95.153
74 mixed-length cases Graph, 64 entries 47.112 47.050 / 47.174

Same exact H200 GPU 2, affinity, BF16, model, CUDA library, frontend and requests. Both capacities are rebuilt from frozen 202c0e1 sources with identical options and dependencies; their model source differs only in the cache limit. Every configuration validates its first long inference after real readiness to preallocate the maximum workload scratch size, then validates the short request. One server is reused for one excluded feasibility pass and exactly two measured passes per workload: 32 repeated short requests per pass, then all 74 mixed cases per pass. HTTP measurement is unprofiled.

All 74 probabilities and decisions remain exactly unchanged (maximum delta 0.0). Both graph-64 pass means improve over eager; the mixed aggregate narrowly passes the prespecified 2% gate. Two passes are observed variability, not confidence intervals. After mixed passes, graph-64 uses 122 MB more observed device memory than graph-eight (scheduler samples, not peak). The excluded graph-64 mixed feasibility mean is 110.255 ms; it records capture cost and is not a controlled cold-latency result.

A separate preceding Nsight run proves one graph launch, zero recaptures and zero individual runtime kernel-launch calls per warm short request, with all 834 kernels retained. Node-level graph tracing shows larger gaps despite lower unprofiled latency, so it supplies replay proof rather than a speedup estimate. Full Cua-S1 checkpoint graph inference is not revalidated; its shared opt-in cache bound also increases.

Format, workspace Clippy/tests/release build, strict MkDocs and diff checks pass. The PR model source matches the measured 64-entry candidate exactly; CUDA kernels and tests are unchanged. Both GPU jobs clean up their owned processes and release GPU 2 to 0 MB used. Raw plans, source copies, commands, responses, traces, memory samples and verification are preserved in profile/jev-cuda-graph-20261003-074613/ and profile/jev-cuda-graph-cache64-20261003-075238/ outside the PR. These native A/B results are separate from the matched raw HF/Fast comparison above. See the validation details.

Earlier isolated packed-SiLU A/B — 2026-10-02

With cached RMSNorm fixed in both native variants, packed SiLU lowered native mean 49.527→48.297 ms (2.483%), with disjoint observed pass ranges and all 74 native probabilities and decisions exactly unchanged. Fast means in that earlier run were 57.345 / 50.908 ms, with 6.436 ms spread associated mainly with preparation variability in two long-policy requests. These earlier timings used an inactive Nsight session with possible CUPTI instrumentation and are separate from the new HF comparison.

Separate CUDA traces showed long-request SiLU 21.217→6.380 ms (69.93%), versus Fast 5.466 ms, closing 94.19% of that measured kernel-family gap. Kernel totals are separate from HTTP latency; Nsight Systems supplied timelines because Nsight Compute counters were denied. See the October 2 revisions, controls and reproduction details.

Self-review

Contributor checklist from CONTRIBUTING.md, left for completion before requesting review:

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
@hsliuustc0106
hsliuustc0106 marked this pull request as ready for review October 3, 2026 06:09
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[New Model]: Native Rust/CUDA support for Open-Jev-27B-v1.1

1 participant