open_jev: add native Rust/CUDA text worker - #55
Open
hsliuustc0106 wants to merge 11 commits into
Open
hsliuustc0106 wants to merge 11 commits into
hsliuustc0106 wants to merge 11 commits into
Conversation
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
5 of 7 tasks
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
1 of 8 tasks
4 tasks
1 task done
4 tasks
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
4 tasks done
hsliuustc0106
marked this pull request as ready for review
October 3, 2026 06:09
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
1 of 10 tasks
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Closes #54
7.47× faster than raw HF Transformers on one H200: mean warm HTTP latency drops 362.21→48.50 ms (86.61% lower). The matched comparison covers 74 single-candidate JevBench
noulrequests per pass, with two measured passes per backend (BF16, concurrency 1).Serve Open-Jev-27B-v1.1 through the existing Rust frontend with a native Rust/CUDA text worker. It compiles and tokenizes independent candidate prompts, applies the trained scalar decision head and saved temperature, and returns typed
choice,scoreandnoulanswers.The CUDA ABI is now 4; rebuild the shared library and both native workers together.
Test Plan
System1-Omni Version / Commit: target
mainat58b8cbe9d4738f6c217f6fe584852dc9361b2c24; merge base1be7d41eb74d5b49fd25194042041319950e5dce; head3685d2037c6910a10ffa033db2262f734ba3fc11.Open-Jev tests and fixtures live under the repository-level
tests/open_jev/; shared Qwen JSON/configuration and CUDA tests live undertests/qwen3_5/. Cargo explicitly registers the integration tests, and private unit tests load their files from that same top-level tree. Frontend CPU mock-worker coverage is tracked separately in issue #46 and PR #58.Check prompts, typed answers and token IDs against pinned Open-Jev fixtures. Run the workspace checks and reserved-GPU kernel/worker validation. The October 2 comparison isolates scalar versus packed SiLU with cached RMSNorm fixed. The October 3 comparison adds the raw HF Transformers / unmerged-LoRA / PyTorch-fallback baseline and newly times native and original Fast. Both use the same 74 single-candidate JevBench requests, one excluded feasibility pass and two measured passes per configuration; their timing instrumentation and results are reported separately below.
Test Result
Passed locally after cleanup and the CUDA Graph cache update:
cargo fmt --all --check cargo clippy --workspace --locked --offline --all-targets -- -D warnings cargo test --workspace --locked --offline cargo build --workspace --release --locked --offlinemkdocs build --strictpassed. The October 2 evidence is preserved separately. Runtime code is unchanged by these documentation updates.61b83b3. The CUDA source and GPU test contents are unchanged; their location and registration changed, and device checks were not repeated. SM89 compilation passed previously; SM89 device execution and full Cua-S1 checkpoint inference were not checked.Latest warm HTTP latency — H200, 2026-10-03
74 real JevBench
noulrequests, one candidate each, BF16, max length 16384 and concurrency 1 on exact H200 GPU 2 (UUIDGPU-cbf66259-f4ab-0ede-1811-82037dde5924; NUMA 0, CPUs 0–15). All three backends were newly timed, with one reused server, one excluded feasibility pass and two measured passes each: 148 measured requests per configuration.Native's mean latency gives a 7.47× speedup over raw HF (86.61% lower latency) on this subset. Native and HF agree on all 74 thresholded decisions and both score 64/74, with maximum probability difference 0.020423. Fast scores 63/74, with one different decision (
hard-opus-a-temporal_numeric-09); native/Fast maximum probability difference is 0.034353. Each backend's probabilities are exactly unchanged across its feasibility and measured passes. These counts do not establish statistical accuracy superiority or full numerical parity.Raw HF baseline: original Open-Jev server and
DecisionModelwith the unmerged PEFT LoRA adapter, trained scalar head and saved temperature. Full attention uses stock SDPA; runtime checks verify all 48 linear-attention layers usetorch_chunk_gated_delta_rule, stock convolution andQwen3_5RMSNormGated, with 160 unmerged LoRA modules. The baseline disables the optional FLA/causal-conv1d availability checks before importing model classes, while reusing the same prepared environment. It uses no custom Fast model, torch.compile, CUDA Graph replay or prefix cache. Native retains its accepted merged-LoRA CUDA path; original Fast retains its custom kernels and CUDA Graph stack, reporting 30 retained graphs.Timing includes localhost HTTP through the same frozen Rust frontend, tokenization, worker execution and UTF8 response decoding. Client body serialization and response JSON parsing are excluded. Downloads, preparation, process-to-readiness, warmup, first inference after readiness and the complete feasibility pass are excluded. No Nsight launcher or trace collection was used. P50 is the median; P95 uses JevBench's
sorted[int(0.95*N)-1]rank. Pass ranges are observed variability, not confidence intervals.Native's observed mean is 4.78% below Fast in this run; Fast has a lower median. The mean pass ranges are disjoint, but the fixed two-pass budget and single-candidate workload do not establish a general winner. Multi-candidate prefix sharing and broader JevBench coverage remain unvalidated. Native graph replay is validated separately below. The author's 17.3 ms B300 result uses different hardware/workload.
Frozen native worker/frontend
202c0e1and the accepted packed-SiLU library were reused; PR source at measurement wasad1cb81. Open-Jev3308a15, Fastc52b8bb, JevBenchf8ce713; pinned base1d4bf0f, adapter28cf730, temperature 2.5343690298472983. Torch 2.13.0+cu130, Transformers 5.10.2 and PEFT 0.19.1. Raw commands, rows, token IDs, source snapshots and 48 input hashes are archived locally inprofile/jev-hf-transformers-comparison-20261003/.A collector cleanup assertion treated native's intentional SIGTERM exit as failure after all its measurements were saved. Only the remaining Fast configuration continued on a second reservation of the same GPU/affinity; no extra measured passes were added. All task-owned processes exited and GPU 2 returned to 0 MB used.
Native CUDA Graph replay — H200, 2026-10-03
Warm mixed-length HTTP: 48.086→47.112 ms (2.03% lower) with the new bounded 64-entry exact-length cache; fixed 107-token requests improve 3.77%. The former eight-entry cache regresses the mixed workload to 95.289 ms through repeated capture. Graph mode stays opt-in with
CUA_S1_GRAPH=1; new lengths and scratch growth still incur preparation costs.Same exact H200 GPU 2, affinity, BF16, model, CUDA library, frontend and requests. Both capacities are rebuilt from frozen
202c0e1sources with identical options and dependencies; their model source differs only in the cache limit. Every configuration validates its first long inference after real readiness to preallocate the maximum workload scratch size, then validates the short request. One server is reused for one excluded feasibility pass and exactly two measured passes per workload: 32 repeated short requests per pass, then all 74 mixed cases per pass. HTTP measurement is unprofiled.All 74 probabilities and decisions remain exactly unchanged (maximum delta 0.0). Both graph-64 pass means improve over eager; the mixed aggregate narrowly passes the prespecified 2% gate. Two passes are observed variability, not confidence intervals. After mixed passes, graph-64 uses 122 MB more observed device memory than graph-eight (scheduler samples, not peak). The excluded graph-64 mixed feasibility mean is 110.255 ms; it records capture cost and is not a controlled cold-latency result.
A separate preceding Nsight run proves one graph launch, zero recaptures and zero individual runtime kernel-launch calls per warm short request, with all 834 kernels retained. Node-level graph tracing shows larger gaps despite lower unprofiled latency, so it supplies replay proof rather than a speedup estimate. Full Cua-S1 checkpoint graph inference is not revalidated; its shared opt-in cache bound also increases.
Format, workspace Clippy/tests/release build, strict MkDocs and diff checks pass. The PR model source matches the measured 64-entry candidate exactly; CUDA kernels and tests are unchanged. Both GPU jobs clean up their owned processes and release GPU 2 to 0 MB used. Raw plans, source copies, commands, responses, traces, memory samples and verification are preserved in
profile/jev-cuda-graph-20261003-074613/andprofile/jev-cuda-graph-cache64-20261003-075238/outside the PR. These native A/B results are separate from the matched raw HF/Fast comparison above. See the validation details.Earlier isolated packed-SiLU A/B — 2026-10-02
With cached RMSNorm fixed in both native variants, packed SiLU lowered native mean 49.527→48.297 ms (2.483%), with disjoint observed pass ranges and all 74 native probabilities and decisions exactly unchanged. Fast means in that earlier run were 57.345 / 50.908 ms, with 6.436 ms spread associated mainly with preparation variability in two long-policy requests. These earlier timings used an inactive Nsight session with possible CUPTI instrumentation and are separate from the new HF comparison.
Separate CUDA traces showed long-request SiLU 21.217→6.380 ms (69.93%), versus Fast 5.466 ms, closing 94.19% of that measured kernel-family gap. Kernel totals are separate from HTTP latency; Nsight Systems supplied timelines because Nsight Compute counters were denied. See the October 2 revisions, controls and reproduction details.
Self-review
Contributor checklist from CONTRIBUTING.md, left for completion before requesting review: