Run Laya encoder and decision layers on CUDA - #49
linear3735 wants to merge 28 commits into
Conversation
|
can we push this PR faster? |
|
Please add a controlled H800 latency comparison against upstream Laya 0.3.20's CUDA fast path. The 208 hidden-state comparisons establish numerical parity; we also want to know whether this Rust encoder improves latency. Keep the comparison scoped to the encoder plus decision transformer layers:
This would let us assess the performance benefit before the later scorer, HTTP and CUDA Graph integration. |
Added the H800 comparison, CUDA-event timings and reproduction steps to Test Result. Latency is close to upstream; all 30 final-hidden parity checks passed with zero error. |
Purpose
Run Laya's 28 encoder layers and two decision transformer layers from Rust. Validate the checkpoint and kernel bundle before loading, then leave hidden states on the GPU for the scorer.
Part of #14. Depends on #48 and the existing Hopper kernel build. Uses original RoPE and eager execution. Scorer, HTTP and CUDA Graph integration are separate steps.
Kept in draft until #48 merges. The diff against
maincurrently includes its dependencies. Review this increment: 419 core/configuration lines for encoder execution.Tests now live under the repository-root
tests/directory. Cargo target names and test coverage are unchanged.Test Plan
cargo fmt --all --check cargo clippy --workspace --locked --all-targets -- -D warnings cargo test --workspace --locked cargo build --workspace --release --lockedCompare with official Laya 0.3.20 on H800: 13 real request cases, including 16 distinct rows at 16×512. Check intermediate and final hidden states, reverse request order, and repeat each workspace without intermediate readbacks. Reproduction commands are in recipe/laya/README.md. The residency/workspace recipe covers the inherited resource checks.
System1-Omni Version / Commit:
b7c9384. Encoder increment:2f5ff4b→ff4d8fa; GPU run:684470a. Later changes add pre-load rotary-size validation, device selection and test/documentation layout; model execution and kernels are unchanged.Test Result
Current CPU CI: 45 tests passed; 9 checkpoint/GPU tests skipped. fmt, strict Clippy and release build passed. This documentation-only follow-up also passed local strict MkDocs and rendered-anchor checks. The historical GPU validation and benchmark below were not rerun; encoder and CUDA source are unchanged.
All 208 valid-token comparisons had zero measured error. Repeated runs produced identical output bytes. Candidate outputs were finite; empty-key padding was excluded from equality checks. These checks validate implementation parity.
CI for
b7c9384: Rust CI, Docs build, benchmark harness tests passed.H800 comparison, 2026-10-03: original RoPE in #49 vs Laya 0.3.20 eager fast path. Same GPU, checkpoint, token IDs, lengths and question types; BF16, concurrency 1, CUDA Graphs disabled. Scope: embedding + 28 encoder + 2 decision layers; scorer and HTTP are outside this benchmark.
Host p50 is 0.3–1.6% lower for short/long inputs and 0.4–0.5% higher at
16×512.Each cell is run 1 / run 2; latencies are in ms. Order: upstream→Rust, then Rust→upstream. Each case/backend/run uses 20 warmup calls and 100 measured calls, after a feasibility pass. Host timing includes uploads and completion sync; req/s counts encoder calls. Rust includes validation/serialization and per-upload sync; upstream uses prebuilt CPU tensors and a final sync. Loading, compilation, warmup and readback are excluded.
30/30 final-hidden comparisons passed:
nRMS=0,max_abs=0; existing thresholds unchanged. These are implementation-parity checks.CUDA-event timings and reproduction
CUDA events were measured in separate passes after uploads and before completion sync. They cover the compute-stream interval, including launch gaps. Same two rounds and units as above.
Reproduction, frozen inputs, measurement scripts and all raw timings. The two rounds remain separate; no samples were removed or pooled.
Measurement source:
ff4d8fa13b2c8d52027aa2c565a0b97940d1c0ca; the 12 Laya model/runtime files match #49 headb7c9384. Only the benchmark harness and event wrapper were added for measurement. Checkpoint:convaiinnovations/laya@55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851.H800 UUID:
GPU-9e874df2-5775-c81e-0730-d23a5f091c3d; driver 580.159.03. Python 3.10.21, Laya 0.3.20, Torch 2.11.0+cu128, TileLang 0.1.14, transformers 5.17.0, Rust 1.98.1, nvcc 13.0.88. Source, checkpoint, library and input hashes are in the evidence files.Self-review
Before marking this PR ready for review or requesting maintainer review, complete
the self-review checklist.
Keep the PR in draft while this work is incomplete.
For agent assistance, use the optional precheck-pr skill.