Pre-resolve Laya encoder resources and kernel handles - #71
Draft
Levius-Fubuki wants to merge 1 commit into
Draft
Levius-Fubuki wants to merge 1 commit into
Levius-Fubuki wants to merge 1 commit into
Conversation
4 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Resolve encoder weights, rotary tables, attention choices, and kernel handles at load time. Encoder dispatch then uses fixed argument arrays without per-launch string lookup or pointer-vector allocation. Shared buffer ownership and resolved handles retain the CUDA context and native code for their required lifetimes.
Stack 1/4 for #66, dependent on #49. The temporary base is a pinned copy of #49’s head. Keep this PR in draft until #49 merges; then rebase if needed, retarget to
main, and recheck the diff before merging. Do not merge this PR into the temporary base.Test Plan
System1-Omni Version / Commit:
d235783328a5fea9b5cb5c771583779997ba63ebIncremental base:
b7c9384ae79ed666b22092e178efe16e91635f64(codex/laya-pr49-base)Review this incremental diff, run the workspace checks, and validate eager/graph correctness and cache behavior with the intended Hopper kernel bundle. This PR contains core source only; added regression tests and validation artifacts remain local.
Test Result
git diff --check b7c9384 d235783passed.Local fixture checks passed with
cargo test -p omni-cuda -p omni-laya --offlineandcargo clippy -p omni-cuda -p omni-laya --all-targets --offline -- -D warnings. Coverage included buffer and library lifetime, argument ordering, foreign contexts, and the legacy slice API. These checks used local regression additions that are not included in this PR.All four standard checks passed on an isolated export of the assembled core-only stack at
d249c5e244de4e42544d6fa3837aef40700cd7dc, using the committed test suite:cargo fmt --all --check cargo clippy --workspace --locked --all-targets -- -D warnings cargo test --workspace --locked cargo build --workspace --release --lockedThe assembled-stack results do not establish independent full-workspace validation of every intermediate commit. Environment-dependent tests were ignored as reported by the suite.
Hopper validation
Fresh validation on 2026-10-04 compared unchanged
b7c9384with assembled stackd249c5eon one H800 PCIe 80 GB (driver 595.71.05, CUDA 13.0.88, Rust 1.98.1). Python 3.12.3 / Torch 2.12.1+cu130 were used with Laya 0.3.20, TileLang 0.1.14 and Transformers 5.17.0. This is a new same-environment comparison, not a reproduction of the earlier Torch 2.11 report.Checkpoint:
convaiinnovations/laya@55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851; the same native operator bundle, original RoPE and dynamic attention served every native configuration.The existing intermediate check was run for both source trees with
cargo test --release --locked -p omni-laya --lib real_encoder_matches_official_hidden_states -- --ignored --nocapture, using the pinned checkpoint, bundle and frozen oracle. Local persistent-worker checks usedcompare_workers.pyandcheck_cache_gpu.py; their added harnesses, raw samples, identities and logs are retained locally rather than added to this core-only PR.One excluded feasibility pass preceded two measured rounds in opposite configuration order: 10 warmups and 50 samples per shape/configuration, 2,000 measured calls. Warm host latency includes validation, packing, uploads, execution and synchronization. Loading, capture, warmup, readback and IO are excluded; first-call costs were recorded separately. This covers serial encoder+decision-transformer execution, not HTTP or full task decisions.
The combined prepared-eager changes measured 3.065–3.067 ms at 1×48 versus baseline 3.092–3.094 ms. Most other shapes were inconclusive or nearly unchanged; at 16×512 the combined path was 0.04–0.18% slower. These results do not isolate the gain of this individual increment.
Self-review
Keep this PR in draft until contributor self-review is complete, following CONTRIBUTING.md.