Skip to content

Pre-resolve Laya encoder resources and kernel handles - #71

Draft
Levius-Fubuki wants to merge 1 commit into
codex/laya-pr49-basefrom
codex/laya-execution-plan
Draft

Levius-Fubuki wants to merge 1 commit into
codex/laya-pr49-basefrom
codex/laya-execution-plan

Conversation

@Levius-Fubuki

Copy link
Copy Markdown
Collaborator

Purpose

Resolve encoder weights, rotary tables, attention choices, and kernel handles at load time. Encoder dispatch then uses fixed argument arrays without per-launch string lookup or pointer-vector allocation. Shared buffer ownership and resolved handles retain the CUDA context and native code for their required lifetimes.

Stack 1/4 for #66, dependent on #49. The temporary base is a pinned copy of #49’s head. Keep this PR in draft until #49 merges; then rebase if needed, retarget to main, and recheck the diff before merging. Do not merge this PR into the temporary base.

Test Plan

System1-Omni Version / Commit: d235783328a5fea9b5cb5c771583779997ba63eb
Incremental base: b7c9384ae79ed666b22092e178efe16e91635f64 (codex/laya-pr49-base)

Review this incremental diff, run the workspace checks, and validate eager/graph correctness and cache behavior with the intended Hopper kernel bundle. This PR contains core source only; added regression tests and validation artifacts remain local.

Test Result

git diff --check b7c9384 d235783 passed.

Local fixture checks passed with cargo test -p omni-cuda -p omni-laya --offline and cargo clippy -p omni-cuda -p omni-laya --all-targets --offline -- -D warnings. Coverage included buffer and library lifetime, argument ordering, foreign contexts, and the legacy slice API. These checks used local regression additions that are not included in this PR.

All four standard checks passed on an isolated export of the assembled core-only stack at d249c5e244de4e42544d6fa3837aef40700cd7dc, using the committed test suite:

cargo fmt --all --check
cargo clippy --workspace --locked --all-targets -- -D warnings
cargo test --workspace --locked
cargo build --workspace --release --locked

The assembled-stack results do not establish independent full-workspace validation of every intermediate commit. Environment-dependent tests were ignored as reported by the suite.

Hopper validation

Fresh validation on 2026-10-04 compared unchanged b7c9384 with assembled stack d249c5e on one H800 PCIe 80 GB (driver 595.71.05, CUDA 13.0.88, Rust 1.98.1). Python 3.12.3 / Torch 2.12.1+cu130 were used with Laya 0.3.20, TileLang 0.1.14 and Transformers 5.17.0. This is a new same-environment comparison, not a reproduction of the earlier Torch 2.11 report.

Checkpoint: convaiinnovations/laya@55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851; the same native operator bundle, original RoPE and dynamic attention served every native configuration.

  • Baseline and candidate each passed 208 intermediate/final comparisons across 13 frozen input cases in forward/reverse order, including encoder27: maximum nRMS and absolute error were both 0.
  • Eager, fixed Graph and cached Graph passed changed-input A→B→A, padding and shape-boundary checks. All 112 comparison records, including first-call outputs, matched the official reference with zero error; native outputs were byte-identical.
  • Six cache policies passed 48 requests covering count/byte eviction, surviving entries, exact budgets, oversized fallback and disabled caching. H800 runtime round-trip and Graph failure/lifetime checks also passed.

The existing intermediate check was run for both source trees with cargo test --release --locked -p omni-laya --lib real_encoder_matches_official_hidden_states -- --ignored --nocapture, using the pinned checkpoint, bundle and frozen oracle. Local persistent-worker checks used compare_workers.py and check_cache_gpu.py; their added harnesses, raw samples, identities and logs are retained locally rather than added to this core-only PR.

One excluded feasibility pass preceded two measured rounds in opposite configuration order: 10 warmups and 50 samples per shape/configuration, 2,000 measured calls. Warm host latency includes validation, packing, uploads, execution and synchronization. Loading, capture, warmup, readback and IO are excluded; first-call costs were recorded separately. This covers serial encoder+decision-transformer execution, not HTTP or full task decisions.

The combined prepared-eager changes measured 3.065–3.067 ms at 1×48 versus baseline 3.092–3.094 ms. Most other shapes were inconclusive or nearly unchanged; at 16×512 the combined path was 0.04–0.18% slower. These results do not isolate the gain of this individual increment.

Self-review

Keep this PR in draft until contributor self-review is complete, following CONTRIBUTING.md.

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant