Skip to content

Reuse Laya input staging and group CUDA uploads - #72

Draft
Levius-Fubuki wants to merge 1 commit into
codex/laya-execution-planfrom
codex/laya-grouped-uploads
Draft

Levius-Fubuki wants to merge 1 commit into
codex/laya-execution-planfrom
codex/laya-grouped-uploads

Conversation

@Levius-Fubuki

Copy link
Copy Markdown
Collaborator

Purpose

Reuse workspace-owned host staging for IDs, lengths, and types, and submit their copies with one synchronization. Validate the entire upload group before entering CUDA and drain submitted work on copy failure while host memory remains borrowed. This removes repeated input packing allocations and two upload synchronizations per request; it does not provide transfer/compute overlap.

Stack 2/4 for #66, dependent on #71 and #49. The base is codex/laya-execution-plan, so this diff contains only the upload increment. After its predecessor lands in main, rebase if needed and retarget to main before merging.

Test Plan

System1-Omni Version / Commit: d97c21fff90ada1926084d530fba3a053e13d3fe
Incremental base: d235783328a5fea9b5cb5c771583779997ba63eb (codex/laya-execution-plan)

Review this incremental diff, run the workspace checks, and validate eager/graph correctness and cache behavior with the intended Hopper kernel bundle. This PR contains core source only; added regression tests and validation artifacts remain local.

Test Result

git diff --check d235783 d97c21f passed.

Local fixture checks passed with cargo test -p omni-cuda -p omni-laya; after the staging-loop Clippy repair, cargo test -p omni-laya --lib workspace::tests and cargo clippy -p omni-cuda -p omni-laya --all-targets -- -D warnings passed. Coverage included staging reuse, group prevalidation, one-sync transfers, and copy/sync failures. These checks used local regression additions that are not included in this PR.

All four standard checks passed on an isolated export of the assembled core-only stack at d249c5e244de4e42544d6fa3837aef40700cd7dc, using the committed test suite:

cargo fmt --all --check
cargo clippy --workspace --locked --all-targets -- -D warnings
cargo test --workspace --locked
cargo build --workspace --release --locked

The assembled-stack results do not establish independent full-workspace validation of every intermediate commit. Environment-dependent tests were ignored as reported by the suite.

Hopper validation

Fresh validation on 2026-10-04 compared unchanged b7c9384 with assembled stack d249c5e on one H800 PCIe 80 GB (driver 595.71.05, CUDA 13.0.88, Rust 1.98.1). Python 3.12.3 / Torch 2.12.1+cu130 were used with Laya 0.3.20, TileLang 0.1.14 and Transformers 5.17.0. This is a new same-environment comparison, not a reproduction of the earlier Torch 2.11 report.

Checkpoint: convaiinnovations/laya@55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851; the same native operator bundle, original RoPE and dynamic attention served every native configuration.

  • Baseline and candidate each passed 208 intermediate/final comparisons across 13 frozen input cases in forward/reverse order, including encoder27: maximum nRMS and absolute error were both 0.
  • Eager, fixed Graph and cached Graph passed changed-input A→B→A, padding and shape-boundary checks. All 112 comparison records, including first-call outputs, matched the official reference with zero error; native outputs were byte-identical.
  • Six cache policies passed 48 requests covering count/byte eviction, surviving entries, exact budgets, oversized fallback and disabled caching. H800 runtime round-trip and Graph failure/lifetime checks also passed.

The existing intermediate check was run for both source trees with cargo test --release --locked -p omni-laya --lib real_encoder_matches_official_hidden_states -- --ignored --nocapture, using the pinned checkpoint, bundle and frozen oracle. Local persistent-worker checks used compare_workers.py and check_cache_gpu.py; their added harnesses, raw samples, identities and logs are retained locally rather than added to this core-only PR.

One excluded feasibility pass preceded two measured rounds in opposite configuration order: 10 warmups and 50 samples per shape/configuration, 2,000 measured calls. Warm host latency includes validation, packing, uploads, execution and synchronization. Loading, capture, warmup, readback and IO are excluded; first-call costs were recorded separately. This covers serial encoder+decision-transformer execution, not HTTP or full task decisions.

The combined prepared-eager changes measured 3.065–3.067 ms at 1×48 versus baseline 3.092–3.094 ms. Most other shapes were inconclusive or nearly unchanged; at 16×512 the combined path was 0.04–0.18% slower. These results do not isolate the gain of this individual increment.

Self-review

Keep this PR in draft until contributor self-review is complete, following CONTRIBUTING.md.

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant