Skip to content

laya: bound native workspace cache before allocation - #74

Open
Levius-Fubuki wants to merge 41 commits into
mainfrom
codex/laya-graph-cache
Open

Levius-Fubuki wants to merge 41 commits into
mainfrom
codex/laya-graph-cache

Conversation

@Levius-Fubuki

@Levius-Fubuki Levius-Fubuki commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Configure retained limits through CacheConfig and Model::load_with_cache, preserving four-shape/512 MiB defaults. Estimate the allocator’s same 17-buffer layout and evict LRU workspaces before allocating a miss. Admit exact fits; zero limits or oversized shapes reuse one eager workspace. Model::clear_cache releases scratch while retaining loaded weights. Budget excludes weights, graph/driver overhead, host staging and eager fallback storage.

Stack 4/4, dependent on #73; merge in order. This PR targets main to run current-head CI and includes predecessor increments until they merge. Review this increment against 231cf358770e1464efbbce86ec6ce47cb5746056.

Test Plan

Published head: ae8e34fcd292d4c8f21c35d4d143033a8f56a796
Validated source revision: 307380eef58133851c614f0a5e0f7dc20041c070 (source-identical tree; the published head adds only an empty CI-trigger commit)
Incremental baseline: 231cf358770e1464efbbce86ec6ce47cb5746056
Current-main baseline: 7f39ac40902c374803992407bb26eeba29c8a588

cargo fmt --all --check
cargo clippy --workspace --locked --all-targets --offline --features omni-laya/serve -- -D warnings
cargo test --workspace --locked --offline --features omni-laya/serve
cargo build --workspace --release --locked --offline --features omni-laya/serve

Review incremental and cumulative diffs. Existing history is preserved: the previous PR head remains an ancestor. Current full-request processing/heads/HTTP/runtime integration replaces obsolete encoder-only modules.

Test Result

All four checks passed on the exact individual validated source revision above. The published head has the same tree, confirmed with Git; its current-head CI is checked separately. 83 passed, 0 failed, 10 ignored. Ignored checkpoint/GPU tests remain unverified. Incremental git diff --check passed. The PR remains core source only; external fixtures/logs/evidence are retained locally.

Seven external model fixtures passed all 160 supported-shape size estimates; exact/count/byte limits; allocation-peak verification proving preallocation eviction; LRU and validation before mutation; disabled/oversized eager reuse; clear/release/recapture and failed-capture admission cleanup/retry; unchanged launch trace and A→B→A staging updates. The complete backend equals real-driver-validated #73 byte-for-byte.

Fresh generic CUDA validation passed using exact #73 backend source 231cf358770e1464efbbce86ec6ce47cb5746056 on RTX 4090 (sm89), driver 580.105.08, nvcc 13.0.88 and Rust 1.99.0. An external init-only sm89 bootstrap and scalar-add kernel exercised production resource/copy/capture/replay/free operations. The production Hopper gate was preserved. #71 resolved dispatch matches tested methods byte-for-byte; #72 group copy matches outside the later capture guard; #74 backend source equals #73. Two setup attempts stopped before GPU work because Python was missing from PATH; corrected PATH passed with unchanged source/assertions.

Validation limit: This is generic host/backend FFI evidence, not Laya support on RTX 4090. Full optimized Laya eager/Graph/scorer/action-head/HTTP numerical parity and performance were not rerun on Hopper after this port. No new Laya accuracy or speedup claim is made. The earlier 2026-10-04 H800 encoder-stack comparison remains historical.

Numerical CUDA/TileLang/RoPE implementation, operator export/build sources, weights/conversion, preprocessing, padding, decoding, executor/serving and Cargo metadata match current main. Only host execution/resources/cache and scoped C capture cleanup change.

Self-review

Incremental and cumulative diffs were reviewed against current architecture/runtime contracts, supported inputs, precision, pointer/drop lifetimes, failure paths and claims. Local review found no remaining actionable defect in this host scope. Fresh full-Hopper validation remains unverified; contributor and maintainer review are separate from this assistance.

Demo / evidence

Raw current-source workspace, CPU trace/copy/capture/cache fixture logs, exact source/ancestry hashes and real-driver output are retained under artifacts/all-pr-followups-20261005/laya/, outside this core-code diff. The deliberately caught panic in the driver log is an expected cleanup test; the suite exit code is zero. No video or measured speedup is claimed. Current-head GitHub Rust, benchmark and docs CI all passed; deploy skipped. CI run.

Rebuild liblaya_cuda.so and regenerate its build-manifest hash to include the C capture cleanup change. The C ABI signatures remain compatible; the old library does not gain the C-side fix merely by rebuilding the Rust worker.

linear3735 and others added 30 commits September 30, 2026 11:37
@hsliuustc0106 hsliuustc0106 mentioned this pull request Oct 5, 2026
2 of 4 tasks
@Levius-Fubuki
Levius-Fubuki changed the base branch from codex/laya-shape-graphs to main October 5, 2026 15:58
@Levius-Fubuki Levius-Fubuki changed the title Bound Laya encoder graph caching by shape count and workspace bytes laya: bound native workspace cache before allocation Oct 5, 2026
@Levius-Fubuki
Levius-Fubuki marked this pull request as ready for review October 5, 2026 16:08
Copilot AI balanced review requested due to automatic review settings October 5, 2026 16:08

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants