cuda: clean up failed and unwound Laya graph captures - #73
Open
Levius-Fubuki wants to merge 39 commits into
Open
Levius-Fubuki wants to merge 39 commits into
Levius-Fubuki wants to merge 39 commits into
Conversation
This was referenced Oct 5, 2026
Levius-Fubuki
marked this pull request as ready for review
October 5, 2026 16:08
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Harden the existing fixed-shape Encoder/Decision Graph path: resolve the complete optional API before capture; reject nested capture and prohibited allocation/copy/readback/sync/replay; end capture on callback error or unwind; reject null success and own partial handles. The C capture-end path releases temporary/partial graphs on every outcome.
Stack 3/4, dependent on #72; merge in order. This PR targets main to run current-head CI and includes predecessor increments until they merge. Review this increment against
6ecf4a11bff5a71e81215a980b4389a88c54600e.Test Plan
Published head:
30b63b7791c81458c786d8ec6ffd254b4b90cb83Validated source revision:
231cf358770e1464efbbce86ec6ce47cb5746056(source-identical tree; the published head adds only an empty CI-trigger commit)Incremental baseline:
6ecf4a11bff5a71e81215a980b4389a88c54600eCurrent-main baseline:
7f39ac40902c374803992407bb26eeba29c8a588cargo fmt --all --check cargo clippy --workspace --locked --all-targets --offline --features omni-laya/serve -- -D warnings cargo test --workspace --locked --offline --features omni-laya/serve cargo build --workspace --release --locked --offline --features omni-laya/serveReview incremental and cumulative diffs. Existing history is preserved: the previous PR head remains an ancestor. Current full-request processing/heads/HTTP/runtime integration replaces obsolete encoder-only modules.
Test Result
All four checks passed on the exact individual validated source revision above. The published head has the same tree, confirmed with Git; its current-head CI is checked separately. 83 passed, 0 failed, 10 ignored. Ignored checkpoint/GPU tests remain unverified. Incremental
git diff --checkpassed. The PR remains core source only; external fixtures/logs/evidence are retained locally.Five external CPU ABI fixtures passed resolved lifetime, grouped copies/errors, capture error/panic recovery and operation guards, missing optional Graph symbols preserving eager use, and replay rejection inside capture. The real-driver suite passed callback error, deliberately caught unwind, nested/prohibited-operation guards and successful same-stream recapture. Model trace/staging fixtures also passed.
Fresh generic CUDA validation passed using exact #73 backend source
231cf358770e1464efbbce86ec6ce47cb5746056on RTX 4090 (sm89), driver 580.105.08, nvcc 13.0.88 and Rust 1.99.0. An external init-only sm89 bootstrap and scalar-add kernel exercised production resource/copy/capture/replay/free operations. The production Hopper gate was preserved. #71 resolved dispatch matches tested methods byte-for-byte; #72 group copy matches outside the later capture guard; #74 backend source equals #73. Two setup attempts stopped before GPU work because Python was missing from PATH; corrected PATH passed with unchanged source/assertions.Validation limit: This is generic host/backend FFI evidence, not Laya support on RTX 4090. Full optimized Laya eager/Graph/scorer/action-head/HTTP numerical parity and performance were not rerun on Hopper after this port. No new Laya accuracy or speedup claim is made. The earlier 2026-10-04 H800 encoder-stack comparison remains historical.
Numerical CUDA/TileLang/RoPE implementation, operator export/build sources, weights/conversion, preprocessing, padding, decoding, executor/serving and Cargo metadata match current main. Only host execution/resources/cache and scoped C capture cleanup change.
Self-review
Incremental and cumulative diffs were reviewed against current architecture/runtime contracts, supported inputs, precision, pointer/drop lifetimes, failure paths and claims. Local review found no remaining actionable defect in this host scope. Fresh full-Hopper validation remains unverified; contributor and maintainer review are separate from this assistance.
Demo / evidence
Raw current-source workspace, CPU trace/copy/capture/cache fixture logs, exact source/ancestry hashes and real-driver output are retained under
artifacts/all-pr-followups-20261005/laya/, outside this core-code diff. The deliberately caught panic in the driver log is an expected cleanup test; the suite exit code is zero. No video or measured speedup is claimed. Current-head GitHub Rust, benchmark and docs CI all passed; deploy skipped. CI run.Rebuild
liblaya_cuda.soand regenerate its build-manifest hash to include the C capture cleanup change. The C ABI signatures remain compatible; the old library does not gain the C-side fix merely by rebuilding the Rust worker.