Try budgeted cache recovery before splitting a minimum wave - #1053
Draft
bradhilton wants to merge 1 commit into
Draft
bradhilton wants to merge 1 commit into
bradhilton wants to merge 1 commit into
Conversation
bradhilton
deployed
to
trainer-rank-gpu-validation
September 30, 2026 01:36 — with
GitHub Actions
Active
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
When a minimum microbatch wave fits only after splitting, give its rejected flat demand one cache-recovery opportunity under the existing budget before accepting the split. The existing WORLD fallback vote carries this opportunity uniformly across split, flat and empty local shares. Recovery runs ordinary fresh search once; denied or ineffective recovery keeps the split only if a fresh admission check still fits.
This intentionally changes the prior “a split fits, return immediately” behavior. It preserves the shared lifetime first attempt and subsequent 5% ledger, physical-free admission and error handling. It adds no ordinary-path collective, new budget, public API or
art.megatronchange.Validation: 62 focused CPU tests and 9 subtests passed, with CUDA hidden and actual cleanup. The changed real-estimator denied-upgrade test also passed in the earlier run; its runtime/test bytes are unchanged. Ruff and diff checks passed. Rebased onto current main
c1b99e36with identical candidate/runtime/dependency files; the intervening change concerns trajectory-capture cleanup. Actual two-rank collective and candidate-native qualification remain pending.Motivation comes from a separate baseline GPU diagnostic, not qualification of this candidate: two approximately 135 ms releases restored the exact natural 64,902-row flat plan. Its same-plan allocated peak stayed 93.849 GiB. Quiet flat forward/backward took 8.332 s versus 10.506 s for a historical split, with higher peak memory; this separate-run comparison is not a throughput guarantee.
Related to #848 and #870. This does not close their broader qualification scope.