fix: release GDN LoRA components before output addition - #1048
Conversation
|
Agent validation summary for exact head
The PR remains draft. Merge of the art.megatron change requires Brad's decision. |
|
Exact-head addendum for The earlier maintained old-red/new-green and component results remain applicable: 1.493 GiB lower allocated peak, 18 bitwise comparisons, and no improvement in reserved peak or synchronized physical free memory. There is no whole-model or planner claim. No local GPU tests were repeated for the annotation. Normal CI is running on the new head; this PR remains draft until applicable checks pass, and merging the runtime change remains held for Brad. |
|
Ready for review at exact head The measured benefit remains 1.493 GiB lower component allocated peak, with unchanged reserved peak and synchronized physical free memory. There is no whole-model or planner claim. The CI cluster's termination is recorded in its job log. This PR is unmerged; the |
|
Whole-model qualification update for unchanged head The genuine 64,902-token A+B forward comparison completed with zero allocated or reserved peak reduction. Both arms reached 94,780,217,856 bytes (88.271 GiB) peak allocated. Corresponding target hashes, starting allocator counters, natural plans, model state, RNG and inputs match. Warm2 was compiler-quiet in both arms. This was gradient-enabled forward only; it did not run backward. This limits the earlier 1.493 GiB component result: that local saving has not translated into a whole-model forward benefit on this workload. The exact CP1 source route reaches the patched projection under a recursively compiler-disabled call, but no local allocation trace establishes why the global peak is unchanged. A separate full backward comparison is being prepared; no backward, physical-memory or larger-batch admission benefit is claimed. Both retained operation histories are closed, including the prior candidate plan-serialization refusal; the final candidate reused the completed immutable baseline after the full canonical-plan guard correction. The final 63 recorded process identities/groups are absent and GPU2 is returned. Independent result receipt: |
|
Full-model forward/backward qualification is now complete for reviewed head
Corresponding baseline/candidate source, input, initial weights, RNG, plans and allocator starts matched. Warm 2 alone had no observed compilation-counter changes. The roughly 9 GiB cold-to-warm reduction is not evidence that warming makes the same execution plan cheaper: the planner changed the execution plan. All 660 gradients were present and finite. Warm target outputs/losses matched, but gradients varied within unchanged arms as well as across arms: relative L2 was 2.6285% between baseline repeats, 1.1323% between candidate repeats and 2.6255% between the final baseline/candidate passes. The first repeat still compiled. These small numbers of observations establish neither a causal numerical effect of this patch nor numerical equivalence. Cold loss differed by 0.03095%; objective formula, masks and denominator were unchanged. Both fresh processes completed with actual exit status 0 and no optimizer updates. All recorded owned identities and process groups were absent at closure; the GPU was released. No larger-batch admission, physical-memory peak or production-throughput improvement is claimed. Independent report: |
GDN LoRA projection keeps qkv, z, beta and alpha tensors alive after concatenating them. Release those four Python owners before the final output addition to reduce temporary component storage; tensor math and the out-of-place output stay unchanged.
A real TE component comparison on one H200, with fixed synthetic BF16 values at the 64,902-row Qwen geometry, measured a 1,603,339,264-byte (1.493 GiB) reduction in peak allocated memory across all 12 pairs with equal starting allocations. All 18 output/input-gradient/four-LoRA-gradient comparisons were bitwise exact, and the designated tensor state was unchanged. Peak reserved memory and synchronized physical-free readings were unchanged. This is component evidence; it establishes no whole-model savings or planner allowance.
Validation: changed-file Ruff lint and formatting checks pass. The maintained regression uses the real GDN fixture in no_grad and gradient paths. On the same H200, its two cases fail on the unmodified runtime at the intended storage-lifetime assertion, then both pass with this change, including input and all four LoRA gradient checks. Setup and teardown pass in both arms; normal pytest artifacts, JUnit and process exit statuses agree.
Draft for review. Merge of this art.megatron change is held for Brad's decision.