Price the adapter gradients a recompute backward holds at its peak - #1002
Open
bradhilton wants to merge 19 commits into
Open
bradhilton wants to merge 19 commits into
bradhilton wants to merge 19 commits into
Conversation
A full-recompute backward allocates each layer's adapter gradients as it passes, and the step's optimizer frees them. Recomputing layer i still holds the saved boundaries of layers 0..i, so a short first wave peaks at layer 0 with nearly every layer's gradients live; the checkpoint floor priced only the last layer's end (all boundaries). On Qwen3.6-35B-A3B CP2, 2k and 4k token first waves were admitted 11.5% and 2.0% under their peaks. The floor now adds the largest excess of pending gradients over released boundaries across the real layers, while a slot's gradients are unallocated, and an unprofiled wave's 64 MiB of first-execution transients. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton
had a problem deploying
to
trainer-rank-gpu-validation
September 27, 2026 02:15 — with
GitHub Actions
Failure
Name gradient slots with sorted kind/name JSON instead of a hash. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton
had a problem deploying
to
trainer-rank-gpu-validation
September 27, 2026 02:45 — with
GitHub Actions
Error
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton
deployed
to
trainer-rank-gpu-validation
September 27, 2026 03:04 — with
GitHub Actions
Active
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton
deployed
to
trainer-rank-gpu-validation
September 28, 2026 06:43 — with
GitHub Actions
Active
Autograd drains the last-forwarded group's chain before an earlier one's, and separate backward calls may come in either order, so a short group's gradients can peak beside another group's unreleased boundaries. Price each gradient group against its own boundaries, with groups not yet run holding theirs and groups already run holding their gradients, over every order. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Any set of the other groups can have run first, so the worst order adds every other group whose gradients outweigh its boundaries to one group's own walk. Exact for any number of groups, without walking permutations. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton
deployed
to
trainer-rank-gpu-validation
September 28, 2026 08:45 — with
GitHub Actions
Active
bradhilton
added a commit
that referenced
this pull request
Sep 28, 2026
…ed-mixer floor Resolves _subforward_cost/_estimate for the stack's one-layer gradient and adds _checkpoint_layer_boundaries (uniform). Still to do: #978 per-rank, per-layer layout boundaries; merge #1002's later commits; suites; reviews. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton
added a commit
that referenced
this pull request
Sep 28, 2026
bradhilton
added a commit
that referenced
this pull request
Sep 28, 2026
bradhilton
added a commit
that referenced
this pull request
Sep 28, 2026
bradhilton
added a commit
that referenced
this pull request
Sep 28, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This was referenced Sep 28, 2026
bradhilton
marked this pull request as ready for review
September 28, 2026 09:36
bradhilton
had a problem deploying
to
trainer-rank-gpu-validation
September 28, 2026 09:36 — with
GitHub Actions
Error
This was referenced Sep 28, 2026
…floor # Conflicts: # src/art/trainer_rank/_impl.py
Grouped planner replay (#1028) recomputes each subforward's cost from primitive runtime facts, with group indices standing in for slot refs. The adapter-gradient floor reads the gradient slot's unallocated LoRA gradients from the live model, so replay could neither resolve the slot nor reproduce the term. Capture each gradient group's slot kind and name and its pending gradient bytes per decoder layer with the selection (runtime facts version 2), and have ReplayRank answer the floor's slot and pending-gradient readers from those facts. The capture's stock-estimator check now covers the floor's readers too. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Replay sizes per-layer boundary tuples from the report's recorded num_layers, which only capture bounded; refuse counts outside capture's 1024-layer limit before any estimator runs. A slot without a kind (a megatron-less reference) has no pending gradients, so reject kindless facts that claim some. Test forged adapter facts, the layer bound, and an instance-overridden pending-gradient reader. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton
marked this pull request as draft
September 29, 2026 04:38
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton
deployed
to
trainer-rank-gpu-validation
September 29, 2026 04:44 — with
GitHub Actions
Active
bradhilton
marked this pull request as ready for review
September 29, 2026 05:19
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A full-recompute backward allocates each layer's adapter gradients as it passes, and the step's optimizer frees them. While it recomputes layer i it still holds the saved boundaries of layers 0..i. So a short wave peaks near layer 0, with almost every layer's gradients live. The checkpoint floor priced only the other end (all boundaries), so short first waves were admitted below their peaks.
Change:
Before/after (Qwen3.6-35B-A3B, 40 layers, CP2, one real sequence per wave; raw = estimate ÷ 1.1 over the peak, admitted = estimate over the peak):
Testing: 21 new CPU tests, including brute-force checks over random per-layer sizes and every backward order of up to six gradient groups, mutation-checked. The trainer-rank suite passes. The GPU runs above are single-group waves, whose pricing the sequential-groups change leaves unchanged.
CI status: GPU validation passed on this exact head in run 36399223939 ('Run on 2x H200' in 20m58s). The later run 36404599822 was started automatically by the ready-for-review event, and I cancelled it before it provisioned a GPU to keep shared H200 jobs one at a time. Its failed status is that cancellation, not a code failure.
Limits:
--allow-source-drift) as mismatched.🤖 Generated with Claude Code