Skip to content

Price the adapter gradients a recompute backward holds at its peak - #1002

Open
bradhilton wants to merge 19 commits into
mainfrom
dalinar/adapter-grad-floor
Open

bradhilton wants to merge 19 commits into
mainfrom
dalinar/adapter-grad-floor

Conversation

@bradhilton

@bradhilton bradhilton commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

A full-recompute backward allocates each layer's adapter gradients as it passes, and the step's optimizer frees them. While it recomputes layer i it still holds the saved boundaries of layers 0..i. So a short wave peaks near layer 0, with almost every layer's gradients live. The checkpoint floor priced only the other end (all boundaries), so short first waves were admitted below their peaks.

Change:

  • The floor adds the largest excess of pending gradients over released boundaries, while a gradient group's slot still has unallocated adapter gradients (the step's first gradient wave).
    • It is computed over the real per-layer sizes, with no uniform-layer assumption.
    • A parameter shared by several layers counts once, at its highest layer; only local shards count.
    • Parameters outside the decoder, including any a head shares with a layer, and a checkpoint's other trainable parameters count as live throughout.
  • Gradient groups run their backward one after another, not layer by layer together. Autograd drains the last-forwarded group first, and separate backward calls can come in either order. While one group runs, groups already run hold their gradients and groups not yet run hold their boundaries. The price is the worst order, in closed form. A single group prices exactly as before.
  • Split children training the same slots charge it once; children with different slots each charge their own.
  • Unprofiled gradient waves also carry 64 MiB of measured first-execution transients. The non-full-recompute path already has an equivalent allowance.
  • The cheap width estimate defers to the exact plan while a slot has pending gradients, since only slot metadata can price them.

Before/after (Qwen3.6-35B-A3B, 40 layers, CP2, one real sequence per wave; raw = estimate ÷ 1.1 over the peak, admitted = estimate over the peak):

before raw / admitted after raw / admitted
EP1 cold, 2,047 tokens -19.5% / -11.5% +18.7% / +30.6%
EP1 cold, 4,095 tokens -10.9% / -2.0% +3.8% / +14.2%
EP1 cold, 8,337 tokens +7.0% / +17.7% +10.0% / +21.0%
EP1 cold, 20,245 tokens +14.9% / +26.4% +16.3% / +27.9%
EP2 cold, 2,047 tokens +11.3% / +22.4% +28.1% / +40.9%

Testing: 21 new CPU tests, including brute-force checks over random per-layer sizes and every backward order of up to six gradient groups, mutation-checked. The trainer-rank suite passes. The GPU runs above are single-group waves, whose pricing the sequential-groups change leaves unchanged.

CI status: GPU validation passed on this exact head in run 36399223939 ('Run on 2x H200' in 20m58s). The later run 36404599822 was started automatically by the ready-for-review event, and I cancelled it before it provisioned a GPU to keep shared H200 jobs one at a time. Its failed status is that cancellation, not a code failure.

Limits:

  • Measured on one MoE model at CP2 (EP1 and EP2). Dense models, TP>1, and waves with two or more trainable slots (priced one group's backward after another) are covered by CPU tests only.
  • Short waves now over-estimate more once their gradients are pending: a warm 2k EP1 wave goes from +17% to +34%, and a 2k EP2 wave from +35% to +50%. Waves above ~8k tokens change by 1–3 points.
    • Allocator snapshots put most of the remaining EP2 excess in main's own terms: the MoE workspace's 1.5x routing allowance and the full-copy input gradient.
    • Pricing each backward layer from only its own live parts would tighten this, but on main it would under-admit EP1 4k (its live MoE stage and mixer exceed the priced workspace), so that belongs with a more complete workspace model.
  • The 64 MiB allowance is the transients measured on this model; other models' first executions weren't measured.
  • The cost components gain two fields, so reports recorded before this change replay (with --allow-source-drift) as mismatched.

🤖 Generated with Claude Code

A full-recompute backward allocates each layer's adapter gradients as it
passes, and the step's optimizer frees them. Recomputing layer i still holds
the saved boundaries of layers 0..i, so a short first wave peaks at layer 0
with nearly every layer's gradients live; the checkpoint floor priced only
the last layer's end (all boundaries). On Qwen3.6-35B-A3B CP2, 2k and 4k
token first waves were admitted 11.5% and 2.0% under their peaks.

The floor now adds the largest excess of pending gradients over released
boundaries across the real layers, while a slot's gradients are unallocated,
and an unprofiled wave's 64 MiB of first-execution transients.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 27, 2026 02:15 — with GitHub Actions Failure
Name gradient slots with sorted kind/name JSON instead of a hash.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 27, 2026 02:45 — with GitHub Actions Error
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 27, 2026 03:04 — with GitHub Actions Active
bradhilton and others added 2 commits September 28, 2026 06:33
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 28, 2026 06:43 — with GitHub Actions Active
bradhilton and others added 3 commits September 28, 2026 08:08
Autograd drains the last-forwarded group's chain before an earlier one's, and
separate backward calls may come in either order, so a short group's
gradients can peak beside another group's unreleased boundaries. Price each
gradient group against its own boundaries, with groups not yet run holding
theirs and groups already run holding their gradients, over every order.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Any set of the other groups can have run first, so the worst order adds
every other group whose gradients outweigh its boundaries to one group's own
walk. Exact for any number of groups, without walking permutations.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 28, 2026 08:45 — with GitHub Actions Active
bradhilton added a commit that referenced this pull request Sep 28, 2026
…ed-mixer floor

Resolves _subforward_cost/_estimate for the stack's one-layer gradient and
adds _checkpoint_layer_boundaries (uniform). Still to do: #978 per-rank,
per-layer layout boundaries; merge #1002's later commits; suites; reviews.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton added a commit that referenced this pull request Sep 28, 2026
Uses _gradient_slots so base-model groups own no adapter gradients.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton added a commit that referenced this pull request Sep 28, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton added a commit that referenced this pull request Sep 28, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton added a commit that referenced this pull request Sep 28, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton marked this pull request as ready for review September 28, 2026 09:36
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 28, 2026 09:36 — with GitHub Actions Error
bradhilton and others added 6 commits September 28, 2026 17:01
…floor

# Conflicts:
#	src/art/trainer_rank/_impl.py
Grouped planner replay (#1028) recomputes each subforward's cost from
primitive runtime facts, with group indices standing in for slot refs. The
adapter-gradient floor reads the gradient slot's unallocated LoRA gradients
from the live model, so replay could neither resolve the slot nor reproduce
the term. Capture each gradient group's slot kind and name and its pending
gradient bytes per decoder layer with the selection (runtime facts version
2), and have ReplayRank answer the floor's slot and pending-gradient readers
from those facts. The capture's stock-estimator check now covers the floor's
readers too.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Replay sizes per-layer boundary tuples from the report's recorded
num_layers, which only capture bounded; refuse counts outside capture's
1024-layer limit before any estimator runs. A slot without a kind (a
megatron-less reference) has no pending gradients, so reject kindless facts
that claim some. Test forged adapter facts, the layer bound, and an
instance-overridden pending-gradient reader.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton and others added 4 commits September 29, 2026 04:16
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton marked this pull request as draft September 29, 2026 04:38
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 29, 2026 04:44 — with GitHub Actions Active
@bradhilton
bradhilton marked this pull request as ready for review September 29, 2026 05:19

This branch was successfully deployed

1 active deployment
trainer-rank-gpu-validation — c41148ec Deployed Sep 29, 2026 by bradhilton via Run on 2x H200 #965
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant