Skip to content

Price the recomputed layer's attention or GDN activations in the checkpoint floor - #963

Draft
bradhilton wants to merge 27 commits into
mainfrom
dalinar/recompute-mixer-floor
Draft

bradhilton wants to merge 27 commits into
mainfrom
dalinar/recompute-mixer-floor

Conversation

@bradhilton

@bradhilton bradhilton commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

Part of #949. This makes TrainerRank's cold memory estimate count what is actually live at the backward peak under full one-layer recompute. The expert-parallel routing allowance also now depends on the EP size.

Builds on #1002, which is merged into this branch; review #1002 first. This PR prices #1002's adapter-gradient term against its own one-incoming-gradient floor. On the full stack with #1002 (#963 + #978 + #981; Qwen3.6-35B-A3B, 40 layers, CP2, one real sequence per wave), the raw estimate against the measured peak is:

  • EP1 cold first waves (2k to 20k tokens): +0.9% to +2.0%;
  • EP2 cold first waves: +3.0% to +13.4%;
  • warm waves: +2.8% to +24.9%.

None are under. The table below predates #1002.

What changes:

  • Recomputed mixer. The recomputed layer's attention or GDN activations stay live beside its MoE stage, and the estimate now counts them. Leaving them out under-predicted CP2/EP1 by 23%.
  • Incoming gradient. When every layer is a MoE layer whose expert stage, FC1 included, the estimate prices, it charges the one incoming gradient live at the peak, not one per layer. The old charge over-counted about 17 GB on the random-weight 40-layer run and 8.4 GB on the real-data one. Other models keep the old charge, which also covers recompute work the estimate doesn't price.
  • Items the old charge hid. These are now priced: the recomputed layer's residual and norm rows, a sixth GDN value tensor, the shared expert's FC1/GLU outputs, router and permutation state, and Transformer Engine's cuBLAS workspaces (first call only).
  • HybridEP dispatch. HybridEP keeps one dispatched copy of each routed row, not two. It dispatches the whole EP group's tokens, so when the EP group is the CP group, routed rows are now priced on the group's even share, not on the busiest CP rank's rows. For example, on a real-data CP2 batch one rank held 52,480 rows, but each rank's even share of routed tokens was 48,397.
  • EP routing allowance. It was a flat 1.5 and is now 1.4 / 1.6 / 2.0 at EP2 / 4 / 8. The values come from a forward-only routing measurement on 3.5M tokens of real retail trajectories. In 200k-token batches the worst layer reached up to 1.22 / 1.40 / 1.62 with pretrained weights and 1.24 / 1.41 / 1.61 with a trained policy; a small rollout sample reached 1.95 at EP8, close to its 2.0 allowance. The old flat 1.5 was too low for EP8. EP4 and EP8 have no memory runs.

Raw estimate vs measured PyTorch peak on the most-loaded rank, Qwen3.6-35B-A3B on two H200s. Random-weight runs use 194,753 tokens. Real-data runs use about 206k tokens of real retail agent trajectories (97k after prefix sharing).

case before (cold) after (cold / warm)
CP1 +2.1% +0.5% / +15.9% ¹
CP2/EP1 −23% (under) +0.3% / +6.8% ¹
EP2/CP2, random weights, 8 layers +38.5% +26.5% / +23.6%
EP2/CP2, random weights, 40 layers not run +19.7% / +17.5%
EP2/CP2, pretrained, real data, 40 layers not run +18.2% / +19.2%
EP2/CP2, trained policy (059, step 225), real data, 40 layers not run +18.3% / +19.4%

¹ Measured on an earlier commit of this PR. The later changes don't alter the estimates for these two configurations.

In the pretrained EP2 cold case, the estimate is 26.81 GB against a 22.67 GB PyTorch peak. 1.05 GB of that estimate is HybridEP's communication buffer, which lives outside PyTorch, so the PyTorch peak can't show it. Leaving it out, the estimate is +13.6%. The GPU's memory outside PyTorch grew 2.76 GB during that call; the estimate prices only the HybridEP buffer part of that.

Limitation. EP2 is still well above the 10% target. With this formula and allowance, on these measured workloads, most of the remaining gap is the routing allowance: 1.4 is priced, while the layer at the peak saw about 1.0. The allowance covers the worst imbalance measured: 1.16 to 1.24 per layer on real data, and about 1.35 inferred from one production run. A cold estimate can't know which layer will be imbalanced, and these routing samples bound only what was measured, not all possible routing. Warm calls could price the routing they have already seen; that is a separate proposal.

Testing: all trainer-rank unit tests pass. The table's runs are on local H200s.

🤖 Generated with Claude Code

@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 06:42 — with GitHub Actions Failure
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 07:09 — with GitHub Actions Failure
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 25, 2026 07:20 — with GitHub Actions Active
@bradhilton
bradhilton force-pushed the dalinar/recompute-mixer-floor branch from 1ffff5e to 4b3d1e0 Compare September 25, 2026 15:47
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 15:47 — with GitHub Actions Failure
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 15:55 — with GitHub Actions Failure
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 25, 2026 15:59 — with GitHub Actions Active
bradhilton and others added 8 commits September 25, 2026 17:17
…kpoint floor

Full one-layer recompute replays a layer with gradients, so its mixer's
saved activations stay live beside that layer's MoE stage. The checkpoint
floor priced boundaries and the MoE stage only, which left context-parallel
runs short: Qwen3.6-35B-A3B at CP2 peaked 9-11 GB above the floor on the
most loaded rank. Price the larger of the model's attention and GDN mixers
per recomputed row, with context-parallel stage buffers and GDN exchange
copies, from allocator traces at CP1 and CP2.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Price GDN from its saved tensors (norm output, q/k with fp32 l2norm copies,
v, z, segment-layout tensors, gated norm and the chunk decay matrix) instead
of a ratio fit, and its context-parallel exchanges from hidden and value
widths rather than the key width. Divide CP attention extras by TP like the
retained widths, and say CP above 2 reuses the CP2 allowance.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The l2-normalized q and k are expanded to the value heads before they are
saved, so their width follows value_heads * key_head_dim, not twice the key
width. Qwen3.6 is unchanged; geometries with more value than key heads were
under-priced.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The EP1 all-to-all holds its permuted copy and the exchanged rows at the
expert stage; HybridEP permutes while it dispatches and returns one tensor.
A Qwen3.6 CP2/EP2 allocator trace holds exactly one routed H-wide input
beside the FC1 and FC2 stage tensors (9,728 features per routed row), where
the planner charged two (11,776).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…pendent

Backward recomputes the last layer first, so the checkpoint floor's peak meets
every saved boundary but only the one incoming gradient. Where the MoE stage
is priced, charge that gradient instead of one per boundary (39 hidden rows
per token too many at 40 layers), and price what the old allowance was
silently covering, all from Qwen3.6-35B-A3B allocator traces:

- the recomputed layer's residual and pre-MLP norm output (2H per row);
- GDN's sixth value-width tensor (the projected q/k/v includes v);
- the shared expert's saved FC1 gate/up and GLU outputs;
- router scores and map plus the dispatcher's row-id map (EP1) or probability
  copy and handle (HybridEP);
- TE's cuBLAS workspaces, as growth until its GEMMs allocate them.

Without a priced MoE stage the per-boundary allowance stays: it also covers
dense MLP and other recompute work the floor does not price.

The EP>1 routed-row allowance becomes EP-dependent (1.4, 1.6, 2.0 at EP2, 4,
8), from pretrained Qwen3.6 routing of 3.5M tokens of retail agent
trajectories (worst layer 1.21, 1.41, 1.63) and one production EP2 run (1.35).
Routed rows are no longer rounded up to whole rows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton force-pushed the dalinar/recompute-mixer-floor branch from 2b54d0c to c59a692 Compare September 25, 2026 18:50
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 18:51 — with GitHub Actions Failure
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 25, 2026 19:03 — with GitHub Actions Active
HybridEP dispatches the whole EP group's rows. When that group is this
rank's CP group, a balanced rank receives the group's rows over EP, not
the busiest CP rank's share: a CP2/EP2 real-data trace put 52,480 rows on
one rank while each layer dispatched exactly 8 x 96,794 pairs across both.
Price only the routed part (and its converted stages) on that share; the
shared expert, mixer and boundaries stay on local rows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 20:12 — with GitHub Actions Failure
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 20:14 — with GitHub Actions Failure
Charge only the one incoming gradient when every decoder layer is a priced
MoE layer that encloses its FC1 stage, for each gradient group's slot. A
positive FC2-only coefficient, dense layers or a slot that reprices to zero
keep one gradient per boundary, which also covers unpriced recompute work.

Count an empty CP rank's padding row in the EP group's total: dispatch runs
at least one row per rank, and that row is routed too.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 25, 2026 20:48 — with GitHub Actions Active
A layer counts as enclosed only if its FC1 converted stages are priced too,
unless FC1 has no adapter or the selected slot has no FC1 tensors. A slot
with FC1 adapters but no FC2 adapter prices FC2 rows from the original
metadata yet skips the whole converted-stage block, so it now keeps one
gradient per boundary. A slot's walk must enclose as many layers as the
constructor's, which already matched every decoder layer.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 21:08 — with GitHub Actions Error
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Main added #988's TP x SP floor, #971's HybridEP combine extent, #991's warm
profile, #992 and #998 since this branch was cut. _checkpoint_memory_floor
keeps #988's traced pricing at TP > 1 (the recomputed mixer is not validated
there), prices TP 1 with this branch's mixer floor, and applies #971's
combine-extent floor after either. Tests that passed routed rows positionally
use the keyword; #971's fixture asserts this branch's coefficient.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton added a commit that referenced this pull request Sep 26, 2026
Picks up main (0e0c31b) through the updated #963 and #978 branches.
_checkpoint_floor_decoder stays TP1 by default, so the dense widths and
their discounts remain TP1/CP2 only; with sequence_parallel it applies main's
#988 checks for the traced TP x SP floor. _checkpoint_memory_floor prices
TP > 1 with #988's floor, then the layout and generic floors, and #971's
combine-extent floor after all three.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton added a commit that referenced this pull request Sep 26, 2026
Picks up main (0e0c31b) through the updated #963 and #978 branches. Clean
merge; the source equals the composition measured in M9 plus #998.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 26, 2026 23:14 — with GitHub Actions Active
Review of the rebased stack: the MoE one-gradient allowance was traced at
CP1 and CP2 only, while above CP2 a rank runs more remote attention stages
than the mixer's CP2 allowance prices, so larger CP keeps one gradient per
boundary. #971's combine-extent floor now includes the TE cuBLAS workspaces
that are live beside it, instead of taking the maximum after they were added
to the stage.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton and others added 13 commits September 27, 2026 02:04
A full-recompute backward allocates each layer's adapter gradients as it
passes, and the step's optimizer frees them. Recomputing layer i still holds
the saved boundaries of layers 0..i, so a short first wave peaks at layer 0
with nearly every layer's gradients live; the checkpoint floor priced only
the last layer's end (all boundaries). On Qwen3.6-35B-A3B CP2, 2k and 4k
token first waves were admitted 11.5% and 2.0% under their peaks.

The floor now adds the largest excess of pending gradients over released
boundaries across the real layers, while a slot's gradients are unallocated,
and an unprofiled wave's 64 MiB of first-execution transients.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Name gradient slots with sorted kind/name JSON instead of a hash.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ed-mixer floor

Resolves _subforward_cost/_estimate for the stack's one-layer gradient and
adds _checkpoint_layer_boundaries (uniform). Still to do: #978 per-rank,
per-layer layout boundaries; merge #1002's later commits; suites; reviews.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Uses _gradient_slots so base-model groups own no adapter gradients.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Autograd drains the last-forwarded group's chain before an earlier one's, and
separate backward calls may come in either order, so a short group's
gradients can peak beside another group's unreleased boundaries. Price each
gradient group against its own boundaries, with groups not yet run holding
theirs and groups already run holding their gradients, over every order.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Any set of the other groups can have run first, so the worst order adds
every other group whose gradients outweigh its boundaries to one group's own
walk. Exact for any number of groups, without walking permutations.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 28, 2026 09:13 — with GitHub Actions Active
bradhilton added a commit that referenced this pull request Sep 28, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton added a commit that referenced this pull request Sep 28, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton added a commit that referenced this pull request Sep 28, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
trainer-rank-gpu-validation — fd5ef3a1 Deployed Sep 28, 2026 by bradhilton via Run on 2x H200 #856
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant