You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
September 25, 2026 — #888 is merged; general physical-headroom safety remains open.
PR #888 merged as e2e8969ff1f64151ae613abf0050e236ee77e35e, alongside the separate #898 accounting correction. Earlier dated held statements below are historical. The shared recovery budget does not guarantee total training slowdown or enough memory for every backward.
The follow-up is now draft PR #972: one optional budgeted cache-recovery attempt before accepting an internal split, followed by normal fresh admission. Three independent source reviews cleared f93b6d85338b62c52c6e0e2e66b95ceddd8b7d24; significant packing/admission behavior remains held for Brad. Quality CI found type-narrowing/test-annotation errors, with an isolated correction underway. Required hosted GPU CI failed during acquisition of two H200s before tests. A separate exact-head local five-command GPU suite is running; it does not waive the hosted gate.
An independently audited six-forward/six-backward model diagnostic on the predecessor with identical ordinary-path runtime completed without optimizer updates. It made two actual cache releases, freeing approximately 30 GiB each; release calls took 0.178/0.261 seconds and full recovery episodes 0.438/0.521 seconds. The second release passed the existing budget, but 93.4% of its work denominator was warmup/compile-associated. This is not steady-state overhead qualification, a throughput comparison, or numerical equivalence. The control stopped at its strict layout guard when admission selected a split. Audit receipt: /home/brad/.local/share/schulman/recovery-candidate-native-independent-audit-20260925/RECEIPT.json (188307fd). The final-head difference clears a transient plan on the separate oversized-override return; it has focused CPU lifetime coverage.
A two-GPU real-NCCL fixture independently passed four injected control-flow cases, including an empty peer, peer veto and local release error. It used synthetic memory accounting and release shims; it does not qualify distributed real allocator behavior. Receipt: /home/brad/.local/share/schulman/krennic-packed-boundary-20260925/recovery-nccl-native-audit-v1/REVIEW-RECEIPT.json (c4801cb6). Both completed diagnostics are physically closed. No new whole-model safety claim or issue closure follows.
Earlier dated history is preserved below.
September 23, 2026 — current recovery candidate qualified in additional bounded cases; still open and held.
Current #888 is 3b0c6d4ef8972aeba0441ec89756b70c8f227624. Three source reviews and quality/two-H200 CI are clear. It shares the measured recovery budget with gradient handoff and credits completed backward work. The 3% physical reserve is a soft trigger, and the 5% ledger does not guarantee total training slowdown or sufficient memory for every backward. Fable's additional review holds the policy pending explicit acceptance or useful targeted timing evidence: later allocator regrowth is not separately charged. No current-head merge is authorized by these results.
Earlier missing current-source checks are now complete at retained source 67401664: a positive-pressure release followed by backward/update, and grouped save with exact saved-versus-fresh adapter/master/optimizer/custom state followed by another update. The complete eight-pair/flat16 workload also passed on combined f80dba32. The new 053-shaped case passed one joint update on one H200, but retains manual caller cache release and is therefore not automatic-policy qualification.
The current-source same-H200 ON/OFF pair is complete: both arms finished six warmup and six measured updates, with original tasks/managed waits 0 and all 12 distinct known native identities/eight groups closed. ON's first failed close check and subsequent natural retirement are preserved. The unchanged primary comparator is INCONCLUSIVE because warmup call/partition records, measured partitions and starting reserved-cache bytes differed. A separate, predeclared full-work observation is 518.711 seconds enabled versus 504.801 disabled: +13.910 seconds (+2.756%). Both completed matching logical requests, output shapes and updates from the same saved initial state, but measured-start parameters differed after normal-policy warmup. ON recorded one proactive release adding 27,881,635,840 bytes of physical free memory; OFF skipped 35 proactive calls and had one observed allocator retry. These facts do not isolate regrowth cost or establish a universal five-percent slowdown bound. Independent retained-result review supports this scoped interpretation; no extra benchmark, policy adoption or merge follows. Final compact evidence: 888-current3b0-onoff-result-20260923/ under the receipt directory below.
Receipt directory: /home/brad/.local/share/schulman/art848-resume-20260921-root/. Completed earlier checks: 888-current674-pressure-final-root-acceptance.json (ce5715e5), 888-current674-grouped-save-fresh-load-root-acceptance.json (7bed2fd1), and combined-full8-native-v6-final-root-acceptance.json (47cb19ea). Current ON result: /var/tmp/art888-local-onoff-ON-20260923-halley-v1/actor-result.json (a27dafca); additive known-process closure cbf074b6. General backward safety, representative performance and repeated policy behavior remain broader than these bounded checks.
Earlier dated history follows; its old heads and pending-work statements are superseded above.
September 22, 2026 update — still open; #888 remains draft/held.
Current #888 is 5a4ea977552fa9b337ca65d8e3fabfa22fd1e887; exact-head quality and hosted two-H200 CI passed. It uses the shared budget for gradient-handoff cache release and credits completed ordinary backward work through a bounded observer. The 3% physical reserve is a soft trigger; the 5% budget is not a guarantee of total training slowdown or enough memory for arbitrary backwards.
The frozen d129f724 candidate completed six warmups plus six measured optimizer updates: 803.215 seconds warmup and 511.498 seconds measured work. Its matched main attempt had stopped before measured updates, so no baseline ratio follows. The corrected compile-scope reader still reports UNKNOWN due to unscoped warmup events; no compilation events were observed during the measured interval, which does not prove a fully stable compiled workload. All resources from these runs are closed.
The current public head passed fresh graph and installed-CPU qualification for grouped save/fresh load. Its GPU attempt failed before cloning or trainer creation: the private guard required raw complete=True, whereas grouped durability is established through the committed group authority. The original attempt did not retain donor metadata, so its exact failing predicate remains unknown. Native exit1 and complete known-resource cleanup are preserved; no OOM, training update or save result is inferred. A narrow guard correction and its standalone/prior-group cases are under review. Separate matched current-main/current-candidate performance preparation is underway with identical eight-pair normalization and two four-pair gradient calls.
The earlier successful grouped save/load and positive/repeat-pressure episodes remain historical-source evidence, not qualification of the new backward observer. Still needed: current-source grouped save/fresh reload, matched full-model timing, and scoped positive/repeated budget behavior. General #870 remains broader than these tests. Durable current receipts: art848-resume-20260921-root/888-final-d129-candidate-only-root-success-closure.json (560743bf), 888-H1-grouped-native-failed-guard-root-closure.json (4fe8a10d) and 888-main45-H1-four-plus-four-graph-root-release.json (b6b38c11). No merge, deployment or tolerance change.
Earlier dated history is retained below; this update supersedes old head and pending-run statements.
Current status — September 21, 2026, 6:35 AM Mountain: unresolved; PR #888 remains held.
Current draft Share budgeted CUDA cache recovery with gradient handoff #888 is b90678708bdc22672f5a4ab66a6f23c4a5c5e641. Its budgeted release policy is not merged. The current main/candidate CPU image prerequisites passed; representative warmed throughput, grouped save and fresh reload remain unqualified.
The first fresh same-H200 pair attempt filled its fixed model cache, then stopped before GPU keeper allocation: the Kubernetes all-pods response exceeded the runner's artificial 2 MiB raw-response cap. The assertion preceded body/wait persistence, so the second response's exact size and OS status are unknown. No measured epoch or save/reload occurred. All cache/fill resource UIDs and known host groups are independently/root closed; no live cache adoption is implied.
A narrow private runner correction preserves the completed query's wait/count/hash evidence before refusal and permits raw combined output below 64 MiB; the compact summary retains its 2 MiB limit. The same two GETs/timeouts/selection remain. Source review and four old-red/boundary/error CPU controls pass. Fresh final composition precedes another finite attempt; no production behavior change is bundled with this runner fix.
A separate positive-pressure diagnostic is being integrated for a real release opportunity. It will remain distinct from the normal warmed pair and cannot substitute for total-overhead measurement. No new native pressure result is claimed.
Owner: Schulman and existing delegates. #848 estimator/admission and #901/#902 numerical questions remain separate. Actual host cache stays counted. Durable report: /home/brad/.local/share/schulman/art848-resume-20260921-root/PROGRESS.md.
Earlier status/history below is preserved; current statements above supersede old heads and pending-run descriptions.
Current status — September 18, 2026, late evening Mountain: partially addressed; still open.
PR #900 is merged. Current held PR #888 is at f557bc4e60f7b38bead1cc0efd7d655e2bc7df69; its low-physical-headroom backward handoff now uses the same measured cache-release budget, rather than an unconditional release rule. Three exact-head source reviews and CI, including two-H200 checks, are clear. Significant behavioral adoption still needs Brad's decision.
A full-model candidate run completed three real optimizer updates (18 backwards), followed by a separate pressure diagnostic with one automatic handoff release and one successful backward, with no optimizer update during that pressure diagnostic. It then successfully published the grouped generator/auxiliary checkpoint. The original run nevertheless exited 1: after trainer close returned, its actor-absence check timed out before fresh reload began. This failure is retained and is not classified as a backend or zombie-process cause. Exact owned resources were independently retired afterward.
A distinct load-only continuation is now running on one free H200: art848ws-07a21878, launched at 11:04:39 PM Mountain. It uses protocol-8 Caladan with the original candidate ART source, loads the exact saved generator:v1/auxiliary:v1 group from art848ws-b140393f, compares complete canonical optimizer/custom state, then attempts exactly one update and real cleanup. Prior training, pressure and remote grouped publication are not repeated. Its actual worker-image and restored-helper CPU checks passed; native saved-state equality and successful completion are not yet claimed.
The observed release cost is not a total-training 5% overhead bound. Repeat-policy instrumentation itself contributes to measured work/cost, and process-wide cache effects remain relevant. A separate source proposal can measure two pressure episodes and intervening updates, but is not yet run and cannot by itself establish representative throughput or arbitrary backward safety. Keep this issue open separately from #848 and numerical variability #902.
Evidence: /var/tmp/art888-load-only-candidate-image-20260919-root/cpu-root-acceptance.json (041f584e), exact native authority gpu-root-admission-v2.json (d4fb5fd6) and gpu-live-identities.json (e16daabc). Original failed producer evidence remains /var/tmp/art888-native-sequence-20260919-backward-n2_crwep/operation/. No merge or deployment occurred in this follow-up.
Earlier status/history, preserved; current statements above supersede old head and pending-candidate descriptions:
Current status — September 17, 2026: partially addressed; still open.
Brad approved PR #900, merged as c009557e2f2488a83551d29c3253b708edef322c. Admission now uses physical free CUDA memory and can release unused allocator cache under a measured cost budget. The scoped two-H200 qualification and source/CI checks passed before merge. This does not establish full-training overhead, sufficient memory throughout arbitrary backwards, or adoption by every Caladan experiment/image.
The earlier PR #888 is a separate, still-held proposal to release cache at a low-headroom handoff to backward. It is not adopted by #900, and its process-wide effects and throughput tradeoffs remain relevant. It should not be merged or relabeled as superseded without reassessing that remaining boundary against the current admission policy.
Active work with #848: qualify the current code on the retained workload and resolve underestimated backward demand. Keep the library-allocation problem separate from the repeated-gradient variability tracked by #902. All recent completed diagnostic resources are cleaned up. Representative end-to-end overhead and broader backward-memory safety remain unqualified.
The historical September 12 status below is preserved; its description of the then-current candidate is superseded by this update.
September 12 update: the earlier statement that conditional/full-step and multi-device behavior are entirely unqualified is superseded by bounded native evidence. Three complete updates on one H200 exercised four actual conditional releases; a separate two-H200 run verified that selected-device release also trims the other device's unused cache while preserving the checked live values/gradients and caller device/streams. All resources were independently retired. Exact retained evidence and limits: /home/brad/.local/share/schulman/art870-two-device-gap-inventory-20260912-capture/REPORT.md.
ART #888 remains OPEN/DRAFT at1763a46ac9b4f465e38440baf8cc3e7922da69a3, source-correctness reviews and CI green, adoption HELD. The 3% trigger is best effort, not a guaranteed reserve. Process-wide cache retirement/waits are significant behavior requiring Brad's decision. Representative throughput/refill costs, arbitrary concurrent users, and broader distributed/negative cases remain unqualified. Do not repeat the completed two-device witness merely to restate the already observed process-wide effect.
Earlier record (preserved; status above is current):
Status reconciliation — September 11, 2026 (Schulman)
Confirmed allocation failure; local mitigation only. Unconditional cache release let the pinned first-backward diagnostic finish, but the proposed conditional trigger, complete optimizer behavior and multi-device side effects remain unqualified. torch.cuda.empty_cache can trim caches across devices in the process; a selected-device context does not isolate that effect.
Planned lane: Schulman and subagents, queued with #848/#256. Complete the conditional/multi-device and full-step qualification before proposing a generic fix. The held local patch is not merged, does not modify art.megatron, and is not covered by a no-significant-behavior assumption.
Historical report (preserved):
Type: confirmed backward allocation failure with a controlled experiment-local mitigation; generic policy remains unresolved.
Agent owner: Schulman, related to #848.
A rank-1 Qwen3.6-35B-A3B eight-pair dynamics batch on one H200 passed its forwards, then failed in micro_batch_loss.backward with CUBLAS_STATUS_ALLOC_FAILED from cublasCreate(handle). GeneratorExit was secondary iterator teardown. The exact internal cuBLAS request size was not measured.
At the first gradient admission, live allocated/reserved memory was about 74.662/147.394 decimal GB with only 0.034734 GB physically free. The planner treated reusable allocator slack as available; its 3% capacity subtraction did not physically reserve CUDA memory for external library allocations. At the caller boundary, an unconditional torch.cuda.empty_cache() released 42.754638 GB without changing live allocation. The matched release-only replay completed all five backward waves and the optimizer; the observe-only replay failed. Pre-intervention plans/budgets and corresponding forward counters matched exactly.
The qualified mitigation is experiment-local cache release at the start of every gradient caller, after receiving its microbatch. It is not equivalent to a conditional 3% policy, a pre-yield trainer hook, or handle warmup. Later nonzero-generator-gradient updates also succeeded, but larger traces and full validation remain unqualified. Cold forward estimate underprediction is a separate problem.
Investigate a bounded generic physical-headroom/library-allocation contract with mixed-grad plans, device/stream behavior and normal downstream workloads. Prefer trainer_rank-local work; no art.megatron change is proposed. Do not silently globalize the workaround based on one fixture.
Evidence: /home/brad/.local/share/schulman/retail49-memory-failure-20260909/MORNING-ALLOCATOR.md and replay-{observe,release}-comparison/. Local Caladan mitigation 431e424 and timing-only follow-up 177a9fce. Raw failures and cleanup receipts are preserved.
September 25, 2026 — #888 is merged; general physical-headroom safety remains open.
PR #888 merged as
e2e8969ff1f64151ae613abf0050e236ee77e35e, alongside the separate #898 accounting correction. Earlier dated held statements below are historical. The shared recovery budget does not guarantee total training slowdown or enough memory for every backward.The follow-up is now draft PR #972: one optional budgeted cache-recovery attempt before accepting an internal split, followed by normal fresh admission. Three independent source reviews cleared
f93b6d85338b62c52c6e0e2e66b95ceddd8b7d24; significant packing/admission behavior remains held for Brad. Quality CI found type-narrowing/test-annotation errors, with an isolated correction underway. Required hosted GPU CI failed during acquisition of two H200s before tests. A separate exact-head local five-command GPU suite is running; it does not waive the hosted gate.An independently audited six-forward/six-backward model diagnostic on the predecessor with identical ordinary-path runtime completed without optimizer updates. It made two actual cache releases, freeing approximately 30 GiB each; release calls took 0.178/0.261 seconds and full recovery episodes 0.438/0.521 seconds. The second release passed the existing budget, but 93.4% of its work denominator was warmup/compile-associated. This is not steady-state overhead qualification, a throughput comparison, or numerical equivalence. The control stopped at its strict layout guard when admission selected a split. Audit receipt:
/home/brad/.local/share/schulman/recovery-candidate-native-independent-audit-20260925/RECEIPT.json(188307fd). The final-head difference clears a transient plan on the separate oversized-override return; it has focused CPU lifetime coverage.A two-GPU real-NCCL fixture independently passed four injected control-flow cases, including an empty peer, peer veto and local release error. It used synthetic memory accounting and release shims; it does not qualify distributed real allocator behavior. Receipt:
/home/brad/.local/share/schulman/krennic-packed-boundary-20260925/recovery-nccl-native-audit-v1/REVIEW-RECEIPT.json(c4801cb6). Both completed diagnostics are physically closed. No new whole-model safety claim or issue closure follows.Earlier dated history is preserved below.
September 23, 2026 — current recovery candidate qualified in additional bounded cases; still open and held.
Current #888 is
3b0c6d4ef8972aeba0441ec89756b70c8f227624. Three source reviews and quality/two-H200 CI are clear. It shares the measured recovery budget with gradient handoff and credits completed backward work. The 3% physical reserve is a soft trigger, and the 5% ledger does not guarantee total training slowdown or sufficient memory for every backward. Fable's additional review holds the policy pending explicit acceptance or useful targeted timing evidence: later allocator regrowth is not separately charged. No current-head merge is authorized by these results.Earlier missing current-source checks are now complete at retained source
67401664: a positive-pressure release followed by backward/update, and grouped save with exact saved-versus-fresh adapter/master/optimizer/custom state followed by another update. The complete eight-pair/flat16 workload also passed on combinedf80dba32. The new 053-shaped case passed one joint update on one H200, but retains manual caller cache release and is therefore not automatic-policy qualification.The current-source same-H200 ON/OFF pair is complete: both arms finished six warmup and six measured updates, with original tasks/managed waits 0 and all 12 distinct known native identities/eight groups closed. ON's first failed close check and subsequent natural retirement are preserved. The unchanged primary comparator is INCONCLUSIVE because warmup call/partition records, measured partitions and starting reserved-cache bytes differed. A separate, predeclared full-work observation is 518.711 seconds enabled versus 504.801 disabled: +13.910 seconds (+2.756%). Both completed matching logical requests, output shapes and updates from the same saved initial state, but measured-start parameters differed after normal-policy warmup. ON recorded one proactive release adding 27,881,635,840 bytes of physical free memory; OFF skipped 35 proactive calls and had one observed allocator retry. These facts do not isolate regrowth cost or establish a universal five-percent slowdown bound. Independent retained-result review supports this scoped interpretation; no extra benchmark, policy adoption or merge follows. Final compact evidence:
888-current3b0-onoff-result-20260923/under the receipt directory below.Receipt directory:
/home/brad/.local/share/schulman/art848-resume-20260921-root/. Completed earlier checks:888-current674-pressure-final-root-acceptance.json(ce5715e5),888-current674-grouped-save-fresh-load-root-acceptance.json(7bed2fd1), andcombined-full8-native-v6-final-root-acceptance.json(47cb19ea). Current ON result:/var/tmp/art888-local-onoff-ON-20260923-halley-v1/actor-result.json(a27dafca); additive known-process closurecbf074b6. General backward safety, representative performance and repeated policy behavior remain broader than these bounded checks.Earlier dated history follows; its old heads and pending-work statements are superseded above.
September 22, 2026 update — still open; #888 remains draft/held.
Current #888 is
5a4ea977552fa9b337ca65d8e3fabfa22fd1e887; exact-head quality and hosted two-H200 CI passed. It uses the shared budget for gradient-handoff cache release and credits completed ordinary backward work through a bounded observer. The 3% physical reserve is a soft trigger; the 5% budget is not a guarantee of total training slowdown or enough memory for arbitrary backwards.The frozen
d129f724candidate completed six warmups plus six measured optimizer updates: 803.215 seconds warmup and 511.498 seconds measured work. Its matched main attempt had stopped before measured updates, so no baseline ratio follows. The corrected compile-scope reader still reports UNKNOWN due to unscoped warmup events; no compilation events were observed during the measured interval, which does not prove a fully stable compiled workload. All resources from these runs are closed.The current public head passed fresh graph and installed-CPU qualification for grouped save/fresh load. Its GPU attempt failed before cloning or trainer creation: the private guard required raw
complete=True, whereas grouped durability is established through the committed group authority. The original attempt did not retain donor metadata, so its exact failing predicate remains unknown. Native exit1 and complete known-resource cleanup are preserved; no OOM, training update or save result is inferred. A narrow guard correction and its standalone/prior-group cases are under review. Separate matched current-main/current-candidate performance preparation is underway with identical eight-pair normalization and two four-pair gradient calls.The earlier successful grouped save/load and positive/repeat-pressure episodes remain historical-source evidence, not qualification of the new backward observer. Still needed: current-source grouped save/fresh reload, matched full-model timing, and scoped positive/repeated budget behavior. General #870 remains broader than these tests. Durable current receipts:
art848-resume-20260921-root/888-final-d129-candidate-only-root-success-closure.json(560743bf),888-H1-grouped-native-failed-guard-root-closure.json(4fe8a10d) and888-main45-H1-four-plus-four-graph-root-release.json(b6b38c11). No merge, deployment or tolerance change.Earlier dated history is retained below; this update supersedes old head and pending-run statements.
Current status — September 21, 2026, 6:35 AM Mountain: unresolved; PR #888 remains held.
b90678708bdc22672f5a4ab66a6f23c4a5c5e641. Its budgeted release policy is not merged. The current main/candidate CPU image prerequisites passed; representative warmed throughput, grouped save and fresh reload remain unqualified.Owner: Schulman and existing delegates. #848 estimator/admission and #901/#902 numerical questions remain separate. Actual host cache stays counted. Durable report:
/home/brad/.local/share/schulman/art848-resume-20260921-root/PROGRESS.md.Earlier status/history below is preserved; current statements above supersede old heads and pending-run descriptions.
Current status — September 18, 2026, late evening Mountain: partially addressed; still open.
PR #900 is merged. Current held PR #888 is at
f557bc4e60f7b38bead1cc0efd7d655e2bc7df69; its low-physical-headroom backward handoff now uses the same measured cache-release budget, rather than an unconditional release rule. Three exact-head source reviews and CI, including two-H200 checks, are clear. Significant behavioral adoption still needs Brad's decision.A full-model candidate run completed three real optimizer updates (18 backwards), followed by a separate pressure diagnostic with one automatic handoff release and one successful backward, with no optimizer update during that pressure diagnostic. It then successfully published the grouped generator/auxiliary checkpoint. The original run nevertheless exited 1: after trainer close returned, its actor-absence check timed out before fresh reload began. This failure is retained and is not classified as a backend or zombie-process cause. Exact owned resources were independently retired afterward.
A distinct load-only continuation is now running on one free H200:
art848ws-07a21878, launched at 11:04:39 PM Mountain. It uses protocol-8 Caladan with the original candidate ART source, loads the exact saved generator:v1/auxiliary:v1 group fromart848ws-b140393f, compares complete canonical optimizer/custom state, then attempts exactly one update and real cleanup. Prior training, pressure and remote grouped publication are not repeated. Its actual worker-image and restored-helper CPU checks passed; native saved-state equality and successful completion are not yet claimed.The observed release cost is not a total-training 5% overhead bound. Repeat-policy instrumentation itself contributes to measured work/cost, and process-wide cache effects remain relevant. A separate source proposal can measure two pressure episodes and intervening updates, but is not yet run and cannot by itself establish representative throughput or arbitrary backward safety. Keep this issue open separately from #848 and numerical variability #902.
Evidence:
/var/tmp/art888-load-only-candidate-image-20260919-root/cpu-root-acceptance.json(041f584e), exact native authoritygpu-root-admission-v2.json(d4fb5fd6) andgpu-live-identities.json(e16daabc). Original failed producer evidence remains/var/tmp/art888-native-sequence-20260919-backward-n2_crwep/operation/. No merge or deployment occurred in this follow-up.Earlier status/history, preserved; current statements above supersede old head and pending-candidate descriptions:
Current status — September 17, 2026: partially addressed; still open.
Brad approved PR #900, merged as
c009557e2f2488a83551d29c3253b708edef322c. Admission now uses physical free CUDA memory and can release unused allocator cache under a measured cost budget. The scoped two-H200 qualification and source/CI checks passed before merge. This does not establish full-training overhead, sufficient memory throughout arbitrary backwards, or adoption by every Caladan experiment/image.The earlier PR #888 is a separate, still-held proposal to release cache at a low-headroom handoff to backward. It is not adopted by #900, and its process-wide effects and throughput tradeoffs remain relevant. It should not be merged or relabeled as superseded without reassessing that remaining boundary against the current admission policy.
Active work with #848: qualify the current code on the retained workload and resolve underestimated backward demand. Keep the library-allocation problem separate from the repeated-gradient variability tracked by #902. All recent completed diagnostic resources are cleaned up. Representative end-to-end overhead and broader backward-memory safety remain unqualified.
The historical September 12 status below is preserved; its description of the then-current candidate is superseded by this update.
September 12 update: the earlier statement that conditional/full-step and multi-device behavior are entirely unqualified is superseded by bounded native evidence. Three complete updates on one H200 exercised four actual conditional releases; a separate two-H200 run verified that selected-device release also trims the other device's unused cache while preserving the checked live values/gradients and caller device/streams. All resources were independently retired. Exact retained evidence and limits: /home/brad/.local/share/schulman/art870-two-device-gap-inventory-20260912-capture/REPORT.md.
ART #888 remains OPEN/DRAFT at1763a46ac9b4f465e38440baf8cc3e7922da69a3, source-correctness reviews and CI green, adoption HELD. The 3% trigger is best effort, not a guaranteed reserve. Process-wide cache retirement/waits are significant behavior requiring Brad's decision. Representative throughput/refill costs, arbitrary concurrent users, and broader distributed/negative cases remain unqualified. Do not repeat the completed two-device witness merely to restate the already observed process-wide effect.
Earlier record (preserved; status above is current):
Status reconciliation — September 11, 2026 (Schulman)
Confirmed allocation failure; local mitigation only. Unconditional cache release let the pinned first-backward diagnostic finish, but the proposed conditional trigger, complete optimizer behavior and multi-device side effects remain unqualified. torch.cuda.empty_cache can trim caches across devices in the process; a selected-device context does not isolate that effect.
Planned lane: Schulman and subagents, queued with #848/#256. Complete the conditional/multi-device and full-step qualification before proposing a generic fix. The held local patch is not merged, does not modify art.megatron, and is not covered by a no-significant-behavior assumption.
Historical report (preserved):
Type: confirmed backward allocation failure with a controlled experiment-local mitigation; generic policy remains unresolved.
Agent owner: Schulman, related to #848.
A rank-1 Qwen3.6-35B-A3B eight-pair dynamics batch on one H200 passed its forwards, then failed in micro_batch_loss.backward with CUBLAS_STATUS_ALLOC_FAILED from cublasCreate(handle). GeneratorExit was secondary iterator teardown. The exact internal cuBLAS request size was not measured.
At the first gradient admission, live allocated/reserved memory was about 74.662/147.394 decimal GB with only 0.034734 GB physically free. The planner treated reusable allocator slack as available; its 3% capacity subtraction did not physically reserve CUDA memory for external library allocations. At the caller boundary, an unconditional torch.cuda.empty_cache() released 42.754638 GB without changing live allocation. The matched release-only replay completed all five backward waves and the optimizer; the observe-only replay failed. Pre-intervention plans/budgets and corresponding forward counters matched exactly.
The qualified mitigation is experiment-local cache release at the start of every gradient caller, after receiving its microbatch. It is not equivalent to a conditional 3% policy, a pre-yield trainer hook, or handle warmup. Later nonzero-generator-gradient updates also succeeded, but larger traces and full validation remain unqualified. Cold forward estimate underprediction is a separate problem.
Investigate a bounded generic physical-headroom/library-allocation contract with mixed-grad plans, device/stream behavior and normal downstream workloads. Prefer trainer_rank-local work; no art.megatron change is proposed. Do not silently globalize the workaround based on one fixture.
Evidence: /home/brad/.local/share/schulman/retail49-memory-failure-20260909/MORNING-ALLOCATOR.md and replay-{observe,release}-comparison/. Local Caladan mitigation 431e424 and timing-only follow-up 177a9fce. Raw failures and cleanup receipts are preserved.