Repository navigation
Fix/fora v14 recipe - #10
Merged
Merged
Conversation
The MR pipeline was adapted from the CT DINOv3-7B continuation recipe that failed in FORA. Bring over every fix the working FORA ViT-L run uses, on by default: - cosine prototypes on both heads (plain std-0.02 prototypes kept Sinkhorn targets flat) - DINOv2 ViT-L recipe: lr 1e-3 @1024 sqrt-scaled, warmup 2.5k, teacher temp warmup 10k, wd 0.04->0.4, EMA 0.992, beta2 0.999, clip 3, layer-wise lr decay 0.9, patch-embed lr x0.2, prototype layers frozen for 1,250 steps (GroupedAdamW param groups) - ViT-L default (16k/16k prototypes, 2048/256 heads), batch 8/GPU, FP8 and activation checkpointing off; 7B/H+ launchers kept as legacy - global crops >=25% overlap snapped to the patch lattice, locals inside one global - cross-crop twin-token patch objective against last-block position leakage (pairs only same-sequence globals by default, preserving the no cross-sequence patch target rule) - position-binned (2x2x2) Sinkhorn for patch targets - target entropy / top-1 / usage / KL / cross metrics in metrics.jsonl Tests: tests/test_fora_recipe.py (8 new); all 18 pass on CPU.
The frozen prototype layers get no gradient during the first --freeze-last-layer-steps steps, so AdamW holds no state for them and checkpoints from that phase lack their optimizer keys. The strict DCP resume then aborted (GPU smoke 941555). Allow exactly those keys to be missing; any other missing model/optimizer key still raises.
A DINO target shared by a T1 and a FLAIR crop rewards dropping contrast-specific findings, and in atlas space the cheapest shared signal is the crop position. Keep a small rate so MIL bags that mix sequences share one embedding space; 0 vs 0.25 to be settled by the validation MIL probe. Also fix the stale 75% global-overlap line in the README.
Coreg (default) keeps every sequence on the study's T1w center grid; atlas resamples into the 1 mm MNI box. The model input is 1.0x0.5x0.5 mm, so from atlas the 0.5 mm in-plane detail is interpolation, and the MNI box crops skull base and neck (-12% head volume on a checked study). Both spaces are rigidly aligned across sequences; native_space is rejected. Checkpoints record their space and resume refuses to cross spaces (older checkpoints are atlas-space). mr_dino.probe scores a checkpoint like FORA CT-DINO: export the EMA teacher, extract frozen last-block tokens of every labelled val study (phase-1-size tiles, 2x2x2 foreground pooling, fp16 shards on node-local /tmp), patient-level 5-fold CV of MR-RATE's ClassifyThenAggregate head on the 14 merged groups, out-of-fold macro/per-class AUROC and AUPRC. Launchers: probe_node.sh (one node or inside a training allocation), probe_checkpoint.sbatch, probe_watch.sh (probes listed steps as checkpoints appear). The GPU smoke now also runs the probe on its own checkpoint.
At 4 nodes (global batch 128) 30k steps is ~6 passes over the ~635k training sequences; the same ViT-L recipe on 4 nodes peaked at ~23.5k steps in FORA. 125k (~25 epochs at 4 nodes, ~200 at 32) was far past that. The cosine decay ends at STEPS, so the default is the real run length; the probe confirms it.
~10 passes over the training sequences at 4 nodes (batch 128).
- train_mrdino.sbatch (replaces train_32n_vitl): ViT-H+ 0.84B default, 4 nodes. Batch/checkpointing from a 1-node H+ benchmark at real crop sizes: phase-1 crops batch 8 OOM without checkpointing, 0.39 steps/s at 32 GiB with it; highres crops batch 4 + checkpointing 0.10 steps/s at 81 GiB. - Phases 2/3 restart the schedule, so they get FORA CT-DINO's tested settings (fixed teacher temperature 0.07, EMA 0.996, short warmup, no prototype freeze; gram lr 5e-4, highres lr 4e-4 batch 4 with 4 locals) instead of re-running the phase-1 warmups. - A requeued phase-2/3 job continued from the previous phase (RESUME stays in its environment); its own complete checkpoint now wins. Phases 2/3 refuse to start without RESUME. - GPU smoke also runs pretrain -> gram -> highres; README describes the full training (model, phase table, memory/speed, time estimates, commands, rules).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.