Skip to content

Fix/fora v14 recipe - #10

Merged
sezginerr merged 7 commits into
mainfrom
fix/fora-v14-recipe
Oct 6, 2026
Merged

sezginerr merged 7 commits into
mainfrom
fix/fora-v14-recipe

Conversation

@sezginerr

Copy link
Copy Markdown
Member

No description provided.

The MR pipeline was adapted from the CT DINOv3-7B continuation recipe that failed in FORA.
Bring over every fix the working FORA ViT-L run uses, on by default:

- cosine prototypes on both heads (plain std-0.02 prototypes kept Sinkhorn targets flat)
- DINOv2 ViT-L recipe: lr 1e-3 @1024 sqrt-scaled, warmup 2.5k, teacher temp warmup 10k,
  wd 0.04->0.4, EMA 0.992, beta2 0.999, clip 3, layer-wise lr decay 0.9, patch-embed lr x0.2,
  prototype layers frozen for 1,250 steps (GroupedAdamW param groups)
- ViT-L default (16k/16k prototypes, 2048/256 heads), batch 8/GPU, FP8 and activation
  checkpointing off; 7B/H+ launchers kept as legacy
- global crops >=25% overlap snapped to the patch lattice, locals inside one global
- cross-crop twin-token patch objective against last-block position leakage (pairs only
  same-sequence globals by default, preserving the no cross-sequence patch target rule)
- position-binned (2x2x2) Sinkhorn for patch targets
- target entropy / top-1 / usage / KL / cross metrics in metrics.jsonl

Tests: tests/test_fora_recipe.py (8 new); all 18 pass on CPU.
The frozen prototype layers get no gradient during the first
--freeze-last-layer-steps steps, so AdamW holds no state for them and
checkpoints from that phase lack their optimizer keys. The strict DCP
resume then aborted (GPU smoke 941555). Allow exactly those keys to be
missing; any other missing model/optimizer key still raises.
A DINO target shared by a T1 and a FLAIR crop rewards dropping contrast-specific findings,
and in atlas space the cheapest shared signal is the crop position. Keep a small rate so
MIL bags that mix sequences share one embedding space; 0 vs 0.25 to be settled by the
validation MIL probe. Also fix the stale 75% global-overlap line in the README.
Coreg (default) keeps every sequence on the study's T1w center grid; atlas
resamples into the 1 mm MNI box. The model input is 1.0x0.5x0.5 mm, so from
atlas the 0.5 mm in-plane detail is interpolation, and the MNI box crops skull
base and neck (-12% head volume on a checked study). Both spaces are rigidly
aligned across sequences; native_space is rejected. Checkpoints record their
space and resume refuses to cross spaces (older checkpoints are atlas-space).

mr_dino.probe scores a checkpoint like FORA CT-DINO: export the EMA teacher,
extract frozen last-block tokens of every labelled val study (phase-1-size
tiles, 2x2x2 foreground pooling, fp16 shards on node-local /tmp), patient-level
5-fold CV of MR-RATE's ClassifyThenAggregate head on the 14 merged groups,
out-of-fold macro/per-class AUROC and AUPRC. Launchers: probe_node.sh (one
node or inside a training allocation), probe_checkpoint.sbatch, probe_watch.sh
(probes listed steps as checkpoints appear). The GPU smoke now also runs the
probe on its own checkpoint.
At 4 nodes (global batch 128) 30k steps is ~6 passes over the ~635k training
sequences; the same ViT-L recipe on 4 nodes peaked at ~23.5k steps in FORA.
125k (~25 epochs at 4 nodes, ~200 at 32) was far past that. The cosine decay
ends at STEPS, so the default is the real run length; the probe confirms it.
~10 passes over the training sequences at 4 nodes (batch 128).
- train_mrdino.sbatch (replaces train_32n_vitl): ViT-H+ 0.84B default, 4 nodes.
  Batch/checkpointing from a 1-node H+ benchmark at real crop sizes: phase-1
  crops batch 8 OOM without checkpointing, 0.39 steps/s at 32 GiB with it;
  highres crops batch 4 + checkpointing 0.10 steps/s at 81 GiB.
- Phases 2/3 restart the schedule, so they get FORA CT-DINO's tested settings
  (fixed teacher temperature 0.07, EMA 0.996, short warmup, no prototype
  freeze; gram lr 5e-4, highres lr 4e-4 batch 4 with 4 locals) instead of
  re-running the phase-1 warmups.
- A requeued phase-2/3 job continued from the previous phase (RESUME stays in
  its environment); its own complete checkpoint now wins. Phases 2/3 refuse
  to start without RESUME.
- GPU smoke also runs pretrain -> gram -> highres; README describes the full
  training (model, phase table, memory/speed, time estimates, commands, rules).
@sezginerr
sezginerr merged commit d379859 into main Oct 6, 2026
3 of 4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant