-
Notifications
You must be signed in to change notification settings - Fork 1.6k
Pull requests: modelscope/ms-swift
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Add fused operator patch for Qwen3-VL dense models on NPU
#9987
opened Aug 26, 2026 by
addsubmuldiv
Collaborator
Loading…
1 of 4 tasks
fix: materialize state_dict for SentenceTransformer full-parameter save
#9986
opened Aug 25, 2026 by
Lossfull
Loading…
1 of 4 tasks
fix(npu): reuse fused AdamW group caches
#9984
opened Aug 25, 2026 by
addsubmuldiv
Collaborator
•
Draft
4 tasks
feat(fsdp2): load model on meta device for non-rank0 ranks (0 CPU RAM per worker)
#9982
opened Aug 24, 2026 by
cben484
Loading…
[megatron] Fix router replay CP partition and expert index overflow
#9981
opened Aug 24, 2026 by
taking-lying-flat
Contributor
Loading…
fix(fsdp): set ACCELERATE_USE_FSDP so device_map and cpu_ram_efficient_loading take effect
#9980
opened Aug 24, 2026 by
cben484
Loading…
3 of 7 tasks
Fix the compatibility of py3.14 and fix qwen3-vl training bug
#9979
opened Aug 24, 2026 by
tastelikefeet
Collaborator
Loading…
1 of 4 tasks
fix: align channel loss with sequence parallel labels
#9977
opened Aug 24, 2026 by
taking-lying-flat
Contributor
Loading…
1 of 4 tasks
Fix Megatron-FSDP DTensor checkpoint save/load compatibility
#9976
opened Aug 24, 2026 by
addsubmuldiv
Collaborator
•
Draft
1 of 4 tasks
Preserve Qwen3.5 linear attention FP32 weights
#9972
opened Aug 24, 2026 by
liuhao-labs
Loading…
1 of 4 tasks
perf(megatron): defer loss all-reduce to log time
#9966
opened Aug 22, 2026 by
gakkiri
Contributor
Loading…
1 task done
fix(train): fail fast on incompatible DLRover flash checkpoint APIs
#9964
opened Aug 21, 2026 by
Excelius-Wang
Contributor
Loading…
Fix colocated vLLM cleanup ordering on rollout failures
#9952
opened Aug 20, 2026 by
pureoxygen123
Loading…
Fix zero multimodal learning rates in Megatron
#9948
opened Aug 19, 2026 by
taking-lying-flat
Contributor
Loading…
Fix Megatron GRPO vocab-parallel log-prob gradients
#9947
opened Aug 19, 2026 by
taking-lying-flat
Contributor
Loading…
Add --dataloader_multiprocessing_context to work around Python 3.14 incompatibility
#9941
opened Aug 18, 2026 by
sliedes
Loading…
1 of 4 tasks
[Megatron] Preserve RNG state across checkpoint resume
#9935
opened Aug 17, 2026 by
taking-lying-flat
Contributor
Loading…
[Megatron] Fix reference adapter loading for LoRA RLHF
#9934
opened Aug 17, 2026 by
taking-lying-flat
Contributor
Loading…
Fix reward model margin broadcasting and alignment
#9927
opened Aug 16, 2026 by
taking-lying-flat
Contributor
Loading…
Fix DPO IPO log-prob normalization
#9925
opened Aug 16, 2026 by
taking-lying-flat
Contributor
Loading…
Fix BatchSamplerShard tail sampling
#9907
opened Aug 14, 2026 by
taking-lying-flat
Contributor
Loading…
[Train] Reduce padding-free embedding output memory
#9893
opened Aug 11, 2026 by
taking-lying-flat
Contributor
Loading…
Previous Next
ProTip!
no:milestone will show everything without a milestone.