Skip to content

Pull requests: modelscope/ms-swift

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

Add fused operator patch for Qwen3-VL dense models on NPU
#9987 opened Aug 26, 2026 by addsubmuldiv Collaborator Loading…
1 of 4 tasks
fix: materialize state_dict for SentenceTransformer full-parameter save
#9986 opened Aug 25, 2026 by Lossfull Loading…
1 of 4 tasks
fix(npu): reuse fused AdamW group caches
#9984 opened Aug 25, 2026 by addsubmuldiv Collaborator Draft
4 tasks
Fix the compatibility of py3.14 and fix qwen3-vl training bug
#9979 opened Aug 24, 2026 by tastelikefeet Collaborator Loading…
1 of 4 tasks
fix: align channel loss with sequence parallel labels
#9977 opened Aug 24, 2026 by taking-lying-flat Contributor Loading…
1 of 4 tasks
Fix Megatron-FSDP DTensor checkpoint save/load compatibility
#9976 opened Aug 24, 2026 by addsubmuldiv Collaborator Draft
1 of 4 tasks
Preserve Qwen3.5 linear attention FP32 weights
#9972 opened Aug 24, 2026 by liuhao-labs Loading…
1 of 4 tasks
perf(megatron): defer loss all-reduce to log time
#9966 opened Aug 22, 2026 by gakkiri Contributor Loading…
1 task done
feat(grpo): add M2PO for stale-rollout training
#9965 opened Aug 21, 2026 by primorLee Contributor Draft
fix: preserve incomplete flash checkpoints
#9963 opened Aug 21, 2026 by LeterTao Loading…
Fix zero multimodal learning rates in Megatron
#9948 opened Aug 19, 2026 by taking-lying-flat Contributor Loading…
Fix Megatron GRPO vocab-parallel log-prob gradients
#9947 opened Aug 19, 2026 by taking-lying-flat Contributor Loading…
[Megatron] Preserve RNG state across checkpoint resume
#9935 opened Aug 17, 2026 by taking-lying-flat Contributor Loading…
[Megatron] Fix reference adapter loading for LoRA RLHF
#9934 opened Aug 17, 2026 by taking-lying-flat Contributor Loading…
fix: 修复 Megatron GRPO 多模态缓存失配
#9933 opened Aug 17, 2026 by tutao0123 Contributor Loading…
Fix reward model margin broadcasting and alignment
#9927 opened Aug 16, 2026 by taking-lying-flat Contributor Loading…
Fix DPO IPO log-prob normalization
#9925 opened Aug 16, 2026 by taking-lying-flat Contributor Loading…
Fix BatchSamplerShard tail sampling
#9907 opened Aug 14, 2026 by taking-lying-flat Contributor Loading…
[Train] Reduce padding-free embedding output memory
#9893 opened Aug 11, 2026 by taking-lying-flat Contributor Loading…
ProTip! no:milestone will show everything without a milestone.