-
Notifications
You must be signed in to change notification settings - Fork 579
Pull requests: NVIDIA/Model-Optimizer
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
fp32 master weights for the DFlash draft, and keep them across a resume
#2322
opened Sep 3, 2026 by
h-guo18
Contributor
Loading…
Reject unsupported partial-block INT4/W4A8 AWQ export
#2320
opened Sep 2, 2026 by
realAsma
Contributor
Loading…
Fix TEGroupedMLP quantizer checkpoint resharding
#2319
opened Sep 2, 2026 by
jenchen13
Contributor
Loading…
[6410139] Fix ONNX AutoCast for large external initializers
cherry-pick-0.47.0
Upcoming release
#2317
opened Sep 2, 2026 by
ajrasane
Contributor
Loading…
docs: document speculation profiles and how to produce them
#2316
opened Sep 2, 2026 by
yeyu-nvidia
Contributor
Loading…
ar_validate: emit a speculation profile from in-training AR validation
#2315
opened Sep 2, 2026 by
yeyu-nvidia
Contributor
Loading…
export: attach speculation_profile.json to exported draft checkpoints
#2313
opened Sep 2, 2026 by
yeyu-nvidia
Contributor
Loading…
Add the NVFP4 experts-only PTQ recipe for zai-org/GLM-5.3-Flash
#2312
opened Sep 2, 2026 by
shengliangxu
Collaborator
Loading…
refactor(export): let the caller own the layerwise exporter
#2303
opened Sep 1, 2026 by
Fridah-nv
Contributor
Loading…
Add the NVFP4 PTQ recipe for Qwen/Qwen3.8-2.4T-A95B
#2302
opened Sep 1, 2026 by
shengliangxu
Collaborator
Loading…
Forward kv_cache_free_gpu_memory_fraction to the TensorRT-LLM engines (NVBug 6701763)
#2300
opened Sep 1, 2026 by
cjluo-nv
Collaborator
Loading…
Fix fsdp2_aware_weight_update masking setup errors with UnboundLocalError
#2295
opened Sep 1, 2026 by
harshal-96
Loading…
2 of 4 tasks
specdec: config_overrides for nested text_config checkpoints + load VLM-capable bases in merge_lora
#2289
opened Aug 31, 2026 by
yeyu-nvidia
Contributor
Loading…
ar_validate: fail loudly when every sample fails
#2288
opened Aug 31, 2026 by
yeyu-nvidia
Contributor
Loading…
[chore]: weekly bump of uv.lock on main (2026-08-31)
#2285
opened Aug 31, 2026 by
github-actions
Bot
Loading…
[OMNIML-5570] 2/2 Compose GEMM and KV-cache AutoQuant workflows
#2273
opened Aug 27, 2026 by
meenchen
Contributor
Loading…
[OMNIML-5570, OMNIML-5569] 1/2 Add layer-wise KV-cache AutoQuant with forward KL
#2272
opened Aug 27, 2026 by
meenchen
Contributor
Loading…
Gkarch/sync main 449a3992
puzzletron_v2
Related to feature/puzzletron_v2 branch
#2266
opened Aug 27, 2026 by
grzegorz-k-karch
Contributor
Loading…
Docs: Add WOA documentation
cherry-pick-0.46.1
#2264
opened Aug 27, 2026 by
haoxiz-nvidia
Contributor
Loading…
Previous Next
ProTip!
What’s not been updated in a month: updated:<2026-08-03.