-
Notifications
You must be signed in to change notification settings - Fork 830
Pull requests: NVIDIA/TransformerEngine
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[PyTorch] Enable LayerNormLinear UB AG overlap under no_grad
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3546
opened Sep 18, 2026 by
ravimajeti
Loading…
7 of 8 tasks
[PyTorch] Keep DistributedWeight objects on ctx across saved-tensor hooks
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3545
opened Sep 18, 2026 by
xrennvidia
Collaborator
Loading…
1 of 13 tasks
feat: vmm slot for local cuda graph offload
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3544
opened Sep 18, 2026 by
Xdydy
Loading…
1 of 13 tasks
Use PyTorch CUTLASS for BF16 grouped GEMM on SM100
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
Fix distributed Newton-Schulz correctness
2.20
#3541
opened Sep 17, 2026 by
vcherepanov-nv
Collaborator
Loading…
4 of 13 tasks
[PyTorch][CP] Support compact metadata with FA4
#3540
opened Sep 17, 2026 by
sudhakarsingh27
Member
Loading…
13 tasks
fix: add missing NVTE_BHSD labels to fused_attn to_string
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3538
opened Sep 17, 2026 by
andrewwhitecdw
Contributor
Loading…
Prototype of green context + VMM localization
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
test(attention): anchor the context-parallel suite to an independent reference
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3536
opened Sep 17, 2026 by
nvegesna-netizen
Contributor
Loading…
[PyTorch] [torch.compile] torch.compile support for LayerNormLinear and LayerNormMLP
#3534
opened Sep 17, 2026 by
pggPL
Collaborator
Loading…
[PyTorch] Retain NVFP4 RNG tensors through quantization
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3533
opened Sep 17, 2026 by
Connor-XY
Loading…
fix(attention): stop handing FlashAttention 4 the -1 window sentinel
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3532
opened Sep 17, 2026 by
nvegesna-netizen
Contributor
Loading…
[PyTorch] Extend no-load-balance CP to the a2a comm type
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3530
opened Sep 16, 2026 by
Rudin6
Loading…
13 tasks
feat(attention): cuDNN FROST attention backend for head_dim in (256, 512], with context parallelism
2.21.0
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3527
opened Sep 16, 2026 by
nvegesna-netizen
Contributor
Loading…
[Pytorch][NCCL EP] Create tokens_per_expert on pinned CPU memory when running eager
#3526
opened Sep 16, 2026 by
YangFei1990
Collaborator
Loading…
8 of 13 tasks
[JAX] WAR for XLA associative scan rewriter crash
2.20
#3525
opened Sep 16, 2026 by
KshitijLakhani
Collaborator
Loading…
2 of 13 tasks
Remove the view() calls from Linear/LNLinear/LNMLP modules
#3523
opened Sep 15, 2026 by
ptrendx
Member
Loading…
1 of 13 tasks
[JAX] Support single-process-multi-device EP via XLA-borrowed communicator
#3522
opened Sep 15, 2026 by
phu0ngng
Collaborator
Loading…
8 of 13 tasks
Accept compact MXFP8 scales in GEMM preparation
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
[PyTorch] Support distributed weights in GroupedLinear's grouped-tensor path
org-contribution
#3517
opened Sep 15, 2026 by
fanshiqing
Member
Loading…
7 of 8 tasks
[JAX] Align fused attn backward output gradient sharding and other test strengthening
2.20
attention
bug
Something isn't working
#3516
opened Sep 14, 2026 by
KshitijLakhani
Collaborator
Loading…
1 of 13 tasks
Support building docs for publishing via GHA
#3515
opened Sep 14, 2026 by
fheinecke
Collaborator
Loading…
1 of 13 tasks
Previous Next
ProTip!
Find all pull requests that aren't related to any open issues with -linked:issue.