-
Notifications
You must be signed in to change notification settings - Fork 799
Pull requests: NVIDIA/TransformerEngine
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
mxfp8: add swizzled-scale fast path for cast-only quantization
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3338
opened Aug 10, 2026 by
WanZzzzzz
Contributor
Loading…
13 tasks
[common] Improved performance of Group MXFP8 kernels
#3337
opened Aug 10, 2026 by
Oleg-Goncharov
Collaborator
Loading…
6 of 13 tasks
[PyTorch] Fine-grained recipe docs
documentation
Improvements or additions to documentation
#3336
opened Aug 10, 2026 by
negvet
Collaborator
Loading…
13 tasks
Add a backwards linear function to be used with the fused mla q up-proj
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3330
opened Aug 7, 2026 by
chaseblock
Contributor
Loading…
13 tasks
[PyTorch] Fix deferred initialization in fusible ops
#3327
opened Aug 7, 2026 by
denera
Collaborator
Loading…
8 of 13 tasks
Refactor GroupedLinear quantization dispatch
#3326
opened Aug 7, 2026 by
negvet
Collaborator
Loading…
13 tasks
[PyTorch] Enable NVFP4 row-scaled (per-token) backward for GroupedLinear
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3324
opened Aug 7, 2026 by
cael-ling
Contributor
Loading…
1 of 13 tasks
[Pytorch] Enable TE Op to consume extra_outputs from a previously run Op in TE Sequential
#3320
opened Aug 5, 2026 by
vthumbe1503
Collaborator
Loading…
13 tasks
[PyTorch] Advance FusedAdam step counter for empty param groups
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3318
opened Aug 5, 2026 by
adityasingh2400
Loading…
[Common][PyTorch] Fuse the RHT into grouped NVFP4 quantize on non-SM100 architectures
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3317
opened Aug 5, 2026 by
davidkny22
Contributor
Loading…
6 of 13 tasks
[Common/PyTorch] Grouped weighted-SwiGLU MXFP8 kernel
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3315
opened Aug 4, 2026 by
cael-ling
Contributor
Loading…
3 of 13 tasks
Log when thd with dropout falls to the composite cuDNN engine
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3313
opened Aug 4, 2026 by
bzantium
Loading…
[CI] Publish GB200 aarch64 wheel artifacts
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3311
opened Aug 4, 2026 by
bvolpato
Loading…
5 of 13 tasks
nvrtc MXFP8 kernels
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3302
opened Aug 3, 2026 by
CarlosGomes98
Contributor
•
Draft
13 tasks
NVRTC NVFP4 quantization kernels
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3301
opened Aug 3, 2026 by
CarlosGomes98
Contributor
Loading…
8 of 13 tasks
Add NVFP4 RHT Support for SM120 and SM121
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3300
opened Aug 3, 2026 by
new-TonyWang
Loading…
[PyTorch] Scope the quantized-param caching flag to its own graph capture
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3298
opened Aug 1, 2026 by
xiuhu17
Contributor
Loading…
6 tasks done
[Pytorch][Common] Option to Disable 2nd level Scale in NVFP4
#3297
opened Jul 31, 2026 by
vthumbe1503
Collaborator
Loading…
13 tasks
[PTQ] Store FP32 global scaling factors (absmax or scale_inv) for all quantized activations and weights.
org-contribution
#3296
opened Jul 31, 2026 by
cspades
Member
Loading…
7 of 13 tasks
[Common] Fix pointer arithmatic to generate correct LDS/STS instructions
#3291
opened Jul 30, 2026 by
kainzhong
Collaborator
Loading…
8 of 13 tasks
Fix cuDNN SDPA score_mod cache collision, sm_12x arch gates, per-device plan cache
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3289
opened Jul 30, 2026 by
YangXu1990uiuc
Loading…
[PyTorch][torch.compile] Support for DotProductAttention on flash and unfused backends
#3286
opened Jul 30, 2026 by
pggPL
Collaborator
Loading…
9 tasks done
Previous Next
ProTip!
Adding no:label will show everything without a label.