PyTorch 2.14 ships NVGEMM, bringing CuTeDSL-generated CUTLASS kernels to Inductor with epilogue fusion and scaled/NVFP4 GEMM support autotuned alongside Triton and ATen. The release introduces a new nccl2 backend for PyTorch Distributed ported from torchcomms, and makes fault tolerance a first-class concept in c10d with in-place process-group reconfiguration and a backend-agnostic Flight Recorder. Apple Silicon gains native linear algebra operations including SVD, eigh, QR, and Cholesky, while torch.compile adds experimental support for complex-valued tensors.
