Back Home

GitHub Repo

PyTorch Community Reports Compiler Dropping a Coefficient, Potentially Amplifying 2.14 CUDA Momentum Updates Tenfold

A PyTorch community reproduction on October 5 found that Inductor may drop an addition’s alpha coefficient under a specific combination of clipping and scaling. A draft fix was proposed on October 6 but has not been merged; the impact on full training runs remains to be verified.

Pytorch Deepdream (https://github.com/gordicaleksa/pytorch-deepdream) by gordicaleksa (https://github.com/gordicaleksa/pytorch-deepdream/commits?author=gordicaleksa) · MIT · Image source
zh-Hant

A regression in PyTorch’s CUDA compilation path may silently change optimizer values. On October 5, user rwightman reported that while handling timm’s Adafactor momentum update, results from PyTorch 2.14 with torch.compile were too large. The attached minimal reproduction showed compiled output at ten times the result of eager execution. The same case produced the correct ratio with the PyTorch 2.13 CUDA and 2.14 CPU paths. Issue report

The triggering sequence first clips the update tensor by its root mean square, then multiplies it by a scaling coefficient, and finally runs m.mul_(0.9).add_(u, alpha=0.1). According to the official API definition, alpha must be multiplied by the tensor being added, so this step should compute 0.9m + 0.1u. The attached Triton code retained only the earlier scaling and omitted the final 0.1; when the test initialized momentum to zero, this produced a tenfold difference. This multiplier cannot be assumed to apply to all nonzero momentum values or full training results. Official addition documentation, in-place addition documentation

A draft fix proposed on October 6 points to Inductor’s pattern matching: reject a replacement when a node has a keyword argument that the pattern does not declare and that differs from its default, preventing a fused operation from losing the argument’s semantics. However, the author stated that the proposal description was AI-generated and had not yet been reviewed by them; a subsequent bot review also flagged issues with the tests and description. The draft should not be treated as a verified fix. Fix proposal

For engineering teams working on compiled optimizers, this case shows that throughput checks should be accompanied by checks of update tensors and momentum state. The lerp_ rewrite in the report produced a correct result, but that is only an observation from this reproduction. Next steps are to track the fix review, regression tests, and inclusion in an official release, and to compare numerical results before and after compilation with the team’s own optimizer. Public evidence is currently insufficient to determine the scope of impact across all GPUs, data types, or long-running training jobs.

Sources

  1. PyTorch Issue #199839:Inductor 2.14 FMA lowering drops alpha
  2. PyTorch PR #199871:Don't pattern-match nodes that set undeclared non-default kwargs
  3. torch.add — PyTorch 2.14 documentation
  4. torch.Tensor.add_ — PyTorch 2.14 documentation