GitHub Repo
Transformers Community Proposes MiniMax Precision Fix as bf16 Recurrence Decay May Shift Long-Sequence Decoding
A community reproduction posted on October 4 found that bf16 decay calculations in the development version of Transformers’ MiniMax Lightning Attention may prevent some attention heads from forgetting older information. A fix using fp32 has been proposed but not yet merged; its impact on model quality and performance remains to be validated.

On October 4, the Transformers community reported that low-precision arithmetic in MiniMax Lightning Attention may cause numerical drift that grows with decoding length. The case used a development branch, PyTorch 2.11, and a CPU; the issue points to the per-token decay coefficient and recurrent state updates. A fix has been proposed and is awaiting upstream review. Issue report
This computation is especially important for long-context models. The official documentation says MiniMax-Text-01 combines Lightning Attention, Softmax Attention, and MoE, with one Softmax Attention layer after every seven Lightning Attention layers. Decay in the recurrent state controls how older information is retained, so deployments need to check the computation precision of state updates as well as weight precision. Architecture documentation
The report says bf16 lacks sufficient precision near 1: decay coefficients around 0.998 or higher may round to 1, causing the corresponding attention heads to stop decaying. Using the public configuration, the author estimated that 388 of 4,480 Lightning Attention heads meet this condition. This is an analysis of the coefficients and should not be interpreted directly as an error rate or a drop in quality on real-world tasks. Reproduction and limitations
The proposed fix promotes decay-related operations and cached states to fp32, then converts back to the original precision before the output projection. Using the same small-model reproduction, the author reported that at 3,000 tokens, hidden-state error relative to fp32 fell from about 7.39% to 0.39%. This result supports the direction of the fix, but is not a long-context evaluation of the full MiniMax model. Proposed fix
The engineering tradeoff is that using fp32 for the linear-attention state roughly doubles the capacity required for that portion of the cache compared with bf16; it does not double memory use for the entire model. Deployment teams should next track which released versions are affected, GPU validation with real weights, and the throughput and memory costs after the fix. Acceptance testing should also evaluate prompt prefill and token-by-token decoding separately, so short-output tests do not mask accumulated error. The author explicitly said that a full MiniMax checkpoint has not been tested, so for now this should be treated as a numerical issue report with a reproduction script. Validation scope