GitHub Repo
Transformers Community Proposes Length-Control Fix After Parameter Combination May End Generation Early
Community reproductions show that using a minimum generation length together with an exponential length penalty may produce NaNs, causing generation to end early or sampling to fail. A proposed fix appeared on October 2, but its status in an official release remains unconfirmed.

The Transformers community has proposed a fix for an interaction bug in generation length controls. A report published on October 1 said that when exponential_decay_length_penalty takes effect before the minimum generation length is reached, the end-of-sequence (EOS) token score may become NaN. The public PR list showed a proposal on October 2 to “preserve masked EOS scores”; it is still listed as open. Issue report, public PR list
The two controls serve different purposes. The official documentation says min_new_tokens sets the EOS score to negative infinity until the specified length is reached, making the model continue generating. The exponential length penalty, meanwhile, progressively increases the EOS score after a starting position so the output ends naturally. The first sets a minimum length; the second encourages shorter outputs, so deployers may configure both. Official generation utilities documentation
According to the reporter’s analysis, the problem occurs when the later score adjustment is still applied to the masked EOS score: negative infinity plus positive infinity yields NaN. As a result, a score that was meant to represent “temporarily prohibit ending” no longer behaves as intended. The community reproduced the issue with small GPT-2 and T5 configurations, observing that greedy decoding and beam search ended before the minimum length, while sampling raised an error because the probability tensor contained invalid values. Reproduction and analysis
The reporter’s local fix applies the score increase only to finite EOS scores, preserving the existing mask. This offers a concrete direction for a fix, but test results in the report were provided by the proposal’s author; they do not establish that all models or deployment paths have been validated, and the proposal should not be treated as a released fix. Fix description
For engineering teams, this case shows why generation parameters need to be validated in combination. Based on the mechanism described above, checks should cover the relative positions of the penalty start and minimum length, the actual number of newly generated tokens, and errors under different decoding methods. Teams should track whether the proposal is merged and which version includes the fix, then rerun tests with their own models and parameters to confirm that the minimum-length constraint holds.