GitHub Repo
Transformers Community Reports Memory Bloat in Sparse Attention Indexer; Long DeepSeek-V3.2 Prompts May Exhaust GPU Memory
A reproduction on the October 5 development build found that the Lightning Indexer retains full per-attention-head scores, potentially increasing temporary GPU memory substantially for long prompts. The report tests the indexer in isolation; the impact on the full model and stable releases remains unconfirmed.

On October 5, the Transformers community reported a performance issue: in 5.19.0.dev0, DeepSeek-V3.2’s Lightning Indexer may allocate oversized intermediate tensors. Using an NVIDIA L4, random inputs, and the indexer in isolation, the reporter measured about 2.09 GiB of additional peak GPU memory at 2,048 tokens, rising to 8.16 GiB at 4,096 tokens; at 8,192 tokens, the test ran out of memory. These figures come from a component-level test and should not be treated as deployment requirements for the full model. Community reproduction
The issue concerns the implementation cost of sparse attention’s “score first, then select” approach. According to the report’s analysis, the code first creates per-head fp32 scores with shape [B,S,H,T], then sums across attention heads. The scaling operation creates another copy. With a single prompt and equal query and key lengths, the two tensors take about 2×S²×H×4 bytes. The default of 64 index heads amplifies temporary memory use, so keeping only the top-k results at the end does not remove the earlier memory peak. Issue analysis
The official documentation explains that DeepSeek Sparse Attention uses an indexer to score each query against previous tokens, then selects 2,048 tokens by default for the main attention operation. The reference implementation uses FP8, while the Transformers port computes in bf16/fp32. The documentation also notes that the sparse computation path with flash_mla is not yet supported. This means deployment teams need to assess indexer temporary memory and main attention costs separately; the actual savings cannot be estimated from the model architecture alone. Official technical documentation
As of the time of review, the issue remained open, with no merged fix in view. The report also points to similar code blocks in models such as GLM-5, but provides no measurements for those full models. Engineering teams can pin the package version, increase prompt length in stages, and measure peak memory during prefill. They should also track whether chunked or fused operations avoid materializing the full per-head tensor, and check top-k results and numerical consistency after any fix.