GitHub Repo
Transformers Community Reports Memory Growth in Sparse Attention Indexer; Long DeepSeek-V3.2 Prompts May Exhaust GPU Memory
A reproduction on the October 5 development build found that the Lightning Indexer retains full per-attention-head scores, which may sharply increase temporary GPU memory use for long prompts. The report provides an isolated indexer test; effects on the full model and official releases remain unconfirmed.

The Transformers community reported a performance issue on October 5: in 5.19.0.dev0, DeepSeek-V3.2’s Lightning Indexer may allocate oversized intermediate tensors. The reporter tested the standalone indexer on an NVIDIA L4 with random inputs, measuring about 2.09 GiB of additional peak GPU memory at 2,048 tokens, rising to 8.16 GiB at 4,096 tokens; at 8,192 tokens, the test ran out of memory. These figures come from a component test and should not be treated as deployment requirements for the full model. Community reproduction
The issue concerns the implementation cost of sparse attention’s “score first, select afterward” approach. According to the report’s analysis, the code first creates fp32 per-head scores with shape [B,S,H,T], then sums across attention heads. The scaling operation creates another copy. When query and key lengths are equal for a single prompt, the two tensors occupy about 2×S²×H×4 bytes. The default 64 indexer heads amplify temporary memory use, so retaining only the top-k results at the end does not eliminate the earlier memory peak. Issue analysis
The official documentation explains that DeepSeek Sparse Attention uses an indexer to score queries against previous tokens, then selects 2,048 tokens by default for the main attention operation. The reference implementation uses FP8, while the Transformers port computes in bf16/fp32; the documentation also notes that the sparse path with flash_mla is not yet supported. This means deployers need to assess the indexer’s temporary memory and the main attention operation separately; the actual savings cannot be estimated from the model architecture alone. Official technical documentation
As of the time of review, the issue remained open, with no merged fix visible. The report also notes similar code blocks in models such as GLM-5, but provides no measurements for those full models. Engineering teams can first pin the package version, then increase prompt length step by step while measuring peak memory during prefill. They should track whether later chunked or fused operations avoid creating the full per-head tensor, and whether the corrected top-k results remain numerically consistent.