Back Home

GitHub Repo

vLLM Community Reports Suspected SM120 Speculative Decoding Issue; GLM Draft Token Acceptance Falls to Zero

Tests on eight RTX PRO 6000 GPUs show that with a vLLM nightly build and the native SM120 attention backend, all MTP draft tokens were rejected and decoding slowed noticeably. The issue awaits upstream confirmation; the existing comparison also changes software dependencies, so a single cause has not been established.

Cepice · CC BY-SA 4.0 · Image source
zh-Hant

On October 2, the vLLM community reported a speculative decoding issue: with eight RTX PRO 6000 GPUs, FP8 GLM-5.3-Flash, and an eight-way tensor-parallel configuration, that day’s nightly build continued to generate MTP draft tokens, but the acceptance rate was zero. In one reported test, 822 draft tokens were produced and none were accepted; decoding ran at about 80 tokens per second. Issue report

MTP uses a model’s multi-token prediction capability to propose candidates, then a verification process decides which tokens to accept, reducing the cost of step-by-step decoding. The vLLM documentation describes speculative decoding as a way to reduce inter-token latency in memory-bandwidth-bound workloads with low to medium request volumes, and cautions that gains depend on the model, hardware, traffic, and sampling settings. So a service generating text normally does not by itself show that the acceleration path is working. Official documentation

The reporter tested the default CUDA Graph mode, forced eager execution, and reducing the number of draft tokens from three to one; the acceptance rate remained zero. These results narrow the scope of the investigation, but do not fully rule out effects from the execution mode or draft-token workflow. Logs show that the system automatically selected FLASHINFER_MLA_SPARSE_SM120 and marked the old attention-backend environment variable as an unknown setting. The reporter therefore suspects an issue in the draft inference path. Tests and logs

A community image on the same hardware and model, using a backported SM90 sparse MLA path, reportedly achieved about 60% draft-token acceptance and 163 to 190 tokens per second. However, the comparison also changes the vLLM build and FlashInfer version, so the speed difference cannot be attributed directly to the SM120 kernel or generalized as the extent of degradation across all deployments. Comparison configuration

For engineering teams using nightly builds, this case suggests that upgrade acceptance checks should include draft-token acceptance rate, the actual attention backend, and inter-token latency, while retaining a baseline with speculative decoding disabled. The next step is to wait for upstream to reproduce the issue with fixed dependency versions and verify the draft outputs and verification process before assessing a fix and its impact on stable releases. As of this review, the issue remains open and the page lists no related fix.

Sources

  1. vLLM issue #59724:SM120 後端下 MTP 草稿接受率為零
  2. vLLM 官方推測解碼文件