GitHub Repo
vLLM Community Reports Multi-Node Decoding Latency; Even an Empty KV Connector May Reduce Performance
In a two-node H100 test, adding a KV connector that moves no data increased per-token latency by about 87% under a specific workload. The investigation points to response aggregation and thread waiting; the cause and any fix still await upstream confirmation.

On September 27, the vLLM community reported a multi-node decoding performance issue: in version 0.30.0, using Model Runner V2 with asynchronous scheduling, even a KV cache connector that performs no reads or writes increased latency. The test used two nodes with 16 H100 80GB GPUs in total, running GLM-5.2 FP8 with eight-way pipeline parallelism and two-way tensor parallelism. The input was about 12,000 tokens, with generation fixed at 1,024 tokens. With eight requests running concurrently, per-token latency rose from 32.1 ms to 60.0 ms. [Community report](https://github.com/vllm-project/vllm/issues/58920)
The official code provides a way to verify a key difference: when the multiprocess executor receives an output aggregator, it sets `output_rank` to `None` and collects responses from all workers. A path that previously needed results from only a designated process therefore becomes one that reads multiple response queues in turn. Even if each message is small, the number of responses may add overhead. [Version 0.30.0 executor documentation](https://docs.vllm.ai/en/v0.30.0/api/vllm/v1/executor/multiproc_executor/)
On the writer side of the shared-memory queue, while a block remains unread, the code repeatedly checks its status and calls `sched_yield`. The reporter attributed the bottleneck to contention for Python’s Global Interpreter Lock (GIL) during the wait, which interfered with the main thread’s preparation of inputs. In a diagnostic change, the reporter switched to a wait method that releases the lock, reducing the added latency to about 13%. This attribution and measurement come from the reporter. [Queue source code](https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/distributed/device_communicators/shm_broadcast.py), [diagnostic experiment](https://github.com/vllm-project/vllm/issues/58920)
This type of bottleneck bears directly on V2’s design goals. The official documentation says asynchronous scheduling lets the CPU prepare inputs for the next step while the GPU executes the current step, and requires CPU operations to remain non-blocking. It follows that if a response thread interferes with input preparation, the overall decoding pace can still be constrained even when GPU compute kernels are no slower. [V2 design documentation](https://docs.vllm.ai/en/latest/design/model_runner_v2/)
KV connectors are also an important interface in disaggregated prefill and decode architectures. The official documentation describes disaggregated deployment as a way to tune time to first token and subsequent-token latency separately, and notes that the feature does not increase throughput. When evaluating a deployment, teams should therefore measure both the benefits of cache transfer and the costs of the control path, rather than estimating performance from network bandwidth alone. [Disaggregated deployment documentation](https://docs.vllm.ai/en/latest/features/disagg_prefill/)
Engineering teams can compare an unconfigured connector, an empty connector, and a real connector while holding the model, parallelism topology, and workload constant, and monitor both inter-token latency and CPU threads. The available evidence is limited to a single deployment and cannot be generalized to all hardware. As of this check, the issue remains open, with no maintainer confirmation or related fix in view. Follow-up should track upstream changes to response aggregation, queue capacity, and waiting strategies. [Issue status](https://github.com/vllm-project/vllm/issues/58920)