GitHub Repo
vLLM Community Reports Silent Loss of Image Input; Per-Request Thinking Settings May Affect Visual Understanding
A community report involving vLLM 0.30.0 suggests that adding `chat_template_kwargs` to image requests may leave the model with text only, while the server still returns success. The issue has yet to be confirmed upstream; deployments combining visual input with thinking-mode controls should verify the actual input.

On October 3, the vLLM community reported a multimodal input issue: after adding chat_template_kwargs to a /v1/chat/completions request, images may disappear from the model input, while the endpoint still returns HTTP 200 with no error or warning. The case used vLLM 0.30.0, an NVFP4-quantized Gemma-4-26B-A4B, and an RTX PRO 4500 Blackwell. This is a user report involving a specific configuration; it does not establish that all vision models are affected. Issue report
The reporter provided two comparison requests with the same image and text. Without per-request template parameters, the prompt contained 276 tokens and the model described the image. After adding enable_thinking: true, the prompt token count fell to 22, and the response said it had not received an image. The same behavior appeared when the setting was changed to false or when using one or three images. Notably, the server was already configured with thinking enabled, so specifying the same value again per request still changed the input. Comparison tests
The official documentation describes this parameter as a supported interface for controlling thinking: deployers can set a global default with --default-chat-template-kwargs, while clients can override it per request. The documentation also states that Gemma 4 can enable thinking through enable_thinking. This makes the report practically relevant: applications that adjust reasoning behavior by task may also change whether visual data reaches the model. The documentation itself, however, does not confirm this defect or explain its cause. Official documentation
For image analysis and visual agents, this kind of failure could evade monitoring that checks only status codes. An engineering implication is that acceptance testing should include recognition tests with fixed images and compare input usage and whether responses match the images, both with and without the parameter. An unusual token count can be a clue for investigation, but cannot alone prove that an image was lost. Temporarily removing per-request settings based on this report would also require confirming that the global configuration still meets the application's needs.
As of the time of review, the issue remained open, and the page listed no related fix or maintainer confirmation. Next steps include monitoring full logs, reproduction results across other models and quantization configurations, and whether upstream can locate the input-processing path and add a regression test. The behavior should not currently be attributed directly to quantization or a particular GPU. Issue status