GitHub Repo
Ollama Community Reports Qwen GGUF Thinking Levels Not Taking Effect; Low-Intensity Settings May Still Use the Default
A community report involving Ollama 0.35.0 says a specific Qwen3.8 GGUF may retain the template’s default `xhigh` setting after receiving a lower-intensity thinking request. Official documentation says to choose a supported value based on model metadata; whether template capabilities and API controls align remains unconfirmed.

A new report about thinking controls was posted to Ollama’s popular repository on October 3. A user loaded a Qwen3.8 27B GGUF with version 0.35.0 and passed think: "low", "medium", or "high" through /api/chat. The observed behavior still resembled enabled thinking; only false clearly changed the output. This is a community report from a specific environment, and there is not yet enough evidence to conclude that all similar models are affected. Issue report
The chat template attached to the report uses reasoning_effort to set the intensity. Its default is xhigh, and it accepts xhigh, medium, and low. The author suspects that the request’s think level is not passed to the template. After rendering a prompt manually with reasoning_effort="low" and sending it through the /api/generate path with raw: true, they observed shorter thinking content in a similar prompt. However, this comparison does not establish the cause, and it does not show that the API’s high is equivalent to the template’s xhigh. Template and comparison test
The current official documentation sets an important boundary: thinking controls vary by model. It recommends calling /api/show, checking thinking.values and the default, then using an exact matching name. The documentation also says that models which resolve levels from metadata use the model’s default when given an unsupported name. Therefore, the presence of level controls in a GGUF template is not enough to prove Ollama exposes them as supported request settings. Thinking controls documentation
For engineering teams, this kind of mismatch could cause agent workflows to exceed expected latency and generation volume. That is a deployment inference, not a broadly measured impact. To verify behavior, retain the model metadata, template version, and actual thinking output; repeat comparisons across levels after warm-up; and record generated token counts and elapsed time. Ollama’s chat API provides eval_count, eval_duration, and load_duration fields to help distinguish model loading from generation cost. API field documentation
Next, watch for upstream confirmation of a metadata recognition or template parameter mapping issue, and for supported-value guidance and regression validation. As of this check, the issue remains open and the page lists no related fix; the scope of the impact and the remedy are still unconfirmed. Issue status