Back Home

GitHub Repo

Ollama Community Reports VRAM Reservation Failure; Fix Targets llama-server Auto-Configuration

Community reports for Ollama 0.34.4 indicate that llama-server may retain its original model-layer configuration after a VRAM reservation is specified. A proposed fix has been submitted but not merged, and its full scope remains unclear.

Mattruffoni · CC BY-SA 4.0 · Image source
zh-Hant

On September 27, the Ollama community reported that the llama-server backend in 0.34.4 did not apply the VRAM reservation setting to model-layer configuration. The reporter tested two models on Windows with an RTX 4070, requested a 6 GiB VRAM reservation, and said the loading results were the same as when no reservation was set. Maintainers have yet to confirm the full scope of the issue. [Issue report](https://github.com/ollama/ollama/issues/18679)

The setting has a specific purpose. Ollama’s Go package documentation defines `OLLAMA_GPU_OVERHEAD` as the amount of VRAM to reserve on each GPU; its default is zero. On machines running models alongside other GPU workloads, whether this value actually affects configuration decisions directly impacts capacity planning. A startup log showing that the parameter was read is not enough to prove that the resources were reserved. [Package documentation](https://pkg.go.dev/github.com/ollama/ollama/envconfig)

The proposed fix points to a gap in parameter passing: when the default `num_gpu=-1` is used, llama.cpp’s `--fit` automatically configures the model layers, and the reserved space must be passed through `LLAMA_ARG_FIT_TARGET`. The existing path accounted only for the extra space needed by the multimodal projector; it did not add the user-specified VRAM reservation. This offers a technical explanation for how the setting could be read by the higher-level application while leaving backend configuration unchanged. [Fix proposal](https://github.com/ollama/ollama/pull/18680)

Submitted the same day, the proposed fix would convert the reservation to MiB and add it to the space required by the projector. As of September 28, the proposal remained open with no review recorded, so it cannot be considered a fix in an officially released version. [Proposal status](https://github.com/ollama/ollama/pull/18680) The report also notes that some scheduler memory estimates do not subtract the reservation. Further validation is therefore needed to determine whether a fix to this single parameter-passing path would cover the entire loading process. [Code analysis in the report](https://github.com/ollama/ollama/issues/18679)

The technical impact is that both multi-model loading and parallel requests rely on available-memory estimates. Official documentation explains that parallel requests increase context memory requirements, and insufficient capacity can also affect queuing and model unloading. When testing an update, engineering teams should compare actual available VRAM, model-layer distribution, and concurrent workload before and after setting the reservation, and validate text and vision models separately. These deployment risks are inferred from the resource-management mechanisms; no data is available to quantify the service failure rate caused by this issue. [Official memory management documentation](https://docs.ollama.com/faq)

Sources

  1. OLLAMA_GPU_OVERHEAD is ignored by the llama-server backend — Issue #18679
  2. llm: include OLLAMA_GPU_OVERHEAD in llama-server fit target — PR #18680
  3. Ollama envconfig 套件文件
  4. Ollama FAQ:並行請求與 GPU 記憶體管理