Back Home

GitHub Repo

Ollama Maintainer Proposes MLX State Fix to Reduce Hidden Memory Use

Proposed on October 5, the fix targets recursive states that continue referencing large buffers after speculative decoding rollback, preventing unnecessary memory from being retained with the prefix cache. The author reports 3.52 GiB less memory use in the final round of a 12-call tool test, but the proposal has not yet been merged.

Mattruffoni · CC BY-SA 4.0 · Image source
zh-Hant

Ollama maintainer dhiltgen opened PR #18805 on October 5 to address oversized buffers retained by the MLX inference path after speculative decoding rollback. The author reports that, in an A/B test involving 12 tool calls, the fix reduced MLX-held memory by 3.52 GiB in the final round, while logical cache usage remained comparable. At the time of review, the proposal was still open, targeting the release_v0.40.0 branch. Proposed fix

The issue lies in the state lifecycle: after rollback and selection of one state, a per-token recursive state may still point to a larger underlying buffer. If the request ends before the next model operation, that buffer can remain alive with the cached state; the cache’s reported usage accounts only for the selected state, while the actual allocation may be larger. The proposal marks restored states that may share buffers, then compacts them in a batch when a request closes, before they are retained in the prefix cache. It clears the marker when a normal model step replaces the state and adds tests for rollback and buffer separation. Mechanism details

The new fix is linked to a community report from September 24. The reporter ran qwen3.6:27b-mlx with Ollama 0.34.2 and 0.34.4 on a 32 GB M1 Max Mac Studio, enabled MTP speculative decoding, and made repeated tool calls. They observed an additional ~0.43 GiB for each request ending in a tool call, not reflected in the prefix-cache count; a regular chat control did not show the same increase. This was reproduced under a specific configuration and cannot be generalized directly to all models and backends. Original report and script

For local agent deployments, this means capacity planning should track both reported cache usage and runtime allocations. The official MLX documentation says get_active_memory() reports the number of bytes in active use and excludes cache buffers, so its value may not match the memory usage shown by the operating system. MLX memory documentation

Engineers should next track the merge and release, then compare memory curves, swap usage, and latency before and after the fix under repeated tool calls. The 3.52 GiB figure comes from a single test by the author; the degree of improvement for other models, the cost of compaction, and whether the fix fully addresses related accumulation issues remain to be verified.

Sources

  1. mlx: compact restored recurrent state after speculative rollback — PR #18805
  2. MLX runner: tool-call requests retain memory outside the prefix-cache budget — Issue #18620
  3. mlx.core.get_active_memory — MLX documentation