Back Home

GitHub Repo

Ollama Community Reports MLX Mixed-Quantization Loading Defect: Successful Import May Still Fail at Inference

A community report involving Ollama 0.35.1 suggests that per-layer quantization settings for MLX models may not be applied correctly, causing tensor shape mismatches during loading. The model configuration confirms mixed-precision settings, but the root cause and the scope of impact on other models remain unconfirmed upstream.

Mattruffoni · CC BY-SA 4.0 · Image source
zh-Hant

On October 4, the Ollama community reported an MLX model compatibility issue: when importing mixed-precision weights on Apple Silicon, the model creation process appeared to succeed, but an actual generation request returned HTTP 500. The case used Ollama 0.35.1 on an M4 Pro with mlx-community/Qwen3-Coder-Next-4bit, creating the model through an experimental import path. The error surfaced only during loading. Issue report

Although the model is named “4bit,” its configuration does not use the same precision for every layer. The public config.json specifies 4-bit quantization globally, with a group size of 64 elements, but sets mlp.gate and shared_expert_gate to 8-bit in 48 layers, for a total of 96 overrides. The model configuration also lists 512 experts and a hidden size of 2048, so the loader must preserve the per-layer settings to interpret the packed weights correctly. Model configuration

The MLX documentation explains that quantized matrices pack multiple weight values into unsigned 32-bit integers, while scale factors are grouped according to group_size. Based on this, if the same packed width is misread as 4-bit instead of 8-bit, the inferred logical width would double, and the number of required scale factors would increase accordingly. The report gives weight dimensions of (512,512) and scale dimensions of (512,32); interpreting them as 4-bit would imply 64 scale columns. This is consistent with the reporter’s proposed root cause, but does not mean maintainers have confirmed a software defect. MLX operation documentation

For local inference deployments, this case highlights a gap between successful import and actual execution. Engineering teams can add a minimal generation request to deployment acceptance checks and verify per-layer quantization metadata against tensor dimensions, avoiding incompatibilities discovered only after large weights have been written. Follow-up should track whether upstream confirms how overrides are handled, adds pre-import checks, and creates regression tests using mixed-precision models. As of this review, the issue remains open; the available evidence is insufficient to generalize to all MLX models, and it has not been shown to cause silent degradation in inference quality. Current status and reproduction steps

Sources

  1. Ollama issue #18789:MLX 逐層量化覆寫與載入形狀錯配
  2. Qwen3-Coder-Next-4bit 的量化與模型配置
  3. mlx.core.quantized_matmul API 文件