模型推論與執行期
llama.cpp Adds Recurrent State Rollback for Kimi K3, Freeing Speculative Decoding from Host-Only Checkpoints
A new prerelease saves KDA state and convolution windows at each candidate position so rejected draft tokens can be rolled back correctly. Initial single-node tests show faster low-concurrency inference, but state memory grows substantially with rollback depth and the number of serving slots.

llama.cpp prerelease b10853 enables “bounded recurrent state rollback” for Kimi K3, closing a critical gap in speculative decoding for hybrid Transformer/recurrent architectures. Speculative decoding uses a smaller draft model to propose multiple tokens before the main model verifies them in a batch. If the main model rejects some of the draft, the runtime must restore its caches to the last accepted position. Conventional Transformers primarily roll back the KV cache, but Kimi K3’s 69 Kimi Delta Attention layers also maintain fixed-size recurrent state and short Q/K/V convolution windows. Saving only the final state would restore a nonexistent snapshot after a draft rejection, causing subsequent computation to diverge from the correct sequence.
The merged implementation saves convolution windows at every rollback-capable position and reuses llama.cpp’s recurrent-attention helpers to create KDA state snapshots. Tests cover checkpoint restoration, split replay across multiple sequences, and sequence isolation; targeted tests passed on both macOS CPU and NVIDIA Vulkan. Nonzero-fill testing was particularly important: when Kimi K3 was merely added to the allowlist, zero-initialized caches could mask errors, whereas filling them with `0x3e` exposed logit differences beyond the tolerance threshold.
Developers tested a quantized Kimi K3 with four serving slots, seven draft tokens, and eight RTX PRO 6000 Blackwell GPUs. Compared with host checkpoints, the rollback approach was approximately 13% to 53% faster across rewriting, knowledge, and coding workloads. However, this was not a placement-controlled A/B test: model configurations and available capacity differed. Compared with running without speculative decoding, TPOT improved at a concurrency of one but worsened at a concurrency of four. The cost was also clear: with four slots and seven rollback positions, the recurrent-state cache grew from about 1.77 GiB to 14.18 GiB. Deployers should next measure draft acceptance rate, concurrency, and additional state memory rather than looking only at single-request speed.