Back Home

GitHub Repo

vLLM Community Reports Ignored Per-Layer LoRA Scaling, Potentially Degrading Fine-Tuning Results in Deployment

A report says vLLM 0.30.0 does not apply PEFT’s per-layer rank and alpha settings, which could silently change adapter behavior. The community has provided a reproduction and a weight-conversion comparison; an official fix and the scope of impact remain unconfirmed.

VamosSandor · CC BY-SA 4.0 · Image source
zh-Hant

On October 2, the vLLM community reported a LoRA compatibility issue: adapters trained with Hugging Face PEFT may use incorrect scaling when deployed on vLLM 0.30.0 if individual modules are configured through rank_pattern or alpha_pattern. The report says the loading process reads only the global r and lora_alpha values, ignoring per-layer settings without issuing a warning. At the time of review, the issue remained open and listed no related fix. Issue report

PEFT’s official documentation confirms that these two fields can override settings by module name or regular expression. Standard LoRA updates are scaled by alpha/r. So even when the weights load successfully, inconsistent scaling semantics can change the model’s computations. For example, with a global rank of 16 and alpha of 32, a layer changed to rank 4 should have a scaling factor of 8; using the global factor instead gives 2. This also means that the mere presence of these fields does not establish an error: what matters is whether the per-layer ratio changes. PEFT documentation

The reporter tested fine-tuning OLMoE-1B-7B on wikitext-2 for 200 steps, reporting perplexity of 10.06 with PEFT and 11.28 with vLLM; lower perplexity is better. After folding each module’s scaling into lora_B and removing the patterns, vLLM returned to 10.06. The report also includes code for a small random-model case to compare token log probabilities. These figures are the author’s own tests and do not establish a quality loss for all LoRA adapters. Tests and reproduction code

The issue is especially worth tracking for MoE fine-tuning: PEFT’s documentation demonstrates lowering the rank of expert layers to control the adapter parameter budget, and points to vLLM for accelerated inference. The engineering implication is that teams should compare the actual updates in addition to checking file formats between training and serving. MoE configuration example

Deployers can first inspect adapter settings, compare log probabilities from PEFT and the serving system using the same token inputs, and then evaluate weight conversion. Follow-up areas include per-module scaling support, capacity checks for the maximum rank in patterns, and upstream validation across models. Public information has not yet confirmed an official fixed release. Potential fixes

Sources

  1. vLLM Issue #59799:PEFT rank_pattern/alpha_pattern 縮放相容性回報
  2. PEFT LoRA 文件:逐層 rank、alpha 與 MoE 專家設定