Back Home

GitHub Repo

Transformers Community Reports FSDP2 Loading Memory Bug; Multi-GPU Startup May Create Duplicate Full Models

A community report involving Transformers 5.12.0 says that after enabling the CPU memory-efficient loading option, other training processes may still allocate a full model, exhausting host memory. Official documentation says the option should prevent duplicate loading across processes, but the root cause and fix status await upstream confirmation.

Marcus Qwertyus · Public domain · Image source
zh-Hant

On October 4, the Hugging Face Transformers community reported an FSDP2 loading issue: even with cpu_ram_efficient_loading=True enabled, non-rank 0 processes may still build a full model on the CPU. The reported setup used Transformers 5.12.0, Accelerate 1.15.0, and PyTorch 2.10. It launched 16 processes on a single node to load a dense BF16 model with about 27 billion parameters, and reportedly ran out of host memory. Issue report

The reporter traced the problem to the from_pretrained() stage: while handling parameters on the meta device, a relevant code path calls torch.zeros_like(..., device="cpu") for non-main processes, creating real CPU tensors. Meta tensors normally retain only information such as shape and do not allocate storage for parameters. If they are materialized too early, each process may use about 54 GB. Based on the figures in the report, 16 model copies would require about 864 GB, excluding other loading overhead. This is a capacity estimate for the reported configuration, not a fixed threshold for all deployments. Code path and reproduction steps

This behavior appears to diverge from the option’s intended purpose. Accelerate’s official documentation says that with CPU RAM-efficient loading enabled, only the first process should load the pretrained checkpoint; the other processes should retain empty weights and then obtain parameters through synchronization. The documentation also requires initializing the distributed process group before calling from_pretrained(). FSDP parameter sharding can reduce memory requirements during training, but if full copies are created before sharding, startup can still become a bottleneck. Official FSDP documentation

For engineering teams, this report is a reminder to include the model-loading phase in capacity checks. One approach is to use a smaller model to observe each rank’s resident memory before and after from_pretrained(), then check the process-group initialization order and package versions to see whether memory use scales with the number of processes. The public evidence currently consists mainly of a single user report, which is not enough to establish the impact on other versions, quantized models, or multi-node deployments. Follow-up should track whether upstream confirms the FSDP2 behavior in this code path and whether a fix can keep weights empty until synchronization is complete.

Sources

  1. FSDP2 CPU RAM efficient loading 問題回報 #49297
  2. Accelerate:Fully Sharded Data Parallel 官方文件