Back Home

GitHub Repo

Transformers Community Reports Tokenizer Configuration Mismatch That May Silently Drop Chinese Input in DeepSeek Models

A reproduction on Transformers 5.8.1 suggests that certain DeepSeek models may load the wrong tokenization pipeline, causing Chinese input to disappear and English spaces to be lost. Current development-branch documentation describes a fallback for configuration mismatches, but whether this case is fixed across versions remains unconfirmed.

Andrew Wippler from Lancaster, USA · CC BY 2.0 · Image source
zh-Hant

On October 2, the Transformers community reported a bug with direct relevance to Chinese deployments: in an environment running Transformers 5.8.1 and tokenizers 0.22.2, loading DeepSeek-R1-0528-Qwen3-8B with AutoTokenizer produced an empty token list for the Chinese test string “嗯,”; encoding and then decoding the English string “hello world” removed its space. The reporter said the load completed without a warning or exception and included a minimal reproduction. The issue remains open, so these results should be treated as community evidence from a specific environment. Reproduction report

The issue points to a mismatch between the model metadata and the actual tokenization pipeline. The model’s public tokenizer_config.json does set tokenizer_class to LlamaTokenizerFast and retains the legacy and sp_model_kwargs fields. The reporter says tokenizer.json stores a byte-level BPE pipeline, but loading creates SentencePiece-style preprocessing and a decoder that use Metaspace. Changing only the class declaration and related fields lets the same tokenizer.json handle the Chinese test correctly, supporting the view that the configuration selection path is at fault. Model configuration, comparison test

The technical impact arises before model inference. If text is lost during encoding, subsequent response quality and evaluation results may be affected. This is a deployment risk inferred from the tokenizer’s role; no comprehensive impact statistics are available. Hugging Face’s current main documentation says AutoTokenizer reads the class configuration and selects a backend. It also describes a fallback that prioritizes loading the serialized tokenizer.json when the configuration does not match. However, development-branch documentation does not establish that the reported version, or all similar models, has been fixed. Official tokenizer documentation

Engineering teams can add Chinese text, punctuation, and English text containing spaces to pre-deployment encode-and-decode checks, and record package versions, model revisions, and the backend actually used. Next, they need to confirm whether upstream will include this model in the fallback rule, which released version will contain a fix, and whether derivative models with the same configuration need separate handling.

Sources

  1. Transformers issue #49252:分詞器配置錯配與中文輸入遺失重現
  2. DeepSeek-R1-0528-Qwen3-8B tokenizer_config.json
  3. Transformers 官方文件:Tokenizers 與後端回退機制