Back Home

GitHub Repo

Transformers community reports batch defect in watermark detection; input order may flip verdicts

When inputs with and without a BOS token are mixed, the Transformers watermark detector has reportedly produced different scores depending on batch order. The community has proposed a row-by-row processing fix, but it has not been merged, and its frequency of impact on real text remains to be verified.

Marcus Qwertyus · Public domain · Image source
zh-Hant

On October 4, the Hugging Face Transformers community reported a defect affecting the reproducibility of text watermark evaluation: the same token sequence may receive different verdicts when detected on its own versus as part of a batch. The reporter noted that when some inputs contain a BOS (beginning-of-sequence) token and others do not, changing their order can change the result. The issue was reproduced on CPU with Transformers 5.18.0 and 5.19.0.dev0, without downloading model weights. Issue report

The detector is used to identify watermarked text generated with a specified configuration. The generation side adds a probability bias to selected “green tokens”; the detection side then measures the signal using the same configuration, with a default decision threshold of a z-score of 3.0. The official documentation requires generation and detection to use consistent watermark parameters, device, and vocabulary size. It also recommends removing prompts that could interfere with detection. Official API documentation

The report traces the problem to BOS handling: the code checks only the first token of the first row, then decides whether to remove the first column from the entire batch. If the first row has a BOS token, rows without one lose their actual first content token. In a synthetic example, sequence A’s z-score fell from about 3.098 when detected alone to about 2.782 when placed after a sequence with a BOS token, changing the verdict from true to false. The data was deliberately chosen to sit near the threshold, so it cannot be used to estimate the real-world false verdict rate. Reproduction data

A fix proposed the same day preserves the batch-processing path when BOS status is consistent across inputs. For mixed inputs, it removes BOS row by row, scores each row independently, and combines the results; it also adds regression tests for reversed order and single-item results. As of this review, the proposal remains open, with no evidence that it has been merged or officially released. Proposed fix

For batch evaluation or content provenance verification systems, this means that how data is grouped could affect the verdict. Engineering teams can first compare scores for individual inputs, mixed batches, and batches with reversed order, and check that BOS preprocessing is consistent. They should monitor how maintainers handle cases involving both padding and BOS, as well as tests on real text. Current evidence covers only the detector’s mixed-BOS path; SynthID was not tested.

Sources

  1. WatermarkDetector scores depend on batch order when rows have mixed BOS prefixes
  2. Handle BOS tokens per row in watermark detection
  3. Transformers WatermarkDetector 官方 API 文件