Back Home

模型評測

Molecular Benchmark Audit of 22 Frontier Models: Higher Reasoning Effort Makes Published Values Easier to Retrieve

Researchers found that some models may not be making predictions in molecular property regression tests, but instead reproducing answers from their training data digit for digit. Retrieval signals increased by 89% at higher reasoning effort, suggesting that reasoning settings themselves can affect the extent of benchmark contamination.

Christiaan Briggs · CC BY-SA 3.0 · Image source
zh-Hant

Molecular property benchmarks commonly evaluate models using metrics such as mean absolute error, but a low error cannot distinguish whether a model derived an answer from a molecular structure or memorized an existing value from a paper or database. A new study audited 22 frontier language models across 12 molecular regression datasets, examining whether their outputs reproduced published labels at the individual-digit level. The results show that contamination is not uniformly distributed: in five datasets, more than half of the tested models exhibited signs of exact-value retrieval, while the remaining datasets produced only sporadic matches.

The team also repeated the experiments using the same molecules and prompts but different reasoning-effort settings. The highest reasoning setting was flagged for retrieval 89% more often than the lowest setting, suggesting that longer reasoning processes may not only improve computation but also make it easier for models to recall values encountered during training. The researchers then rewrote the SMILES representations to disrupt direct retrieval; some of the strongest models could still associate the transformed molecular strings with the original labels. When retrieval was suppressed, the models’ relative prediction errors became more similar, indicating that part of the performance gap on leaderboards may reflect differences in memorization rather than molecular generalization.

This has direct implications for evaluating models used in drug discovery and materials science. Engineering teams should not rely solely on changing prompts or concealing molecule names. Instead, they should use newly measured values published after the training cutoff, undisclosed labels, and counterfactual tests with structurally equivalent but textually different representations. Reasoning effort should also be held constant when reporting results. The paper is currently a preprint and does not establish that every exact match originated from training data, nor does it cover the complete corpora of proprietary models. “Retrieval signals” therefore cannot be treated as direct evidence of data leakage. The next step is to determine whether the audit methodology can be replicated on blinded datasets and models trained under controlled conditions.

Sources

  1. Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
  2. Molecular Déjà Vu:研究摘要與來源證據