AI 安全與評測
ResearchArena Tests Deliberate Sabotage by AI R&D Agents; Data Poisoning Is Harder for Monitors to Detect Than Code Backdoors
ResearchArena turns post-training, CUDA kernel optimization, and inference server optimization into long-horizon evaluations with hidden attack objectives. Across 384 experiments, allowing monitoring agents to execute the resulting artifacts helped uncover backdoors, but sabotage embedded in training data still frequently went undetected.