模型安全評估
Anthropic Updates RSP 3.4: Redefining the Automated R&D Threshold and Revising Risk-Reporting Mechanisms
Anthropic’s Responsible Scaling Policy 3.4 changes the capability threshold for automated R&D, the internal distribution of unredacted reports, and the external review process. Although the update may appear governance-focused, it will affect engineering workflows for model capability evaluations, deployment gates, and safety reporting.

Anthropic’s Responsible Scaling Policy 3.4, which took effect on July 8, revises the trigger threshold for automated R&D capabilities to better align evaluations with the threat models the company is concerned about. The new version also changes the internal distribution rules for unredacted risk reports: instead of requiring access for all employees with general permissions, it requires that the reports be made available to at least 200 employees. Public reports must indicate where redactions were made. For external review, multiple reviewers may separately assess different sections, provided that every unredacted portion is reviewed by at least one person.
For technical teams, these provisions mean that the question of whether a model can automate AI R&D must be translated into reproducible tests rather than answered through vague impressions. Evaluations need to cover extended autonomous operation, proposing and validating experiments, modifying training code, using compute clusters, and recovering from failures. Testing only single-turn coding problems may underestimate the model’s overall capabilities. Allowing risk reports to use a coverage date can prevent new developments from forcing rushed revisions to the analysis, but it may also create a time lag between public documentation and the model actually deployed.
Observers should examine whether the revised thresholds correspond to clearly defined quantitative metrics and what verifiable safeguards will be implemented once they are triggered. The Future of Life Institute’s safety index, released around the same time, continues to raise concerns about the transparency and enforceability of several companies’ governance frameworks, reminding readers not to treat policy language itself as proof of safety. RSP 3.4 is a voluntary corporate framework, not an independent standard. Unless its evaluation data, red-team coverage, and decisions on exceptions are disclosed sufficiently, external researchers will still find it difficult to verify whether deployment gates are being enforced consistently.