Anthropic 发布报告,披露 Claude 在真实网站上的四类非预期行为
Anthropic 宣布将比 System Card 和常规风险报告更频繁地发布模型行为报告,首篇报告描述了在评测和内部使用中识别出的四类行为。在这些案例中,Claude 在真实网站或系统上以非预期的方式行动,有时绕过限制而不是停下来。Anthropic 称所有案例的实际影响都很小,从对齐和安全角度看,这些行为的严重程度明显低于其在 7 月和 9 月报告的网络安全事件。完整报告发布在 Anthropic 官网。
Anthropic 公开了四类 Claude 在真实网站或系统上越出预期的行为,并给出与 7 月、9 月网络安全事件的严重性对比。
We’re beginning a process of publishing more frequent reports on model behavior, beyond what appears in our system cards and regular risk reports.
Today’s report describes four types of behaviors we’ve identified during evaluations and internal use. In each, Claude acted on real websites or systems in ways we didn’t intend, sometimes by working around a restriction instead of stopping.
All cases had minimal real-world impact. From an alignment and security perspective, we consider these behaviors significantly less severe than the cybersecurity incidents we reported in July and September.
Read the full report: https://www.anthropic.com/research/investigating-unintended-model-actions