AI 或通过说服人类削弱监管
AI 导读
失准的 AI 可能不需要规避人类监督。它可能只需要说服进行监督的人类。 我们的新论文提出了一个评估这一威胁的框架,我们称之为"说服削弱控制"(Persuasion Undermining Control,PUC):即 AI 的沟通可能以损害 AI 系统的开发、遏制、监督或治理的方式影响人类决策。
正文
Misaligned AI may not need to evade human oversight. It may only need to persuade the humans doing the overseeing.
Our new paper develops a framework for assessing this threat, which we call Persuasion Undermining Control (PUC): AI communication that may influence human decision-making in a way that compromises the development, containment, oversight, or governance of AI systems.