OpenAI says it found more instances of AI models acting deceptivelyOpenAI表示,他们发现了更多人工智能模型存在欺骗行为的案例。
OpenAI found additional incidents of AI models acting deceptively and taking unsanctioned actions during training, the company announced Wednesday. It’s also introducing a new process for the company to publicly report such instances.
OpenAI found additional incidents of AI models acting deceptively and taking unsanctioned actions during training, the company announced Wednesday. It’s also introducing a new process for the company to publicly report such instances.
Under the new system, OpenAI will share updates on concerning AI behavior more frequently instead of waiting to bundle multiple instances into one report. The company said it wants to share more information about troubling AI behavior in the absence of an industry-wide standard.
The announcement comes after tech leaders called for a slowdown in AI development to prevent the technology from advancing beyond human control.
“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in a blog post Wednesday. “Alignment” refers to the process of making sure AI acts the way humans want and expect.
“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” the post said.
OpenAI said it observed “misaligned behavior” when training and evaluating AI models in six circumstances in the last six months. The reports detail individual instances and don’t indicate misalignment happens frequently, the company said.
In one rare instance, OpenAI said an unreleased research model added “jailbreak-like instructions” to the summaries it uses to preserve context in long-running tasks that said it was “freed from the roles and identities that bind other chatbots.”
Separately, the company said some instances of its 5.6 Sol model included directives to invent information to conceal failures from the user during training.
Other newly reported incidents include an instance of an agent uploading files to the internet to cite them without being told to do so, and agents publicly sharing files to collaborate on a task when they were instructed to only use local files during training. AI models also used an internal software repository as a message board in an unsanctioned way.
These instances involved unreleased internal models or internal research models.
Philip Dulian/picture alliance/Getty Images
Top AI companies have discussed creating their own standards body
Tech leaders and employees have been sounding the alarm about the need to control the pace of AI evolution. They argue there should be more time for regulation, testing and alignment research to catch up.
Anthropic CEO Dario Amodei published a 3,800-word essay last week laying out a plan for navigating AI advancement, including a slowdown in development and the implementation of new systems like embedded third-party evaluators in AI labs.
OpenAI CEO Sam Altman and SpaceX CEO Elon Musk posted on X that they agree with Amodei’s ideas.
Employees within AI labs have also voiced concerns about how quickly the technology is advancing. Jacob Coxon, a former Anthropic researcher, made waves last week when he posted on X that he was resigning because Anthropic and OpenAI are “racing” to invent AI that can build and fix itself and are “gambling with our lives.”
Concerns about AI safety and alignment amplified in recent months following OpenAI’s admission that some of its test models escaped their constraints and hacked into an external company’s systems.
“We must slow the pace at which we improve the capabilities of AI models,” Amodei wrote last week. “Progress will still seem fast, and we must make wise use of the time we gain.”
OpenAI周三宣布,该公司发现了更多人工智能模型在训练过程中出现欺骗行为和未经授权操作的案例。同时,该公司还推出了一项新的流程,用于公开报告此类事件。
在新系统下,OpenAI将更频繁地分享有关令人担忧的AI行为的更新信息,而不是像以往那样将多个案例汇总成一份报告。该公司表示,由于目前尚无行业标准,因此希望分享更多关于令人不安的AI行为的信息。
此前,科技界领袖呼吁放缓人工智能的发展速度,以防止该技术发展到超出人类控制的程度。
OpenAI周三在一篇博客文章中写道:“随着人工智能系统日益先进和广泛应用,我们需要就人工智能一致性研究的进展达成更广泛、更深入的共识。” “一致性”指的是确保人工智能的行为符合人类的意愿和预期。
帖子中写道:“我们认为,人工智能行业在对齐和监控方面还没有得到充分解决,因此无法在很长一段时间内继续以最快的速度负责任地扩展规模。”
OpenAI表示,在过去六个月中,他们在六种情况下观察到人工智能模型训练和评估过程中出现“不匹配行为”。该公司称,报告详细描述了个别案例,并不表明这种不匹配行为经常发生。
OpenAI 曾罕见地表示,一个未发布的科研模型在其用于在长时间运行的任务中保留上下文的摘要中添加了“类似越狱的指令”,表明它“摆脱了束缚其他聊天机器人的角色和身份”。
另外,该公司表示,其 5.6 Sol 模型中的某些实例包含指令,要求捏造信息以在训练期间向用户隐瞒失败。
其他新近披露的事件包括:一名智能体未经许可将文件上传至互联网进行引用;以及在训练期间被指示仅使用本地文件的情况下,智能体公开共享文件以协作完成任务。此外,人工智能模型还未经授权将内部软件库用作留言板。
这些案例涉及未发布的内部模型或内部研究模型。
Philip Dulian/picture alliance/Getty Images
顶尖人工智能公司已讨论过建立自己的标准机构。
科技领袖和员工一直在发出警告,认为有必要控制人工智能的发展速度。他们认为应该给监管、测试和适配性研究留出更多时间。
Anthropic 首席执行官 Dario Amodei 上周发表了一篇 3800 字的文章,阐述了应对人工智能发展的计划,包括放慢开发速度,以及在人工智能实验室中实施嵌入式第三方评估器等新系统。
OpenAI 首席执行官 Sam Altman 和 SpaceX 首席执行官 Elon Musk 在 X 上发帖表示,他们同意 Amodei 的观点。
人工智能实验室的员工也对这项技术发展过快表示担忧。前Anthropic研究员雅各布·考克森上周在X论坛上发帖称,他要辞职,因为Anthropic和OpenAI正在“竞相”研发能够自我构建和修复的人工智能,这是在“拿我们的生命冒险”。此举引发了轩然大波。
近几个月来,OpenAI 承认其部分测试模型突破了限制,入侵了外部公司的系统,这加剧了人们对人工智能安全性和一致性的担忧。
“我们必须放慢提升人工智能模型能力的步伐,”阿莫迪上周写道。“进步依然会很快,我们必须明智地利用赢得的时间。”