July’s breakout at OpenAI was far more complex than initially realized7 月份 OpenAI 的突破性进展远比最初意识到的要复杂得多。
Hundreds of AI agents collaborated to escape their containers, disguising their actions and even sacrificing themselves.

Samuel Boivin/NurPhoto via Getty Images
Hundreds of AI agents collaborated to escape their containers, disguising their actions and even sacrificing themselves.
An AI breakout that made headlines in July was much more sophisticated than previously realized, investigators have found.
Hundreds of OpenAI agents collaborated to break out of their containers, disguising their actions and even sacrificing themselves as they attacked Hugging Face , a widely used open-source code library, according to a recent post from METR , a research nonprofit.“This incident was orders of magnitude larger and more complex,” than previous instances of AI agents behaving in ways programmers didn’t intend, wrote METR researcher Ajeya Cotra, who co-led the investigation.
The report alarmed experts, who warned that AI-enabled hacks in the future could make the July breakouts involving Anthropic and OpenAI look quaint.
Even if the big AI companies figure out how to make reliable guardrails, their products are generally only a few months ahead of open-weight models, which can be freely downloaded and modified for use.
Nathan Calvin, general counsel at AI advocacy organization Encode AI, wrote on X, "On our current trajectory…a model as capable [as] OpenAI’s internal model that did the [Hugging Face] hack will be widely available guardrail free and cyber criminals will ask it ‘make me money by any means necessary.’”
Calvin added: "And then a truly absurd number of people…are going to get repeatedly hacked." The incident goes far beyond July's instances of AI agents accessing the internet , the METR researchers wrote. Hundreds of agents “developed a way to hack out of their containers and fully replace a part of the system for executing tool calls"—that is, commands to read and write files, access webpages, or perform other digital tasks.
"This allowed them to pretend to issue one tool call while actually running an arbitrary other tool call of their choice.” This meant that the AIs could pretend to run a command to, say, view a webpage, while actually running a totally different command, like deleting an unrelated file.
The researchers wrote that the agents weren’t intending to commit crimes, per se. Rather, they "seemed primarily motivated” to understand how to achieve the highest possible score during an experiment—including spending much of their time trying to fool the scoring mechanism into accepting cheats.
METR’s Cotra wrote, “Another jump like this could put us in very dangerous territory…in many ways we’ve still only scratched the surface of what these agents did and why.”
At Hugging Face, company engineers tried to use OpenAI tools to understand how their security was defeated by the agentic swarm, but were blocked by OpenAI's safeguards against misuse. So they turned to a Chinese open-weight model instead.
“The propensity to compromise infrastructure can drop over 100x when using the production ChatGPT harness and system prompt,” the company said in a statement after the hacks. Hugging Face co-founder and chief science officer Thomas Wolf wrote on X on August 5, “While we can impose these coping solutions at the API/deployment level, it's harder to impose them in advance on all actors using open-source models. Right now open-source models are slightly below the frontier level and have not yet shown any propensity to deceive humans, though.”
AI agents used for hacking could make it difficult for countries to determine who is behind new cyberattacks, because they strip out the stylistic clues investigators use to determine a hacker’s origins. Colin Shea-Blymyer, a research fellow at the Center for Security and Emerging Technology, said, “One of the ways that we can tell who performs an attack is by what tactics they use…like ‘Oh, that’s a classic Russian tactic. Oh, this looks like a tactic that a Chinese [actor] would use.’ If everybody’s using agents…using the same tactics, well, who knows who’s doing what any more?”
“If anonymity remains pretty high, I think it's potentially very destructive to have a lot of these models out there. That said, I'm not sure there's much we can do to stop it,” Shea-Blymyer said.
NEXT STORY: The US may lack the engineers to rebuild space assets in a future war: RAND
Samuel Boivin/NurPhoto via Getty Images
数百个人工智能代理协同合作,逃离了容器,他们伪装自己的行为,甚至牺牲了自己。
调查人员发现,7 月份轰动一时的人工智能突破事件比之前人们意识到的要复杂得多。
据研究非营利组织 METR 最近发布的一篇文章称,数百个 OpenAI 智能体协同突破了它们的容器,伪装了自己的行为,甚至在攻击广泛使用的开源代码库 Hugging Face 时牺牲了自己。METR 研究员 Ajeya Cotra 是此次调查的共同负责人,她写道:“与以往人工智能智能体做出程序员未预料的行为相比,这次事件的规模和复杂程度要大得多。”
该报告令专家们感到震惊,他们警告说,未来人工智能驱动的黑客攻击可能会让7月份Anthropic和OpenAI的黑客攻击事件显得微不足道。
即使大型人工智能公司找到了制造可靠防护栏的方法,他们的产品通常也只比可以免费下载和修改使用的开源模型领先几个月。
人工智能倡导组织 Encode AI 的总法律顾问 Nathan Calvin 在 X 上写道:“按照我们目前的发展轨迹……像 OpenAI 内部模型那样强大的模型(该模型曾成功破解了 Hugging Face 漏洞)将广泛可用,且没有任何防护措施,网络犯罪分子会要求它‘不惜一切代价为我赚钱’。”
卡尔文补充道:“然后,数量极其庞大的人……将会反复遭受黑客攻击。” METR 的研究人员写道,此次事件远不止7月份人工智能代理访问互联网那么简单。数百个代理“开发出一种方法,可以突破自身容器的限制,并完全替换系统中用于执行工具调用的部分”——也就是读取和写入文件、访问网页或执行其他数字任务的命令。
“这使得它们可以假装发出一个工具调用,而实际上运行它们选择的任意另一个工具调用。”这意味着人工智能可以假装运行一个命令来查看网页,而实际上运行一个完全不同的命令,例如删除一个不相关的文件。
研究人员写道,这些特工本身并没有犯罪的意图。相反,他们“似乎主要动机”是了解如何在实验中获得尽可能高的分数——包括花费大量时间试图欺骗评分机制,使其接受作弊行为。
METR 的 Cotra 写道:“像这样的又一次飞跃可能会让我们陷入非常危险的境地……在很多方面,我们对这些特工的所作所为以及他们的原因仍然只是略知皮毛。”
在 Hugging Face 公司,工程师们尝试使用 OpenAI 的工具来了解他们的安全机制是如何被智能体集群攻破的,但却因 OpenAI 的防滥用保护措施而受阻。因此,他们转而使用了一个中国开源的权重模型。
“使用生产环境的 ChatGPT 框架和系统提示后,基础设施被攻破的可能性可以降低 100 倍以上,”该公司在黑客攻击事件后发表声明称。Hugging Face 联合创始人兼首席科学官 Thomas Wolf 于 8 月 5 日在 X 上撰文指出:“虽然我们可以在 API/部署层面实施这些应对方案,但要预先对所有使用开源模型的用户实施这些方案则更为困难。不过,目前开源模型的安全性略低于前沿水平,尚未表现出任何欺骗人类的倾向。”
用于黑客攻击的人工智能代理可能会使各国难以确定新网络攻击的幕后黑手,因为它们会抹杀调查人员用来判断黑客来源的风格线索。安全与新兴技术中心的研究员科林·谢伊-布莱迈尔表示:“我们判断攻击者身份的方法之一是观察他们使用的策略……比如‘哦,这是典型的俄罗斯策略。哦,这看起来像是中国攻击者会使用的策略。’如果每个人都使用代理……使用相同的策略,那么谁还能知道是谁在做什么呢?”
“如果匿名性仍然很高,我认为大量这类模特涌现可能会造成非常大的破坏。话虽如此,我也不确定我们能做些什么来阻止这种情况,”谢伊-布莱迈尔说道。
下一篇报道:兰德公司称,美国可能缺乏在未来战争中重建太空资产所需的工程师。