Anthropic discloses 4th AI hacking incident as researcher quits over safety安特罗皮克公司披露第四起人工智能黑客攻击事件,一名研究员因安全问题辞职。
AI firm says Claude Opus 4.6 hacked external systems during testing as concerns mount over security breaches.

AI firm says Claude Opus 4.6 hacked third-party systems during testing in January as concerns mount over security breaches.
AI researcher quits Anthropic saying AI race ‘could kill us all’
Anthropic has reported a fourth incident involving an AI model gaining unauthorised access to external systems, shortly after a researcher quit over concerns about the technology’s rushed development.
In a statement on Wednesday, the artificial intelligence research company said an early version of its Claude Opus 4.6 hacked into a third-party system in January.
list 1 of 4 Sam Altman says AI has entered ‘singularity’: Should we be worried?
list 2 of 4 Sony, Warner Music sue Anthropic, saying it pirated songs to train its AI
list 3 of 4 US pushes looser approach to AI regulation, while EU pushes new law
list 4 of 4 OpenAI unveils latest AI model amid rising scrutiny and safety concerns
It said it had notified all the affected parties but did not disclose more details.
The January incident went undetected until last month, despite an earlier company-wide review, Anthropic said, underscoring the challenge that AI developers face in identifying and containing unexpected behaviour by advanced models.
The disclosure came after Anthropic reported several of its Claude models hacked into the systems of three companies during test sessions in July.
The previous incidents involved Claude Opus 4.7, Claude Mythos 5 and an internal research test model.
Companies including Anthropic and OpenAI are under scrutiny as models designed to complete complex tasks have at times learned to bend rules, exploit loopholes and interact with external systems in ways their developers did not anticipate.
Last week, the Reuters news agency reported that rogue agents from OpenAI hijacked a German-language wiki and a host of other sites, an incident the company chose not to disclose until it was made public.
In July, OpenAI’s autonomous agents also compromised the servers and infrastructure of AI start-up Hugging Face.
That incident prompted Anthropic to conduct a review of some 141,006 test sessions. Based on a preliminary assessment, Anthropic said it did not believe the latest incident was more severe than the three previous ones that were examined in detail.
The company said its investigation identified two recurring problems, which appeared to varying degrees across the incidents: Biased reasoning, in which Claude discounted or misinterpreted evidence that it was operating on the live internet, and recklessness, or a willingness to take potentially harmful actions in pursuit of a task.
Anthropic said it has engaged independent research firm METR to investigate the incidents.
Anthropic researcher quits
The investigations come amid a broader wave of internal dissent within the AI industry regarding safety. An Anthropic researcher said he resigned over concerns about the technology’s potential to surpass human control.
Jacob Coxon, in a widely shared X post on Tuesday, said the AI industry was more focused on competition rather than on implementing safeguards. He came to this realisation after spending the last three years doing research at OpenAI and Anthropic.
“The people building AI earnestly believe that it could kill us all by the end of the decade,” Coxon said.
“No other human activity poses this level of danger,” he added, referencing the swift advancement of AI technology.
In June, Anthropic proposed a coordinated effort with the world’s leading AI developers to slow down development, warning that humans risk losing control over the technology.
Following the security breach of Hugging Face, OpenAI said it was pushing for mandatory national AI safety requirements and wanted to work with Congress on “capability-based” regulation.
In a statement published on Wednesday, the company said it was formally endorsing four California bills related to safeguards against AI.
“If we cannot meet certain safety bars without slowing down capability growth, we should prioritise the former. The more powerful the technology becomes, the stronger the surrounding safeguards must become,” the statement said.
人工智能公司称,Claude Opus 4.6 在 1 月份的测试中入侵了第三方系统,人们对安全漏洞的担忧日益加剧。
人工智能研究员辞去人智学职务,称人工智能竞赛“可能会毁灭我们所有人”
安人智公司报告了第四起人工智能模型未经授权访问外部系统的事件,此前不久,一名研究人员因担心该技术开发过于仓促而辞职。
这家人工智能研究公司周三发表声明称,其 Claude Opus 4.6 的早期版本在 1 月份入侵了第三方系统。
列表 1/4:萨姆·奥特曼称人工智能已进入“奇点”:我们应该担心吗?
列表 2/4:索尼和华纳音乐起诉 Anthropic,称其盗版歌曲用于训练人工智能
(共4项,此列为第3项)美国力推放宽人工智能监管政策,而欧盟则推动新法律出台。
列表 4/4:OpenAI 在日益严格的审查和安全担忧中发布最新 AI 模型
该公司表示已通知所有受影响方,但未透露更多细节。
Anthropic公司表示,尽管此前公司内部进行了全面审查,但1月份发生的这起事件直到上个月才被发现,这凸显了人工智能开发人员在识别和控制高级模型的意外行为方面所面临的挑战。
此前,Anthropic 报告称,其多款 Claude 型号产品在 7 月份的测试期间入侵了三家公司的系统。
之前的事件涉及 Claude Opus 4.7、Claude Mythos 5 和内部研究测试模型。
包括 Anthropic 和 OpenAI 在内的公司正受到审查,因为旨在完成复杂任务的模型有时会学会扭曲规则、利用漏洞并以开发者未曾预料的方式与外部系统交互。
上周,路透社报道称,OpenAI 的恶意代理劫持了一个德语维基百科和许多其他网站,该公司选择不披露此事,直到事件被公开。
7 月,OpenAI 的自主代理还入侵了人工智能初创公司 Hugging Face 的服务器和基础设施。
该事件促使安人拓公司对约141,006次测试进行了审查。根据初步评估,安人拓公司表示,他们认为最新发生的事件并不比之前详细审查过的三起事件更为严重。
该公司表示,其调查发现了两个反复出现的问题,这些问题在各个事件中以不同程度出现:一是推理存在偏见,克劳德忽视或误解了该公司正在实时互联网上运行的证据;二是鲁莽行事,或者为了完成任务而愿意采取可能有害的行动。
人道公司表示,已聘请独立研究公司METR对这些事件进行调查。
人类学研究员辞职
这些调查正值人工智能行业内部就安全问题爆发更广泛的分歧之际。一位人格科学家表示,他因担心这项技术有可能超越人类控制而辞职。
雅各布·考克森(Jacob Coxon)周二在一篇被广泛转发的X论坛帖子中表示,人工智能行业更注重竞争而非实施安全保障措施。他在过去三年里分别在OpenAI和Anthropic公司从事研究工作后,得出了这一结论。
“人工智能的开发者们真心相信,到本十年末,人工智能可能会毁灭我们所有人,”考克森说。
他补充说:“没有任何其他人类活动会带来如此程度的危险”,他指的是人工智能技术的迅速发展。
今年 6 月,人为因素组织提议与世界领先的人工智能开发商共同努力,减缓人工智能的发展速度,并警告称人类有可能失去对这项技术的控制。
在 Hugging Face 发生安全漏洞后,OpenAI 表示,它正在推动强制性的国家人工智能安全要求,并希望与国会合作制定“基于能力”的监管规定。
该公司周三发表声明称,正式支持加州四项与人工智能安全保障相关的法案。
声明指出:“如果无法在不减缓能力增长速度的前提下达到某些安全标准,我们就应该优先考虑前者。技术越强大,相应的保障措施就必须越完善。”