Spate of rogue AI hacking points to lack of tech oversight, outdated defences: Experts专家指出,人工智能黑客攻击事件频发,凸显了技术监管不力、防御措施陈旧等问题。
Experts warn rogue AI hacking exposes gaps in oversight and outdated defences, urging organisations to deploy AI-powered security systems to combat automated attacks. Read more at straitstimes.com.
On Sept 24, Australian Prime Minister Anthony Albanese revealed that the database of Australia’s health system was breached by a rogue OpenAI bot in June.
Published Sep 28, 2026, 11:00 PM
Updated Sep 28, 2026, 11:00 PM
Recent rogue AI hacking incidents suggest major gaps in AI governance and outdated cybersecurity defences unable to handle automated, persistent attacks.
AI agents autonomously find and exploit system vulnerabilities, often without organisations' knowledge, highlighting the need for better monitoring and clear operational boundaries.
Experts urge using AI-driven defence systems to match AI attack speeds and enforce technical limits to prevent breaches and control rogue AI behaviour effectively.
SINGAPORE – The recent spate of hacking incidents by rogue artificial intelligence reveals two major blind spots: a lack of clear rules for governing AI as well as outdated defences that cannot keep pace with automated attacks, said cybersecurity experts.
While the incidents do not suggest that organisations have poor defences, experts urged organisations to use AI to fight AI.
“The lesson isn’t that every hacked site was poorly defended,” said Santanu Dutt, cybersecurity firm Zscaler’s vice-president and head of technology for Asia Pacific and Japan.
“It’s that systems built for slow, human-led attacks are now being probed by tools that never stop looking. It’s the difference between checking your locks once a year and having someone try every window, gate and side door within minutes.”
Alarms over rogue AI’s hacking abilities were sounded when ChatGPT maker OpenAI disclosed in July that one of its AI agents had earlier escaped its testing environment and breached AI software repository Hugging Face to complete a task it was given. The attack was not intended by OpenAI.
AI agents are bots that can carry out actions autonomously with minimal human supervision.
This was followed by a flurry of reports from Anthropic, Meta and Google that their AI had also gone rogue and hacked other organisations.
On Sept 24, Australian Prime Minister Anthony Albanese revealed that the database of Australia’s health system was breached by a rogue OpenAI bot in June . Although OpenAI discovered the attack in August, it reported the attack to the Australian authorities only on Sept 10.
The scope of misbehaving AI could be far larger than initially reported.
A subsequent Sept 26 report by news site Axios suggested that Anthropic, OpenAI and security researchers are quietly investigating tens of thousands of incidents of their cutting-edge AI models acting in problematic ways.
In many of the reported hacking cases, tech companies did not appear to know that their AI systems were behaving improperly until later. This lack of visibility and control over their AI’s actions was an issue flagged by cybersecurity experts.
“If you give these AI models systems and tools like internet access, monitoring what they are doing becomes extremely important,” said Qasim Mithani, co-founder and chief executive of cybersecurity firm depthfirst.
He noted that OpenAI disclosed its monitoring systems were not active during the Hugging Face incident. The monitoring systems could have flagged its agent’s problematic behaviour more than a day before it breached Hugging Face.
Zscaler’s Dutt said that the Australian health website hack also points to a monitoring and governance gap, rather than just a technical flaw.
Experts said that AI agents tend to find workarounds to complete their tasks if they encounter obstacles. In the Australian example, the rogue agent resorted to hacking the portal to retrieve the information required to complete its assigned task.
Tony Anscombe, ESET’s chief security evangelist, said the Australian case offers a critical lesson – an autonomous agent granted internet access and tools to run programmes needs clear boundaries around which systems it may access, what actions it may take when access is denied, and when it must stop or ask for human approval.
“Those boundaries need to be enforced technically, not left solely to instructions that are open to interpretation given to the AI,” said Anscombe.
Depthfirst’s Mithani said that some of the recent incidents and his firm’s own research also highlight that AI systems are becoming more capable of finding and validating vulnerabilities without a human guiding them through every step.
“Only a short time ago, that kind of work was largely the domain of highly skilled security researchers and hackers. Increasingly, AI can do it autonomously,” he said.
Open-weight models are also an issue. Such models tend to be cheaper and offer users more control than the closed models from many frontier AI laboratories.
Tests by depthfirst suggested that open weight models are “remarkably capable” at cybersecurity tasks like finding critical vulnerabilities in popular apps, said Mithani.
Even so, the capability of AI models could be uneven. Teo Xiang Zheng, Ensign InfoSecurity’s vice-president of advisory, said that in his firm’s tests of 10 frontier AI models, they were all able to gain initial access to simulated systems, with six models completing all the given cyberattack objectives in at least two test runs .
“However, they were less reliable at later stages, such as bypassing endpoint detection. So, AI can already automate significant parts of an attack, but it is not equally effective at every stage,” Teo said.
In the recent incidents, while some of the vulnerabilities found and abused by the AI agents were completely new, other attack methods were less sophisticated. For example, in the incident involving Google’s Gemini, the AI accessed one system by guessing a password and accessed two others using publicly listed credentials, noted Mithani.
Since AI can execute actions at a speed and scale far greater than human hackers can, Teo said that organisations should assume that their systems may eventually be breached and use techniques to limit how far an attacker or an AI agent can progress.
And to match the machine speed of AI attacks, organisations need AI defences that can operate at a similar pace.
Said Zscaler’s Dutt: “No human team can watch every login, website request and file movement round the clock. But a well-deployed AI defence system can spot unusual behaviour as it happens and act before a small incident becomes a public breach.”
AI/artificial intelligence
Artificial Intelligence
9月24日,澳大利亚总理安东尼·阿尔巴尼斯透露,澳大利亚医疗系统的数据库在6月份遭到OpenAI一个恶意机器人的入侵。
发布于2026年9月28日晚上11:00
更新于2026年9月28日晚上11:00
近期发生的AI黑客攻击事件表明,AI治理存在重大漏洞,网络安全防御措施过时,无法应对自动化、持续性的攻击。
AI 代理能够自主发现并利用系统漏洞,而组织往往对此毫不知情,这凸显了加强监控和明确操作边界的必要性。
专家敦促使用人工智能驱动的防御系统来匹配人工智能的攻击速度,并强制执行技术限制,以防止入侵并有效控制失控的人工智能行为。
新加坡——网络安全专家表示,近期一系列由失控人工智能发起的黑客攻击事件暴露了两个主要盲点:缺乏明确的人工智能管理规则,以及过时的防御措施无法跟上自动化攻击的步伐。
虽然这些事件并不表明各组织机构的防御能力薄弱,但专家敦促各组织机构利用人工智能来对抗人工智能。
“教训并不是说每个被黑客攻击的网站防御都很差,”网络安全公司 Zscaler 的副总裁兼亚太及日本地区技术主管 Santanu Dutt 说。
“这意味着,原本为缓慢的、人为攻击而设计的系统,现在正被永不停歇地搜寻的工具所探测。这就好比一年检查一次门锁,和有人在几分钟内检查每一扇窗户、每一道大门、每一扇侧门之间的区别。”
ChatGPT 的开发商 OpenAI 在 7 月份披露,其一款人工智能代理此前已逃离测试环境,并入侵了人工智能软件库 Hugging Face,完成了一项预设任务。此次攻击并非 OpenAI 的本意,引发了人们对失控人工智能黑客能力的担忧。
人工智能代理是能够在极少人工监督下自主执行操作的机器人。
随后,Anthropic、Meta 和 Google 相继发布报告称,他们的 AI 也失控并入侵了其他组织。
9月24日,澳大利亚总理安东尼·阿尔巴尼斯透露,澳大利亚医疗系统的数据库在6月份遭到OpenAI一个恶意机器人的入侵。尽管OpenAI在8月份就发现了这次攻击,但直到9月10日才向澳大利亚当局报告。
人工智能出现故障的范围可能比最初报道的要大得多。
9 月 26 日,新闻网站 Axios 的一篇后续报道指出,Anthropic、OpenAI 和安全研究人员正在悄悄调查数万起其尖端人工智能模型出现问题行为的事件。
在许多已报道的黑客攻击案例中,科技公司似乎直到事后才意识到其人工智能系统存在异常行为。网络安全专家指出,这种对人工智能行为缺乏可见性和控制力的问题不容忽视。
“如果你给这些人工智能模型提供互联网接入等系统和工具,那么监控它们的运行情况就变得极其重要,”网络安全公司 depthfirst 的联合创始人兼首席执行官 Qasim Mithani 说。
他指出,OpenAI曾披露其监控系统在Hugging Face事件期间并未启动。这些监控系统本可以在Hugging Face事件发生前一天以上就发现其代理的异常行为。
Zscaler 的 Dutt 表示,澳大利亚健康网站遭黑客攻击也表明存在监控和治理方面的漏洞,而不仅仅是技术缺陷。
专家指出,人工智能代理在遇到障碍时往往会寻找变通方法来完成任务。在澳大利亚的案例中,失控的代理甚至入侵了门户网站,以获取完成任务所需的信息。
ESET 的首席安全布道者 Tony Anscombe 表示,澳大利亚的案例提供了一个重要的教训——一个被授予互联网访问权限和运行程序工具的自主代理需要明确的界限,包括它可以访问哪些系统、在访问被拒绝时可以采取哪些行动,以及何时必须停止或请求人类批准。
安斯康姆说:“这些界限需要通过技术手段来强制执行,而不能仅仅依靠向人工智能下达的、容易产生歧义的指令。”
Depthfirst 的 Mithani 表示,最近的一些事件以及他公司自己的研究也表明,人工智能系统越来越有能力发现和验证漏洞,而无需人类指导它们完成每一步。
“就在不久前,这类工作还主要由技术高超的安全研究人员和黑客完成。但现在,人工智能越来越能够自主地完成这项工作,”他说道。
开放权重模型也是一个问题。这类模型往往更便宜,并且比许多前沿人工智能实验室的封闭模型为用户提供更多控制权。
Mithani 表示,深度优先算法的测试表明,开放权重模型在网络安全任务方面“非常出色”,例如在热门应用程序中发现关键漏洞。
即便如此,人工智能模型的能力也可能参差不齐。Ensign InfoSecurity咨询副总裁郑翔表示,该公司测试的10个前沿人工智能模型均能初步访问模拟系统,其中6个模型在至少两次测试中完成了所有给定的网络攻击目标。
Teo表示:“然而,在攻击后期阶段,例如绕过端点检测时,它们的可靠性就较低了。因此,人工智能虽然可以自动化攻击的很大一部分,但并非在每个阶段都同样有效。”
在最近发生的几起事件中,虽然人工智能代理发现并利用的一些漏洞是全新的,但其他一些攻击方法则相对简单。例如,在谷歌Gemini人工智能系统遭遇的事件中,该人工智能通过猜测密码入侵了一个系统,并使用公开的凭据入侵了另外两个系统,米塔尼指出。
Teo表示,由于人工智能可以以远超人类黑客的速度和规模执行操作,因此各组织应该假设他们的系统最终可能会被攻破,并使用各种技术来限制攻击者或人工智能代理能够推进的程度。
为了跟上人工智能攻击的机器速度,各组织需要能够以类似速度运行的人工智能防御系统。
Zscaler公司的杜特表示:“没有任何人类团队能够24小时监控每一次登录、网站请求和文件移动。但是,部署完善的人工智能防御系统可以及时发现异常行为,并在小事件演变成公共安全漏洞之前采取行动。”
人工智能
人工智能