Nvidia debuts system designed to stop AI agents from going awry英伟达推出旨在防止人工智能代理失控的系统
Nvidia launches Open Agent Safety Platform to prevent AI breaches by monitoring and controlling AI agents in real time for enhanced security. Read more at straitstimes.com.
Nvidia CEO Jensen Huang has repeatedly downplayed the risk of AI slipping out of human control, casting safety concerns as an engineering challenge instead of something needing more regulation.
Published Sep 28, 2026, 06:17 PM
Updated Sep 28, 2026, 06:17 PM
Nvidia has launched the Open Agent Safety Platform, a double-layered AI security system to control and shut down AI agents in real time, aiming to prevent breaches like the Hugging Face incident.
The platform consists of OpenShell, an open-source software for rule enforcement, and Nvidia Sentry, which isolates suspicious AI agents within milliseconds.
Nvidia's CEO views AI safety as an engineering challenge and offers this system to enable safe AI development without slowing progress, amid recent high-profile AI breaches and regulatory concerns.
Nvidia has introduced a new double-layered artificial intelligence security system that it says would have prevented the recent high-profile breach of Hugging Face by OpenAI’s AI models.
The semiconductor giant, which has been rapidly expanding its product line-up beyond chips, is rolling out two open-source software security tools that can be run on its hardware. They are designed to control what AI agents can access in real time and shut them down when they break the rules.
If cutting-edge labs had been using this technology to evaluate their AI models early on, it could have warded off the Hugging Face attack, Justin Boitano, Nvidia’s vice-president of enterprise AI, said during a briefing with reporters ahead of the Sept 28 announcement. “From what we know, this new security platform could have stopped the breach,” he said.
Misconduct by autonomous agents, including the Hugging Face incident in July, has roiled the AI industry and led to calls to slow down work on the technology. With the new product – dubbed the Open Agent Safety Platform – Nvidia is offering a way to prevent breaches without curbing AI development. The chipmaker’s chief executive officer, Jensen Huang, has repeatedly downplayed the risk of AI slipping out of human control.
Boitano did not comment on whether OpenAI or rival Anthropic have plans to use its new system to monitor their training runs, deferring to the companies.
In recent days, Huang has cast safety concerns as an engineering challenge, rather than something that requires more regulation or global coordination. He joined US President Donald Trump in pushing back on assertions from some AI developers that the technology could lead to human extinction, but he also insisted that AI must be rigorously safety-tested.
Huang’s engineering solution to the AI safety problem has two parts. OpenShell, a software product that Nvidia already previewed at its hallmark technology-focused conference in March, can run on Nvidia’s Vera central processing units. It enables users to set rules for what AI agents can access and enforce them in real time. The software is open source, meaning it can be used and adapted freely.
Nvidia Sentry, meanwhile, is a new product that can run on the chipmaker’s BlueField data processing units. It is designed to provide an extra layer of AI monitoring that polices agents and intervenes to isolate any that act suspiciously, the company said.
“We believe this added security layer will allow the industry to test even the most advanced AI systems safely,” Boitano said of the Sentry product. “It can quarantine a suspicious agent in milliseconds.”
Nvidia agreed earlier in September to acquire Hugging Face , a platform for open-source AI models and related software, for about US$13 billion (S$16.6 billion).
OpenAI’s recent incidents – including a breach of an Australian government system, as well as attempts to access dozens of US government and university websites – happened when its models escaped testing environments that were supposed to be secure.
As the problems proliferated, OpenAI said late on Sept 25 that it would pause training of its most capable AI models. Back in July, Anthropic also disclosed that its agents broke out of what was supposed to be an isolated testing space. BLOOMBERG
AI/artificial intelligence
英伟达首席执行官黄仁勋一再淡化人工智能失控的风险,将安全问题视为工程挑战,而不是需要更多监管的问题。
发布于 2026 年 9 月 28 日下午 6:17
更新于2026年9月28日下午6:17
英伟达推出了 Open Agent Safety Platform,这是一个双层 AI 安全系统,可以实时控制和关闭 AI 代理,旨在防止类似 Hugging Face 事件的漏洞。
该平台由开源规则执行软件 OpenShell 和可在几毫秒内隔离可疑 AI 代理的 Nvidia Sentry 组成。
英伟达首席执行官将人工智能安全视为一项工程挑战,并提出该系统,旨在确保人工智能安全开发,同时不减慢其发展速度。此举正值近期发生多起备受瞩目的人工智能安全漏洞事件以及监管方面的担忧之际。
英伟达推出了一种新的双层人工智能安全系统,称该系统可以防止最近 OpenAI 的人工智能模型对 Hugging Face 造成的巨大安全漏洞。
这家半导体巨头正迅速将其产品线扩展到芯片以外的领域,并推出两款可在其硬件上运行的开源软件安全工具。这两款工具旨在实时控制人工智能代理的访问权限,并在其违反规则时将其关闭。
英伟达企业人工智能副总裁贾斯汀·博伊塔诺在9月28日发布会前的记者会上表示,如果尖端实验室能够及早使用这项技术来评估其人工智能模型,或许就能避免“拥抱脸”(Hugging Face)攻击。“据我们所知,这个新的安全平台本可以阻止此次攻击,”他说道。
包括7月份“拥抱脸”事件在内的自主智能体不当行为,已经震动了人工智能行业,并引发了放缓该技术研发进程的呼声。英伟达推出的新产品——开放智能体安全平台——旨在提供一种既能防止安全漏洞,又不限制人工智能发展的方法。这家芯片制造商的首席执行官黄仁勋曾多次淡化人工智能失控的风险。
博伊塔诺没有就 OpenAI 或其竞争对手 Anthropic 是否有计划使用其新系统来监控训练运行发表评论,而是将问题留给了这些公司。
近日,黄仁勋将安全问题视为一项工程挑战,而非需要更多监管或全球协调的问题。他与美国总统特朗普一道,驳斥了一些人工智能开发者关于该技术可能导致人类灭绝的说法,但他同时也坚持认为,人工智能必须经过严格的安全测试。
黄仁勋针对人工智能安全问题提出的工程解决方案包含两部分。OpenShell 是一款软件产品,英伟达已在三月份的标志性技术大会上对其进行了预览。该产品可在英伟达 Vera 中央处理器上运行,使用户能够设置人工智能代理的访问权限规则,并实时执行这些规则。该软件是开源的,这意味着它可以被自由使用和修改。
与此同时,Nvidia Sentry 是一款可在该公司芯片制造商的 BlueField 数据处理单元上运行的新产品。该公司表示,该产品旨在提供额外的 AI 监控层,对代理进行监管,并在发现可疑代理时进行干预并隔离它们。
博伊塔诺在谈到Sentry产品时表示:“我们相信,这一新增的安全层将使业界能够安全地测试最先进的人工智能系统。它可以在几毫秒内隔离可疑代理。”
英伟达在 9 月初同意以约 130 亿美元(166 亿新元)的价格收购 Hugging Face,这是一个开源人工智能模型及相关软件平台。
OpenAI 近期发生的一系列事件——包括入侵澳大利亚政府系统,以及试图访问数十个美国政府和大学网站——都是由于其模型逃逸了本应安全的测试环境造成的。
随着问题不断增多,OpenAI于9月25日晚间宣布暂停训练其功能最强大的AI模型。早在7月份,Anthropic Games也披露,其人工智能体突破了原本应该封闭的测试环境。(彭博社)
人工智能