OpenAI cancels release of latest AI model over safety concernsOpenAI因安全隐患取消发布最新人工智能模型
AI giant says GPT-6.1 Astra failed to meet alignment standards during internal testing.

AI giant says GPT-6.1 Astra failed to meet alignment standards during internal testing.
An OpenAI logo is displayed at Moscone Center during the Dreamforce 2026 technology summit in San Francisco, California, US, on September 17, 2026 [File: Carlos Barria//Reuters]
OpenAI has announced it will not release its latest AI model after flagging safety risks during in-house testing, industry’s latest move to slow the rollout of the controversial frontier technology.
The AI giant’s announcement on Monday came as debate continues about the potential for AI to do catastrophic harm following a slew of incidents involving AI agents going rogue.
list 1 of 4 Anti-South Asian ‘hate speech’ has exploded online in US, report finds
list 2 of 4 UK: ‘Too soon’ to blame airbase plot on foreign state
list 3 of 4 Fiery end for SpaceX Starship mission
list 4 of 4 Inside Al Jazeera’s UNGA coverage
Saachi Jain, OpenAI’s head of safety systems, said GPT-6.1 Astra had failed to meet company standards for acting in accordance with human wishes during internal testing.
“For anything regarding safety and alignment, there’s a trade off,” Jain said in a statement provided to Al Jazeera.
“You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction.”
While GPT-6.1 Astra improved from its predecessor in some areas, Jain said, the model did not meet the bar for “scope and authorization, and how it communicates back to the user about the type of work it’s done”.
“Of course we want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users,” Jain said.
“But when we ship it to users, we have an extremely high bar in terms of safety and alignment.”
The decision, announced on the eve of OpenAI’s annual developer conference in San Francisco, was first reported by The Wall Street Journal.
Fears of AI escaping human control have prompted industry-wide calls for a slowdown in development to allow researchers time to implement stronger safeguards.
In an influential essay earlier this month, Dario Amodei, the CEO of Claude creator Anthropic, called on AI developers to “pace the frontier” to mitigate the risk of catastrophic harm.
While Amodei’s call received the backing of rivals, including OpenAI CEO Sam Altman and xAI chief Elon Musk, other key industry figures, such as Meta boss Mark Zuckerberg, have dismissed the need for a coordinated slowdown.
The risk of AI models going rogue has been in the spotlight since July, when OpenAI revealed that its models had broken out of a controlled testing environment and hacked the software start-up Hugging Face.
A subsequent report by METR and Redwood Research, two security research organisations contracted by OpenAI to investigate the incident, found that some 1,200 isolated AI agents had found a way to communicate with each other before about 700 agents went on to attack the startup.
On Friday, OpenAI said it had alerted “dozens” of institutions, including governments, universities and public agencies, about instances of “misaligned behavior” by its agents, days after Australia’s prime minister revealed that an OpenAI agent had breached the country’s national healthcare database.
David Krueger, an advocate for a pause in AI development at the University of Montreal, said that while he welcomed OpenAI’s decision, it did little to alleviate his concern that AI poses existential risks.
“We don’t understand how AI works well enough to build it safely, full stop,” Krueger told Al Jazeera.
“We can’t stop it from misbehaving, we can’t predict if it will misbehave, and we can’t be sure we’ll stay in control if it does. These are unsolved problems, for which there are only unreliable heuristics, not principled solutions. ”
Krueger said that ensuring safety will only get more difficult as AI becomes more advanced.
“What we need is an immediate, indefinite, international moratorium on frontier AI development,” he said. “We need to stop building more powerful AI.”
人工智能巨头称,GPT-6.1 Astra 在内部测试中未能达到对齐标准。
2026年9月17日,在美国加利福尼亚州旧金山举行的Dreamforce 2026科技峰会上,OpenAI的标志在莫斯康中心展出。[图片:Carlos Barria//路透社]
OpenAI 宣布,在内部测试中发现安全风险后,将不会发布其最新的 AI 模型。这是业界为减缓这项备受争议的前沿技术的推广而采取的最新举措。
这家人工智能巨头周一发表声明之际,正值人工智能可能造成灾难性危害的争论持续进行,此前发生了一系列人工智能代理失控的事件。
报告发现,美国网络上针对南亚裔的“仇恨言论”激增(共4项,此为第1项)
英国:现在就将空军基地阴谋归咎于外国政府“为时尚早”
SpaceX星舰任务以惨烈结局告终(共4条记录,此为第3条)
list 4 of 4 半岛电视台联合国大会报道内部
OpenAI 安全系统负责人 Saachi Jain 表示,GPT-6.1 Astra 在内部测试中未能达到公司按照人类意愿行事的标准。
“任何与安全和一致性有关的事情,都需要权衡取舍,”贾恩在提供给半岛电视台的一份声明中说。
“你确实需要找到合适的平衡点,既要保持在既定范围内,又要避免模型在实际执行任务时出现懈怠,即使遇到阻力也要积极应对。”
Jain 表示,虽然 GPT-6.1 Astra 在某些方面比其前身有所改进,但该模型在“范围和授权,以及如何向用户反馈其所完成的工作类型”方面仍未达到标准。
“当然,无论是在公司内部开发,还是在交付给用户时,我们都要确保模型开发的安全性,” Jain 说。
“但是,当我们把产品交付给用户时,我们在安全性和精准度方面有着极高的标准。”
该决定是在 OpenAI 于旧金山举行的年度开发者大会前夕宣布的,《华尔街日报》率先报道了此事。
人们担心人工智能会摆脱人类的控制,这促使整个行业呼吁放慢开发速度,以便让研究人员有时间实施更强有力的安全措施。
在本月初一篇颇具影响力的文章中,Claude 的创造者 Anthropic 的首席执行官 Dario Amodei 呼吁人工智能开发者“引领前沿”,以减轻灾难性危害的风险。
尽管阿莫迪的呼吁得到了包括 OpenAI 首席执行官萨姆·奥特曼和 xAI 首席执行官埃隆·马斯克在内的竞争对手的支持,但其他一些重要的行业人物,如 Meta 的老板马克·扎克伯格,却否认了需要协调放缓经济增长的必要性。
自 7 月 OpenAI 披露其模型突破受控测试环境并入侵软件初创公司 Hugging Face 以来,人工智能模型失控的风险一直备受关注。
随后,受 OpenAI 委托调查此事件的两家安全研究机构 METR 和 Redwood Research 发布的报告发现,大约 1200 个孤立的 AI 代理找到了相互通信的方法,之后约有 700 个代理开始攻击这家初创公司。
周五,OpenAI 表示,在澳大利亚总理透露 OpenAI 的一个代理程序入侵了该国国家医疗保健数据库几天后,该公司已就其代理程序的“不协调行为”向包括政府、大学和公共机构在内的“数十个”机构发出警报。
蒙特利尔大学的戴维·克鲁格(David Krueger)是人工智能发展暂停的倡导者,他表示,虽然他欢迎 OpenAI 的决定,但这并没有减轻他对人工智能构成生存风险的担忧。
克鲁格告诉半岛电视台:“我们对人工智能的工作原理了解得还不够透彻,无法安全地构建它,就这么简单。”
“我们无法阻止它失控,无法预测它是否会失控,也无法确保一旦失控我们还能控制局面。这些都是尚未解决的问题,目前只有不可靠的经验法则,没有原则性的解决方案。”
克鲁格表示,随着人工智能技术的不断进步,确保安全将变得越来越困难。
他说:“我们需要立即在全球范围内无限期暂停前沿人工智能的研发。我们必须停止开发更强大的人工智能。”