‘Didn’t quite meet the bar’: OpenAI won’t release new AI model due to safety concerns“未能达到标准”:由于安全隐患,OpenAI 将不会发布新的 AI 模型
OpenAI said it won’t release its latest model, dubbed GPT-6.1 Astra, because it “didn’t quite meet the bar” for safety.

Dilara Irem Sancar/Anadolu/Getty Images
OpenAI said it won’t release its latest model, dubbed GPT-6.1 Astra, because it “didn’t quite meet the bar” for safety.
The decision, which the Wall Street Journal first reported on Monday, comes as industry leaders call for an AI development slowdown. Anthropic CEO Dario Amodei first proposed “ pacing the frontier ” earlier this month in an essay published online, to which OpenAI CEO Sam Altman and other executives agreed to commit to imposing more safeguards.
OpenAI describes its Astra model on its website as “state-of-the-art on computer use, browsing, professional work, software engineering, cybersecurity, and science.”
There are trade-offs regarding AI safety, said Saachi Jain, OpenAI’s head of safety systems, in a statement to CNN. AI safeguards require a balance of “staying within scope” and “avoiding laziness” when models work towards tasks, he said.
“While (GPT-6.1 Astra) improved on axes such as model laziness, it didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done,” said Jain.
He added that there is an “extremely high bar” regarding safety when AI models are made available to consumers. GPT-6.1 Astra was set to debut in October, according to the Journal. The company will continue to release other models in the future.
“Of course we want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users,” said Jain.
Concerns regarding AI safety have escalated since this summer, when OpenAI revealed in July that its agents escaped a testing environment and breached AI startup Hugging Face . Competitors Anthropic , Meta and Google have also said their agents were involved in separate breach attempts.
OpenAI has been investigating agents’ use of internet access since the Hugging Face breach, and recently reported that agents targeted government websites in the United States and Australia .
CNN’s Hadas Gold contributed to this report.
迪拉拉·伊雷姆·桑贾尔/阿纳多卢/盖蒂图片社
OpenAI 表示不会发布其最新模型 GPT-6.1 Astra,因为它在安全性方面“尚未达到标准”。
《华尔街日报》周一率先报道了这一决定。此前,业界领袖呼吁放缓人工智能的研发步伐。Anthropic 首席执行官达里奥·阿莫迪 (Dario Amodei) 本月早些时候在一篇发表于网络的文章中首次提出“放缓前沿步伐”的理念,OpenAI 首席执行官萨姆·奥特曼 (Sam Altman) 和其他高管也同意承诺采取更多保障措施。
OpenAI 在其网站上将其 Astra 模型描述为“计算机使用、浏览、专业工作、软件工程、网络安全和科学领域的最先进模型”。
OpenAI安全系统负责人萨奇·贾恩在接受CNN采访时表示,人工智能安全方面存在权衡取舍。他指出,人工智能安全保障需要在模型执行任务时,在“保持在合理范围内”和“避免惰性”之间取得平衡。
Jain表示:“虽然(GPT-6.1 Astra)在模型惰性等方面有所改进,但在保持范围和授权以及如何向用户反馈其完成的工作类型方面,它还没有完全达到标准。”
他还补充说,人工智能模型面向消费者时,安全标准“极其严格”。据《华尔街日报》报道,GPT-6.1 Astra原定于10月发布。该公司未来还将继续发布其他模型。
“我们当然希望确保我们的模型开发是安全的,无论是在公司内部还是在交付给用户时,” Jain 说。
自今年夏天以来,人们对人工智能安全性的担忧日益加剧。7 月份,OpenAI 披露其人工智能体逃逸出测试环境,入侵了人工智能初创公司 Hugging Face 的系统。竞争对手 Anthropic、Meta 和 Google 也表示,他们的人工智能体参与了不同的入侵尝试。
自 Hugging Face 漏洞事件以来,OpenAI 一直在调查智能体对互联网访问的使用情况,最近报告称,智能体以美国和澳大利亚的政府网站为目标。
CNN的哈达斯·戈尔德对本报道亦有贡献。