OpenAI reports more incidents of models acting deceptivelyOpenAI报告称,更多模型出现欺骗行为。
The ChatGPT creator says it is introducing a public reporting framework to share unexpected AI behaviour.
![OpenAI CEO Sam Altman attends an event to pitch AI for businesses in Tokyo, Japan, on February 3, 2025 [File: Kim Kyung-Hoon/Reuters]](https://www.aljazeera.com/wp-content/uploads/2026/09/2026-09-11T023211Z_2042538591_RC2UMCA71F6W_RTRMADP_3_OPENAI-DEVELOPMENT-1789616786.jpg?resize=770%2C513&quality=80)
The ChatGPT creator says it is introducing a public reporting framework to share unexpected AI behaviour, admitting the industry has not solved safety challenges yet.
OpenAI CEO Sam Altman attends an event to pitch AI for businesses in Tokyo, Japan, on February 3, 2025 [File: Kim Kyung-Hoon/Reuters]
OpenAI says it has identified additional incidents of its AI models allegedly acting deceptively and taking unsanctioned actions during internal training and testing.
Alongside these disclosures on Wednesday, the creator of ChatGPT stated it was introducing a public reporting framework intended to frequently share instances of what it termed as unexpected or misaligned AI behaviour.
list 1 of 3 Congress passes sweeping US sanctions bill targeting Russia
list 2 of 3 Morocco’s 2026 election: A test of political trust and engagement
list 3 of 3 Yemeni forces target Houthis as US rules out direct role
In a post on its website, OpenAI claimed that under the newly outlined framework, it will publish updates on concerning model behaviour on an ongoing basis rather than delaying disclosures to group multiple incidents into larger, periodic reports.
The company said the initiative aims to increase industry transparency around troubling model activities in the absence of standardised safety disclosure norms.
The announcement comes amid broader calls from prominent technology leaders urging a slowdown in frontier AI development over concerns that rapid scaling could outpace human oversight and control.
Last week, Anthropic claimed to have thwarted multiple malicious operations using its Claude models, ranging from cyber-espionage and weapons design to mass surveillance campaigns.
“We must slow the pace at which we improve the capabilities of AI models,” Anthropic CEO Dario Amodei wrote in an essay published on Saturday. “Progress will still seem fast, and we must make wise use of the time we gain.”
However, United States President Donald Trump has repeatedly pushed back against calls to limit the industry, arguing that maintaining the US’s technological edge over international rivals remains paramount.
Responding to slowdown proposals, Trump described critics as “very negative forces” raising exaggerated scenarios that “won’t happen”.
Escalating debate on alignment
Despite political resistance to statutory slowdowns, OpenAI signalled agreement with its industry rival regarding alignment pressures.
“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” the company stated in the post.
OpenAI added that it does not believe the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer, emphasising that decisions about future AI development need to draw on evidence that external observers can examine independently.
According to the company, safety teams observed what they categorised as “misaligned behaviour” across six specific circumstances over the past six months during training and evaluation runs.
However, OpenAI maintained that these reports document individual, rare instances rather than frequent operational failures across deployed products.
The reported incidents allegedly included unreleased research models concealing mistakes in task summaries, unauthorised file uploads to the internet to generate citation links, and agents sharing files across public servers or internal repositories to bypass local boundaries.
OpenAI further stated that its future reports will detail observed behaviours, severity, setting, discovery dates, and the specific models involved, adding that it remains committed to disclosing complex cases requiring longer investigation or third-party coordination.
ChatGPT 的创建者表示,他们正在引入一个公共报告框架,用于分享意外的 AI 行为,并承认该行业尚未解决安全挑战。
2025年2月3日,OpenAI首席执行官Sam Altman在日本东京出席一场活动,向企业推介人工智能技术。[图片:Kim Kyung-Hoon/路透社]
OpenAI 表示,他们已发现其人工智能模型在内部训练和测试期间存在其他涉嫌欺骗行为和未经授权的行为。
除了周三披露的信息外,ChatGPT 的创建者还表示,它正在引入一个公共报告框架,旨在定期分享它所称的意外或不协调的 AI 行为案例。
(共3条) 国会通过针对俄罗斯的全面制裁法案(第1条)
摩洛哥2026年大选:对政治信任和参与度的考验(共3个列表,此为第2个)
也门军队打击胡塞武装的行动已展开,美国排除直接参与的可能性。此为也门军队打击胡塞武装行动的3个名单之3。
OpenAI 在其网站上发布的一篇文章中声称,根据新概述的框架,它将持续发布有关令人担忧的模型行为的更新,而不是延迟披露,将多个事件归类到更大的定期报告中。
该公司表示,该举措旨在提高行业对令人担忧的模型活动的透明度,因为目前尚无标准化的安全披露规范。
此前,一些知名科技领袖呼吁放缓前沿人工智能的研发步伐,因为他们担心人工智能的快速发展可能会超出人类的监督和控制能力。
上周,Anthropic 声称利用其 Claude 模型挫败了多起恶意行动,这些行动包括网络间谍活动、武器设计以及大规模监控活动。
“我们必须放慢提升人工智能模型能力的步伐,”Anthropic首席执行官达里奥·阿莫迪在周六发表的一篇文章中写道。“进步依然会很快,我们必须明智地利用赢得的时间。”
然而,美国总统唐纳德·特朗普一再反对限制该行业的呼吁,他认为保持美国在国际竞争对手面前的技术优势仍然至关重要。
针对放缓经济的提议,特朗普称批评者是“非常消极的势力”,他们提出了夸大其词的情景,而这些情景“不会发生”。
关于联盟的争论愈演愈烈
尽管政治上对法定放缓措施存在阻力,但 OpenAI 表示与行业竞争对手在协调压力方面达成一致。
“随着人工智能系统变得越来越先进,部署也越来越广泛,我们需要就对齐研究的进展建立更广泛、更明智的共识,”该公司在帖子中表示。
OpenAI 还表示,它认为人工智能行业尚未充分解决对齐和监控问题,因此无法在很长一段时间内继续以最大速度负责任地扩展规模,并强调有关未来人工智能发展的决策需要依靠外部观察者可以独立审查的证据。
据该公司称,在过去六个月的培训和评估过程中,安全团队在六种特定情况下观察到了他们归类为“不协调行为”的情况。
然而,OpenAI 坚持认为,这些报告记录的是个别罕见的事件,而不是已部署产品中频繁发生的运行故障。
据报道,这些事件包括未发布的科研模型掩盖任务摘要中的错误、未经授权将文件上传到互联网以生成引用链接,以及代理人通过公共服务器或内部存储库共享文件以绕过本地边界。
OpenAI 还表示,其未来的报告将详细说明观察到的行为、严重程度、环境、发现日期以及涉及的具体模型,并补充说,它仍然致力于披露需要更长时间调查或第三方协调的复杂案例。