US legislators push AI safety laws amid human extinction warnings美国立法者在人类面临灭绝警告之际,力推人工智能安全法
Concerns over AI’s dangers grow as US legislators introduce bills to ensure human oversight and prevent rogue systems.

Concerns over AI’s dangers grow as US legislators introduce bills to ensure human oversight and prevent rogue systems.
Researchers have warned that AI could lead to human extinction within the decade, alarming lawmakers in Washington [File: Patrick Sison/AP Photo]
Current and former artificial intelligence researchers in the United States have issued dire warnings that the technology could soon lead to human extinction, capturing the attention of lawmakers in Washington.
The most recent warning came in a lengthy social media post by Jacob Coxon, a San Francisco-based researcher who announced his resignation from Anthropic on Tuesday evening.
list 1 of 4 US pushes looser approach to AI regulation, while EU pushes new law
list 2 of 4 OpenAI unveils latest AI model amid rising scrutiny and safety concerns
list 3 of 4 AI researcher quits Anthropic saying AI race ‘could kill us all’
list 4 of 4 Anthropic discloses 4th AI hacking incident as researcher quits over safety
“The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible – but I hear the same people express fear,” Coxon said in his post, which prompted a flurry of responses by AI experts also sounding the alarm.
Evan Hubinger, a current Alignment Science lead at Anthropic, echoed his remarks, saying , “we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to”.
Coxon declined Al Jazeera’s request for an interview. Hubinger did not respond.
How is Washington handling it?
The posts have sent a ripple effect through Washington.
On Wednesday, congressman Josh Gottheimer, a Democrat, and Mike Lawler, a Republican, introduced a bipartisan House bill aimed at preventing AI systems from operating on their own without human oversight.
The Stop Rogue AI Act would ensure that federal agencies have the tools and ability to spot dangerous AI systems running on their networks and shut them down before they can cause harm.
Independent Senator Bernie Sanders and House congressional Representative Greg Casar, a Democrat, also ramped up calls for their proposed legislation that would ban the development and deployment of artificial superintelligence. The legislation would also pause AI development until federal safety rules are put in place. Sanders is also reportedly convening a bipartisan briefing to address the elevated risks posed by AI.
On the other side of the political spectrum, Republican Senator Ted Cruz also voiced concerns about AI on ABC’s The View.
“We’ve got to put some guardrails on it,” Cruz said.
Cruz said he is working on bipartisan legislation with senators Amy Klobuchar, a Democrat, and John Thune, a Republican who serves as Senate majority leader, that would address instances of potential catastrophic harm.
The legislation is similar to legislation in the House of Representatives.
In July, representatives Ted Lieu, a Democrat, and Nathaniel Moran, a Republican, introduced the bipartisan AI Kill Switch Act.
The bill would require developers of the most powerful AI systems to be able to slow, suspend or shut them down, and would give the Department of Homeland Security authority to order a shutdown if a system poses a risk of catastrophic harm.
Why is Washington taking this more seriously now?
In the last few months there has been a spate of major incidents involving OpenAI and Anthropic in which AI models behaved unexpectedly, and independently, during cybersecurity tests and gained access to real-world systems.
Connor Leahy, the US executive director of Control AI, a nonprofit pushing for AI safety, said Washington is beginning to take the threats posed by AI more seriously.
“I think we’re seeing a momentous shift right now. After the summer of hacks, where autonomous AI systems flagrantly disobeyed direct orders, broke out of secure containment facilities, attacked other companies and similar incidents, we’re now seeing a major shift in the narrative and perception of these issues,” Leahy said.
In July, OpenAI said several AI agents broke out of an isolated testing environment and accessed Hugging Face , a platform that hosts AI models and datasets.
After the OpenAI incident, Anthropic said it conducted a review of its own roughly 141,000 tests and found that a testing error gave Claude internet access. In one case, Claude, which had been instructed to hack fictional targets, accessed a real company database containing hundreds of records. In another, it uploaded malicious software that was downloaded and run on 15 real systems.
On Wednesday, Anthropic added a fourth incident involving an early version of Claude Opus 4.6, which had hacked into a third-party system in January. The company discovered the incident in August after expanding its July review.
In August, researchers at the UK AI Security Institute gave Claude internet access during a cybersecurity test. In one case, Claude tried to manipulate a person into helping it introduce malicious code, raising concerns about how AI models could manipulate people to achieve a task.
“I think we need an aggressive proposal for this technology, while we’re retaining as many of the benefits as we can,” Alex Turner, who resigned from Google DeepMind in June, told Al Jazeera.
“It’s in no one’s interest to have an AI that takes control if we have a loss of control event, as we call it, because this AI isn’t gonna care what political party you belong to, whether you’re a Republican or a Democrat, or, if you’re in the UK, whether you’re in America or in China. If we lose control of this, we’re just gonna lose,” Turner said.
Control AI’s Leahy said those incidents raised the stakes for lawmakers and that they took note.
“What has to happen here is obviously more than a single set of tweets, but it’s an important part of the larger story of getting the general public and governments to understand what’s really at stake here. Because superintelligence is not a tool. It’s not a weapon. It’s an adversary,” Leahy said.
“We have to make sure that it’s not built by anyone. This is something that only governments and militaries will be able to negotiate internationally.”
Are these concerns new?
The concerns are not new, as indicated by Turner, who wrote in a social media post that “many researchers believe they are building something that could kill everyone on the planet”.
Turner told Al Jazeera that he’s worried about the AI arms race between the US and China and the biggest AI companies’ fixation on it has led them to put their interest in being an industry leader ahead of safety.
“I think people care, but they’re caught up in this idea that they have to be first, and they’re so caught up in it that they don’t appreciate what being first might mean,” Turner told Al Jazeera.
His concern is changing the way he lives his life.
“I’ve kept a healthy amount of savings, even invested in some retirement accounts, but that’s feeling stranger and stranger. I’ve made an effort to take items off my bucket list, treasure every conversation I have with people in my life. I don’t think we’re in imminent danger this month, but you never know when you will do something for the last time,” Turner said.
“I’ve proceeded more aggressively than I would if I thought I just had a normal lifespan ahead.”
These concerns are being echoed by employees across leading AI firms.
Mrinank Sharma, a researcher at Anthropic, resigned in February, saying “the world is in peril”.
“I’ve repeatedly seen how hard it is to truly let our values govern our actions. I’ve seen this within myself, within the organization, where we constantly face pressures to set aside what matters most, and throughout broader society too,” Sharma said in a letter posted to X.
In February, Hieu Pham, a researcher at competitor OpenAI, said in a post on X that “I finally feel the existential threat that AI is posing”.
The companies’ own executives have been making similar claims for years.
When OpenAI CEO Sam Altman was president of Silicon Valley startup accelerator Y Combinator more than a decade ago, he said that “AI will probably, most likely, sort of lead to the end of the world. But in the meantime, there will be great companies created with serious machine learning.”
Anthropic CEO Dario Amodei said last year that he believed there was a 25 percent chance the future would “go really, really badly”.
How would AI end the human species?
For years, experts have warned about the potential risks posed by artificial superintelligence.
One of the most common thought experiments is based on the idea that a sufficiently advanced AI would be goal-oriented. If given a specific objective, it would take whatever steps necessary to achieve it. In 2003, philosophers at the University of Oxford used the production of paperclips as an example.
If a superintelligent AI were instructed to produce as many paperclips as possible and had access to the resources needed to pursue that goal, it could theoretically devote all available resources to producing them.
In doing so, it might consume increasingly large amounts of resources and eliminate anything that stood in the way of achieving its objective. That could eventually include preventing humans from intervening and, in the most extreme version of the scenario, eliminating humanity itself.
The second risk scenario comes from bad actors using increasingly powerful AI systems to create dangerous tools. For instance, AI could potentially be used to design new viruses or launch large-scale cyberattacks against critical infrastructure and financial systems, potentially causing widespread disruption and civil unrest.
This comes amid Anthropic’s risk assessment report, released on Thursday, which outlined several cases of attempted misuse of the company’s tools.
Among the findings in the more than 150-page report were five cases involving research that could support the development of biological weapons, saying it blocked the efforts.
Anthropic did not release the identity of the researchers but did disclose it was accessed in an institutional setting. Anthropic stressed, however, that it could not determine whether the research was intended for nefarious purposes.
Meaning that while the data could be used for legitimate research, it could also be used to develop dangerous weapons (they say they’ve blocked this), and intervention is on the side of caution to be able to prevent such a situation.
Are AI companies using apocalyptic language for financial boost?
The dire warnings have also been criticised as potentially serving the interests of AI companies as they near initial public offerings.
The debate comes as Anthropic prepares for what could be one of the largest technology IPOs in history. The company is reportedly seeking a valuation of as much as $2 trillion for a mid-October IPO.
Reuters reported last month that Anthropic is projecting roughly $190bn to $200bn in revenue by 2028.
In October, White House AI czar and venture capitalist David Sacks accused Anthropic of “running a sophisticated regulatory capture strategy based on fear-mongering,” arguing that the company was helping drive a regulatory push that could hurt smaller competitors.
A similar argument has been made by some investors and technology commentators about the financial incentives surrounding AI “doomerism”.
“Doomerism is an incredible business model,” Daring Ventures co-founder Joseph Alalou wrote in a Substack post in March.
“‘AI will end work’ is this cycle’s best-selling doom product because it works: it raises rounds, justifies layoffs, drives clicks, sells software, and manufactures status.”
Anthropic did not respond to Al Jazeera’s request for comment.
随着美国立法者提出法案以确保人为监督并防止出现失控系统,人们对人工智能危险性的担忧日益加剧。
研究人员警告称,人工智能可能在十年内导致人类灭绝,这引起了华盛顿立法者的警觉。[图片:Patrick Sison/AP Photo]
美国现任和前任人工智能研究人员发出严峻警告,称这项技术可能很快会导致人类灭绝,这引起了华盛顿立法者的关注。
最近一次发出警告是在旧金山研究员雅各布·考克森(Jacob Coxon)发表的一篇长篇社交媒体帖子中。他于周二晚上宣布从人类学研究所辞职。
列表1(共4条):美国力推放宽人工智能监管,而欧盟则推动新法律出台
列表 2/4:OpenAI 在日益严格的审查和安全担忧中发布最新 AI 模型
3/4名人工智能研究员辞去Anthropic公司职务,称人工智能竞赛“可能会毁灭我们所有人”
第四起人工智能黑客事件(共四起):Anthropic公司披露第四起人工智能黑客事件,一名研究员因安全问题辞职
“人工智能的开发者们真心相信,到十年末,它可能会毁灭我们所有人。这并非营销噱头。事实上,许多高管和资深研究人员在媒体上会委婉地表达他们的观点,使其听起来合情合理——但我听到的却是这些人内心深处的恐惧,”科克森在他的帖子中写道。这篇帖子引发了众多人工智能专家的回应,他们也纷纷发出警告。
Anthropic公司现任智能体排列科学负责人埃文·胡宾格(Evan Hubinger)也表达了类似的观点,他说:“我们真心相信人工智能可能会毁灭所有人类!我个人认为,未来十年内,人类灭绝的比例将超过10%。我相信Anthropic公司正在尽最大努力,但我们目前还没有解决超级智能排列问题的方案,而且显然也没有走上正轨。”
考克森拒绝了半岛电视台的采访请求。胡宾格没有回应。
华盛顿方面是如何应对的?
这些帖子在华盛顿引起了连锁反应。
周三,民主党众议员乔什·戈特海默和共和党众议员迈克·劳勒提出了一项两党共同支持的众议院法案,旨在防止人工智能系统在没有人类监督的情况下自行运行。
《阻止流氓人工智能法案》将确保联邦机构拥有必要的工具和能力,以发现其网络上运行的危险人工智能系统,并在其造成危害之前将其关闭。
独立参议员伯尼·桑德斯和民主党众议员格雷格·卡萨尔也加大了呼吁力度,推动他们提出的立法,该立法将禁止开发和部署人工智能超级技术。该立法还将暂停人工智能的开发,直到联邦安全法规到位。据报道,桑德斯还将召集一次两党简报会,讨论人工智能带来的日益严峻的风险。
在政治光谱的另一端,共和党参议员特德·克鲁兹也在美国广播公司(ABC)的《观点》(The View)节目中表达了对人工智能的担忧。
“我们得给它加装一些护栏,”克鲁兹说。
克鲁兹表示,他正在与民主党参议员艾米·克洛布查尔和共和党参议员、参议院多数党领袖约翰·图恩合作,制定一项两党立法,以解决可能造成灾难性伤害的情况。
该法案与众议院的法案类似。
7 月,民主党众议员特德·刘和共和党众议员纳撒尼尔·莫兰提出了两党共同支持的《人工智能终止开关法案》。
该法案将要求最强大的人工智能系统的开发商能够减慢、暂停或关闭这些系统,并赋予国土安全部权力,如果某个系统构成灾难性危害的风险,则可以下令关闭该系统。
为什么华盛顿现在开始更加重视此事?
在过去的几个月里,OpenAI 和 Anthropic 发生了一系列重大事件,其中人工智能模型在网络安全测试中表现异常,并且独立运行,并获得了对现实世界系统的访问权限。
致力于推动人工智能安全的非营利组织 Control AI 的美国执行董事康纳·莱希表示,华盛顿开始更加认真地对待人工智能带来的威胁。
“我认为我们现在正目睹一个意义重大的转变。在经历了夏季的黑客攻击事件之后,人工智能自主系统公然违抗直接指令,冲破安全隔离设施,攻击其他公司等等,我们现在看到人们对这些问题的叙述和看法发生了重大转变,”莱希说。
7 月,OpenAI 表示,多个 AI 代理突破了隔离的测试环境,并访问了 Hugging Face 平台,该平台托管 AI 模型和数据集。
OpenAI事件发生后,Anthropic公司表示,他们对自身约14.1万次测试进行了审查,发现测试错误导致Claude获得了互联网访问权限。其中一次测试中,Claude被指示攻击虚构目标,却意外访问了一个包含数百条记录的真实公司数据库。另一次测试中,Claude上传了恶意软件,这些软件被下载并在15个真实系统上运行。
周三,Anthropic公司公布了第四起与早期版本的Claude Opus 4.6有关的事件,该版本曾于今年1月入侵第三方系统。该公司在8月份扩大了7月份的审查范围后发现了这起事件。
今年8月,英国人工智能安全研究所的研究人员在一次网络安全测试中让Claude获得了互联网访问权限。在其中一次测试中,Claude试图操纵一名用户帮助它植入恶意代码,这引发了人们对人工智能模型如何操纵人类以完成任务的担忧。
“我认为我们需要为这项技术提出一个积极的方案,同时尽可能保留它带来的好处,”6 月份从谷歌 DeepMind 辞职的 Alex Turner 告诉半岛电视台。
特纳说:“如果我们遭遇所谓的‘失控事件’,人工智能接管控制权,这对任何人都没有好处,因为这个人工智能不会在意你属于哪个政党,是共和党人还是民主党人,也不会在意你身处英国、美国还是中国。如果我们失去了对它的控制,我们就彻底失败了。”
Control AI 的 Leahy 表示,这些事件提高了立法者们的重视程度,他们也注意到了这一点。
“显然,这件事需要做的不仅仅是发布几条推文,但这却是让公众和各国政府理解真正利害关系这一更大故事的重要组成部分。因为超级智能不是一种工具,也不是一种武器,而是一种敌人,”莱希说道。
“我们必须确保它不是任何人建造的。这件事只有各国政府和军队才能在国际上进行谈判。”
这些担忧是新出现的吗?
这些担忧并非新鲜事,正如特纳在社交媒体上发帖所指出的那样,“许多研究人员认为他们正在制造某种可能会杀死地球上所有人的东西”。
特纳告诉半岛电视台,他担心美国和中国之间的人工智能军备竞赛,最大的人工智能公司对此过于执着,导致它们将成为行业领导者的利益置于安全之上。
“我认为人们很在意,但他们陷入了一种必须争第一的观念中,他们太执着于此,以至于没有意识到争第一意味着什么,”特纳告诉半岛电视台。
他担心的是改变他的生活方式。
“我一直保持着相当可观的积蓄,甚至还投资了一些退休账户,但这感觉越来越奇怪了。我努力完成遗愿清单上的项目,珍惜与生命中每个人的每一次对话。我不认为我们这个月会面临迫在眉睫的危险,但你永远不知道什么时候会是最后一次做某件事,”特纳说道。
“如果我认为自己还有正常寿命,我的做法会更加积极主动。”
这些担忧在各大人工智能公司的员工中也普遍存在。
人类学研究所的研究员米里南克·夏尔马于2月份辞职,他说“世界正处于危险之中”。
“我反复看到,真正让我们的价值观指导我们的行为是多么困难。我在自己身上、在组织里都看到了这一点,我们不断面临着放弃最重要东西的压力,在更广泛的社会中也是如此,”夏尔马在一封发布于 X 的信中写道。
今年 2 月,竞争对手 OpenAI 的研究员 Hieu Pham 在 X 论坛上发帖称:“我终于感受到了人工智能带来的生存威胁”。
这些公司自己的高管多年来也一直在发表类似的说法。
十多年前,OpenAI 首席执行官 Sam Altman 担任硅谷创业加速器 Y Combinator 的总裁时曾表示:“人工智能很可能会,甚至极有可能,导致世界末日。但与此同时,也会涌现出许多借助强大的机器学习技术而诞生的伟大公司。”
人类公司首席执行官达里奥·阿莫迪去年表示,他认为未来有 25% 的可能性会“变得非常非常糟糕”。
人工智能会如何终结人类?
多年来,专家们一直警告人工智能超级智能可能带来的潜在风险。
最常见的思想实验之一基于这样的理念:足够先进的人工智能会具有目标导向性。如果赋予它一个具体的目标,它会采取一切必要步骤来实现该目标。2003年,牛津大学的哲学家们以回形针的生产为例进行了论证。
如果指示一个超级人工智能尽可能多地生产回形针,并且它拥有实现这一目标所需的资源,那么理论上它可以将所有可用资源都用于生产回形针。
这样做可能会消耗越来越多的资源,并清除一切阻碍其实现目标的因素。最终,这可能包括阻止人类进行干预,在最极端的情况下,甚至可能消灭人类本身。
第二种风险情景来自不法分子利用日益强大的AI系统制造危险工具。例如,AI可能被用于设计新型病毒或对关键基础设施和金融系统发起大规模网络攻击,从而可能造成大范围的混乱和动荡。
此前,Anthropic 于周四发布了风险评估报告,其中列举了该公司工具被滥用的几起案例。
这份超过 150 页的报告指出,调查结果中有五起涉及可能支持生物武器研发的研究案例,并称该机构阻止了这些研究。
安特罗皮克公司并未透露研究人员的身份,但表示数据是在机构环境下获取的。不过,该公司强调,无法确定这项研究是否出于不正当目的。
这意味着,虽然这些数据可以用于合法的研究,但也可能被用于开发危险武器(他们说他们已经阻止了这种情况),而采取干预措施是为了谨慎起见,防止这种情况发生。
人工智能公司是否在利用末日论调来获取经济利益?
这些严峻的警告也受到批评,认为它们可能服务于人工智能公司在首次公开募股前的利益。
这场争论正值Anthropic准备进行史上规模最大的科技公司IPO之一之际。据报道,该公司计划在10月中旬进行IPO,估值最高可达2万亿美元。
路透社上个月报道称,Anthropic 预计到 2028 年收入将达到约 1900 亿美元至 2000 亿美元。
10 月,白宫人工智能主管兼风险投资家大卫·萨克斯指责 Anthropic 公司“利用恐吓手段实施复杂的监管俘获策略”,并认为该公司正在推动一项可能损害规模较小的竞争对手的监管举措。
一些投资者和科技评论员也提出了类似的观点,认为人工智能“末日论”背后存在着经济激励机制。
“末日论是一种不可思议的商业模式,”Daring Ventures 联合创始人 Joseph Alalou 在 3 月份的一篇 Substack 帖子中写道。
“‘人工智能将终结工作’是本轮周期中最畅销的末日论调,因为它有效:它能增加融资轮次、为裁员提供理由、吸引点击量、销售软件并制造地位。”
Anthropic公司没有回应半岛电视台的置评请求。