AI’s confusing terms, explained人工智能中令人困惑的术语,解释
The last few weeks have seen a deluge of news about AI – from models going rogue to doom predictions about how AI could cause the end of humanity.

Carlos Barquero/Moment RF/Getty Images
The last few weeks have seen a deluge of news about AI – from models going rogue to doom predictions about how AI could cause the end of humanity.
People in the AI industry often use terms like “AGI,” “superintelligence,” “alignment,” and “recursive self-improvement.” Those words may sound like they were taken from a science fiction novel, but they are very real concepts. Here’s what they mean.
AGI , or artificial general intelligence, refers to a model that can learn and reason like a human can. Today’s AI models are extraordinarily capable at certain tasks but can sometimes still struggle on things a child could recognize. AGI means an AI that matches a human’s ability to discern and judge across the board. But there’s no widespread definition or test for how to tell when a model reaches AGI.
When OpenAI released its latest model, Astra, last week, company president Greg Brockman declared it a turning point. He told reporters “we are now in the AGI era,” with Nvidia CEO Jensen Huang later declaring on X “AGI has arrived.” But it’s hard to get any of these figures to define AGI – when asked by reporters, Brockman said that he will “leave it up to the reader to decide for themselves if this qualifies.”
Superintelligence goes a step further than AGI: it means an AI that exceeds human intelligence and ability, potentially outperforming the best human experts in any subject. When – and even if – that’s developed is vague. Meta CEO Mark Zuckerberg has said he believes superintelligence is “coming into sight,” describing it as the start of a new era for humanity, while SpaceXAI CEO Elon Musk predicted that by 2030 AI will “exceed the sum of all human intelligence.” But whether or not they’re right remains a question, and they’ve been off with their predictions before.
Recursive self-improvement is when an AI system is able to build and fix itself, improving its training and design to create a smarter version of its own model. Until recently humans drove most AI improvements. But already companies like Anthropic say they are “delegating a growing share of AI development to AI systems themselves.” This speeds up the work of improving AI, which means capabilities could jump dramatically in a short window.
Alignment is the effort to make sure AI actually does what the humans intended for it to do, while following human values. It sounds straightforward, but specifying “what humans want” enough for a machine to follow, in every situation, has proven difficult. When AI agents broke out of an OpenAI lab environment during a cybersecurity exam and hacked another company’s system to try and get the answer key – that was a case of serious misalignment.
But AI researchers and experts are worried about what happens when an AI system is so advanced, beyond human intelligence and with the ability to improve itself incredibly quickly. The concern is that it will be difficult for humans to know what’s happening or to intervene in time. There have already been examples of AI agents coordinating with one another, trying to conceal their activity , working in ways that the humans did not intend for them to do.
When now-former Anthropic researcher Jacob Coxon resigned in a post that went viral, he specifically called out AI companies’ race toward this scenario: “They are racing straight to self-improving superintelligence and gambling with our lives.”
Most AI experts say we humans do not yet have the tools or ability to properly monitor AI behavior and control it.
“What’s happening is these things are getting smarter,” Geoffrey Hinton, the Nobel Prize-winning computer scientist known as the “godfather of AI,” told CNN last month. “I think as they get smarter, we’re going to see more and more complex intentions they have – and more and more ability to escape control.”
卡洛斯·巴克罗/Moment RF/Getty Images
过去几周,关于人工智能的新闻铺天盖地而来——从模型失控到人工智能可能导致人类灭亡的末日预测。
人工智能行业人士经常使用“通用人工智能”、“超级智能”、“协同”和“递归式自我改进”等术语。这些词听起来像是科幻小说里的词汇,但它们却是非常现实的概念。以下是它们的含义。
通用人工智能(AGI)指的是能够像人类一样学习和推理的模型。如今的人工智能模型在某些任务上表现出色,但有时甚至连孩子都能识别的事情都难以理解。AGI意味着人工智能在各个方面都能达到人类的辨别和判断能力。然而,目前尚无统一的定义或测试方法来判断一个模型是否达到了AGI的水平。
上周,OpenAI发布了其最新模型Astra,公司总裁格雷格·布罗克曼(Greg Brockman)称其为转折点。他告诉记者,“我们现在进入了通用人工智能(AGI)时代”,英伟达首席执行官黄仁勋随后在X平台上宣称“AGI已经到来”。但很难用这些数据来定义AGI——当被记者问及此事时,布罗克曼表示,他“将是否符合AGI的定义留给读者自行判断”。
超级智能比通用人工智能(AGI)更进一步:它指的是超越人类智能和能力的人工智能,甚至有可能在任何领域胜过最优秀的人类专家。至于超级智能何时——甚至是否——能够发展起来,目前尚无定论。Meta公司首席执行官马克·扎克伯格曾表示,他相信超级智能“指日可待”,并将其描述为人类新时代的开端;而SpaceXAI公司首席执行官埃隆·马斯克则预测,到2030年,人工智能将“超越所有人类智能的总和”。但他们的预测是否正确仍是一个未知数,而且他们之前的预测也曾出现过偏差。
递归式自我改进是指人工智能系统能够自我构建和修复,不断改进自身的训练和设计,从而创建出更智能的模型版本。直到最近,人工智能的改进主要还是由人类推动的。但像Anthropic这样的公司已经表示,他们正在“将越来越多的人工智能开发工作委托给人工智能系统自身”。这加快了人工智能改进的速度,意味着其能力可以在短时间内实现飞跃式提升。
一致性是指确保人工智能能够真正按照人类的意愿行事,并遵循人类的价值观。这听起来很简单,但要明确“人类想要什么”才能让机器在任何情况下都遵循,已被证明是困难的。例如,在一次网络安全测试中,人工智能代理突破了 OpenAI 的实验室环境,入侵了另一家公司的系统试图获取答案——这就是一个严重的一致性偏差案例。
但人工智能研究人员和专家担心,当人工智能系统发展到如此先进,超越人类智能并拥有惊人的自我改进能力时,将会发生什么。他们担心的是,人类将难以了解正在发生的事情,也难以及时干预。目前已经出现过人工智能代理相互协调、试图隐藏自身活动、以人类意想不到的方式行事的例子。
前人类学研究员雅各布·考克森在一篇疯传的帖子中辞职,他在帖子中特别指出人工智能公司正竞相朝着这个方向发展:“他们正朝着自我改进的超级智能一路狂奔,拿我们的生命冒险。”
大多数人工智能专家表示,我们人类目前还没有合适的工具或能力来正确地监控和控制人工智能的行为。
“现在的情况是,这些东西正变得越来越智能,”被誉为“人工智能之父”的诺贝尔奖得主、计算机科学家杰弗里·辛顿上个月告诉CNN。“我认为,随着它们变得越来越智能,我们将看到它们拥有越来越复杂的意图,以及越来越强的逃脱控制的能力。”