Inside China, artificial intelligence is a snake eating its own tail在中国,人工智能就像一条吞噬自己尾巴的蛇。
Opinion: AI researchers refer to this as “model collapse,” a phenomenon in which models trained on their own synthetic outputs degrade over successive generations.

China’s greatest technological ambition and its greatest political obsession are quietly destroying each other.
The same censorship apparatus the Party built to control its people is now corrupting the AI systems its leaders depend on. The United States, by leaning into an open marketplace of information and ideas, will gain advantage as it takes a different path.
AI is increasingly training newer, faster AI models . This typically involves scraping the internet for content and then loading it into datasets for new programs. The problem: online content used for training is increasingly generated by AI. As a result, each generation of technology drifts from reality.
AI researchers refer to this as “model collapse,” a phenomenon in which models trained on their own synthetic outputs degrade over successive generations. The only defense is a constant influx of fresh, honest, human-generated information. Without it, the system folds in on itself.
China’s Great Firewall cuts off that influx, expediting the impact of model collapse within its borders. The first generation of large language models worldwide was trained on massive datasets of publicly available human-generated text. These human-derived pieces of information built an algorithmic approximation of how people think, argue, explain, and communicate.
Now, the internet is filling with AI-generated content at a rate inconceivable five years ago. Marketing copy, product descriptions, social media captions and news summaries are increasingly produced by AI systems and published online. They’ve all become generic and detached from human signals.
With each cycle of new AI, the models drift further from their human origins. The newer systems amplify the patterns AI systems favor while losing nuance and amplifying existing biases. Each successive AI generation is one step further removed from the humans these programs were originally built to serve.
In China, the Great Firewall accelerates this problem. The Chinese Ministry of Public Security built the Great Firewall of China in the late 1990s to censor the Chinese people. It is now the most sophisticated information control infrastructure in human history. The Great Firewall does not just restrict what Chinese users can see. It shapes the data used to train Chinese AI systems. By design, it strips out politically sensitive events, dissenting viewpoints, and independent reporting. Left behind is a curated record of reality aligned with the Party’s narrative. That filtered record of information becomes the raw material for LLMs.
The training data fed into LLMs in China does not contain any criticism of the government, fair reportage of controversial topics, or accurate information about Chinese history. Events such as the the Uyghur detention camps exist inside the Firewall only as the state chose to describe them.
Now add model collapse on top of this. Chinese AI companies such as Baidu, Alibaba, ByteDance and dozens of others aggressively deploy AI-generated content across their platforms. This material becomes training data for the next generation of Chinese AI models. With independent reporting locked out, model collapse expedites inside the Great Firewall with no escape valve. The practical consequences of this divergence are already visible and will accelerate.
Chinese LLMs struggle with tasks that require new observation, original synthesis or human complexity that the training data was designed to suppress. If asked about repression in Chinese history or controversial current events such as the Uyghur detentions, these LLMs either do not answer or produce a response indistinguishable from a Party press release. A Chinese trade official relying on domestic AI to model the economic impact of Western sanctions is working from a system incapable of providing an honest account of how those sanctions functioned or failed in comparable historical cases.
The West has a milder version of this problem.
Western AI models are trained on increasingly synthetic content, but human reporters regularly push new information into the ecosystem. Free and open societies have structural advantages to retain the capacity to reason about what is happening in the world rather than what previous AI systems described. When asked about Tiananmen Square, Western accounts typically focus on the 1989 protests and the crackdown. By contrast, Chinese models either refuse to answer or return state-aligned language. This information gap becomes part of the training data for the next generation.
To maintain a competitive AI advantage over China, Washington should treat human-generated data as a strategic asset and invest in journalism, open web archives, and synthetic-content labeling. Preserving the integrity of American training data is a defense imperative, not a tech problem.
Chinese leaders deploying AI products to make decisions about economics, geopolitics, and public health will make those decisions based on systems trained on what China’s information control apparatus wants people to believe. That is not an intelligence system. It is a mirror. And the tragedy of model collapse is that a mirror that has been looking at itself long enough no longer reflects anything.
Joe Buccino is a retired U.S. Army colonel and the author of “When Every Word Counts: How to Earn Trust, Command Attention, and Communicate Clearly in Any Situation.”
中国最大的科技雄心和最大的政治执念正在悄然地相互毁灭。
党用来控制人民的审查机制,如今正在腐蚀其领导人赖以生存的人工智能系统。美国若走上一条不同的道路,拥抱开放的信息和思想市场,必将获得优势。
人工智能正在不断训练更新、更快的模型。这通常涉及从互联网抓取内容,然后将其加载到新程序的数据集中。问题在于:用于训练的在线内容越来越多地由人工智能生成。因此,每一代技术都与现实存在偏差。
人工智能研究人员将此称为“模型崩溃”,指的是模型基于自身合成输出进行训练后,随着迭代次数的增加而性能下降的现象。唯一的应对之策是不断引入新鲜、真实的、由人类生成的信息。否则,系统就会自我崩溃。
中国的防火长城阻断了这种信息流入,加速了模型崩溃在其境内的影响。全球第一代大型语言模型是基于海量的公开人类文本数据集进行训练的。这些信息构建了一个算法,用于近似模拟人们的思考、论证、解释和交流方式。
如今,互联网上人工智能生成的内容正以五年前难以想象的速度涌现。营销文案、产品描述、社交媒体标题和新闻摘要越来越多地由人工智能系统生成并发布到网上。它们都变得千篇一律,脱离了人类的表达。
随着人工智能的不断迭代,其模型与人类原型之间的距离也越来越远。新系统强化了人工智能系统偏好的模式,却丢失了细微差别,并加剧了原有的偏见。每一代人工智能都与它们最初服务的对象——人类——渐行渐远。
在中国,防火长城加剧了这个问题。中国公安部在上世纪90年代末期建立了防火长城,旨在审查中国民众的信息。如今,它已成为人类历史上最复杂的信息控制基础设施。防火长城不仅限制中国用户能够看到的内容,还塑造了用于训练中国人工智能系统的数据。它有意剔除了政治敏感事件、异议观点和独立报道。最终留下的,是一份经过精心筛选、符合党的叙事的“现实记录”。这份经过过滤的信息记录,成为了人工智能(LLM)的原材料。
中国法律硕士(LLM)课程的训练数据中不包含任何对政府的批评、对争议性话题的公正报道,也不包含任何关于中国历史的准确信息。诸如维吾尔族拘留营之类的事件,在网络防火墙内仅仅以官方选择的方式呈现。
现在,再加上模型崩溃的问题。百度、阿里巴巴、字节跳动等数十家中国人工智能公司正积极在其平台上部署人工智能生成的内容。这些内容将成为下一代中国人工智能模型的训练数据。由于独立报道渠道受阻,模型崩溃在防火墙内加速发展,且无处可逃。这种分化的实际后果已经显现,并将加速恶化。
中国人工智能系统在处理需要全新观察、原创性综合或人类复杂性的任务时表现不佳,而这些复杂性恰恰是训练数据旨在抑制的。如果被问及中国历史上的镇压事件或诸如维吾尔族拘留等争议性时事,这些人工智能系统要么不作答,要么给出的回答与官方新闻稿毫无二致。一位依赖国产人工智能来模拟西方制裁经济影响的中国贸易官员,实际上是在使用一个无法真实反映这些制裁在类似历史案例中如何发挥作用或失败的系统。
西方也存在类似的问题,但程度较轻。
西方人工智能模型越来越多地使用合成内容进行训练,但人类记者却不断向生态系统推送新的信息。自由开放的社会拥有结构性优势,使其能够独立思考世界正在发生的事情,而不是仅仅依赖于以往人工智能系统的描述。当被问及天安门事件时,西方媒体的报道通常聚焦于1989年的抗议活动和镇压。相比之下,中国模型要么拒绝回答,要么给出与官方立场一致的解释。这种信息鸿沟成为了下一代人工智能模型的训练数据。
为了保持对中国的AI竞争优势,华盛顿应该将人类生成的数据视为战略资产,并投资于新闻业、开放网络档案和合成内容标注。维护美国训练数据的完整性是国防的当务之急,而非技术问题。
中国领导人利用人工智能产品来制定经济、地缘政治和公共卫生方面的决策,这些决策所依据的系统,其训练内容却是中国信息控制机构希望人们相信的。这并非智能系统,而是一面镜子。而模型崩溃的悲剧在于,一面镜子如果长时间只照镜子,就什么也映照不出来。
乔·布奇诺是美国陆军退役上校,著有《字字珠玑:如何在任何情况下赢得信任、赢得关注并清晰沟通》。