ByteDance founder joins AI elite in race to perfect world models字节跳动创始人加入人工智能精英行列,竞相打造完美世界模型
The new model is slated to be launched as soon as October. Read more at straitstimes.com.
ByteDance’s AI model aims to generate virtual worlds that respond to Pico headset users’ voices or movements.
Published Sep 08, 2026, 10:45 AM
Updated Sep 08, 2026, 10:45 AM
ByteDance founder Zhang Yiming is leading the development of a new AI model for real-time spatial video generation, aiming for launch as early as October.
The model will enable interactive virtual worlds for live streams, games, and dramas, integrating ByteDance’s AI, cloud resources, and Pico hardware.
This move positions ByteDance to compete with Meta and Apple in virtual and mixed reality by offering cloud-based spatial content to reduce VR hardware costs.
ByteDance is readying an AI model geared for real-time spatial video generation, taking on Meta Platforms and Alphabet in an arena with applications in robotics and autonomous systems.
Founder Zhang Yiming is personally overseeing the development of a new model slated for launch as soon as October , according to people familiar with the matter.
The billionaire has been coordinating the efforts of business units across the company and throwing AI resources and computing capacity behind the effort, the people said, asking not to be identified discussing private information.
The timing of the launch is not certain and plans may change, one of the people said.
Built on Seedance – ByteDance’s existing AI model for generating cinematic videos – the model would allow users to create interactive virtual worlds for live streams, short-form dramas and games, the people said. A spokesperson for ByteDance did not respond to a request for comment.
Zhang hopes his latest AI model will launch the TikTok creator into the burgeoning arena of world models, armed with top-tier expertise in video. He would be joining the likes of Fei-Fei Li and Yann LeCun in exploring a visual approach to AI that’s considered crucial for robotics, games and self-driving cars.
The move is part of an ongoing strategic shift within ByteDance, as the Beijing-based company redirects much of its business-to-business operations toward video production and generation.
Zhang sees the spatial-video model – which would place the user in a three-dimensional world much like Google’s Genie – as a flywheel, linking ByteDance’s AI models and cloud-computing resources with its content platforms and extended reality arm Pico’s hardware, the people said.
Both ByteDance and Google’s systems are part of a video-centric push to build models that can simulate physical environments for AI agents. Genie was launched earlier in 2026 to let users interact with a real-time rendered world.
If successful, the new model may open a new front in ByteDance’s rivalry with Meta and Apple, which have invested heavily in virtual- and mixed-reality devices.
Meta has sought to build a mass-market ecosystem around its Quest headsets, while Apple has positioned Vision Pro as a premium spatial-computing platform. Both have struggled to go mainstream.
ByteDance’s latest model aims to generate virtual worlds that respond to Pico headset users’ voices or movements. The model can offer on-demand videos with a latency of about 0.05 seconds at 20 frames per second and serve spatial-computing environments and interfaces, one of the people said.
The Chinese model seeks to lower the upfront cost of VR adoption by moving the computationally intensive process of generating spatial content to the cloud. That would reduce processing power required in headsets and open a path for cheaper and less sophisticated hardware.
Swift content generation might also help shift the battleground for users away from hardware specs to AI models, computing infrastructure and content-distribution platforms.
Seedance has emerged as ByteDance’s signature AI success, even as the company lags rivals in a domestic field crowded with the likes of Alibaba Group Holding, DeepSeek and Moonshot AI. It underpins a slew of ByteDance products, including CapCut and Doubao, China’s most popular AI chatbot. The video tool has also become the go-to choice for independent film studios, creators and emerging AI startups.
ByteDance, which Zhang founded alongside Liang Rubo in 2012, has secured a US$30 billion ( S$38 billion ) loan as it rushes to expand its AI capabilities and amass data centres and other hardware, Bloomberg reported last week. BLOOMBERG
AI/artificial intelligence
Technology and research
字节跳动的人工智能模型旨在生成能够对 Pico 头戴式设备用户的声音或动作做出反应的虚拟世界。
发布于 2026 年 9 月 8 日上午 10:45
更新于2026年9月8日上午10:45
字节跳动创始人张一鸣正在领导开发一种用于实时空间视频生成的新型人工智能模型,目标是最早在10月份推出。
该模型将整合字节跳动的人工智能、云资源和Pico硬件,为直播、游戏和电视剧打造交互式虚拟世界。
此举使字节跳动能够通过提供基于云的空间内容来降低虚拟现实硬件成本,从而在虚拟现实和混合现实领域与 Meta 和苹果展开竞争。
字节跳动正在准备一款面向实时空间视频生成的AI模型,在机器人和自主系统应用领域与Meta Platforms和Alphabet展开竞争。
据知情人士透露,创始人张一鸣正在亲自监督一款新车型的研发,该车型计划最早于10月发布。
知情人士透露,这位亿万富翁一直在协调公司各业务部门的努力,并投入人工智能资源和计算能力支持这项工作。由于涉及私人信息,他们要求匿名。
其中一位知情人士表示,发布时间尚未确定,计划可能会有所变动。
据知情人士透露,该模型基于字节跳动现有的电影级视频生成人工智能模型Seedance构建,将允许用户创建用于直播、短剧和游戏的交互式虚拟世界。字节跳动发言人未对此置评。
张希望他最新的AI模型能帮助这位TikTok创作者跻身蓬勃发展的世界级模型领域,并凭借其顶尖的视频制作技术脱颖而出。他将与李飞飞和闫乐存等人一起,探索一种视觉化的AI方法,这种方法被认为对机器人、游戏和自动驾驶汽车至关重要。
此举是字节跳动正在进行的战略转型的一部分,这家总部位于北京的公司正在将其大部分企业对企业业务转向视频制作和生成。
据知情人士透露,张一鸣将空间视频模型(该模型会将用户置于一个类似于谷歌 Genie 的三维世界中)视为一个飞轮,将字节跳动的 AI 模型和云计算资源与其内容平台和扩展现实部门 Pico 的硬件连接起来。
字节跳动和谷歌的系统都属于以视频为中心的人工智能模型构建浪潮的一部分,旨在为人工智能代理构建能够模拟物理环境的模型。Genie 于 2026 年初推出,让用户能够与实时渲染的世界进行互动。
如果成功,这种新模式可能会在字节跳动与 Meta 和苹果的竞争中开辟新的战线,这两家公司都在虚拟现实和混合现实设备领域投入巨资。
Meta一直试图围绕其Quest头显构建大众市场生态系统,而苹果则将Vision Pro定位为高端空间计算平台。但两者都难以真正走向主流市场。
字节跳动最新推出的这款产品旨在创建能够响应Pico头显用户语音或动作的虚拟世界。一位知情人士透露,该型号产品能够以每秒20帧的速度提供延迟约为0.05秒的按需视频,并可用于空间计算环境和界面。
中国模式旨在通过将计算密集型的空间内容生成过程转移到云端,降低虚拟现实技术的初期应用成本。这将减少头显所需的处理能力,并为更便宜、更简单的硬件铺平道路。
快速的内容生成也可能有助于将用户争夺的焦点从硬件规格转移到人工智能模型、计算基础设施和内容分发平台。
尽管字节跳动在国内人工智能领域落后于阿里巴巴集团、DeepSeek和登月科技等竞争对手,但Seedance已成为字节跳动最具代表性的人工智能产品。它为字节跳动的一系列产品提供了技术支持,包括CapCut和中国最受欢迎的AI聊天机器人豆宝。这款视频工具也已成为独立电影工作室、内容创作者和新兴人工智能创业公司的首选。
据彭博社上周报道,张一鸣与梁如波于2012年共同创立的字节跳动公司已获得300亿美元(约合380亿新元)贷款,用于加速拓展人工智能能力,并扩充数据中心和其他硬件设施。
人工智能
技术与研究