General World Models:
The Bridge Connecting the Digital and Physical Worlds

On March 19, at the 2026 Spring Capital Market Forum hosted by CITIC Securities, Professor Jun Zhu — founder of ShengShu Technology, Deputy Dean of the Institute for AI at Tsinghua University, and ACM/IEEE/AAAI Fellow — delivered a keynote speech titled “General-Purpose World Models: The Bridge Connecting the Digital and Physical Worlds,” systematically outlining the key technical pathways for generative AI’s transition from “content generation” to the “physical world.”

He noted that as unified model architectures take shape and data paradigms continue to improve, world models are reaching a critical inflection point, with general-purpose world models becoming a core technical direction toward AGI (Artificial General Intelligence).

In this direction, ShengShu Technology has been a pioneer in deploying general-purpose world models. In July and December 2025, the company, in collaboration with Tsinghua University, successively released Vidar — the first video-based embodied foundation model — and Motus, a unified-architecture general-purpose foundation world model. Compared to the internationally leading VLA model Pi0.5, these models achieved approximately 40% improvement in success rate and were the first to discover generalization capabilities of general-purpose world models across diverse embodied tasks.

Professor Jun Zhu Keynote Speech

At the forum, Zhu systematically introduced ShengShu Technology’s strategic roadmap for general-purpose world models. At its core lies the Foundation World Model, extending upward into a dual-track technology system covering both digital and physical spaces.

The Foundation World Model is built on the globally pioneering U-ViT architecture, accumulating multimodal information across vision, audio, and touch to construct a unified cognitive and modeling capability for the world, providing a unified intelligence foundation for upper-layer applications.

In digital space, ShengShu Technology has built the video foundation model product Vidu based on the World Generation Model (WGM). Vidu’s generation model focuses on single-timestep world simulation, empowering AI productivity in the digital world. Its streaming generation model focuses on multi-timestep world simulation, enabling real-time companionship and interaction. Vidu has significantly improved the efficiency of digital content production, ultimately aiming to achieve AGI in the digital world.

In physical space, ShengShu Technology has built the unified world model product Motus based on the World Action Model (WAM). As the “brain” of real-world embodied intelligence, Motus is dedicated to solving core pain points of traditional embodied intelligence — fragmented pipelines, data scarcity, and weak generalization. It enables zero-shot generalization and cross-embodiment adaptation in the real world, driving robots from “modular execution” toward “unified intelligent agents,” ultimately achieving AGI in the physical world.

This strategic layout connects the complete pathway from “understanding the world” to “generating the world” to “acting in the world,” making the general-purpose world model truly the bridge connecting the digital and physical worlds.

Generative AI Enters a New Phase: From “Generating Content” to “Understanding the World”

Generative AI is entering a new stage of development. Its core objective is no longer limited to content generation, but rather to characterize and understand the complex distributions of the physical world.

“Generative capability itself is becoming an essential foundation for understanding the world — if you cannot generate it, you cannot truly understand it,” Zhu pointed out.

From probabilistic graphical models to deep learning, and then to the rise of large-scale pre-training, Transformers, and diffusion models, the technical pathway has continued to evolve, steadily approaching the capability boundaries of AGI. Zhu stated that the evolution of generative AI is, in essence, a continuous enhancement of world modeling capabilities.

Evolution of Generative AI

Video: The Key Medium Connecting Digital and Physical Worlds

In this process, the focus of AI development is extending from language further into video.

“Compared to language, video naturally contains richer spatiotemporal information and physical laws, making it a key medium connecting the digital and physical worlds,” Zhu noted. “Video is not just a content format — it is a record of how the world operates.”

At the same time, vision plays a dominant role in human cognition. For machines to truly understand the world, they must also center their learning on vision. However, relying solely on large language models remains insufficient to build a complete intelligence loop. True intelligent systems need the ability to learn from experience, predict the future, and execute actions — a process that depends on continuous interaction with the physical world.

Why Generative AI

Breaking the Data Bottleneck: Building a Data Pyramid Centered on Video

On the data front, embodied intelligence has long faced a “data wall”: real-robot data is scarce, expensive, and difficult to reuse.

To address this, a video-centric data pathway is becoming an industry consensus. By constructing a multi-layered data system covering internet videos, human demonstration videos, simulation data, and robot data, it is possible to systematically mine the physical interaction knowledge embedded in videos.

“Video is currently the largest-scale and most information-rich data modality. Fully leveraging scalable, heterogeneous data with video as the primary source is the most viable path for building general-purpose world models,” Zhu stated.

By introducing methods such as “Latent Action,” models can map motion information from videos into action space, maintaining effective action capabilities even without large amounts of real-robot data.

World Models: From “Modular Assembly” to “Unified Architecture”

Against this backdrop, general-purpose world models are increasingly seen as an important pathway toward achieving AGI.

The core objective is to build a unified intelligent system that enables AI to complete the full closed loop from “observing the world” to “predicting the world” to “acting in the world.” However, the current industry technical approach remains fragmented: VLA models focus on behavior imitation, traditional world models on future prediction, and inverse dynamics models on action generation — each covering only part of the capability chain.

“World models should not be modular assemblies; they need to achieve multiple cognitive capabilities through a unified architecture, just like humans,” Zhu stated. A general-purpose world model must integrate perception, reasoning, prediction, and action capabilities within a single model, building a holistic intelligence structure akin to the human “brain.”

Human World Models

Motus: A Unified World Model Opening a New Paradigm for Multi-Task Generalization and Scaling in Embodied Intelligence

Building on the data and architecture pathways described above, Motus — the unified world model open-sourced by ShengShu Technology and Tsinghua University — achieves systematic integration of multimodal capabilities.

In terms of model architecture, Motus is built on the UniDiffuser unified modeling framework. Through Cross-modal Priors Fusion, it integrates visual-language knowledge (VLM), video dynamics knowledge (Video Generation Model), and action skill knowledge (Action Expert) within a single model, achieving unified expression and generation of language, video, and action — building a truly unified world model.

In data utilization and scaling, Motus demonstrates significant advantages. In the Data Scaling experiments, compared to the internationally leading VLA model Pi0.5, Motus can learn from broader heterogeneous data and effectively incorporate multimodal priors from pre-trained foundation models. On the average success rate across 50 tasks, Motus achieved a 35.1% absolute improvement, while demonstrating 13.55× data efficiency at the same performance level.

In the Task Number Scaling experiments, as the number of tasks increased, Motus’s average success rate continued to improve, while the comparison model Pi0.5 showed performance degradation as task complexity grew. Ultimately, Motus achieved a 37% absolute success rate advantage, demonstrating stronger multi-task generalization.

Motus Scaling Experiment Comparison

Even more notably, Motus was the first to reveal a new dimension in embodied intelligence scaling — the multi-task generalization capability curve. This curve provides a critical “North Star metric” for embodied foundation models. Its evolution path closely mirrors that of language models, echoing the core insight proposed by GPT-2 — “Language Models are Unsupervised Multitask Learners” — and has been described as the “GPT-2 moment” for embodied intelligence.

In long-horizon, multi-step complex real-robot tasks, Motus further exhibits near-human-level decision logic and execution stability. It is important to emphasize that these are not simple single-step commands, but rather typical long-horizon, multi-step tasks completed end-to-end by the model, without relying on traditional “fast-slow dual system” decomposition.

The Inflection Point Is Near: General-Purpose World Model Capabilities Will Continue to Leap Forward

As Turing Award winner Richard Sutton pointed out in “The Bitter Lesson,” “General methods that leverage computation are ultimately the most effective, and by a wide margin.” This judgment is being continuously validated in the development of AI.

Zhu stated that a scalable, heterogeneous data system centered on video is the most viable path for building general-purpose world models, and this view is gradually forming industry consensus. As unified model architectures, data paradigms, and training systems continue to mature, the technical pathway for general-purpose world models is becoming increasingly clear, and the industry is entering a phase of scale-driven capability leaps.

In this trend, moving from video generation to world models is becoming the key pathway for AI to transition from “understanding the world” to “changing the world.” As related technologies continue to evolve, general-purpose world models will accelerate into the physical world, becoming the bridge connecting the digital and physical worlds.