
ShengShu Technology Proposes a Five-Level Roadmap for General World Models
SINGAPORE, Aug. 24, 2026 /PRNewswire/ -- At the 2026 World Robot Conference (WRC), ShengShu Technology unveiled its latest research on General World Models (GWMs), proposing a five-level development roadmap charting their evolution from world generation and real-time interaction to physical action, autonomous agents, and world orchestration.
Jun Zhu, Founder and Chief Scientist of ShengShu Technology and an ACM, IEEE and AAAI Fellow, presented the research during a keynote address at the forum "The Evolution of Foundation Models for Embodied Intelligence: From Technology Levels to Industrial Deployment."
"From the perspective of foundation-model development, our goal is not to build another specialized model for a particular task or setting, but to create a general foundation model that can understand the world, imagine possible futures, and take action," Zhu said.
Defining the General World Model from First Principles
World-model research today spans video generation, environment simulation, robot decision-making and action control. Yet the field still lacks a shared definition of what makes such a model truly general.
When people learn to ride a bicycle or drive a car, their movements gradually become stable and precise. One important reason is that the brain develops an "internal model" through continuous interaction with the external world, allowing it to anticipate the consequences of an action.
A General World Model similarly requires three interdependent capabilities:
- Understanding the world: integrating different sources of information to infer the current state;
- Imagining possible futures: predicting what may happen next, including the consequences of different actions;
- Taking action: influencing a digital or physical environment to realize goals, while using real-world feedback to refine subsequent predictions and decisions.
"A General World Model is not merely a generator, simulator, robot action model or policy model, nor is it a collection of isolated capabilities or a simple linear pipeline," Zhu said. "It is a closed-loop feedback system in which Understanding, Imagination, and Action are tightly connected."
Action is not merely the output of the model. It changes the environment, produces new information, and feeds into the next cycle of understanding, imagination and decision-making.
Data, Architecture and Compute: Three Pillars of General World Models
Building a General World Model requires three fundamental elements of foundation-model development: data, architecture and compute.
On the data side, GWMs require a multi-layer data pyramid that moves progressively from observation toward action: web-scale video, curated domain/instructional video, egocentric human video, action-recording human demonstrations, and real-robot interaction.
Lower layers provide greater scale and broader coverage, while higher layers are scarcer and more costly but more directly connect tasks, actions and physical outcomes. Synthetic data can augment every layer, while "imperfect" data—including failed attempts, corrective actions and recovery trajectories—provides valuable learning signals for recovering from failure.
On architecture, a GWM must process images, video, language and robot actions within a unified framework. Mixture-of-Transformers (MoT) assigns modality-specific parameters while retaining shared attention for cross-modal interaction, enabling environment understanding, world-state prediction and action generation to work together within the same model.
ShengShu Technology's World Action Model Motubrain is built on this architecture, jointly modeling understanding, prediction and action to improve the utilization of heterogeneous data and strengthen transfer across tasks.
Compute affects both large-scale pre-training and real-time deployment. While offline video generation can tolerate some latency, interactive generation and robot control must predict and decide before the environment changes. Training infrastructure, inference acceleration, model distillation and efficient attention are therefore critical to real-world deployment.
How Do General World Models Evolve? A Five-Level Roadmap
GWMs can be divided into five progressively evolving levels. Rather than separate product directions, they represent a continuous progression toward deeper world understanding, stronger interaction and greater autonomy. In the past several years, ShengShu Technology's research and deployment efforts have already covered the first three levels.
L1: World Generation
The first step is generating coherent world trajectories. Video helps models learn objects, motion, spatial relationships and temporal dynamics, laying the foundation for deeper understanding and imagination.
In 2024, ShengShu Technology launched the video generation foundation model Vidu. Through continued improvements in video quality, temporal consistency and world-dynamics modeling, Vidu represents the company's initial implementation of L1.
L2: Interactive World
Building on generation, the model receives real-time input and continuously changes what happens next in response to language, speech or control signals, moving from one-off generation toward continuous interaction.
Released in July 2026, Vidu S1 advances this capability into real-time interaction. Users can intervene through speech and continuously influence what happens next, allowing the generated world to evolve dynamically in response to external input.
L3: Actionable World
At L3, the model moves from the digital into the physical world. Environment understanding, future prediction and action generation work together, enabling the model to produce executable robot actions and refine its decisions based on real-world feedback.
Motus and Motubrain mark ShengShu Technology's progression from video generation toward physical action. Motus was released and fully open-sourced in December 2025. Motubrain, released in April 2026, further unifies environment understanding, world-state prediction and action-trajectory generation within a single model, advancing the GWM from visual simulation toward physical decision-making.
Compared with Motus, Motubrain delivers approximately 10× faster inference and can adapt to new robot embodiments with 50–100 human demonstrations. On the RoboTwin 2.0 benchmark, it achieved a score of 96.1, ranking first on the leaderboard.
Motubrain has also been validated across multiple robot embodiments, including robots from Galaxea AI, demonstrating generalization across long-horizon tasks, multi-task scenarios and different robot embodiments.
L4: Autonomous World Agent
At L4, the model moves toward autonomous decision-making, proactively perceiving its environment, decomposing tasks, exploring unknown states and continuously planning actions around long-term goals.
L5: World Orchestrator
At L5, autonomy extends beyond a single agent. The model coordinates robots, digital agents, humans, tools and other resources, enabling complex task allocation and multi-agent collaboration.
From L1 to L5, capabilities build progressively. Without learning and predicting how the world evolves, stable interaction is difficult; without feedback from real-world action, it is difficult to develop autonomous learning, long-term planning and complex coordination.
Toward L4 and L5: Six Key Challenges Ahead
While L1 through L3 have already seen concrete technical implementations, significant gaps remain before higher levels of world intelligence can be achieved. Video models must continue to improve generation quality, controllability and spatiotemporal consistency, while World Action Models need greater stability, generalization and execution efficiency in complex, open-ended environments.
Advancing toward L4 and L5 will require progress in six key areas:
- joint evaluation across understanding, imagination, action and transfer;
- learning physical dynamics such as contact, friction, force and the consequences of intervention;
- persistent and revisable memory;
- online learning and self-improvement through real-world interaction;
- efficient closed-loop deployment under latency and resource constraints; and
- stronger safety and controllability.
Moving forward, ShengShu Technology will continue to strengthen World Generation, Interactive World and Actionable World, while developing goal formation, active exploration, continual learning and long-term memory to lay the foundation for higher-level Autonomous World Agents.
The longer-term goal is to build a general foundation model capable of continuously understanding the world, imagining the consequences of different actions, acting autonomously under explicit constraints, and learning and evolving from real-world feedback.
When Understanding, Imagination and Action form a true closed loop, world models can evolve beyond content generation or task-specific robot policies to become a foundation connecting the digital and physical worlds—and an important step toward general embodied intelligence.
For more information on the General World Model research, visit: https://www.shengshu.com/en/general-world-model/
SOURCE ShengShu Technology
Share this article