Recently, ShengShu Technology, in collaboration with Tsinghua University, has officially open-sourced Motus — a unified, general-purpose foundation world model built on video foundation models and integrating multiple technical approaches. As a pivotal component of ShengShu Technology’s embodied world model strategy, Motus achieves a 40% improvement in success rate over the internationally leading VLA model Pi0.5 across multiple core tasks, validating an effective path for scaling up embodied foundation models and laying the critical technical groundwork for embodied intelligence to advance toward a unified agent form.
Motus was published and fully open-sourced in December 2025, ahead of the industry by two months. Since then, both industry and academia have produced related research findings, further confirming the foresight and practical value of this technical approach in embodied intelligence. Notably, around the core thesis of “using video models as the unified representation foundation for embodied intelligence,” ShengShu Technology and Tsinghua University had already published related work on the Vidar embodied video model as early as July 2025 — half a year ahead of the industry.
ShengShu Technology has consistently maintained that multimodal models are the core pathway toward Artificial General Intelligence (AGI). Video naturally encodes the physical spatiotemporal dynamics, causal logic, and dynamic evolution of the real world, making it the ideal representational foundation connecting perception and action.
Building on this forward-looking strategy, ShengShu Technology is committed to breaking the fragmented paradigm in embodied intelligence — where perception, reasoning, and action have long been siloed — and constructing a unified, general-purpose foundation “world model” that integrates multimodal generation. By deeply fusing the “imagination” of video generation with the “execution” of robotic manipulation, we aim to build an intelligent hub akin to the “human whole brain.”
The general-purpose embodied agent that ShengShu pursues should, like humans, simultaneously possess the ability to understand scenes and language instructions, reason, predict future states and action consequences, and ultimately execute actions — all within a single unified system. However, existing embodied intelligence models still face two core bottlenecks.
On one hand, multimodal generation capabilities have long been fragmented: VLAs, video generation models, world models, and VLMs each operate independently, with models typically covering only a single link in the perception–reasoning–action chain, making it difficult to form a unified closed loop.
On the other hand, models have limited ability to leverage large-scale heterogeneous data, with training heavily reliant on expert-collected task trajectories, resulting in constrained generalization and transfer capabilities.
Motus unifies the multimodal generation capabilities of different paradigms — world models, VLAs, video generation models — within a single framework, while fully mining physical interaction knowledge from large-scale heterogeneous data. Notably, Motus’s unified world model takes a different approach from World Labs and Genie3. Motus is not merely about rendering simulations — it is an end-to-end unified model, like the human whole brain: capable of perceiving and understanding, imagining and reasoning, and executing actions, thereby truly entering the physical world and helping humans get things done.
At the architectural level, Motus is the first to unify five mainstream embodied foundation model paradigms — VLA, world models, video generation models, inverse dynamics models, and video–action joint generation models — into a single framework, building a unified modeling pathway that connects perception, reasoning, and action.
At the data level, Motus further breaks through the real-robot data bottleneck that has long constrained embodied intelligence development. Faced with data scarcity, heterogeneous sources, and sharing difficulties, Motus unifies the action space across cross-embodiment robot data, task-agnostic data, synthetic simulation data, human demonstration videos, and internet videos. Through large-scale pre-training, it shares rich motion priors, providing the foundation for cross-task and cross-platform transfer.
On LinkedIn, numerous AI and robotics professionals have given Motus highly positive reviews, widely recognizing it as a key paradigm shift from “modular assembly systems” to a “unified intelligent agent architecture” in embodied intelligence. Its Latent Action + unified modeling approach is seen as both clever and pragmatic, effectively bridging the path from the “digital world” to the “physical world.”
This time, robots can finally perceive, reason, and act simultaneously — just like humans!
Both Motus code and model weights are open-sourced:
Paper: https://arxiv.org/abs/2512.13030
Project Page: https://motus-robotics.github.io/motus
Code & Model Weights: GitHub: https://github.com/thu-ml/Motus; Hugging Face: https://huggingface.co/motus-robotics
01 Outperforming Pi0.5 by 40%: Motus Leads a New Paradigm in Multi-Task Generalization and Scale-Up for Embodied Intelligence
Motus demonstrates significant advantages in efficient use of sampled data. In the Scaling Curves experiments, Motus exhibits superior scaling properties across both data scale and task scale dimensions.
Even more notably, Motus was the first to discover a new dimension in embodied intelligence scaling — the multi-task generalization curve. This curve provides a critical North Star metric for the development of embodied foundation models. Its evolution path closely mirrors that of language models, echoing the core insight proposed by GPT-2: Language Models are Unsupervised Multitask Learners.
In the Data Scaling experiments, compared to the internationally leading VLA model Pi0.5, Motus can learn from a broader range of data types and effectively incorporate prior capabilities from additional pre-trained foundation models. On the average success rate across 50 tasks, Motus achieves a 35.1% absolute improvement over Pi0.5, while demonstrating 13.55× data efficiency at the same performance level.
These results indicate that, under the Scaling Law, Motus can more efficiently develop more general intelligence capabilities by introducing richer and more heterogeneous multimodal priors.
In the Task Number Scaling experiments, as the number of tasks continues to increase, Motus’s average success rate shows a steady upward trend, while Pi0.5’s success rate continues to decline as task complexity grows. Ultimately, Motus achieves a 37% absolute improvement in success rate over Pi0.5.
This qualitative leap in multi-task performance demonstrates that Motus’s “unified” modeling approach can learn shareable structural knowledge and a unified world understanding across different tasks; in contrast, traditional VLA models are more prone to overfitting on single-task action patterns.
02 Simulation Evaluation: 88% Success Rate, Setting a New Bar for Complex Embodied Manipulation
To further validate the transferability of Motus’s performance advantages to specific tasks and real-world scenarios, the research team conducted systematic evaluations across simulation environments and real-robot deployments, achieving outstanding results on multiple key metrics.
In the RoboTwin 2.0 simulation environment covering 50 general tasks, using the official 27,500 mixed training episodes (spanning both Clean and Randomized scenarios), Motus achieved approximately 88% average success rate, standing out among models of its class.
Even in the extremely challenging stack bowls three task — which demands high alignment precision and dynamic balance, where slight deviations cause complete collapse and has long plagued robots with frequent “shaky hand” errors — baseline models’ best success rate never exceeded 16%. In contrast, Motus achieved a success rate of 91%–95%, representing an order-of-magnitude performance leap.
This performance is not an isolated case. Across multiple high-difficulty tasks where baseline models are constrained or even fail to converge, Motus consistently demonstrates significant and stable performance improvements, continually pushing the success rate ceiling and validating its strong generalization capability and performance ceiling in complex manipulation scenarios.
03 Real-Robot Evaluation: From “Mechanical Execution” to “End-to-End Intelligence”
In real-robot experiments, Motus was deployed on two different embodiment platforms — AC-One and Agilex-Aloha-2 — and achieved significantly higher task success rates than industry-leading models such as Pi0.5 across multiple tasks. Particularly noteworthy is Motus’s ability to continuously benefit from large-scale pre-training, demonstrating stronger cross-task and cross-embodiment generalization.
Additionally, in tasks such as Touch Instructed Keyboard, Place Cube into Plate, and Put Bread into Oven, Motus demonstrates a deep understanding of spatial relationships between objects, planning coherent and reasonable motion trajectories that further showcase its fine-grained vision–action grounding capability with high coordination.
In a series of long-horizon, multi-step complex real-robot tasks, Motus further exhibits near-human-level decision logic and execution stability. It is important to emphasize that these are not simple single-step commands, but rather typical long-horizon, multi-step tasks completed end-to-end by the model, without relying on traditional “fast-slow dual system” decomposition.
In the Cloudflare CAPTCHA verification scenario, faced with an ergonomic mouse, Motus can precisely locate the contact point on the irregular curved surface of the mouse, leverage its visual understanding to gauge the distance between the mouse and the Cloudflare click target, smoothly and continuously move the mouse, and finally click precisely — successfully solving the Cloudflare challenge from the physical world.
In multi-step tasks involving complex rules such as peg solitaire, Motus achieves a complete closed loop from understanding and generation to execution. Through real-time comprehension and analysis of the ever-changing board state, Motus can imagine rule-compliant moves and precisely project them into the physical world, executing consecutive reasonable moves.
In complex household scenarios involving deformable objects, such as folding clothes and towels, Motus can predict fabric deformation in real time and perform continuous, fine-grained action control under an end-to-end framework. Since fabric continuously undergoes uncertain shape changes under external forces, the robot must constantly perform state inference and action adjustment. Motus can stably interact with deformable objects, completing coherent folding sequences and breaking through the long-standing bottleneck of traditional methods in deformable object manipulation.
Strategic Vision: Open-Sourcing a Unified World Model, a Critical Step Toward AGI in the Physical World
As a core engine in ShengShu Technology’s global strategic roadmap, the open-sourcing of Motus marks a milestone. This is not merely the external release of a cutting-edge technical achievement, but a key inflection point for embodied intelligence — from “single-capability breakthroughs” toward “unified foundation model system construction.”
The release of Motus brings robots closer than ever to human-like foundational cognition and action paradigms: understanding the current environment and language instructions, imagining future states and action consequences, and executing coherent, generalizable actions in the real physical world. This goes beyond leading model metrics — it represents a critical leap for AGI from the digital world into the physical world.
We firmly believe that the true flourishing of embodied intelligence should not be built on closed, fragmented technical leadership, but on scalable, reusable unified foundation models.
Through open-sourcing Motus, ShengShu Technology provides global developers and industry partners with a high-performance, highly compatible “universal physical intelligence foundation”: whether industrial robotic arms, commercial service robots, or legged and mobile robots, all can continuously evolve under the same “perception–reasoning–action” world model paradigm.
From video models to world models; from content generation to physical intelligence. ShengShu Technology is enabling AI to not just understand the world, but to truly change it.