DECOWAM Enhances Legged Robot Mobile Manipulation with Decoupled Model.

Siyuan Ma, Boshi Zhang, Yutian Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Qiaojun Yu· August 21, 2026 View original

Key takeaways

  • DECOWAM is a new whole-body world-action model for legged mobile manipulation.
  • It decouples camera ego-motion from base and arm actions for improved prediction.
  • The model enhances future video and action prediction, leading to better coordination and robustness.
  • This approach is a significant step towards more capable and adaptable mobile manipulation robots.

Who benefits

RoboticsLogisticsManufacturingExplorationDefense

Summary

Researchers introduce DECOWAM, a whole-body world-action model for legged mobile manipulation that decouples camera ego-motion from base and arm actions. This model improves future video and action prediction, achieving better whole-body coordination and robustness in real-robot experiments.

Mobile manipulation for legged robots presents a significant challenge: predicting how the combined movements of locomotion and arm actions will affect future observations and control. Existing world-action models, largely developed for stationary platforms, often fail to explicitly differentiate between the robot's camera ego-motion and the independent actions of its base and arm. To address this, a new model called DECOWAM (Decoupled Whole-Body World-Action Model) has been developed. DECOWAM separates these factors through dedicated conditional interfaces, freezing a FastWAM backbone and training residual adapters for efficiency. The model incorporates an action-equivalent future bottleneck, adversarially separated base and arm latents, and base-velocity conditioning for video prediction. To facilitate this research, a new real-robot dataset, ARMDOG, was created, synchronizing video, whole-body state, action, and language. In experiments, DECOWAM improved both future-video and action prediction over FastWAM, reducing action Mean Squared Error by 21.7% with a relatively small number of trainable parameters. Crucially, it demonstrated superior whole-body coordination and base-displacement robustness in closed-loop real-robot trials, indicating a significant step forward in embodiment-aware factorization for joint visual prediction and whole-body control under moving viewpoints.

Why it matters

This research significantly advances the capabilities of legged mobile manipulation robots, making them more robust and coordinated in complex, dynamic environments, which is critical for real-world deployment in various industries.

How to implement this in your domain

  1. 1Adopt decoupled world-action models like DECOWAM for developing advanced control systems in legged mobile robots.
  2. 2Utilize embodiment-aware factorization techniques to improve joint visual prediction and whole-body control in robotic platforms.
  3. 3Leverage datasets like ARMDOG for training and evaluating mobile manipulation models in real-world scenarios.
  4. 4Explore integrating residual adapters and adversarial separation of latents for efficient model adaptation and improved performance.

Original post by Siyuan Ma, Boshi Zhang, Yutian Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Qiaojun Yu

"arXiv:2608.20114v1 Announce Type: new Abstract: Mobile manipulation requires a robot to predict how locomotion and arm motion jointly alter future observations and control. Existing world-action models, developed largely for fixed-base platforms, do not explicitly distinguish cam…"

View on X

Originally posted by Siyuan Ma, Boshi Zhang, Yutian Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Qiaojun Yu on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses