IMPACT Improves World Models for Physically Plausible Interactions.

Rongze Tang, Jianjie Fang, Zhaolu Wang, Ziyou Wang, Xvyuan Liu, Haisheng Su, Xin Zhang, Wei Wu, Chen Gao, Yong Li, Zhibo Chen· September 2, 2026 View original

Key takeaways

  • World models struggle with physically plausible interactions due to supervision mismatch.
  • IMPACT uses cross-attention to identify and reweight dynamic interaction regions.
  • It improves interaction fidelity and physical plausibility without external representations.
  • The framework is scalable and outperforms standard MSE-trained baselines.

Who benefits

RoboticsGamingVirtual RealityAI/ML DevelopmentManufacturing

Summary

IMPACT is a new training framework that enhances world models' ability to simulate physically plausible interactions for embodied agents. It addresses a supervision-allocation mismatch in standard MSE denoising by using cross-attention to identify and reweight dynamic-object regions, improving interaction fidelity without external representations.

World models have shown significant progress in predicting future actions for embodied agents, but they often struggle to accurately model physically plausible interactions. Existing solutions typically rely on external representations like motion or geometry, which are costly to obtain. This research identifies a core problem: the globally averaged mean squared error (MSE) denoising objective disproportionately emphasizes static content, undersupervising the sparse, dynamic regions critical for interaction generation. To overcome this, the IMPACT framework (Interaction-aware Model training with Prior-guided Attention Calibration and Targeting) is introduced. IMPACT leverages cross-attention associated with manipulated-object tokens as an internal prior for action-conditioned changes. It then samples candidate regions from this prior, calibrates them with local prediction errors, and constructs an "interaction map" to reweight denoising supervision. This approach significantly improves interaction fidelity, physical plausibility, and visual quality across various robot-arm and human-hand manipulation tasks, without requiring external representations or inference-time modifications.

Why it matters

This advancement is crucial for developing more capable and reliable embodied AI agents, enabling them to perform complex manipulation tasks in real-world environments with greater accuracy and physical realism.

How to implement this in your domain

  1. 1Investigate IMPACT's attention-calibration and reweighting mechanism for improving training of your own generative models.
  2. 2Apply similar internal prior-guided supervision techniques to address data sparsity issues in dynamic environments.
  3. 3Evaluate the potential of IMPACT's approach to enhance the physical plausibility of simulations for robotics or virtual agents.
  4. 4Explore how to integrate interaction-aware training into your existing world model development pipelines.

Original post by Rongze Tang, Jianjie Fang, Zhaolu Wang, Ziyou Wang, Xvyuan Liu, Haisheng Su, Xin Zhang, Wei Wu, Chen Gao, Yong Li, Zhibo Chen

"arXiv:2609.00161v1 Announce Type: new Abstract: World models have made remarkable progress in action-conditioned future prediction for embodied agents, yet still struggle to model physically plausible interactions. Existing approaches address this limitation by constraining the g…"

View on X

Originally posted by Rongze Tang, Jianjie Fang, Zhaolu Wang, Ziyou Wang, Xvyuan Liu, Haisheng Su, Xin Zhang, Wei Wu, Chen Gao, Yong Li, Zhibo Chen on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses