ForeTime-VLA Improves Robot Grasping of Moving Objects

Siyuan Ma, Yutian Zhang, Boshi Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Xiaojin Huang· August 24, 2026 View original

Key takeaways

  • ForeTime-VLA improves robot manipulation of moving objects by anticipating future contact events.
  • It uses causal future-token distillation from a world action model to learn predictive dynamics efficiently.
  • The policy achieves significantly higher grasp success rates on conveyor belts, especially at speed.
  • This method enhances dynamic manipulation without the high computational cost of deploying a full world model.

Who benefits

ManufacturingLogisticsE-commerceRoboticsAutomation

Summary

ForeTime-VLA is a new vision-language-action (VLA) policy that significantly improves robot manipulation of moving objects by distilling future-aware, action-equivalent representations from a world action model. This causal future-token distillation allows the robot to anticipate contact events, leading to substantially higher grasp success rates on conveyor belts compared to existing methods.

Manipulating objects in motion, such as those on a conveyor belt, presents a significant challenge for robotic policies, as it requires anticipating future contact events. Traditional vision-language-action (VLA) policies often rely solely on current observations, limiting their ability to react effectively to dynamic environments. While world action models (WAMs) can predict future dynamics, deploying them at scale or explicitly imagining future frames is computationally expensive. ForeTime-VLA addresses this by introducing a dense policy that distills a future-aware, action-equivalent representation from a pre-trained, frozen Fast-WAM-derived teacher. This "causal future-token distillation" allows the policy to learn predictive dynamics offline while remaining efficient during inference. The system uses an eight-frame history encoder to predict future video latents, manipulation phase, and time-to-transition, which then condition the VLM prefix and action expert. In real-robot evaluations on a conveyor-belt dataset, ForeTime-VLA achieved significantly higher grasp success rates—81.1% for stationary and 58.9% for slow-moving objects—outperforming the next-best reference by substantial margins, especially at faster belt speeds. This demonstrates that distilling future-aware information causally is an effective strategy for enhancing dynamic manipulation capabilities without the overhead of deploying a full world model.

Why it matters

For professionals in robotics, manufacturing, and logistics, ForeTime-VLA offers a practical and efficient method to improve the reliability and speed of robotic manipulation in dynamic environments, leading to increased automation and productivity.

How to implement this in your domain

  1. 1Evaluate current robotic manipulation systems for their performance with moving objects.
  2. 2Explore integrating causal future-token distillation techniques into existing VLA policies.
  3. 3Develop or leverage world action models to generate future-aware representations for distillation.
  4. 4Conduct real-world robot trials to validate performance gains in dynamic grasping scenarios.
  5. 5Apply this approach to automate tasks on assembly lines, sorting facilities, or warehouse operations.

Original post by Siyuan Ma, Yutian Zhang, Boshi Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Xiaojin Huang

"arXiv:2608.20735v1 Announce Type: new Abstract: Manipulating moving objects requires a policy to anticipate contact events, yet vision-language-action (VLA) policies are commonly fine-tuned from the current observation alone. World action models (WAMs) learn predictive dynamics,…"

View on X

Originally posted by Siyuan Ma, Yutian Zhang, Boshi Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Xiaojin Huang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Engineering & DevTools