New RL Framework Improves Spatial Reasoning, Controls Response Length

Xingjian Tao, Yiwei Wang, Yujun Cai, Jing Tang· July 21, 2026 View original

Summary

This paper introduces LenGuard-GPC, a dense reward framework for reinforcement learning that enhances multi-view spatial reasoning in vision-language models while controlling the verbosity of chain-of-thought responses. It uses token-wise predictive distribution comparisons and a staged length bonus to improve accuracy and manage response length.

Multi-view spatial reasoning tasks, which require vision-language models to compare visual information across images and infer spatial relationships, often lead to overly verbose chain-of-thought reasoning without a corresponding increase in accuracy. Traditional reinforcement learning approaches, relying on sparse outcome-level feedback, struggle to pinpoint reasoning errors or control response length. To address these challenges, researchers propose LenGuard-GPC, a novel dense reward framework. This framework calculates a token-sum KL divergence by comparing predictive distributions under a standard prompt and a guided prompt, providing a granular reward signal. Furthermore, LenGuard-GPC incorporates a staged length bonus. This mechanism prevents the KL penalty from simply favoring shorter responses, ensuring that reasoning length remains within a controlled, optimal range without sacrificing quality. Experiments on six multi-view spatial reasoning benchmarks show that LenGuard-GPC improves accuracy over vanilla GRPO while simultaneously reducing average response length.

Why it matters

For AI systems requiring complex reasoning, this framework offers a way to achieve more accurate and concise outputs, improving efficiency and user experience in applications like robotics, autonomous systems, and advanced visual assistants.

How to implement this in your domain

  1. 1Experiment with LenGuard-GPC's dense reward mechanism to fine-tune existing vision-language models for spatial reasoning tasks.
  2. 2Integrate the staged length bonus into custom reinforcement learning pipelines to manage output verbosity.
  3. 3Apply the guided-prompt consistency concept to improve reasoning quality in other complex AI tasks.
  4. 4Evaluate the framework's impact on user interaction and computational efficiency in deployed AI agents.

Who benefits

RoboticsAutonomous VehiclesAI/ML DevelopmentGamingDefense

Key takeaways

  • LenGuard-GPC improves spatial reasoning accuracy in vision-language models.
  • It uses a dense reward framework based on token-wise predictive distribution comparisons.
  • The framework effectively controls the length of reasoning responses.
  • This approach can lead to more efficient and user-friendly AI outputs.

Original post by Xingjian Tao, Yiwei Wang, Yujun Cai, Jing Tang

"arXiv:2607.17243v1 Announce Type: new Abstract: Multi-view spatial reasoning requires vision-language models to compare visual evidence across images, align object correspondences, and infer spatial relations over long visual contexts, a setting where chain-of-thought reasoning t…"

View on X

Originally posted by Xingjian Tao, Yiwei Wang, Yujun Cai, Jing Tang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses