VLMs Struggle with Physical Strategic Reasoning in Soccer Decisions

Jasin Cekinmez, Addison J. Wu, Haotian Xia, Akshaya Bharadhwaj, Anay Putty, Anirudh Ravishankar, Jaewoong Lee, Jinglin Xiao, Kyumin Andrew Shim, Mishika Ahuja, Nisarga Patil, Leo Liu, Zhuohan Liu, Weining Shen· July 17, 2026 View original

Key takeaways

  • VLMs currently struggle with physical strategic reasoning in dynamic environments like soccer.
  • They exhibit a bias towards lower-variance, lower-reward actions compared to human experts.
  • SportD provides a valuable benchmark for measuring strategic reasoning in VLMs.
  • Models often imitate suboptimal player actions rather than evaluating optimal alternatives.

Who benefits

Sports AnalyticsRoboticsAutonomous VehiclesGamingDefense

Summary

This research introduces SportD, a benchmark using 2022 FIFA World Cup data to evaluate Vision-Language Models (VLMs) on their ability to make strategic decisions in soccer. Findings show VLMs significantly underperform professional players, exhibiting a preference for lower-variance actions and struggling with optimal strategic choices.

While Vision-Language Models (VLMs) are adept at interpreting visual scenes, their capacity for making strategically effective decisions based on that information remains unclear. This study investigates this by focusing on soccer, specifically on-ball decisions like shooting or passing. The SportD benchmark was created using 478 on-ball decisions from the 2022 FIFA World Cup. Each VLM's chosen action is quantitatively evaluated against a possession-value model that estimates the optimal action for increasing the attacking team's scoring probability. Results show that even frontier VLMs select the highest-valued action significantly less often than professional players, incurring greater "regret" from suboptimal decisions. Models tend to prefer lower-variance, lower-reward actions, shooting less and making less progressive passes. They also partially imitate player actions, even when suboptimal, suggesting a lack of consistent counterfactual evaluation.

Why it matters

For professionals developing AI for complex, dynamic environments like sports, robotics, or autonomous systems, this research highlights current limitations of VLMs in strategic physical reasoning and the need for better models of optimal decision-making.

How to implement this in your domain

  1. 1Recognize the current limitations of VLMs in tasks requiring complex physical strategic reasoning and optimal decision-making.
  2. 2Integrate value-grounded evaluation metrics, similar to SportD's possession-value model, when developing AI for dynamic, strategic environments.
  3. 3Focus AI training on encouraging exploration of higher-reward, higher-variance actions rather than just imitating observed behavior.
  4. 4Develop hybrid AI systems that combine VLM perception with explicit strategic planning or reinforcement learning components.

Original post by Jasin Cekinmez, Addison J. Wu, Haotian Xia, Akshaya Bharadhwaj, Anay Putty, Anirudh Ravishankar, Jaewoong Lee, Jinglin Xiao, Kyumin Andrew Shim, Mishika Ahuja, Nisarga Patil, Leo Liu, Zhuohan Liu, Weining Shen

"arXiv:2607.14616v1 Announce Type: new Abstract: Vision--language models have become increasingly capable of interpreting visual scenes, but it remains unclear whether they can use information to make strategically effective decisions. We investigate this question in soccer, where…"

View on X

Originally posted by Jasin Cekinmez, Addison J. Wu, Haotian Xia, Akshaya Bharadhwaj, Anay Putty, Anirudh Ravishankar, Jaewoong Lee, Jinglin Xiao, Kyumin Andrew Shim, Mishika Ahuja, Nisarga Patil, Leo Liu, Zhuohan Liu, Weining Shen on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Research

AI ResearchAI Engineering & DevTools

New Optimizer Accelerates LLM Pretraining with Curvature-Conditioned Momentum

This research proposes a curvature-conditioned multiscale momentum method with sphere constraints to accelerate large language model pretraining. It addresses challenges from noise-dominant gradients and ill-conditioned loss landscapes by enhancing progress along flat directions, significantly improving upon existing adaptive optimizers like AdamW and Muon.

Shuchen Zhu, Yuxin Fang, Mingze Wang, Kun YuanAug 31, 2026
AI ResearchAI Engineering & DevTools

Euclidean Fourier Neural Operators Enhance Domain Transferability

This paper introduces Euclidean Fourier Neural Operators (EFNOs) as a domain-independent alternative to traditional FNOs, addressing their limitation in transferring across different periodic domains. EFNOs achieve this by parameterizing the spectral kernel as a continuous function of the physical wavevector, enabling consistent operator learning across varying domain shapes and sizes.

Nathanael Bosch, Niklas Frederik Schmitz, Michael F. HerbstAug 31, 2026
AI Engineering & DevToolsAI Research

SymboLLM-FE Boosts Feature Engineering with LLMs and Symbolic Regression

This paper introduces SymboLLM-FE, a novel approach combining symbolic regression and large language models for automated feature engineering on tabular data. It aims to generate highly interpretable and performant features while overcoming the limitations of traditional AutoFE and LLM-based methods.

Zi-Jian Cheng, Zi-Yi Jia, Zhi Zhou, Yu-Feng Li, Lan-Zhe GuoAug 31, 2026