LLMs Show Human-Like Mentalization in Economic Games

Aamir Sohail, Xintong Zhong, Arkady Konovalov, Patricia L. Lockwood, Lei Zhang· August 28, 2026 View original

Key takeaways

  • LLMs exhibit clear mentalization capabilities, inferring others' beliefs.
  • GPT-5 agents can adapt reasoning depth and sometimes outperform humans.
  • Strategic prompting improves LLM performance in social reasoning tasks.
  • Computational modeling is a valuable tool for comparing human and AI intelligence.

Who benefits

Customer ServiceHealthcareEducationSocial RoboticsMarketing

Summary

This research assesses mentalization, the ability to infer others' beliefs, in humans and LLMs using economic games and computational modeling. It found that LLMs, particularly GPT-5, exhibit clear behavioral and computational signatures of mentalizing, adapting their reasoning depth to opponents and sometimes outperforming humans.

This study explores "mentalization," the human cognitive ability to infer others' beliefs and intentions, by comparing it in human participants and various large language models (LLMs). Researchers used two economic games and cognitive computational modeling to uncover the latent strategies employed by individual LLM agents from families like DeepSeek, GPT-4.1, GPT-5, and Gemini 2.0 Flash. These LLMs were tested against opponents of varying sophistication, and the impact of strategic prompting on their performance was also examined. The findings revealed that LLMs demonstrated clear behavioral and computational signatures of mentalizing, though capacities varied significantly across model providers and sizes. Strategic prompting generally improved performance by eliciting more sophisticated reasoning. Notably, GPT-5 agents exhibited a flexible adaptation of their recursive depth of reasoning to increasingly sophisticated opponents, even outperforming human participants in some instances. This research highlights the potential of computational modeling to formally assess comparative intelligence between humans and machines.

Why it matters

Understanding LLMs' capacity for mentalization is crucial for developing more sophisticated and socially aware AI, impacting human-AI collaboration, ethical AI design, and applications requiring nuanced social interaction.

How to implement this in your domain

  1. 1Design AI agents for customer service or negotiation to incorporate basic mentalization principles, inferring user intent and adapting responses.
  2. 2Utilize strategic prompting techniques to elicit more sophisticated reasoning from LLMs in complex decision-making scenarios.
  3. 3Develop evaluation metrics for AI systems that go beyond task completion to assess their ability to understand and adapt to human social cues.
  4. 4Explore the use of advanced LLMs like GPT-5 for applications requiring adaptive social intelligence, such as personalized tutoring or virtual assistants.
  5. 5Conduct internal research to benchmark the mentalization capabilities of different LLMs for specific business use cases.

Original post by Aamir Sohail, Xintong Zhong, Arkady Konovalov, Patricia L. Lockwood, Lei Zhang

"arXiv:2608.26291v1 Announce Type: new Abstract: Mentalization - the ability to infer others' beliefs and intentions to guide one's own choices - is a key cognitive function underlying human social interactions. Large language models (LLMs) demonstrate behaviour consistent with hu…"

View on X

Originally posted by Aamir Sohail, Xintong Zhong, Arkady Konovalov, Patricia L. Lockwood, Lei Zhang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Research

AI ResearchAI Engineering & DevTools

Emotional Preferences Regulate Goal Priorities in Reinforcement Learning Agents

This paper proposes a computational framework where higher-level goals autonomously generate state-dependent emotional preferences to regulate the priorities of competing lower-level objectives in reinforcement learning agents. It demonstrates how this emergent preference function exhibits contextual priority switching and improves performance over fixed-preference strategies in multi-objective exploration environments.

Shiqi Liu, Yihua Tan, Hu Fu, Guanyu QiAug 28, 2026
AI Engineering & DevToolsAI Research

New Framework Unifies Task Detection and Adaptation for Continual Learning

This paper proposes FiUni, a Fisher-guided unified framework for task-free continual learning in LLMs that combines batch-level task detection with parameter-efficient adaptation. FiUni uses Fisher information matrix (FIM) properties to dynamically determine whether to reuse, expand, or create new low-rank adaptation (LoRA) subspaces, effectively mitigating catastrophic forgetting without explicit task boundaries.

Dezheng Han, Anbang Zhang, Zhihao Zhu, Shuaishuai GuoAug 28, 2026
AI Engineering & DevToolsAI Research

Soft EMG Interface Enables Machine Learning-Powered Silent Speech Recognition

This paper introduces a soft, active electromyography (EMG) interface worn on the hand that enables word-level silent speech recognition (SSR) using machine learning. The device acquires stable EMG signals from a fingertip electrode near the lips, achieving 97.2% accuracy on a 30-word vocabulary and demonstrating real-time drone control in noisy environments.

Yuta Kurotaki, Shusuke Yamakoshi, Reitaro Yoshida, Yutaka Isoda, Tamami Takano, Yuji Isano, Yusuke Miyake, Kentaro Kuribayashi, Hiroki OtaAug 28, 2026