LLMs Improve Evidence Use, Not Information Seeking, Under Uncertainty

Hua-Dong Xiong, Xinyuan Yan, Ji-An Li, Jingming Xue, Marcelo G. Mattar, Robert C. Wilson· July 30, 2026 View original

Summary

Research shows that 'thinking' in large language models primarily strengthens their ability to use existing evidence and reduces choice noise under uncertainty, rather than increasing active information-seeking behaviors.

This research investigates how "inference-time thinking" affects large language models (LLMs) when faced with uncertainty. By observing action preferences, thinking length, and reported confidence in a controlled two-armed bandit task, the study aimed to differentiate between improved evidence utilization and active information-seeking. Ten open-weight models were tested in both thinking and non-thinking modes. The findings indicate that thinking primarily enhances the models' ability to act on current evidence more effectively and reduces random choice noise. However, it did not lead to increased UCB-like exploration (preference for less-known options) or stronger Thompson-like exploration (choice variability increasing with uncertainty). The study also noted that thinking length increased with information-imbalanced histories, and reported confidence became more sensitive to decision difficulty, suggesting metacognitive control and monitoring processes at play, though not definitively establishing them as such.

Why it matters

Understanding how LLMs process information and make decisions under uncertainty is crucial for developing more reliable and interpretable AI systems, especially for tasks requiring nuanced judgment or strategic planning.

How to implement this in your domain

  1. 1Design LLM prompts that encourage "thinking" to improve evidence utilization in decision-making tasks.
  2. 2Develop evaluation metrics that specifically assess an LLM's ability to act on available evidence under uncertainty.
  3. 3Explore how to explicitly integrate information-seeking mechanisms into LLM architectures, as current "thinking" doesn't naturally foster it.
  4. 4Consider the implications for building AI agents that need to balance exploitation of known information with exploration of new information.

Who benefits

AI DevelopmentCognitive ScienceRoboticsDecision Support Systems

Key takeaways

  • LLM "thinking" improves evidence use and reduces choice noise under uncertainty.
  • It does not inherently increase information-seeking behaviors like UCB or Thompson sampling.
  • Thinking length and confidence patterns suggest metacognitive processes.
  • Decoder settings like temperature can alter behavior but don't replicate thinking's effects.

Original post by Hua-Dong Xiong, Xinyuan Yan, Ji-An Li, Jingming Xue, Marcelo G. Mattar, Robert C. Wilson

"arXiv:2607.26845v1 Announce Type: new Abstract: Inference-time thinking improves the performance of large language models, but aggregate outcomes do not reveal whether models use available evidence more effectively or seek information that could improve future decisions. We disti…"

View on X

Originally posted by Hua-Dong Xiong, Xinyuan Yan, Ji-An Li, Jingming Xue, Marcelo G. Mattar, Robert C. Wilson on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses