LLMs Struggle with Information Discernment from External Sources

Joshua Ashkinaze, Laura Kurek, Alina Faisal, Tongyuan Miao, Mariam Joseph, Ceren Budak, Eric Gilbert· July 23, 2026 View original

Summary

This research introduces Learn2Discern (L2D), a framework and benchmark to evaluate how LLMs weigh information from external sources. It finds consistent failures in source and truth discernment across 13 models, showing models struggle to prioritize reliable sources or update beliefs appropriately based on new claims.

As large language models (LLMs) increasingly rely on external knowledge sources like the internet, their ability to appropriately weigh information becomes critical. This paper formalizes the concept of "information discernment" and introduces Learn2Discern (L2D), an experimental framework and benchmark designed to evaluate how LLMs assess source reliability and update their beliefs based on new claims. A user study confirmed that human users endorse the three normative axioms underpinning L2D and report reduced trust and usage intent when these axioms are violated by LLMs. Across 13 models and nearly 670,000 trials, the study revealed consistent failures in both source and truth discernment. Models performed near chance, often relying on source popularity over reliability and updating their positions equally regardless of whether a claim improved or worsened their accuracy relative to ground truth. The research highlights that LLMs integrate external knowledge most effectively when their prior beliefs are already accurate. While newer and larger models show some improvement in truth discernment, source discernment remains a significant blind spot. The authors also identify simple inference-time interventions that can improve both forms of discernment, offering practical steps towards more trustworthy LLM behavior.

Why it matters

Professionals relying on LLMs for information synthesis and decision-making need to be aware of these models' limitations in discerning reliable information, emphasizing the need for critical human oversight and validation of AI-generated content.

How to implement this in your domain

  1. 1Implement human-in-the-loop validation for critical information generated by LLMs, especially when external sources are involved.
  2. 2Develop internal guidelines for evaluating the reliability of sources cited or implicitly used by LLMs in your workflows.
  3. 3Explore fine-tuning or prompt engineering techniques to encourage LLMs to prioritize reliable sources and demonstrate better truth discernment.
  4. 4Advocate for and adopt LLM solutions that incorporate mechanisms for source attribution and confidence scoring.

Who benefits

MediaConsultingResearchEducationLegal

Key takeaways

  • LLMs consistently fail at source and truth discernment from external knowledge.
  • Models often prioritize source popularity over reliability.
  • Newer and larger models improve truth discernment but not source discernment.
  • Simple interventions can improve LLM information discernment.

Original post by Joshua Ashkinaze, Laura Kurek, Alina Faisal, Tongyuan Miao, Mariam Joseph, Ceren Budak, Eric Gilbert

"arXiv:2607.19355v1 Announce Type: new Abstract: LLMs are increasingly used with external knowledge sources like the internet. Do they weigh information appropriately -- updating more for reliable sources (source discernment) and more when claims bring priors closer to the truth (…"

View on X

Originally posted by Joshua Ashkinaze, Laura Kurek, Alina Faisal, Tongyuan Miao, Mariam Joseph, Ceren Budak, Eric Gilbert on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses