LLMs Struggle with Real Scientific Discovery, New Benchmark Shows

Taolin Han, Yuchen Zhang, Jinghang Wang, Yun Wu, Wai Yuet Chiu, Zhaohai Li, Yifei Zhang, Jinxin Wang, Yuhao Zhou, Chen Zhao, Jiajia Li, Jiaxin Li, Qile Jin, Kewei Sun, Shuang Wu, Weiqi Zhai, Renquan Lv, Junchao Li, Ruodan Chen, Qingteng Chen, Zhibo Yang, Hu Wei, Lin Qu, Shuai Bai, Bing Zhao· August 10, 2026 View original

Key takeaways

  • Current MLLMs cannot reliably make justified, evidence-bounded inferences from experimental results.
  • The SEE benchmark highlights a significant gap in MLLMs' ability for real scientific discovery.
  • Even with tool use, MLLMs struggle to manage information within experimental evidence boundaries.
  • Future MLLM development needs to focus on deriving novel insights, not just explaining concepts.

Who benefits

PharmaceuticalsBiotechnologyMaterials ScienceChemical EngineeringAcademia

Summary

The Science Edge Evaluation (SEE) benchmark reveals that current multimodal large language models (MLLMs) struggle significantly with complex, evidence-based scientific reasoning required for real laboratory science. Even top models achieve less than 50% accuracy, highlighting a gap in deriving novel insights from experimental data.

A new benchmark called Science Edge Evaluation (SEE) has been introduced to assess the capability of large language models (LLMs) in supporting complex, real-world scientific discovery. This multimodal benchmark features expert-curated questions derived from peer-reviewed literature and experimental practices in chemistry, biology, and materials science. The evaluation of 19 different multimodal LLMs (MLLMs) showed that even the best-performing model achieved only 48.7% accuracy, indicating a significant limitation. Interestingly, general-purpose MLLMs often outperformed science-specialized models. While tool use improved accuracy slightly to 52.7% in visual-agent evaluations, it did not fundamentally solve the core challenge: MLLMs struggle to manage tool-derived information within the bounds of original experimental evidence. The findings suggest that current MLLMs are not yet capable of reliably making justified, evidence-bounded inferences from experimental results, a crucial step for genuine scientific discovery, and need to evolve from explaining established concepts to deriving novel insights.

Why it matters

Professionals relying on AI for scientific research, R&D, or complex problem-solving must recognize the current limitations of MLLMs in generating truly novel, evidence-based scientific insights.

How to implement this in your domain

  1. 1Integrate human expert review as a mandatory step for any AI-generated scientific hypotheses or conclusions.
  2. 2Focus AI applications in scientific discovery on tasks like data synthesis or hypothesis generation, rather than final inference.
  3. 3Develop specialized training datasets and architectures for MLLMs that emphasize evidence-bounded reasoning.
  4. 4Collaborate with AI researchers to bridge the gap between current MLLM capabilities and the demands of real scientific discovery.

Original post by Taolin Han, Yuchen Zhang, Jinghang Wang, Yun Wu, Wai Yuet Chiu, Zhaohai Li, Yifei Zhang, Jinxin Wang, Yuhao Zhou, Chen Zhao, Jiajia Li, Jiaxin Li, Qile Jin, Kewei Sun, Shuang Wu, Weiqi Zhai, Renquan Lv, Junchao Li, Ruodan Chen, Qingteng Chen, Zhibo Yang, Hu Wei, Lin Qu, Shuai Bai, Bing Zhao

"arXiv:2608.06931v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce Science Edge Evaluation (SEE), a multimodal benchmark…"

View on X

Originally posted by Taolin Han, Yuchen Zhang, Jinghang Wang, Yun Wu, Wai Yuet Chiu, Zhaohai Li, Yifei Zhang, Jinxin Wang, Yuhao Zhou, Chen Zhao, Jiajia Li, Jiaxin Li, Qile Jin, Kewei Sun, Shuang Wu, Weiqi Zhai, Renquan Lv, Junchao Li, Ruodan Chen, Qingteng Chen, Zhibo Yang, Hu Wei, Lin Qu, Shuai Bai, Bing Zhao on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses