LLMs Streamline Systematic Literature Reviews for Disease Models

Orhan Yagizer Cinar, Timur Emre Ozkose, Emma Von Hoene, Amira Roess, Taylor Anderson, Hamdi Kavak· August 28, 2026 View original

Key takeaways

  • LLMs can significantly streamline systematic literature reviews, achieving high paper-level accuracy.
  • Accuracy varies by field complexity, with subjective fields performing less reliably.
  • LLM agreement can indicate output quality, helping identify hallucinations or data errors.
  • Human oversight remains essential for validating and refining LLM-generated content in SLRs.

Who benefits

AcademiaPharmaceuticalsPublic HealthResearch & DevelopmentConsulting

Summary

This study develops an LLM pipeline for extracting information from agent-based modeling papers for systematic literature reviews (SLRs) on disease spread models. It achieved paper-level accuracies of 77.95% for GPT-4.1 and 81.67% for GPT-5.0, highlighting LLMs' potential to automate research processes while noting limitations in complex or subjective fields.

Recent advances in Large Language Models (LLMs) offer new avenues for automating research processes, including systematic literature reviews (SLRs). This study focused on developing an LLM pipeline to extract model-relevant information from 536 peer-reviewed papers on agent-based disease spread modeling. The results were then compared against a human-conducted SLR. The LLM pipeline achieved paper-level accuracies of approximately 77.95% with GPT-4.1 and 81.67% with GPT-5.0. While field-level accuracy varied significantly (32.40% to 100.00%), more complex or subjective data fields proved less reliable for LLMs. A key finding was that agreement between different LLMs could serve as an indicator of output quality: low agreement might signal hallucinations, while high agreement with low accuracy could point to errors in the human-annotated dataset. The study provides practical insights into prompt engineering and underscores both the potential and limitations of using LLMs for full-scale SLRs in modeling and simulation.

Why it matters

Researchers and analysts can significantly accelerate the initial stages of systematic literature reviews by leveraging LLMs, freeing up human experts for more complex synthesis and critical evaluation. Understanding LLM limitations is crucial for effective implementation.

How to implement this in your domain

  1. 1Develop structured prompts for LLMs to extract specific information from research papers.
  2. 2Pilot LLM-based information extraction on a small subset of papers to refine prompts and evaluate accuracy.
  3. 3Implement a multi-LLM approach to cross-validate outputs and identify potential hallucinations or errors.
  4. 4Integrate human review and correction workflows for complex or subjective data fields.
  5. 5Utilize LLMs to generate initial summaries or data points, then have human experts perform the final synthesis and critical analysis.

Original post by Orhan Yagizer Cinar, Timur Emre Ozkose, Emma Von Hoene, Amira Roess, Taylor Anderson, Hamdi Kavak

"arXiv:2608.26150v1 Announce Type: new Abstract: Recent advancements in Large Language Models (LLMs) have created new opportunities to streamline and potentially automate many research processes, including systematic literature reviews (SLRs). This study reports an LLM pipeline de…"

View on X

Originally posted by Orhan Yagizer Cinar, Timur Emre Ozkose, Emma Von Hoene, Amira Roess, Taylor Anderson, Hamdi Kavak on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Research

AI ResearchAI Engineering & DevTools

Emotional Preferences Regulate Goal Priorities in Reinforcement Learning Agents

This paper proposes a computational framework where higher-level goals autonomously generate state-dependent emotional preferences to regulate the priorities of competing lower-level objectives in reinforcement learning agents. It demonstrates how this emergent preference function exhibits contextual priority switching and improves performance over fixed-preference strategies in multi-objective exploration environments.

Shiqi Liu, Yihua Tan, Hu Fu, Guanyu QiAug 28, 2026
AI Engineering & DevToolsAI Research

New Framework Unifies Task Detection and Adaptation for Continual Learning

This paper proposes FiUni, a Fisher-guided unified framework for task-free continual learning in LLMs that combines batch-level task detection with parameter-efficient adaptation. FiUni uses Fisher information matrix (FIM) properties to dynamically determine whether to reuse, expand, or create new low-rank adaptation (LoRA) subspaces, effectively mitigating catastrophic forgetting without explicit task boundaries.

Dezheng Han, Anbang Zhang, Zhihao Zhu, Shuaishuai GuoAug 28, 2026
AI Engineering & DevToolsAI Research

Soft EMG Interface Enables Machine Learning-Powered Silent Speech Recognition

This paper introduces a soft, active electromyography (EMG) interface worn on the hand that enables word-level silent speech recognition (SSR) using machine learning. The device acquires stable EMG signals from a fingertip electrode near the lips, achieving 97.2% accuracy on a 30-word vocabulary and demonstrating real-time drone control in noisy environments.

Yuta Kurotaki, Shusuke Yamakoshi, Reitaro Yoshida, Yutaka Isoda, Tamami Takano, Yuji Isano, Yusuke Miyake, Kentaro Kuribayashi, Hiroki OtaAug 28, 2026