SciTrue Achieves Top Performance in Scientific Claim Validation

Qiming Bao, Ne\c{s}et \"Ozkan Tan, Siyuan Wang, Mark Gahegan· September 2, 2026 View original

Key takeaways

  • Combining multiple frontier and open LLMs with light post-processing excels in scientific claim validation.
  • Instruction-tuned models are highly competitive for complex text understanding tasks.
  • Data pairing structures and avoiding data leaks are crucial for achieving high accuracy.
  • Current models are approaching the limits of what noisy datasets allow for this task.

Who benefits

Scientific ResearchPublishingPharma & BiotechLegalTechAI/ML Development

Summary

The SciTrue team achieved first place in most categories of the NTCIR-19 SciClaimEval task by benchmarking multiple frontier and open multimodal models with transparent post-processing. Key findings include the strong performance of instruction-tuned models and the significant impact of a leak-free pair prior.

The SciTrue team participated in the NTCIR-19 SciClaimEval task, which challenges systems to verify scientific claims against tables and figures within research papers. Instead of focusing on a single model, SciTrue benchmarked eleven different frontier and open multimodal models using a consistent, per-sample protocol, then combined their outputs with minimal, transparent post-processing. This approach led to SciTrue securing first place in three out of four evidence-category/subtask combinations and tying for first in the fourth on the official leaderboard. Three main factors contributed to their success. Firstly, powerful instruction-tuned models like Claude Opus 4.8 and Gemma-4-31B proved highly competitive, with GPT-5.5 and Claude Fable 5 leading both subtasks. Secondly, the most impactful factor was a "leak-free pair prior" that could infer Supported/Refuted pairings directly from the claim text, significantly boosting accuracy. Lastly, a detailed audit revealed that most remaining errors were due to visually undetectable label-mapping issues or dataset noise, suggesting that current models are already performing near the upper bound of what the dataset allows, with limited room for further improvement through modeling alone.

Why it matters

This research demonstrates the current state-of-the-art in automated scientific claim validation, highlighting the effectiveness of combining multiple advanced LLMs and the critical role of data handling in achieving high accuracy.

How to implement this in your domain

  1. 1Explore ensemble methods combining multiple frontier and open-source LLMs for complex information extraction tasks.
  2. 2Prioritize robust data preprocessing and "leak-free" feature engineering to maximize model performance.
  3. 3Conduct thorough error analysis to distinguish between model limitations and dataset quality issues.
  4. 4Consider using advanced instruction-tuned models for scientific text analysis and verification.

Original post by Qiming Bao, Ne\c{s}et \"Ozkan Tan, Siyuan Wang, Mark Gahegan

"arXiv:2609.00654v1 Announce Type: new Abstract: We describe the SciTrue team's participation in both subtasks of the NTCIR-19 SciClaimEval task~\cite{sciclaimeval}, which asks systems to verify scientific claims against the tables and figures of a paper. Rather than tuning a sing…"

View on X

Originally posted by Qiming Bao, Ne\c{s}et \"Ozkan Tan, Siyuan Wang, Mark Gahegan on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses