New Benchmark Reveals LLMs Struggle with Legal Contract Review

Yejin Bang, Kirsty Fielding, Brandan Oliver, Brian Birke, Nabeel Seedat, Andrew M. Bean· August 21, 2026 View original

Key takeaways

  • LLMs currently struggle with the nuanced task of legal contract "scrubbing."
  • Domain-specific benchmarks are essential for accurately assessing LLM performance in specialized fields.
  • High performance on general benchmarks does not translate directly to complex legal tasks.
  • Further research and development are needed to make LLMs reliable for critical legal review.

Who benefits

LegalTechLaw FirmsFinancial ServicesReal EstateCorporate Legal Departments

Summary

ContractScrub, a new benchmark, evaluates LLMs on the final review of legal contracts for errors and inconsistencies, a task known as "scrubbing." Despite expectations, frontier models perform surprisingly poorly, highlighting the need for domain-specific benchmarks to assess real-world legal AI capabilities.

A new benchmark called ContractScrub has been introduced to specifically test the capabilities of large language models (LLMs) in performing the final review of legal contracts, a process known as "scrubbing." This task involves meticulously checking transactional agreements for various errors, such as misuses of defined terms, incorrect references, and inconsistent language, which is typically a routine but painstaking job for lawyers. Despite the general perception that LLMs excel at text processing and tasks requiring long-context reasoning and consistency checks, the results on ContractScrub were unexpectedly poor. Even frontier models achieved a macro average recall of only 0.75, significantly underperforming compared to their strong results on more general benchmarks. This outcome underscores the critical importance of developing and utilizing narrowly targeted, domain-specific benchmarks to accurately gauge the practical limits and real-world impact of LLMs in specialized fields like law. It suggests that while LLMs have broad capabilities, their application in highly precise and critical domains still requires substantial improvement and tailored evaluation.

Why it matters

Legal professionals and AI developers in LegalTech need to recognize that current LLMs are not yet proficient in complex, high-stakes tasks like contract scrubbing, necessitating caution and further specialized development.

How to implement this in your domain

  1. 1Adopt domain-specific benchmarks like ContractScrub to rigorously evaluate LLM performance for legal applications.
  2. 2Invest in fine-tuning LLMs with extensive, high-quality legal datasets for specific tasks like contract review.
  3. 3Develop hybrid AI solutions that combine LLMs with rule-based systems or human-in-the-loop processes for critical legal tasks.
  4. 4Educate legal teams on the current limitations of LLMs in complex document review to manage expectations.
  5. 5Collaborate with AI researchers to contribute to the development of more robust legal AI models and benchmarks.

Original post by Yejin Bang, Kirsty Fielding, Brandan Oliver, Brian Birke, Nabeel Seedat, Andrew M. Bean

"arXiv:2608.20204v1 Announce Type: new Abstract: Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs. Contract ``scrubbing,'' the final review of transactional agreements for errors and inc…"

View on X

Originally posted by Yejin Bang, Kirsty Fielding, Brandan Oliver, Brian Birke, Nabeel Seedat, Andrew M. Bean on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Research

AI ResearchAI Engineering & DevTools

Decoding Silent Reading from Non-Invasive EEG

This research demonstrates that open-vocabulary word-level and semantic information can be reliably decoded from non-invasive EEG during silent reading. Using a contrastive decoder and a large dataset from a single participant, the study shows decoding scales log-linearly with training data and extends to rare words.

Ingo Marquardt, Anthilia Alchanat, Priyanka JainAug 21, 2026
AI ResearchAI Engineering & DevTools

Exact Learning Coefficients for Singular Models

This paper presents the first deterministic algorithm for exactly computing local learning coefficients (Real Log Canonical Thresholds) for two-dimensional singular models. This breakthrough provides ground truth for calibrating sampling-based estimators and reveals algebraic structure in learning coefficients, outperforming sampling in shallow regimes.

Gr\'egoire Sergeant-Perthuis (CQSB, Sorbonne Universit\'e), Elias Tsigaridas (Ouragan Team, INRIA), Jules Tsukahara (Ouragan Team, INRIA)Aug 21, 2026
AI Engineering & DevToolsAI Research

Standardized ML Evaluation for Power System Protection

This paper proposes a standardized framework for evaluating machine learning applications in power system protection, addressing inconsistencies in current research. It defines seven critical study dimensions and instantiates the framework with a case study on fault classification and localization using a public benchmark.

Julian Oelhaf, Georg Kordowich, Paula Andrea P\'erez-Toro, Christian Bergler, Johann J\"ager, Andreas Maier, Siming BayerAug 21, 2026