Open Models Struggle with High-Risk Information Extraction.

Elias Schubert, Felix Bie{\ss}mann· August 20, 2026 View original

Key takeaways

  • Open-source LLMs and VLMs struggle with high-risk structured information extraction in zero-shot settings.
  • VLMs generally outperform OCR+LLM pipelines, but overall reliability remains low.
  • Model scale does not guarantee proportional performance improvements.
  • Input quality and OCR accuracy are critical factors for extraction success.

Who benefits

GovernmentPublic SectorBFSIHealthcareLegal

Summary

A benchmark evaluating open-source OCR, LLMs, and VLMs for structured information extraction in a high-risk public sector application (student applications) found that most models struggle. While VLMs generally outperformed OCR+LLM pipelines, only a few configurations achieved acceptable F1 scores, highlighting challenges in reliability for critical tasks.

This paper presents a comprehensive benchmark evaluating the performance of open-source Optical Character Recognition (OCR) engines, Large Language Models (LLMs), and Vision-Language Models (VLMs) in a complex, high-risk public sector application: extracting structured information from student applications for an international study program. This task is particularly critical given the EU AI Act's classification of such applications as high-risk. The study aimed to bridge the gap in systematic evaluations of end-to-end extraction pipelines using open models. The empirical evaluation revealed that while VLMs generally showed better performance than pipelines combining OCR with LLMs, the overall reliability of state-of-the-art open-source models in zero-shot settings remains a significant challenge. Only 4 out of 35 configurations achieved an F1 score above 0.5, with roughly 75% scoring below 0.25. The research also noted that model scale does not guarantee proportionally better results, and input quality, particularly the structural integrity of OCR output, is a crucial factor influencing downstream performance, often more so than the LLM or VLM capabilities themselves.

Why it matters

Professionals deploying AI for critical information extraction in regulated or high-stakes environments must be aware that open-source models, even state-of-the-art ones, may not yet offer sufficient reliability without significant fine-tuning or human-in-the-loop processes.

How to implement this in your domain

  1. 1Conduct rigorous, task-specific benchmarks for any AI system intended for high-risk information extraction, especially with open-source models.
  2. 2Prioritize improving input quality and OCR accuracy as a foundational step for any document processing pipeline.
  3. 3Design human-in-the-loop validation processes for critical extraction tasks to mitigate the unreliability of current open models.
  4. 4Investigate fine-tuning open-source models on domain-specific data rather than relying solely on zero-shot performance for high-risk applications.

Original post by Elias Schubert, Felix Bie{\ss}mann

"arXiv:2608.18289v1 Announce Type: new Abstract: The extraction of structured information from unstructured documents represents a critical component of digital transformations in all sectors. While proprietary solutions dominate commercial applications, a rapidly growing ecosyste…"

View on X

Originally posted by Elias Schubert, Felix Bie{\ss}mann on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses