MLLMs and Visualization: Distinguishing Chart-Supported from Model-Supplied Claims

Ishrat Jahan Eliza, Md Dilshadur Rahman· July 29, 2026 View original

Summary

An exploratory study examines how multimodal large language models (MLLMs) generate claims about visualizations, categorizing them as direct, derived, or speculative. It finds that providing accessible chart context shifts models towards more direct claims and improves numeric agreement, highlighting the need for systems that differentiate evidence-based interpretations from model-supplied knowledge.

Multimodal Large Language Models (MLLMs) possess the capability to connect visual patterns in charts with external knowledge, inferring causes and consequences. However, the underlying evidential basis for these interpretations often remains ambiguous. This exploratory study investigates how MLLMs generate claims when presented with visualizations, categorizing these claims into DIRECT (supported by the chart), DERIVED (inferred from the chart), and SPECULATIVE (model-supplied interpretation). The research analyzed 1,224 descriptions generated by three MLLMs (Gemini, GPT, and another unnamed model) across 102 visualizations, under four different input conditions. These conditions varied the MLLMs' access to the image, source-specific accessible chart context, and a "withheld-context" framing. A key finding was that providing accessible chart context significantly influenced Gemini and GPT, leading them to produce more DIRECT claims and improving numeric agreement in some cases. Interestingly, simply adding the image to the full context did not consistently improve numeric agreement, and prompts designed to encourage cautious language in "withheld-context" scenarios did not reliably increase it. The "Real-World Significance" section, as defined in the prompt, remained predominantly SPECULATIVE. These results underscore the critical need for accessible description systems that clearly distinguish between claims directly supported by the provided visual evidence and those that are purely model-supplied interpretations or external knowledge.

Why it matters

For professionals relying on AI to interpret data visualizations, this research is crucial for understanding the trustworthiness and evidential basis of MLLM-generated insights, enabling more informed decision-making and preventing misinterpretation.

How to implement this in your domain

  1. 1Design MLLM-powered visualization interpretation tools to explicitly label claims as "chart-supported," "derived," or "speculative."
  2. 2Prioritize providing accessible chart context to MLLMs to encourage more evidence-based interpretations.
  3. 3Develop user interfaces that clearly distinguish between data-driven facts and AI-generated inferences or external knowledge.
  4. 4Conduct internal audits of MLLM-generated insights from visualizations to assess their evidential basis and numeric agreement.
  5. 5Train MLLMs with datasets that emphasize grounding claims in visual evidence and using cautious language for speculative interpretations.

Who benefits

Data AnalyticsBusiness IntelligenceJournalismEducationMarketing

Key takeaways

  • MLLMs can interpret visualizations but their claims' evidential basis is often unclear.
  • Claims can be direct (chart-supported), derived, or speculative (model-supplied).
  • Accessible chart context improves MLLM's direct claims and numeric agreement.
  • Systems must distinguish evidence-based claims from model-supplied interpretations.

Original post by Ishrat Jahan Eliza, Md Dilshadur Rahman

"arXiv:2607.25021v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) can connect visualization patterns to external causes, consequences, and domain knowledge, but the evidential basis of these interpretations is often unclear. We present an exploratory study…"

View on X

Originally posted by Ishrat Jahan Eliza, Md Dilshadur Rahman on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses