New Benchmark Evaluates MLLMs for Chart Annotation

Zhenghan Chen, Zekai Shao, Lidan Tan, Xin Lin, Xingchen Zeng, Yi Shan, Ziyue Lin, Xiaoliang Fu, Xinyuan Liu, Yuetong Guo, Fen Wang, Bongshin Lee, Siming Chen· August 5, 2026 View original

Key takeaways

  • ChartAnno is a new benchmark for evaluating MLLMs in chart annotation.
  • Proprietary MLLMs currently lead, but open-source models are catching up.
  • Specific instructions are crucial for high-quality chart annotations.
  • Inferring abstract intent from charts remains a significant challenge for MLLMs.

Who benefits

Business IntelligenceData AnalyticsSoftware DevelopmentEdTechMarketing

Summary

Researchers introduce ChartAnno, a benchmark with 1,200 real-world charts and instructions, to evaluate Multimodal Large Language Models (MLLMs) on their ability to generate chart annotations. The study reveals that proprietary models generally outperform open-source ones, and specific instructions improve quality, while inferring abstract intent remains challenging.

While Multimodal Large Language Models (MLLMs) have shown impressive progress in understanding, generating, and editing charts, their capability to annotate existing charts has remained largely unexplored. Chart annotation is a complex communicative task that requires models to infer the intended message, interpret chart semantics, and appropriately place textual or graphical elements. To address this gap, a new benchmark called ChartAnno has been developed. ChartAnno comprises 1,200 real-world charts, each paired with code and annotation instructions across three levels of specificity. The benchmark was used to evaluate ten representative MLLMs under two main input settings: using chart code alone, and using both chart code and the chart image. An ablation study with image-only input was also conducted. The evaluation results indicate that proprietary MLLMs generally achieve stronger overall performance, though large-scale open-source models are narrowing the gap. The study also found that providing more specific instructions significantly improves annotation quality, while the task of inferring abstract intent from charts remains the most difficult challenge for current MLLMs. Interestingly, providing chart images offered only limited overall gains, with improvements primarily observed in design-related metrics. These findings underscore that chart annotation generation is a demanding task requiring deep semantic grounding and effective design capabilities.

Why it matters

Professionals developing data visualization tools, business intelligence platforms, or AI assistants for data analysis can use this benchmark and its findings to improve MLLM capabilities in generating clear, contextually relevant chart annotations.

How to implement this in your domain

  1. 1Utilize the ChartAnno benchmark to evaluate and fine-tune MLLMs for chart annotation tasks.
  2. 2Prioritize providing specific instructions to MLLMs when requesting chart annotations for better quality.
  3. 3Focus MLLM development efforts on improving abstract intent inference for chart understanding.
  4. 4Consider the limited benefit of chart images alone for semantic understanding, emphasizing code or structured data input.

Original post by Zhenghan Chen, Zekai Shao, Lidan Tan, Xin Lin, Xingchen Zeng, Yi Shan, Ziyue Lin, Xiaoliang Fu, Xinyuan Liu, Yuetong Guo, Fen Wang, Bongshin Lee, Siming Chen

"arXiv:2608.03464v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have made significant progress in chart understanding, generation, and editing, but their ability to annotate existing charts remains underexplored. Annotating charts is a common yet challeng…"

View on X

Originally posted by Zhenghan Chen, Zekai Shao, Lidan Tan, Xin Lin, Xingchen Zeng, Yi Shan, Ziyue Lin, Xiaoliang Fu, Xinyuan Liu, Yuetong Guo, Fen Wang, Bongshin Lee, Siming Chen on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses