DashArena Benchmarks LLMs for Interactive Dashboard Generation

Xiaotong Wang, Dazhen Deng· August 12, 2026 View original

Key takeaways

  • DashArena is the first benchmark for evaluating LLMs on interactive analytic dashboard generation.
  • It uniquely assesses both dashboard output and user interaction trajectories.
  • A VLM judge and browser executor provide robust, reproducible evaluation evidence.
  • Current frontier LLMs still exhibit significant failures in generating realistic, interactive dashboards.

Who benefits

Data AnalyticsSoftware DevelopmentBusiness IntelligenceConsulting

Summary

DashArena is a new benchmark for evaluating large language models' ability to generate interactive analytic dashboards, focusing on both dashboard appearance and the quality of user interaction trajectories. It uses a VLM judge and a browser executor to assess analytical support and interaction, revealing current LLM limitations.

Evaluating the effectiveness of large language models (LLMs) in generating interactive analytic dashboards has been challenging, as existing methods often fail to capture the full scope of analytical support and interaction quality. A new benchmark, DashArena, addresses this by requiring LLMs to generate both a dashboard and a replayable interaction trajectory. DashArena employs a browser executor to replay these trajectories, providing reproducible visual and execution evidence of the LLM's intended analytical workflow. A Vision-Language Model (VLM) judge then compares candidates based on this evidence, with results aggregated using Bradley-Terry. The research also introduces DashJudge-8B, an open-weight distilled judge that effectively reproduces human judgments. Experiments with current frontier models highlight persistent issues in rendering, analytical accuracy, and interaction capabilities, underscoring that realistic, interactive dashboard generation remains a significant challenge for LLMs.

Why it matters

Data professionals and product developers can use this benchmark to understand the current capabilities and limitations of LLMs in creating functional and interactive data visualization tools, guiding future development and adoption strategies.

How to implement this in your domain

  1. 1Explore DashArena and DashJudge-8B to understand the evaluation criteria for interactive dashboards.
  2. 2Benchmark internal LLM-powered dashboard generation tools against DashArena's methodology.
  3. 3Identify specific areas (rendering, analytical accuracy, interaction) where current LLMs fall short.
  4. 4Integrate interaction-aware evaluation metrics into your own dashboard development workflows.
  5. 5Contribute to or leverage the open-weight DashJudge-8B for more robust internal evaluations.

Original post by Xiaotong Wang, Dazhen Deng

"arXiv:2608.10567v1 Announce Type: new Abstract: Analytic dashboards combine coordinated views and interactions for data exploration and decision-making. Recent models can generate them from data and natural-language goals, but evaluating their usefulness remains difficult. Dashbo…"

View on X

Originally posted by Xiaotong Wang, Dazhen Deng on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses