QC-T2I-Bench: Scalable Text-to-Image Evaluation and Routing.

Shaoan Zhao, Fang Zhao, Xueqiang Guo, Xinpei Su, Huanlin Gao, Qiang Hui, Ting Lu, Fuyuan Shi, Chao Tan, Bikun Yang, Kai Wang, Shiguo Lian· August 26, 2026 View original

Key takeaways

  • QC-T2I-Bench offers a scalable, question-centric framework for T2I evaluation.
  • It enables reliable model ranking and fine-grained performance diagnosis.
  • Joint completion rates in T2I models decrease significantly with prompt complexity.
  • The framework supports cost-aware routing for efficient T2I inference.

Who benefits

AI DevelopmentMarketingMedia & EntertainmentE-commerceDesign

Summary

This paper introduces QC-T2I-Bench, a question-centric framework for evaluating text-to-image (T2I) models that converts open prompts into attributed atomic questions and organizes dependencies with Davidsonian Scene Graphs. It enables reliable ranking, fine-grained diagnosis, and cost-aware routing, revealing how joint completion rates drop with increasing prompt complexity.

Evaluating modern text-to-image (T2I) models is challenging because models with similar overall scores often have distinct strengths and weaknesses, making practical selection difficult. Existing fine-grained benchmarks often revert to prompt-level scores or fixed categories, obscuring specific failure points and ignoring prompt complexity. This research presents "QC-T2I-Bench," a question-centric framework designed for scalable T2I evaluation. It transforms open prompts into attributed atomic questions and maps their dependencies using Davidsonian Scene Graphs (DSGs). The framework uses hierarchy-constrained question aggregation to prevent simple and complex prompts from being weighted equally and to exclude downstream questions if prerequisites fail. The DSG structure also allows for measuring joint success within prompts and comparing repeated entities, distinguishing basic realization failures from those under additional requirements. Experiments on open-source T2I models with English and Chinese prompts demonstrated that joint completion rates significantly decrease as prompt complexity increases. Furthermore, the framework's records can be reused for training-free, cost-aware routing, achieving comparable performance to a high-performing model with substantial GPU-s/MP savings.

Why it matters

Professionals working with T2I models need precise evaluation tools to select the best model for specific tasks, diagnose performance issues, and optimize resource allocation. QC-T2I-Bench offers a robust, fine-grained, and cost-effective solution.

How to implement this in your domain

  1. 1Adopt question-centric evaluation frameworks like QC-T2I-Bench for T2I model assessment.
  2. 2Utilize Davidsonian Scene Graphs to break down complex prompts into atomic questions.
  3. 3Implement fine-grained diagnostic methods to identify specific T2I model weaknesses.
  4. 4Explore cost-aware routing strategies for T2I inference to optimize resource usage.
  5. 5Apply this evaluation approach to compare and select T2I models for creative or marketing campaigns.

Original post by Shaoan Zhao, Fang Zhao, Xueqiang Guo, Xinpei Su, Huanlin Gao, Qiang Hui, Ting Lu, Fuyuan Shi, Chao Tan, Bikun Yang, Kai Wang, Shiguo Lian

"arXiv:2608.24112v1 Announce Type: new Abstract: Modern text-to-image (T2I) models often have similar total scores but different strengths, making practical selection difficult. Fine-grained benchmarks decompose prompts into questions, yet often return them to prompt scores and fi…"

View on X

Originally posted by Shaoan Zhao, Fang Zhao, Xueqiang Guo, Xinpei Su, Huanlin Gao, Qiang Hui, Ting Lu, Fuyuan Shi, Chao Tan, Bikun Yang, Kai Wang, Shiguo Lian on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses