ResearchAI Research

C4 Evaluates MLLM Creative Understanding of Cross-Concepts.

Ming Wang, Yuqing Zhang, Tingna Xie, Xiangju Li, Xiaocui Yang, Daling Wang, Shi Feng, Yifei Zhang· August 10, 2026 View original

Key takeaways

  • MLLMs struggle with "receptive creativity" and cross-concept understanding.
  • C4 framework evaluates MLLMs using Chengyu-based conceptual relations.
  • Current MLLMs show a significant gap in decoding creatively encoded meaning.
  • Explicit hints offer limited improvement, suggesting a deeper conceptual challenge.

Who benefits

Content CreationMarketingEdTechDesignAI/ML Research

Summary

This paper introduces C4, a cognition-inspired evaluation framework and dataset for assessing Multimodal Large Language Models' (MLLMs) "receptive creativity" through cross-concept understanding. C4 uses Chengyu (Chinese idiom)-based figures to test how well MLLMs decode non-obvious but meaningful conceptual relations, revealing a substantial gap in their current creative capabilities.

Evaluating the creative capabilities of Multimodal Large Language Models (MLLMs) is challenging due to the lack of clear targets and reward signals. This research proposes C4, a new framework designed to assess "receptive creativity," specifically focusing on cross-concept understanding – the ability to grasp intended meaning from non-obvious conceptual relationships. C4 operationalizes this by encoding cross-concept relations, often inspired by Chinese idioms (Chengyu), into items and then evaluating the MLLM's ability to decode them. The C4 Evaluation Set (C4-Eval) includes both synthetic and human-created items, each with detailed cross-concept relations and reasoning paths. Tested across ten MLLMs in various task settings, the strongest closed models achieved only around 50% accuracy, while open-source models performed significantly worse. The study found that while candidate constraints improved accuracy, hints and explanation requests offered only modest gains. These results highlight a considerable deficiency in how current MLLMs process and understand creatively encoded meanings through complex conceptual connections.

Why it matters

Professionals developing or applying MLLMs for creative tasks (e.g., content generation, design, education) need to understand their current limitations in genuine creative understanding, guiding future research and application development.

How to implement this in your domain

  1. 1Review the C4 framework to understand current MLLM limitations in creative reasoning.
  2. 2Incorporate cross-concept understanding tests into your MLLM evaluation pipelines for creative applications.
  3. 3Design MLLM prompts that explicitly guide models through multi-step conceptual reasoning for creative tasks.
  4. 4Contribute to or explore datasets that focus on non-obvious conceptual relations to improve MLLM training.

Original post by Ming Wang, Yuqing Zhang, Tingna Xie, Xiangju Li, Xiaocui Yang, Daling Wang, Shi Feng, Yifei Zhang

"arXiv:2608.06501v1 Announce Type: new Abstract: Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain difficult to evaluate because explicit targets and reward signals are scarce compared with accuracy-oriented tasks. C…"

View on X

Originally posted by Ming Wang, Yuqing Zhang, Tingna Xie, Xiangju Li, Xiaocui Yang, Daling Wang, Shi Feng, Yifei Zhang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses