Multimodal AI Struggles with Abstract Perceptual Reasoning

Garima Arya Yadav, Nilay Yilmaz, Yezhou Yang· August 18, 2026 View original

Key takeaways

  • Multimodal AI models struggle significantly with abstract perceptual reasoning.
  • "The Unwritten Benchmark" reveals a vast performance gap between humans and AI.
  • Leading models often perform worse with multiple modalities in this task.
  • Current AI lacks robust cross-modal causal reasoning and micro-kinematic understanding.

Who benefits

AI/ML DevelopmentRoboticsHuman-Computer InteractionEducation Technology

Summary

Researchers introduce "The Unwritten Benchmark," a challenge where multimodal models must infer words being written from only audio of pen scratches and video of hand movements, without visible ink. Leading models like GPT-4o and Gemini 2.5-Pro perform significantly worse than humans, often degrading performance when both modalities are provided.

A new challenge, "The Unwritten Benchmark," has been introduced to test the abstract perceptual reasoning capabilities of multimodal machine learning models. Unlike tasks involving static visual or auditory content, this benchmark requires models to infer unseen information from dynamic, generative processes. The core task is acousto-kinematic word inference: models must decipher words being written in three different styles, using only the audio of pen scratches and video of hand movements, with no visible ink trace. The evaluation revealed a substantial performance gap between humans and machines. Human participants achieved over 80% ordered letter accuracy, while leading multimodal models, including GPT-4o and Gemini 2.5-Pro, struggled significantly, failing to surpass 10%. This stark difference highlights a critical limitation in current AI's ability to perform abstract perceptual and cognitive tasks. Furthermore, the study identified a paradoxical "fusion effect" where providing both audio and video modalities often led to degraded performance rather than improvement. This suggests a fundamental breakdown in the models' capacity to synthesize complementary perceptual cues for this complex cognitive task, indicating limitations in cross-modal causal reasoning and understanding micro-kinematics.

Why it matters

This research exposes a significant gap in current multimodal AI capabilities, indicating that models still lack human-like abstract reasoning and cross-modal integration, which is crucial for developing truly intelligent and intuitive AI systems.

How to implement this in your domain

  1. 1Re-evaluate the scope of "multimodal" capabilities in current AI systems, recognizing limitations in abstract reasoning.
  2. 2Invest in research and development focused on improving cross-modal causal reasoning and micro-kinematic understanding in AI.
  3. 3Design new benchmarks that specifically target abstract perceptual reasoning beyond static content recognition.
  4. 4Consider human-in-the-loop approaches for tasks requiring nuanced abstract interpretation where AI currently fails.

Original post by Garima Arya Yadav, Nilay Yilmaz, Yezhou Yang

"arXiv:2608.14558v1 Announce Type: new Abstract: Current multimodal models have demonstrated remarkable proficiency in recognizing static visual and auditory content. However, their capacity for abstract perceptual reasoning, inferring unseen information from dynamic, generative p…"

View on X

Originally posted by Garima Arya Yadav, Nilay Yilmaz, Yezhou Yang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses