New Benchmark Reveals AI Models Struggle with Geometry Diagrams

Hsien Xin Peng, Anthony Kim, Alvin Li, Calvin Supasanya, Shivank Garg, Kevin Zhu· August 20, 2026 View original

Key takeaways

  • AI models proficient in solving geometry problems struggle significantly with constructing accurate diagrams.
  • A new benchmark, 'Solving Is Not Drawing,' measures diagrammatic reasoning in Olympiad geometry.
  • Current foundation models have a low success rate (36.14%) in generating faithful diagrams.
  • Strong mathematical reasoning does not imply the ability to construct accurate geometric diagrams.

Who benefits

AI ResearchEngineeringArchitectureEducationScientific Visualization

Summary

A new open-source benchmark, targeting Olympiad geometry problems, reveals a significant gap between AI models' ability to solve mathematical problems and their capacity to construct accurate diagrams. Current foundation models achieve only a 36.14% compile success rate for diagrams.

While large language models (LLMs) like GPT and Claude have shown impressive proficiency in solving Olympiad-level mathematics, a new benchmark highlights a distinct limitation: their inability to accurately construct geometric diagrams. Solving a geometry problem often relies on a faithful diagram with correct auxiliary constructions and incidences, a skill not measured by existing benchmarks. To address this gap, researchers introduced an open-source benchmark comprising 954 self-contained Olympiad geometry problems, including a hard subset of 297 problems. Each problem is paired with its solution and a human-authored, high-fidelity diagram rendered in Asymptote code. Evaluating current foundation models against this benchmark revealed a pronounced disparity. Despite strong mathematical reasoning capabilities, these models produced diagrams that were markedly less faithful, achieving an average compile success rate of only 36.14%. This indicates that advanced mathematical reasoning does not automatically translate into the ability to construct accurate visual representations.

Why it matters

For professionals developing or deploying AI in fields requiring visual reasoning, such as engineering, architecture, or scientific research, this finding underscores a critical limitation. It suggests that current AI models may struggle with tasks requiring precise diagrammatic understanding and generation, necessitating human oversight or specialized tools.

How to implement this in your domain

  1. 1Assess the reliance of your AI applications on diagrammatic reasoning and visual output generation.
  2. 2Integrate human-in-the-loop validation for AI-generated diagrams or visual representations in critical applications.
  3. 3Explore specialized AI models or techniques focused on geometric reasoning and diagrammatic generation.
  4. 4Develop internal benchmarks to evaluate AI models' ability to produce accurate visual outputs relevant to your domain.
  5. 5Provide AI models with structured data or explicit instructions for diagram construction rather than relying solely on natural language prompts.

Original post by Hsien Xin Peng, Anthony Kim, Alvin Li, Calvin Supasanya, Shivank Garg, Kevin Zhu

"arXiv:2608.18111v1 Announce Type: new Abstract: Foundation models such as GPT and Claude now solve olympiad-level mathematics with remarkable proficiency, so much so that geometry problem solving has become a standard proxy for their mathematical reasoning. Yet solving a geometry…"

View on X

Originally posted by Hsien Xin Peng, Anthony Kim, Alvin Li, Calvin Supasanya, Shivank Garg, Kevin Zhu on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses