LLMs Show Promising, Yet Flawed, Mechanical Engineering Awareness

Johannes Gerstmayr, Sebastian Weyrer, Tobias M\"oltner, Peter Manzl, Michael Pieber· August 18, 2026 View original

Key takeaways

  • LLMs can generate multibody simulation models from text, showing mechanical awareness.
  • Performance is strong for rigid-body systems but weaker for flexible multibody tasks.
  • Proprietary models slightly outperform open-weight models but errors persist.
  • LLM capabilities in mechanical engineering are improving but still require human validation.

Who benefits

ManufacturingAutomotiveAerospaceRoboticsProduct Design

Summary

A new benchmark, MecEng, evaluates LLMs' ability to generate multibody simulation models from text, revealing strong performance on rigid-body tasks but significant challenges with flexible multibody systems. While proprietary models slightly outperform open-weight ones, LLMs still exhibit error-prone mechanical engineering awareness.

While Large Language Models (LLMs) excel in code generation and mathematical reasoning, their proficiency in mechanical engineering, specifically mechanics and spatial geometry, has been less systematically quantified. A new automated benchmark called MecEng addresses this by evaluating LLMs on their capacity to create multibody simulation models from parameterized textual descriptions. The benchmark includes 84 tasks of varying difficulty, from simple rigid-body systems to complex flexible multibody systems requiring 3D geometry generation and finite-element meshing. The evaluation pipeline uses LLMs to generate simulation-ready geometry and build models for the Exudyn code, which are then verified against expert ground truth. The study assessed 32 open-weight and two proprietary LLMs. Results indicate that the best open-weight model achieved an 86.0% success rate on rigid-body tasks, close to the strongest proprietary model's 91.4%. However, flexible multibody tasks proved considerably more difficult for all models. The findings suggest that while LLMs are rapidly improving in mechanical engineering awareness, they remain prone to errors, particularly in more complex scenarios involving deformable bodies.

Why it matters

Engineers and product developers in mechanical design, robotics, and simulation can leverage LLMs for initial model generation, but must implement rigorous validation steps. This research highlights both the potential for automation and the current limitations requiring human oversight.

How to implement this in your domain

  1. 1Experiment with LLMs for generating initial drafts of multibody simulation models from textual specifications.
  2. 2Integrate LLM-generated geometry and simulation code into existing CAD/CAE workflows, using human experts for verification.
  3. 3Develop custom fine-tuning datasets for LLMs using internal mechanical engineering documentation and simulation results.
  4. 4Utilize LLMs for rapid prototyping and concept exploration in mechanical design, acknowledging their current error rates.
  5. 5Invest in tools that can automatically verify the correctness of LLM-generated mechanical models against established engineering principles.

Original post by Johannes Gerstmayr, Sebastian Weyrer, Tobias M\"oltner, Peter Manzl, Michael Pieber

"arXiv:2608.14615v1 Announce Type: new Abstract: Large Language Models (LLMs) perform well on established code-generation and mathematical-reasoning benchmarks, but their capabilities in mechanics and spatial geometry, here denoted as mechanical engineering awareness, has not been…"

View on X

Originally posted by Johannes Gerstmayr, Sebastian Weyrer, Tobias M\"oltner, Peter Manzl, Michael Pieber on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses