PhysicsBench: Unified Leaderboard for Engineering AI Models

Sang Won Lee, Hyogu Jeong, Namwoo Kang· August 26, 2026 View original

Key takeaways

  • PhysicsBench unifies evaluation for generative and predictive AI models in engineering.
  • It uses standardized procedures and metrics across seven tasks and 66 models.
  • The benchmark focuses on realistic, limited data scales, unlike academic evaluations.
  • Academic standing weakly predicts small-data performance, emphasizing the need for real-world testing.

Who benefits

EngineeringAerospaceAutomotiveManufacturingConstruction

Summary

PhysicsBench is a new unified benchmark and leaderboard for generative and predictive AI models in engineering design and simulation, standardizing evaluation across seven tasks and 66 models. It assesses models on realistic, limited data scales and uses a common metric suite with a PageRank-based ranking system, revealing that academic standing weakly predicts small-data performance.

The evaluation of generative and predictive AI models in engineering design and simulation has traditionally been fragmented, with models often assessed in isolation on academic datasets using inconsistent metrics. To address this, PhysicsBench has been introduced as a unified benchmark and leaderboard. This platform standardizes the evaluation procedure for both generative and predictive models. PhysicsBench covers seven generation and prediction tasks across 1D, 2D, and 3D domains, ranking 66 models on nine datasets. These datasets include industrial-scale CAD/CFD/FEA simulations and public references, configured into 28 distinct setups. A crucial aspect of PhysicsBench is its focus on realistic, limited data scales (from S to XL), contrasting with the unlimited training sets common in academic benchmarks. The benchmark employs a common metric suite that captures geometric fidelity using distributional distances, physical-field and scalar accuracy, and engineering-specific field and shape validity. BenchRank, a PageRank-based system over a head-to-head dominance graph, debiases correlated metrics and provides a comprehensive ranking, with computational cost presented separately. Findings indicate that an architecture's large-scale academic standing is a weak predictor of its small-data ranking, with the top model changing with data scale in six of seven tasks, and no single model leading more than one task. PhysicsBench aims to transform "state-of-the-art" claims into openly published, verifiable foundations for model selection.

Why it matters

This benchmark provides a critical tool for professionals in engineering and AI to objectively compare and select the best generative and predictive models for real-world applications, especially when data is limited. It shifts model selection from self-reported claims to data-driven, standardized evaluations.

How to implement this in your domain

  1. 1Consult PhysicsBench when selecting AI models for engineering design or simulation tasks to ensure objective, data-driven choices.
  2. 2Evaluate your internal AI models against PhysicsBench's standardized procedures and metrics to understand their real-world performance.
  3. 3Prioritize AI models that demonstrate strong performance at realistic, limited data scales, as highlighted by PhysicsBench's findings.
  4. 4Contribute your own models or datasets to PhysicsBench to foster transparency and accelerate industry-wide AI adoption.

Original post by Sang Won Lee, Hyogu Jeong, Namwoo Kang

"arXiv:2608.24056v1 Announce Type: new Abstract: Generative and predictive artificial intelligence models are increasingly used to generate geometry and to predict physical fields and scalar quantities in engineering design and simulation. Yet these models are typically evaluated…"

View on X

Originally posted by Sang Won Lee, Hyogu Jeong, Namwoo Kang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses