Standardized ML Evaluation for Power System Protection

Julian Oelhaf, Georg Kordowich, Paula Andrea P\'erez-Toro, Christian Bergler, Johann J\"ager, Andreas Maier, Siming Bayer· August 21, 2026 View original

Key takeaways

  • Current ML evaluations in power system protection lack standardization, hindering comparability.
  • A new framework defines seven critical dimensions for consistent evaluation design.
  • Evaluation assumptions significantly impact reported performance and robustness.
  • Standardization is vital for comparable, auditable, and certifiable ML protection functions.

Who benefits

EnergyUtilitiesCritical InfrastructureAI DevelopmentRegulatory Bodies

Summary

This paper proposes a standardized framework for evaluating machine learning applications in power system protection, addressing inconsistencies in current research. It defines seven critical study dimensions and instantiates the framework with a case study on fault classification and localization using a public benchmark.

Researchers have introduced a standardized framework designed to bring consistency and comparability to the evaluation of machine learning models in power system protection. Currently, studies often report near-perfect scores, but these results are difficult to compare due to wide variations in evaluation settings, including protection tasks, physical scope, measurements, timing, and validation protocols. The proposed framework aims to make evaluation design an explicit part of the scientific contribution. The framework outlines seven essential study dimensions: protection objective, physical scope, observability, timing and decision windows, targets and sample validity, validation protocol, and evaluation outputs. To demonstrate its utility, the paper applies the framework to a case study using the PROTECT-90 benchmark for fault classification and localization. The study highlights how different evaluation assumptions, such as observability and decision horizons, significantly impact performance, emphasizing the need for explicit, reproducible evidence to enable more auditable and certifiable ML protection functions.

Why it matters

Standardizing evaluation for ML in power system protection is crucial for building trust, ensuring reliability, and accelerating the adoption of these technologies in critical infrastructure.

How to implement this in your domain

  1. 1Adopt the proposed seven-dimension framework for designing and reporting ML evaluations in power systems.
  2. 2Ensure all evaluation assumptions are explicitly stated and reproducible in research and development.
  3. 3Benchmark ML protection functions against conventional methods under comparable information sets.
  4. 4Focus on robustness testing, including measurement degradation, beyond clean predictive performance.

Original post by Julian Oelhaf, Georg Kordowich, Paula Andrea P\'erez-Toro, Christian Bergler, Johann J\"ager, Andreas Maier, Siming Bayer

"arXiv:2608.20181v1 Announce Type: new Abstract: Studies of machine-learning-based power-system protection increasingly report near-perfect scores, yet the meaning of those scores depends strongly on the evaluation setting. Protection task, physical scope, measurements, timing, ta…"

View on X

Originally posted by Julian Oelhaf, Georg Kordowich, Paula Andrea P\'erez-Toro, Christian Bergler, Johann J\"ager, Andreas Maier, Siming Bayer on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses