RestoreBench Evaluates AI Agents for Power Grid Restoration

Riccardo Mansutti, Andrea Pomarico, Robert Jakob, Qian Zhang, Alberto Berizzi, Kevin O'Sullivan· September 2, 2026 View original

Key takeaways

  • LLM agents show promise for automating complex engineering workflows, like power grid restoration.
  • RestoreBench is a new benchmark for evaluating AI agents in resolving power flow convergence issues.
  • It tests various agent architectures across multiple power grid scenarios.
  • The benchmark provides a reproducible foundation for developing AI systems in power operations.

Who benefits

EnergyUtilitiesInfrastructureAI Engineering

Summary

RestoreBench is a new benchmark designed to evaluate the ability of LLM agents to diagnose and resolve non-convergent power flow cases in electrical grids. It tests various LLM architectures (chatbot, single-agent, multi-agent) across two power grids and 46 cases per grid, providing a reproducible foundation for developing AI systems in power system operations.

Large Language Model (LLM) agents are increasingly being used to automate complex, multi-step engineering workflows, leveraging their ability to use tools, interpret intermediate results, and plan iteratively. A promising, yet underexplored, application area for these agents is the diagnosis and resolution of non-convergent power flow scenarios within electrical grids. This task demands significant engineering judgment, experimental approaches, and precise decision-making within predefined action spaces. To facilitate research in this domain, a new benchmark called RestoreBench has been introduced. This benchmark specifically evaluates the capabilities of LLM agents in restoring power flow convergence. It assesses different LLM architectures, including chatbot, single-agent, and multi-agent systems. The evaluation covers two distinct power grids, each presenting 46 unique cases that require one or more corrective actions to achieve convergence. RestoreBench clearly defines the simulation environment, the available observation and action spaces, and the metrics for evaluation, thereby establishing a reproducible framework for the development of agentic AI systems tailored for power system planning and operational tasks.

Why it matters

This benchmark is critical for advancing AI applications in vital infrastructure sectors like energy, enabling the development and rigorous testing of AI agents that can enhance the stability and reliability of power grids. Professionals in energy and AI engineering can use this to build and validate intelligent automation solutions.

How to implement this in your domain

  1. 1Access the RestoreBench code and documentation from the provided GitHub repository.
  2. 2Set up the simulation environment for the specified power grids.
  3. 3Evaluate various LLM agent architectures (chatbot, single-agent, multi-agent) against the benchmark's cases.
  4. 4Develop and fine-tune custom AI agents to improve performance on power flow convergence tasks.
  5. 5Collaborate with power system engineers to integrate successful agentic solutions into operational planning.

Original post by Riccardo Mansutti, Andrea Pomarico, Robert Jakob, Qian Zhang, Alberto Berizzi, Kevin O'Sullivan

"arXiv:2609.00384v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly automate multi-step engineering workflows through tool use, interpretation of intermediate results, and iterative planning. Diagnosing and resolving non-convergent power flow cases is a…"

View on X

Originally posted by Riccardo Mansutti, Andrea Pomarico, Robert Jakob, Qian Zhang, Alberto Berizzi, Kevin O'Sullivan on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses