Powerbench evaluates language models on power system data tasks
PowerBench: A Benchmark for Agentic Retrieval and Reasoning in Power Systems
Artificial Intelligence
Summary
Understanding how complex energy systems work is hard because data is huge, interconnected, and often secret. The authors created PowerBench, which makes fake but realistic power system data with many devices and documents linked together. They then test advanced language models to see if these models can find and use the right information to answer tough questions about the system. The tests show current models still struggle, scoring around 74% accuracy, revealing areas where they need to improve. This work helps guide the careful use of language tools in energy and industrial settings.
What this means in practice
- •For power system operators: Evaluate and improve AI tools for automated analysis of large-scale power system data under realistic conditions.
- •For industrial ai developers: Develop and test language model agents that autonomously retrieve and reason over heterogeneous industrial datasets.
Authors
Xijing Wang, Yinsheng Yao, Jinru Ding, Yidong Jiang, Ziwen Xu, Yiwen Jiang, Jie Xu, Dawei Cheng
Abstract
Large language model (LLM) agents offer new opportunities for automated analysis in industry. However, rigorous evaluation of such agents-for example, within power system scenarios-remains hindered: real operational data are confidential, and existing public resources fail to fully capture the chained dependencies and heterogeneous evidence. To address this gap, we propose PowerBench, comprising (1) a generation framework that derives interconnected heterogeneous operational data through a common dependency chain, and (2) a synthetic dataset generated by this framework. The dataset covers 761 devices across 100 device types, with 13.35 million hourly telemetry records spanning two years and 24,939 operational documents. Building on this dataset, we construct 300 questions across three task families that evaluate frontier LLMs' ability to complete analysis tasks that require autonomous evidence retrieval and reasoning across interconnected and heterogeneous data under restricted tool calls and time budgets. Results demonstrate that the evaluated frontier LLMs remain challenged on these tasks: the best model reaches only 74.2% joint accuracy. Our trace analysis further reveals that model performance varies across evidence discovery, content retrieval, tool use, reasoning over evidence, and answer submission. These findings provide detailed insights for evaluating LLM agents and guiding their reliable deployment in industry. The framework, dataset, and benchmark tasks are available at https://github.com/open-compass/PowerBench.