AI Grid Ready Data

20260304 tgr photoshoots elissaiossarmas 0014 elissaios sarmas

Chair: Elissaios Sarmas
Energy Policy Unit (EPU) at the National Technical University of Athens (NTUA)
This session focuses on creating AI-ready data ecosystems for the electricity sector. Topics include semantic interoperability, energy data spaces, multimodal and synthetic datasets, data quality assessment, metadata and provenance, privacy-preserving data sharing, and governance frameworks enabling Foundation Models and Digital Twins. Speakers will present practical experiences from utilities and research projects on transforming heterogeneous operational data into reusable assets for rustworthy Grid AI.

Solving data for European power system foundation models

Speaker: Arsam Aryandoust – E.ON Energy Research Center

Power system foundation models require the exchange of security-sensitive data for training and inference. In this presentation, we discuss the open problems that have been identified and been started to work on by the AI.grids initiative. This include benchmarking on low technology readiness level until creating data exchange infrastructures on high technology readiness level

Arsam

Generating diverse synthetic data for the grid with gridfm-datakit

IMG 4584 Copy Alban Puech 1024x1024

Speaker: Alban Puech – IBM Research & ETH Zürich

GridFMs require large and diverse datasets, yet existing datasets often lack realistic scenario diversity and flexibility. datakit is an open-source Python library for generating synthetic Power Flow (PF) and Optimal Power Flow (OPF) datasets at scale. It addresses three key limitations: insufficiently diverse load and topology perturbations, PF datasets restricted to OPF-feasible operating points, and fixed generator cost functions in OPF datasets. datakit combines global load scaling from real-world profiles with localized noise, supports arbitrary N-k topology perturbations, generates PF scenarios beyond operating limits, and enables varying generator costs for OPF. It also scales efficiently to large networks, with demonstrated generation on grids of up to 10,000 buses. In this presentation, we introduce datakit, showcase its capabilities and scalability, and discuss how it enables diverse and challenging datasets for training and benchmarking machine-learning-based grid solvers.

GridArena: An End-to-End Platform for Simulating, Twinning, and Benchmarking Grid AI

David Lima – INESC-TEC

Developing and evaluating AI for the electric grid is fragmented across disconnected tools: a simulator here, a private dataset there, an ad hoc evaluation script somewhere else. We present GridArena, an open-source platform that brings these pieces together in one environment. GridArena lets users register and manage low- and medium-voltage grid topologies, ingest real measurement data, and run topology-aware power-flow simulation, all through a single API and UI. A live digital-twin mode continuously compares a registered grid’s simulated behaviour against real external feeds, and a built-in reinforcement-learning environment supports training grid-control agents end-to-end. When real data is scarce, the platform’s data layer can be augmented with synthetic measurements (diffusion-model generator trained on real data and a large-scale synthetic MV scenario generator) among the tools available to fill gaps in coverage. GridArena’s core contribution for the grid-AI community is its standardized, anonymised benchmark suite spanning four representative tasks (phase identification, topology discovery, state estimation, and voltage control) where a persistent server-side anonymisation key lets participants train and evaluate models on masked node, phase, and grid identifiers without ever touching identifiable utility data. We present GridArena’s architecture and demonstrate the full workflow, from grid registration through simulation, digital-twin comparison, and benchmark scoring.

david lima
Jochen Cremer

Synthetic Grids, Real Physics: Learning Foundation Models for Power Systems

Jochen Cremer – TU Delft

Large-scale AI for power systems is limited by the scarcity and narrow coverage of real-world grid data. This talk presents complementary approaches to address this challenge through synthetic grid generation and equation-based pretraining. We first introduce a synthesizer that generates realistic grid topologies and associated electrical parameters for modeling, simulation, and machine-learning applications. We then show how grid foundation models can be pretrained directly from the structure of the AC power-flow equations, reducing reliance on large solver-generated datasets and improving generalization across unseen topologies and parameters. Together, these approaches point toward scalable ways of combining synthetic data, grid structure, and physical laws for power-system A