The increasing integration of Distributed Energy Resources (DERs) within modern power systems requires Virtual Power Plants (VPPs) to adopt intelligent, adaptive, and scalable scheduling mechanisms to ensure efficient, reliable, and economically optimal operation. This research proposes a hybrid reinforcement learning–based framework that combines decentralized Deep Q-Networks (DQN) in the first layer with a Graph-Constrained Multi-Agent Proximal Policy Optimization method enhanced by a Safety Layer (GC-MAPPO-SL) in the second layer. In the first layer, decentralized DQN agents independently regulate local DER behavior such as generation, storage, and consumption using only local observations, including demand forecasts, renewable generation estimates, and state-of-charge values. The decentralized organization provides the ability of the system to be scalable to heterogeneous DERs and optimally achieve adaptability to varying operating conditions. The second layer, GC-MAPPO-SL, optimizes the global operational decisions with graph-based embeddings in the second layer, but a lightweight centralized critic is used, which does not require communication delays or privacy invasion. One layer provides operational restrictions dynamically, such as power balance, line-flow limit, state-of-charge limit, and ramping limit, which are necessary to guarantee feasibility and operational safety in any circumstance. The hybrid framework was tested on a 24-hour simulation horizon with 15-minute intervals with real historical data on the load and renewable generation. Findings indicate good performance, an overall operating cost of $812, a renewable energy use of 77, zero constraint violations, convergence in 120 episodes, as well as scalability of three heterogeneous DERs. Besides, the technique had a load balance variation of 7.25 kW and a quick reaction of 0.3 s/step. These results underscore the efficacy, dependability, and applicability of the framework to real-life VPP scheduling.
Gao H, Jin T, Feng C, Li C, Chen Q, Kang C. Review of virtual power plant operations: Resource coordination and multidimensional interaction. Applied Energy. 2024;357:122284.
2.
Ceusters G, Camargo LR, Franke R, Nowé A, Messagie M. Safe reinforcement learning for multi-energy management systems with known constraint functions. Energy and AI. 2023;12:100227.
3.
He S, Cui W, Li G, Xu H, Chen X, Tai Y. Intelligent Scheduling of Virtual Power Plants Based on Deep Reinforcement Learning. Computers, Materials & Continua. 2025;84(1):861–86.
4.
Gao H, Jiang S, Li Z, Wang R, Liu Y, Liu J. A Two-Stage Multi-Agent Deep Reinforcement Learning Method for Urban Distribution Network Reconfiguration Considering Switch Contribution. IEEE Transactions on Power Systems. 2024;39(6):7064–76.
5.
Wang L, Pan T, Wang Y, Jin X, Yu H, Cao W. A virtual power plant scheduling strategy considering cooperative participation of multiple industrial users in demand response. AIMS Energy. 2025;(4).
6.
Klaiber J, Van Dinther C. Deep Learning for Variable Renewable Energy: A Systematic Review. ACM Computing Surveys. 2023;56(1):1–37.
7.
Goeckner A, Sui Y, Martinet N, Li X, Zhu Q. Graph Neural Network-based Multi-agent Reinforcement Learning for Resilient Distributed Coordination of Multi-Robot Systems. 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE; 2024. p. 5732–9.
8.
Mishra S, Bordin C, Wu Q, Manninen H. Resilient expansion planning of virtual power plant with an integrated energy system considering reliability criteria of lines and towers. International Journal of Energy Research. 2022;46(10):13726–51.
9.
Ye Y, Papadaskalopoulos D, Yuan Q, Tang Y, Strbac G. Multi-Agent Deep Reinforcement Learning for Coordinated Energy Trading and Flexibility Services Provision in Local Electricity Markets. IEEE Transactions on Smart Grid. 2023;14(2):1541–54.
10.
Comment on egusphere-2026-1109. Copernicus GmbH; 2026.
11.
Jendoubi I, Bouffard F. Data-driven sustainable distributed energy resources’ control based on multi-agent deep reinforcement learning. Sustainable Energy, Grids and Networks. 2022;32:100919.
12.
Razmi D, Babayomi O, Zhang Z. Reinforcement learning-driven dynamic Model Predictive Control for adaptive real-time multi-agent management of microgrids. International Journal of Electrical Power & Energy Systems. 2025;170:110823.
13.
Akhoundzadeh P, Mirjalily G, Sadeghi MT. Optimal D2D Resource Allocation in Heterogeneous Cellular Networks by Decentralized Multi-Agent Deep Q-Learning. 2024 32nd International Conference on Electrical Engineering (ICEE). IEEE; 2024. p. 1–5.
14.
Wang J, Guo C, Yu C, Liang Y. Virtual power plant containing electric vehicles scheduling strategies based on deep reinforcement learning. Electric Power Systems Research. 2022;205:107714.
15.
Yang X, Liu H, Wu W. Attention-Enhanced Multi-Agent Reinforcement Learning Against Observation Perturbations for Distributed Volt-VAR Control. IEEE Transactions on Smart Grid. 2024;15(6):5761–72.
16.
Xiao D. A Review on Risk-Averse Bidding Strategies for Virtual Power Plants with Uncertainties: Resources, Technologies, and Future Pathways. Technologies. 2025;13(11):488.
17.
Fan Z, Zhang W, Liu W. Multi-Agent Deep Reinforcement Learning-Based Distributed Optimal Generation Control of DC Microgrids. IEEE Transactions on Smart Grid. 2023;14(5):3337–51.
18.
Seyfi M, Mehdinejad M, Mohammadi-Ivatloo B, Shayanfar H. Deep learning-based scheduling of virtual energy hubs with plug-in hybrid compressed natural gas-electric vehicles. Applied Energy. 2022;321:119318.
19.
Hu D, Ye Z, Gao Y, Ye Z, Peng Y, Yu N. Multi-Agent Deep Reinforcement Learning for Voltage Control With Coordinated Active and Reactive Power Optimization. IEEE Transactions on Smart Grid. 2022;13(6):4873–86.
20.
Ghasemi Olanlari F, Amraee T, Moradi‐Sepahvand M, Ahmadian A. Coordinated multi‐objective scheduling of a multi‐energy virtual power plant considering storages and demand response. IET Generation, Transmission & Distribution. 2022;16(17):3539–62.
21.
Xu B, Luan W, Yang J, Zhao B, Long C, Ai Q, et al. Integrated three-stage decentralized scheduling for virtual power plants: A model-assisted multi-agent reinforcement learning method. Applied Energy. 2024;376:123985.
22.
Sun Z, Lu T. Collaborative operation optimization of distribution system and virtual power plants using multi‐agent deep reinforcement learning with parameter‐sharing mechanism. IET Generation, Transmission & Distribution. 2023;18(1):39–49.
The statements, opinions and data contained in the journal are solely those of the individual authors and contributors and not of the publisher and the editor(s). We stay neutral with regard to jurisdictional claims in published maps and institutional affiliations.