×
Home Current Archive Editorial board
Instructions for papers
For Authors Aim & Scope Contact
Original scientific article

SCALABLE RESOURCE SCHEDULING FOR VIRTUAL POWER PLANTS USING MULTI-AGENT DEEP Q-NETWORKS FOR OPTIMIZED DISTRIBUTED ENERGY RESOURCE MANAGEMENT

By
Srinivasa Rao Sureddy Orcid logo ,
Srinivasa Rao Sureddy
Contact Srinivasa Rao Sureddy

Research Scholar, Department Of EECE, GITAM Deemed to be University, Visakhapatnam, Andhra Pradesh, India

Nageswara Rao Pulivarthi Orcid logo
Nageswara Rao Pulivarthi

Assistant Professor, Department Of EECE, GITAM Deemed to be University, Visakhapatnam, Andhra Pradesh, India

Abstract

The increasing integration of Distributed Energy Resources (DERs) within modern power systems requires Virtual Power Plants (VPPs) to adopt intelligent, adaptive, and scalable scheduling mechanisms to ensure efficient, reliable, and economically optimal operation. This research proposes a hybrid reinforcement learning–based framework that combines decentralized Deep Q-Networks (DQN) in the first layer with a Graph-Constrained Multi-Agent Proximal Policy Optimization method enhanced by a Safety Layer (GC-MAPPO-SL) in the second layer. In the first layer, decentralized DQN agents independently regulate local DER behavior such as generation, storage, and consumption using only local observations, including demand forecasts, renewable generation estimates, and state-of-charge values. The decentralized organization provides the ability of the system to be scalable to heterogeneous DERs and optimally achieve adaptability to varying operating conditions. The second layer, GC-MAPPO-SL, optimizes the global operational decisions with graph-based embeddings in the second layer, but a lightweight centralized critic is used, which does not require communication delays or privacy invasion. One layer provides operational restrictions dynamically, such as power balance, line-flow limit, state-of-charge limit, and ramping limit, which are necessary to guarantee feasibility and operational safety in any circumstance. The hybrid framework was tested on a 24-hour simulation horizon with 15-minute intervals with real historical data on the load and renewable generation. Findings indicate good performance, an overall operating cost of $812, a renewable energy use of 77, zero constraint violations, convergence in 120 episodes, as well as scalability of three heterogeneous DERs. Besides, the technique had a load balance variation of 7.25 kW and a quick reaction of 0.3 s/step. These results underscore the efficacy, dependability, and applicability of the framework to real-life VPP scheduling.

References

1.
Gao H, Jin T, Feng C, Li C, Chen Q, Kang C. Review of virtual power plant operations: Resource coordination and multidimensional interaction. Applied Energy. 2024;357:122284.
2.
Ceusters G, Camargo LR, Franke R, Nowé A, Messagie M. Safe reinforcement learning for multi-energy management systems with known constraint functions. Energy and AI. 2023;12:100227.
3.
He S, Cui W, Li G, Xu H, Chen X, Tai Y. Intelligent Scheduling of Virtual Power Plants Based on Deep Reinforcement Learning. Computers, Materials & Continua. 2025;84(1):861–86.
4.
Gao H, Jiang S, Li Z, Wang R, Liu Y, Liu J. A Two-Stage Multi-Agent Deep Reinforcement Learning Method for Urban Distribution Network Reconfiguration Considering Switch Contribution. IEEE Transactions on Power Systems. 2024;39(6):7064–76.
5.
Wang L, Pan T, Wang Y, Jin X, Yu H, Cao W. A virtual power plant scheduling strategy considering cooperative participation of multiple industrial users in demand response. AIMS Energy. 2025;(4).

Citation

This is an open access article distributed under the  Creative Commons Attribution Non-Commercial License (CC BY-NC) License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. 

Article metrics

Google scholar: See link

Issue image
Issue 36, 2026
See full issue

Citations

Crossref Logo

0

The statements, opinions and data contained in the journal are solely those of the individual authors and contributors and not of the publisher and the editor(s). We stay neutral with regard to jurisdictional claims in published maps and institutional affiliations.