Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributional deep reinforcement learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Deep reinforcement learning assisted co-optimization of Volt-VAR grid service in distribution networks

With the increasing penetration of distributed energy resources in distribution networks, Volt-VAR control and optimization (VVC/VVO) have become very important to ensure an acceptable quality of service to all customers. System operators can rely on slow-responding utility devices, including capacitor banks and on-load tap changing transformers, along with fast-responding battery and photovoltaic (PV) inverters for the VVC/VVO implementation. Because of variations in response time of these two classes of devices, and different control actions (discrete versus continuous), coordinated and optimal scheduling and operation have become of utmost importance. Here, this paper develops a look-ahead deep reinforcement learning (DRL)-based multi-objective VVO technique to improve the voltage profile of active distribution networks, decrease network and inverter power loss, and save the operational cost of the grid. It proposes a deep deterministic policy gradient (DDPG)-based approach to schedule the optimal reactive and/or active power set-points of fast-responding inverters, and a deep Q-network (DQN)-based DRL agent to schedule the discrete decisions variables of slow-responding assets. The reactive power output of PV and battery smart inverters are scheduled at 30-minute intervals and the capacitors’ commitment status is scheduled with several hour intervals. The proposed framework is validated on the modified IEEE 34-bus and 123-bus test cases with embedded PV and PV-plus-storage. To validate the efficacy of the proposed VVO, it is compared with several scenarios, including the base case without VVO, localized droop control of DERs, DDPG-only, and twin delayed DDPG (TD3) agent-based DRL techniques. The results justify the superior performance of the proposed method to improve the voltage profile, reduce network power loss, and minimize the look-ahead grid operational cost while minimizing the undesirable power losses in inverters as a result of power factor adjustments.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Dynamic Role-Based Access Control Policy for Smart Grid Applications: An Offline Deep Reinforcement Learning Approach

Role-based access control (RBAC) is adopted in the information and communication technology domain for authentication purposes. However, due to a very large number of entities within organizational access control (AC) systems, static RBAC management can be inefficient, costly, and can lead to cybersecurity threats. In this paper, a novel hybrid RBAC model is proposed, based on the principles of offline deep reinforcement learning (RL) and Bayesian belief networks. The considered framework utilizes a fully offline RL agent, which models the behavioral history of users as a Bayesian belief-based trust indicator. Thus, the initial static RBAC policy is improved in a dynamic manner through off-policy learning while guaranteeing compliance of the internal users with the security rules of the system. By deploying our implementation within the smart grid domain and specifically within a Distributed Energy Resources (DER) ecosystem, we provide an end-to-end proof of concept of our model. Finally, detailed analysis and evaluation regarding the offline training phase of the RL agent are provided, while the online deployment of the hybrid RL-based RBAC model into the DER ecosystem highlights its key operation features and salient benefits over traditional RBAC models.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Hybrid-RL-MPC4CLR (Hybird-Reinforcement-Learning-Model-Predictive-Control-for-Reserve-Policy-Assisted-Critical-Load-Restoration-in-Distribution-Grids)

Hybrid-RL-MPC4CLR was developed as a hybrid controller for active distribution grid critical load restoration, combining deep reinforcement learning (RL) and model predictive control (MPC) aiming at maximizing total restored load following an extreme event. The RL determines a policy for quantifying operating reserve requirements, thereby hedging against uncertainty, while the MPC models grid operations incorporating the RL policy actions (i.e., reserve requirements), renewable (wind and solar) power predictions, and load demand forecasts. The developers formulated the reserve requirement determination problem as a sequential decision-making problem based on the Markov Decision Process (MDP) and design an RL learning environment based on the OpenAI Gym framework and MPC simulation. The RL agent reward and MPC objective function aim to maximize and monotonically increase total restored load and minimize load shedding and renewable power curtailment. The software is developed using various software packages in Python. The MPC's optimal power flow (OPF) model is implemented using the Pyomo package, the RL simulation environment is implemented using the MPC simulation with various scenarios of renewable energy and load demand profiles and power outage beginning times, based on the OpenAI Gym framework. The RL agent training is performed using the RLlib Ray package. The RL algorithm is trained offline using historical forecasts of renewable generation and load demand profiles. Simulation analysis and performance tests are conducted using a modified IEEE 13-bus distribution test feeder containing wind turbine, photovoltaic, microturbine, and battery.

Eseye, Abinet Tesfaye↗

Comprehensive assessment of deep reinforcement learning approaches for economic dispatch in nuclear-driven microgrids

As the electrical grid integrates more variable renewable energy sources such as wind and solar, the demand for distributed and flexible systems to address this increased variability becomes critical. Nuclear-driven microgrids provide a promising solution by offering stable generation to complement intermittent renewables, ensuring grid reliability and operating efficiency. This paper proposes a recurrent deep reinforcement learning framework for optimal economic dispatch in a nuclear-powered microgrid integrating renewable energy sources, small modular reactors, battery storage systems, and balance-of-plant dynamics. A three-agent control architecture is developed, where demand and renewable energy agents act as forecasters, and a reinforcement learning-based dispatch agent performs real-time energy allocation. A nonlinear programming formulation is first used to generate an optimal baseline for benchmarking. The proposed dispatch controller, based on Proximal Policy Optimization enhanced with Long Short-Term Memory networks, exploits temporal correlations in system dynamics by taking advantage of the time series used as inputs to improve policy robustness under uncertainty. Comparative analysis against established deep reinforcement learning methods, including Proximal Policy Optimization with a feedforward architecture, Soft Actor-Critic, and Twin Delayed Deep Deterministic Policy Gradient, demonstrates superior performance. Numerical results indicate that the proposed controller achieves a 0.39% cost reduction relative to the nonlinear programming benchmark and outperforms other learning-based methods by generating additional revenue of up to 0.35%. All reinforcement learning controllers compute dispatch actions in less than 0.3 s, resulting in a computational speedup of more than three orders of magnitude over the nonlinear programming baseline. The findings of this paper highlight their applicability for real-time operation and control in nuclear-integrated microgrids under volatile operating conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Post-Disaster Microgrid Formation for Enhanced Distribution System Resilience

This paper proposes a deep reinforcement learning (DRL) based approach for post-disaster critical load restoration in active distribution systems to form microgrids through network reconfiguration to minimize critical load curtailments. Distribution networks are represented as graph networks, and optimal network configurations with microgrids are obtained by searching for the optimal spanning forest. The constraints to the research question being explored are the radial topology and power balance. Unlike existing analytical and population-based approaches, which necessitate the repetition of entire analyses and computation for each outage scenario to find the optimal spanning forest, the proposed approach, once properly trained, can quickly determine the optimal, or near-optimal, spanning forest even when outage scenarios change. When multiple lines fail in the system, the proposed approach forms microgrids with distributed energy resources in active distribution systems to reduce critical load curtailment. The proposed DRL-based model learns the action-value function using the REINFORCE algorithm, which is a model-free reinforcement learning technique based on stochastic policy gradients. A case study was conducted on a 33-node distribution test system, demonstrating the effectiveness of the proposed approach for post-disaster critical load restoration.

active distribution systems↗

Branching Dueling Q-Network Based Online Scheduling of a Microgrid With Distributed Energy Storage Systems

This letter investigates a Branching Dueling Q-Network (BDQ) based online operation strategy for a microgrid with distributed battery energy storage systems (BESSs) operating under uncertainties. Additionally, the developed deep reinforcement learning (DRL) based microgrid online optimization strategy can achieve a linear increase in the number of neural network outputs with the number of distributed BESSs, which overcomes the curse of dimensionality caused by the charge and discharge decisions of multiple BESSs. Numerical simulations validate the effectiveness of the proposed method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Federated Deep Reinforcement Learning for Decentralized VVO of BTM DERs

The future of grid control requires a hybrid approach combining centralized and decentralized methods to fully utilize the potential of smart edge devices with artificial intelligence (AI) capabilities. This paper aims to develop and evaluate a federated deep reinforcement learning (FDRL) framework for decentralized adaptive volt-var optimization (VVO) of behind-the-meter (BTM) distributed energy resources (DERs). First, this paper models a single deep reinforcement learning (DRL) agent using the Markov Decision Process (MDP) framework for decentralized adaptive VVO of BTM DERs. Two DRL algorithms, soft actor-critic (SAC) and twin-delayed deep deterministic policy gradient (TD3), are compared for their effectiveness in optimizing VVO. Results show that TD3 outperforms SAC, achieving a 71.3% improvement in mean reward. Finally, the DRL agent is deployed within the FDRL framework, using the Flower platform, to enhance learning, provide adaptive control, and ensure data privacy for BTM DERs.

Ravi, Abhijith↗

Optimal Coordination of Distributed Energy Resources Using Deep Deterministic Policy Gradient

Recent studies showed that reinforcement learning (RL) is a promising approach for coordination and control of distributed energy resources (DER) under uncertainties. Many existing RL approaches, including Q-learning and approximate dynamic programming, are based on lookup table methods, which become inefficient when the problem size is large and infeasible when continuous states and actions are involved. In addition, when modeling battery energy storage system (BESS), the loss of life is not reasonably considered into the decision-making process. This paper proposes an innovative deep RL method for DER coordination considering BESS degradation. The proposed deep RL is designed based on an adaptive actor-critic architecture and employs an off-policy deterministic policy gradient method for determining the dispatch operation that minimizes the operation cost and BESS life loss. Case studies were performed to validate the proposed method and demonstrate the effects of incorporating degradation models into control design.

Das, Avijit↗

Soft Actor Critic Based Volt-VAR Co-optimization in Active Distribution Grids

Modern distribution networks are undergoing several technical challenges, such as voltage fluctuations, because of high penetration of distributed energy resources (DERs). This paper proposes a deep reinforcement learning (DRL)-based Volt VAR co-optimization technique for reducing voltage fluctuations as well as power loss under high penetration of DERs. In addition, the proposed approach minimizes the operational cost of the grid. A stochastic policy optimization based soft actor critic (SAC) agent is proposed to configure the optimal set-points of the reactive power outputs of the inverters. The performance of the proposed model is verified on the modified IEEE 34- and 123-bus systems and compared with a base case scenario with no reactive supply by inverters, and a local droop control approach. The results demonstrate that the proposed framework outperforms the conventional droop control method in improving the voltage profile, minimizing the network power loss, and reducing grid operational cost.

—Distribution grids, deep reinforcement learning, ↗

Topology-Aware Reinforcement Learning for Voltage Control: Centralized and Decentralized Strategies

Volt-VAR control (VVC) methods based on deep reinforcement learning (DRL) can effectively control distribution grid voltage and minimize power loss by implementing corrective and preventive control measures on the reactive power output of inverter-based distributed energy resources (DERs). However, model-free DRL-based VVC approaches usually cannot capture the important topological feature of the power system since they use a fully-connected network (FCN) to deliver the action. Therefore, this paper proposes a graph convolutional network (GCN)-based DRL approach that can employ the topological information of the network to take better control action for regulating the voltage. Our implementation allows for both centralized and decentralized configurations, utilizing a single agent and multiple agents respectively. Although the centralized GCN-based DRL approach has its advantages of minimizing voltage fluctuation and power loss, it is not suitable for large scale power systems due to its challenges in terms of scalability, computation speed and potential single points of failure. Therefore, these problems can be resolved using the decentralized GCN-based DRL approach. Moreover, to ensure the safe operation of the model, our proposed approach incorporates an exponential barrier function while formulating the reward function for each agent. To validate performance of the proposed approaches, the proposed model is tested on modified IEEE test systems and the performances are measured in terms on voltage fluctuation reduction, minimization of power loss and computational speed. Finally, the results show that the proposed topology-aware approach outperforms the FCN-based DRL approach in terms of reducing voltage fluctuation and minimizing power loss of the network. Moreover, it is shown that the decentralized GCN-based DRL has faster computational speed than other approaches.

42 ENGINEERING↗

Quantum Reinforcement Learning for Volt-VAR Control in Power Distribution Systems

Volt-VAR control (VVC) is crucial in active distribution networks for optimizing voltage profiles and minimizing network losses. While traditional deep reinforcement learning (DRL) algorithms exhibit promise for VVC, they often require extensive computational resources to handle such a high-dimensional problem. As a potential solution, quantum reinforcement learning (QRL) algorithms integrate the computational capabilities of quantum computing into the DRL framework. However, existing QRL algorithms struggle with complex VVC problems due to the limitations of current quantum hardware. To bridge this gap, this paper proposes an innovative QRL algorithm featuring an end-to-end architecture that integrates a classical autoencoder, variational quantum circuits (VQCs), and classical post-processing layers. This design efficiently compresses high-dimensional grid states, enabling VQCs to leverage quantum advantages while producing multiple control device outputs tailored for VVC tasks. Numerical studies on three representative distribution systems verify the effectiveness and scalability of the proposed QRL algorithm, and demonstrate its enhanced performance over classical approaches with only approximately 1% of the parameters. Additionally, the robustness of our developed algorithm is validated through noisy quantum environments.

97 MATHEMATICS AND COMPUTING↗

Towards intelligent emergency control for large-scale power systems: Convergence of learning, physics, computing and control

Here, this paper has delved into the pressing need for intelligent emergency control in large-scale power systems, which are experiencing significant transformations and are operating closer to their limits with more uncertainties. Learning-based control methods are promising and have shown effectiveness for intelligent power system control. However, when they are applied to large-scale power systems, there are multifaceted challenges such as scalability, adaptiveness, and security posed by the complex power system landscape, which demand comprehensive solutions. The paper first proposes and instantiates a convergence framework for integrating power systems physics, machine learning, advanced computing, and grid control to realize intelligent grid control at a large scale. Our developed methods and platform based on the convergence framework have been applied to a large (more than 3000 buses) Texas power system, and tested with 56 000 scenarios. Our work achieved a 26% reduction in load shedding on average and outperformed existing rule-based control in 99.7% of the test scenarios. The results demonstrated the potential of the proposed convergence framework and DRL-based intelligent control for the future grid.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Decentralized Voltage Control of Large-Scale Distribution System with PVs Based on MADRL

This paper proposes a model-free decentralized control framework for the voltage regulation of large-scale distribution systems through the coordinated control of PV inverters. This is achieved by developing a novel interaction mechanism between the surrogate model and the centralized training and decentralized execution multiagent deep reinforcement learning framework. Specifically, the sparse Gaussian processes regression method is first utilized to develop the surrogate model of the original distribution system for reward calculation during the training stage, where each agent represents a sub-region in the centralized fashion for coordination strategy learning. After that, the learned control rules are used to inform controllers within each sub-region for real-time decisions with only local measurements. Comparative tests among various methods on the EPRI Ckt5 test system demonstrate the effectiveness of the proposed method.

distribution system↗

A Reinforcement Learning Approach to Parameter Selection for Distributed Optimal Power Flow

With the increasing penetration of distributed energy resources, distributed optimization algorithms have attracted significant attention for power systems applications due to their potential for superior scalability, privacy, and robustness to a single point-of-failure. The Alternating Direction Method of Multipliers (ADMM) is a popular distributed optimization algorithm; however, its convergence performance is highly dependent on the selection of penalty parameters, which are usually chosen heuristically. In this work, we use reinforcement learning (RL) to develop an adaptive penalty parameter selection policy for alternating current optimal power flow (ACOPF) problem solved via ADMM with the goal of minimizing the number of iterations until convergence. We train our RL policy using deep Q-learning and show that this policy can result in significantly accelerated convergence (up to a 59% reduction in the number of iterations compared to existing, curvatureinformed penalty parameter selection methods). Furthermore, we show that our RL policy demonstrates promise for generalizability, performing well under unseen loading schemes as well as under unseen losses of lines and generators (up to a 50% reduction in iterations). This work thus provides a proof-of-concept for using RL for parameter selection in ADMM for power systems applications.

alternating current optimal power flow↗

Learning Sequential Distribution System Restoration via Graph-Reinforcement Learning

We report a distribution service restoration algorithm as a fundamental resilient paradigm for system operators provides an optimally coordinated, resilient solution to enhance the restoration performance. The restoration problem is formulated to coordinate distribution generators and controllable switches optimally. A model-based control scheme is usually designed to solve this problem, relying on a precise model and resulting in low scalability. To tackle these limitations, this work proposes a graph-reinforcement learning framework for the restoration problem. We link the power system topology with a graph convolutional network, which captures the complex mechanism of network restoration in power networks and understands the mutual interactions among controllable devices. Latent features over graphical power networks produced by graph convolutional layers are exploited to learn the control policy for network restoration using deep reinforcement learning. The solution scalability is guaranteed by modeling distributed generators as agents in a multi-agent environment and a proper pre-training paradigm. Comparative studies on IEEE 123-node and 8500-node test systems demonstrate the performance of the proposed solution.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Deep Reinforcement Scheduling of Energy Storage Systems for Real-time Voltage Regulation in Unbalanced LV Networks with High PV Penetration

The ever-growing higher penetration of distributed energy resources (DERs) in low-voltage (LV) distribution systems brings both opportunities and challenges to voltage support and regulation. This paper proposes a deep reinforcement learning (DRL)-based scheduling scheme of energy storage systems (ESSs) to mitigate system voltage deviations in unbalanced LV distribution networks. The ESS-based voltage regulation problem is formulated as a multi-stage quadratic stochastic program, with the objective of minimizing the expected total daily voltage regulation cost while satisfying operational constraints. While existing voltage regulation methods are mostly focused on onetime- step control, this paper explores a day-horizon systemwide voltage regulation problem. In other words, the size of action and state spaces are extremely high-dimensional and need to be delicately handled. Furthermore, in order to overcome the difficulty of modeling uncertainties and develop a realtime solution, a learn-to-schedule feedback control framework is proposed by adapting the problem to a model-free DRL setting. The proposed algorithm is tested on a customized 6-bus system and a modified IEEE 34-bus system. Simulation results validate the effectiveness and near-optimality of voltage regulation by ESS in comparison with a deterministic quadratic program solution.

Wang, Shengyi↗

PowerNet: Multi-agent Deep Reinforcement Learning for Scalable Powergrid Control

This paper develops an efficient multi-agent deep reinforcement learning algorithm for cooperative controls in powergrids. Specifically, we consider the decentralized inverter-based secondary voltage control problem in distributed generators (DGs), which is first formulated as a cooperative multi-agent reinforcement learning (MARL) problem. We then propose a novel on-policy MARL algorithm, PowerNet, in which each agent (DG) learns a control policy based on (sub-)global reward but local states and encoded communication messages from its neighbors. Motivated by the fact that a local control from one agent has limited impact on agents distant from it, we exploit a novel spatial discount factor to reduce the effect from remote agents, to expedite the training process and improve scalability. Furthermore, a differentiable, learning-based communication protocol is employed to foster the collaborations among neighboring agents. In addition, to mitigate the effects of system uncertainty and random noise introduced during on-policy learning, we utilize an action smoothing factor to stabilize the policy execution. To facilitate training and evaluation, we develop PGSim, an efficient, high-fidelity powergrid simulation platform. Here, experimental results in two microgrid setups show that the developed PowerNet outperforms the conventional model-based control method, as well as several state-of-the-art MARL algorithms. The decentralized learning scheme and high sample efficiency also make it viable to large-scale power grids.

24 POWER TRANSMISSION AND DISTRIBUTION↗