Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributional deep reinforcement learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Distributional Deep Reinforcement Learning-Based Emergency Frequency Control

Emergency frequency control is one of the most critical approaches to maintain power system stability after major disturbances. With the increasing number of grid-connected renewable energy sources, existing model-based methods of frequency control are facing up with challenges of computational speed and scalability for large-scale systems. In this paper, the emergency frequency control problem is formulated as a Markov Decision Process (MDP) and solved through a novel Distributional Deep Reinforcement Learning (DDRL) method, namely the distributional soft actor critic (DSAC) method. Compared with other RL methods that only estimate the mean value, the proposed DSAC model estimates the distribution of value function over returns. This advancement can lead to more insights and knowledge for the agent, with the benefit of a much faster and more stable learning process, and the improved frequency control performance. Here, the simulation results on IEEE 39-bus and IEEE 118-bus systems demonstrate the effectiveness and robustness of proposed models, as well as the advantage compared to other state-of-the-art DRL algorithms.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Volt-VAR Optimization in Distribution Networks Using Twin Delayed Deep Reinforcement Learning

Modern distribution grids are undergoing new challenges due to the stochastic nature of distributed energy resources (DERs). High penetration of DERs has a significant impact on Volt-VAR profile and system power losses. This work proposes a deep reinforcement learning (DRL)-based Volt-VAR optimization approach for improving voltage profile and reducing system power loss under high penetration of distributed energy resources, such as battery energy storage and solar photovoltaic units in distribution grids. The twin delayed deep deterministic policy gradient (TD3) method-based DRL agent is proposed to configure optimal set-points of reactive power outputs of fast responding smart inverters. The agent schedules the reactive power of inverters according to their physical capabilities, such as minimum allowed power factor, e.g., 0.9 leading/lagging. The reward function of the proposed DRL scheme is designed carefully to ensure a proper voltage profile of the grids with effective scheduling of reactive power outputs from inverters. The performance of the proposed model is verified on modified IEEE 34- and 123-bus systems and compared with base case with no reactive supply by inverters, and local droop Volt-VAR control approach. The results show that the proposed method performs better than the local droop control and deep deterministic policy gradient (DDPG)-based DRL method for reducing voltage fluctuation and minimizing power loss.

Hossain, Rakib↗

Two-Stage Deep Reinforcement Learning for Distribution System Voltage Regulation and Peak Load Management: Preprint

The growing integration of distributed solar photovoltaic (PV) in distribution systems could result in adverse effects during grid operation. This paper develops a soft actor critic-based deep reinforcement learning (SAC-DRL) solution to simultaneously control PV inverters and battery energy storage systems for voltage regulation and peak load demand shaving. The novel two-stage framework, featured with two different control agents, is applied for daytime and nighttime operation to enhance the control performance. Comparison results with other control methods on a real feeder in Western Colorado demonstrate that the proposed method can provide advanced voltage regulation with modest active power curtailment for peak demand reduction.

deep reinforcement learning↗

Two-Stage Deep Reinforcement Learning for Distribution System Voltage Regulation and Peak Load Management

The growing integration of distributed solar photovoltaic (PV) in distribution systems could result in adverse effects during grid operation. This paper develops a two-agent soft actor critic-based deep reinforcement learning (SAC-DRL) solution to simultaneously control PV inverters and battery energy storage systems for voltage regulation and peak demand reduction. The novel two-stage framework, featured with two different control agents, is applied for daytime and nighttime operations to enhance control performance. Comparison results with other control methods on a real feeder in Western Colorado demonstrate that the proposed method can provide advanced voltage regulation with modest active power curtailment and reduce peak load demand from feeder's head.

deep reinforcement learning↗

Deep Reinforcement Learning for Distribution System Restoration Using Distributed Energy Resources and Tie-Switches

Distributed energy resources (DERs), such as solar PVs and energy storage, can be used to restore distribution system critical loads after the extreme weather events to increase grid resilience. However, coordinating multiple DERs together with tie-switches for multi-step restoration process under renewable uncertainty is challenging. This paper proposes a deep reinforcement learning to control discrete actions of switching on/off tie switches and DERs for critical load restoration. The restoration problem is first cast into the Markov decision process suitable for DRL. Then, the original soft actor critic (SAC) method for continuous actions has been extended to handle discrete and continuous actions. Numerical comparison results with other stochastic optimization-based approaches on the modified IEEE 33-bus system show that the proposed method can achieve fast critical load restoration in the presence of substation power outage while maintaining system voltage limit throughout the restoration process.

active distribution systems↗

Deep Reinforcement Learning for Distribution System Cyber Attack Defense with DERs

The use of smart inverter capabilities of distributed energy resources (DERs) enhances the grid reliability but in the meanwhile exhibits more vulnerabilities to cyber-attacks. This paper proposes a deep reinforcement learning (DRL)-based defense approach. The defense problem is reformulated as a Markov decision making process to control DERs and minimizing load shedding to address the voltage violations caused by cyber-attacks. The original soft actor-critic (SAC) method for continuous actions has been extended to handle discrete and continuous actions for controlling DERs' setpoints and loadshedding scenarios. Numerical comparison results with other control approaches, such as Volt-VAR and Volt-Watt on the modified IEEE 33-node, show that the proposed method can achieve better voltage regulation and have less power losses in the presence of cyber-attacks.

active distribution systems↗

Deep Reinforcement Learning for Distribution System Operations: A Tutorial and Survey

Here, the rapid evolution of modern electric power distribution systems into complex networks of interconnected active devices, distributed generation (DG), and storage poses increasing difficulties for system operators. The large-scale integration of distributed energy resources (DERs) and the rapid exchange of measurement data via communication networks present major opportunities for advancing grid operations but also introduce greater uncertainty, higher data dimensionality, more complex network and device models, and challenging control and optimization problems. Deep reinforcement learning (DRL) algorithms are promising in addressing these challenges. However, they have not been effectively adapted for power systems applications, requiring extensive customization for implementation and evaluation. This has resulted in reproducibility challenges and a steep learning curve for researchers new to applying DRL algorithms to the power systems domain. To bridge these gaps, this tutorial aims to serve as a valuable resource for researchers interested in exploring learning-based algorithms to operate active power distribution networks. Specifically, this work presents a generalized process for translating sequential decision-making problems in power distribution systems into Markov decision process (MDP) formulations, illustrated through concrete grid service examples. Additionally, we introduce a simple environment design strategy to develop and evaluate example DRL algorithms for distribution system applications, complete with an included code repository to guide users through environment construction.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Blockchain Enabled Intelligence of Federated Systems (BELIEFS): An attack-tolerant trustable distributed intelligence paradigm

In this article, a Blockchain Enabled Intelligence of Federated Systems (BELIEFS) is proposed to conduct cooperative control for the multi-regional large-scale power system with a multi-agents system (MAS). By establishing a two levels blockchain, each regional AI agent can simultaneously manage intra-regional controllers and cooperate with other AI agents. Under the consensus mechanism, the agents, which respectively conducted distributed deep reinforcement learning (DDRL) algorithm in multi-regions, can have the tolerant capability of malicious attacks in their training process. The demonstration of the proposed approach is within a multi-regional large-scale interconnected power system. Under the mode of "centralized dispatching and hierarchical management", this article aims to definite a mathematical model to deal with the control problem of the power systems. With the comparison experiments, the effectiveness and efficiency of our proposed method in the training process are verified. In addition, malicious attacks are set on the main chain and shard chains to verify the attack-tolerant capability. We expect that such approach and results can suggest a new paradigm of attack-tolerant trustable distributed AI deployment.

97 MATHEMATICS AND COMPUTING↗

Safe Deep Reinforcement Learning for Active Distribution System Model Predictive Control with EVs and DERs

The temporal and spatial mismatch between PV generation and electric vehicle (EV) charging and discharging may cause voltage violations in active distribution networks. Despite the widespread use of deep reinforcement learning (DRL) in power system optimization and control, it lacks guarantees on constraint satisfaction during both training and deployment. This paper proposes a Lagrangian-based safe DRL approach for model predictive control (MPC) of active distribution systems with large-scale integration of PVs, EVs, and energy storage systems (ESSs). A Transformer-LSTM time-series model is proposed to forecast EV charging demand, which is then formulated as a constraint to ensure charging requirements are met. Using this prediction, a Lagrangian-based safe soft actor-critic (SAC) framework is developed for real-time control in a three-phase unbalanced distribution system, enforcing voltage safety constraints while optimizing the cumulative net reward. By integrating the forecasting model with multi-period constraints, the proposed framework jointly coordinates PV systems, EV charging and discharging, and ESS scheduling within the MPC horizon. Numerical experiments on a modified IEEE 123-bus system with real-world data show that, under a high PV penetration scenario, the proposed method increases the net reward by 30.74% and reduces average voltage violations from 0.0011 p.u. to 0.0002 p.u. compared with standard SAC. Compared with the optimal power flow (OPF) approach, it achieves similar voltage security while yielding lower line losses. It also maintains real-time control capability, reducing operation latency to 53.21 ms per 15-minute control interval. The proposed method remains effective under varying PV/EV penetrations and load conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

On the Verification of Deep Reinforcement Learning Solution for Intelligent Operation of Distribution Grids

Capabilities of deep reinforcement learning (DRL) in obtaining fast decision policies in high dimensional and stochastic environments have led to its extensive use in operational research, including the operation of distribution grids with high penetration of distributed energy resources (DER). However, the feasibility and robustness of DRL solutions are not guaranteed for the system operator, and hence, those solutions may be of limited practical value. This paper proposes an analytical method to find feasibility ellipsoids that represent the range of multi-dimensional system states in which the DRL solution is guaranteed to be feasible. Empirical studies and stochastic sampling determine the ratio of the discovered to the actual feasible space as a function of the sample size. In addition, the performance of logarithmic, linear, and exponential penalization of infeasibility during the DRL training are studied and compared in order to reduce the number of infeasible solutions

Hosseini, Mohammad Mehdi↗

Multi-Agent Deep Reinforcement Learning for Realistic Distribution System Voltage Control Using PV Inverters

Over the last few decades, the deployment of distributed solar photovoltaic (PV) systems has increased consistently. High PV penetration could cause adverse effects on the grid, such as voltage violations. This paper proposes a new distributed soft actor-critic based multi-agent deep reinforcement learning (SAC-MADRL) control solution to minimize the PV real power curtailment while keeping the grid voltage in an acceptable range. New reward functions have been designed to coordinate different agents during the learning process, yielding improved convergence. Comparison results with other control methods on a real feeder in western Colorado U.S. with 80% penetration of PVs demonstrate that the proposed method has better capability of effectively regulating voltage while minimizing the PV real power curtailment.

distribution system↗

A Novel Multi-Agent Deep Reinforcement Learning-enabled Distributed Power Allocation Scheme for mmWave Cellular Networks

We consider the power allocation problem over shared spectrum for millimeter-Wave (mmWave) cellular downlink. Existing approaches usually find sub-optimal solutions by solving a non-convex optimization which leads to scalability issues due to centralized control. Therefore, distributed and adaptive approaches are desirable. Recently, model-free Deep Reinforcement Learning (DRL) has achieved success in such wireless resource management tasks. By modeling the radio environment as a Markov Decision Process (MDP) with the base stations (BSs) being the agents, power allocation can be automated at the agent level with comparable throughput performance to conventional centralized schemes. The multi-agent setting presents new challenges as the radio environment is impacted by the joint actions of the agents and is no longer stationary from any individual agent’s perspective. Existing literature bypasses this non-stationarity violation by ignoring it which may cause performance degradation. To tackle this issue, we propose a distributed continuous power allocation scheme based on a modified version of multi-agent Deep Deterministic Policy Gradient (MADDPG) that is tailored for the distributed multiple-agent setting. The proposed scheme employs a centralized-training distributed-execution framework where Q-functions are trained over subsets of BSs while each BS determines its transmit power based only on its own local observation. It admits constant per-BS communication and computation complexity and is thus scalable to large networks. Numerical evaluation shows that the proposed scheme adapts well to a wide range of interference conditions and can achieve comparable or better performance than several state-of-the-art non-learning approaches.

99 GENERAL AND MISCELLANEOUS↗

Controlling distributed energy resources via deep reinforcement learning for load flexibility and energy efficiency

Behind-the-meter distributed energy resources (DERs), including building solar photovoltaic (PV) technology and electric battery storage, are increasingly being considered as solutions to support carbon reduction goals and increase grid reliability and resiliency. However, dynamic control of these resources in concert with traditional building loads, to effect efficiency and demand flexibility, is not yet commonplace in commercial control products. Traditional rule-based control algorithms do not offer integrated closed-loop control to optimize across systems, and most often, PV and battery systems are operated for energy arbitrage and demand charge management, and not for the provision of grid services. More advanced control approaches, such as MPC control have not been widely adopted in industry because they require significant expertise to develop and deploy. Recent advances in deep reinforcement learning (DRL) offer a promising option to optimize the operation of DER systems and building loads with reduced setup effort. However, there are limited studies that evaluate the efficacy of these methods to control multiple building subsystems simultaneously. Additionally, most of the research has been conducted in simulated environments as opposed to real buildings. This paper proposes a DRL approach that uses a deep deterministic policy gradient algorithm for integrated control of HVAC and electric battery storage systems in the presence of on-site PV generation. The DRL algorithm, trained on synthetic data, was deployed in a physical test building and evaluated against a baseline that uses the current best-in-class rule-based control strategies. Performance in delivering energy efficiency, load shift, and load shed was tested using price-based signals. The results showed that the DRL-based controller can produce cost savings of up to 39.6% as compared to the baseline controller, while maintaining similar thermal comfort in the building. The project team has also integrated the simulation components developed during this work as an OpenAIGym environment and made it publicly available so that prospective DRL researchers can leverage this environment to evaluate alternate DRL algorithms.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Data-Driven Distribution System Coordinated PV Inverter Control Using Deep Reinforcement Learning

The deployment of distributed solar photovoltaic (PV) systems has increased consistently over the past decades. High penetrations of PVs could cause a series of adverse grid impacts, such as voltage violations. The recent development of smart inverter technologies rises the incentives of developing PV control solutions that regulate the inverter output power and seeking the optimization on system operational objectives. This paper proposes a data-driven control solution based on deep reinforcement learning (DRL) to optimize PV inverters for voltage regulation. The proposed solution can minimize PV real power curtailment while maintaining network voltage at an acceptable range. Comparison results between the proposed DRL control algorithms with deep deterministic policy gradient (DDPG) and volt-var control on a real feeder in west Colorado highlight the advantage of the proposed framework in controlling the system voltage while minimizing the PV real power curtailment.

deep reinforcement learning↗

Adaptive Deep Reinforcement Learning Algorithm for Distribution System Cyber Attack Defense With High Penetration of DERs

With grid modernization, smart inverters are increasingly used to execute advanced controls for distribution network reliability. However, this also increases the cyber-attack space. Here this paper focuses on the defense approaches to restore the system to normal operation circumstances in the presence of cyber-attacks. A unique deep reinforcement learning (DRL) method is developed to minimize voltage violations and reduce power losses for impacted feeders. The defense problem is reformulated as a Markov decision-making process to dynamically control DERs while minimizing load shedding. This is achieved via an improved soft actor-critic (SAC)-based DRL algorithm, which can govern DER set points and load-shedding scenarios in discrete and continuous modes via the auto-tune entropy and Gaussian policy features. Numerical comparison results on the modified IEEE 123-node system with other control approaches, such as Volt-VAR (VV), Volt-Watt (VW), and model predictive control (MPC) show that the proposed method can eliminate voltage violations and provide feasible control actions that perform complete mitigation of cyber-threats.

24 POWER TRANSMISSION AND DISTRIBUTION↗