Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “soft actor critic”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Soft Actor Critic Based Volt-VAR Co-optimization in Active Distribution Grids

Modern distribution networks are undergoing several technical challenges, such as voltage fluctuations, because of high penetration of distributed energy resources (DERs). This paper proposes a deep reinforcement learning (DRL)-based Volt VAR co-optimization technique for reducing voltage fluctuations as well as power loss under high penetration of DERs. In addition, the proposed approach minimizes the operational cost of the grid. A stochastic policy optimization based soft actor critic (SAC) agent is proposed to configure the optimal set-points of the reactive power outputs of the inverters. The performance of the proposed model is verified on the modified IEEE 34- and 123-bus systems and compared with a base case scenario with no reactive supply by inverters, and a local droop control approach. The results demonstrate that the proposed framework outperforms the conventional droop control method in improving the voltage profile, minimizing the network power loss, and reducing grid operational cost.

—Distribution grids, deep reinforcement learning, ↗

Power System Frequency Dynamics Modeling, State Estimation, and Control using Neural Ordinary Differential Equations (NODEs) and Soft Actor-Critic (SAC) Machine Learning Approaches

With the global energy transition of the electric power system, grid control, supervision, and protection is becoming more challenging. With the increasing integration of renewable energy sources (RES), the system dynamics are changing, causing traditional power system dynamic modeling with swing equation-based modeling approaches to fail. Additionally, the converter-dominated power grid is decreasing the system inertia, making the power system more fragile to the frequency swings. This paper first investigates and compares the application of a model-based Kalman filter state estimation approach with (i) a model-free machine learning approach --- neural ordinary differential equations (NODEs) --- and (ii) a data-driven system identification (SysId) approach to model and infer critical state values of the power system frequency dynamics. Then a model predictive control (MPC) framework is compared to a model-free Soft Actor-Critic (SAC) reinforcement learning (RL) control algorithm in providing efficient fast frequency response (FFR) to the power system frequency dynamics. The approaches are compared in terms of their performance goals as well as their per-timestep computational efficiency. Furthermore, the comparative study for state estimation shows that for the model-free requirement, both NODEs and SysId can provide accurate state estimates; however, with increasing model complexity, NODEs can be a better choice for model identification. Similarly, the results from the FFR comparative study show that the SAC RL-based FFR, once trained, outperforms MPC with better control signals and faster computation time, making the SAC RL-based FFR better option for providing FFR to the power system.

97 MATHEMATICS AND COMPUTING↗

Multi-agent voltage control in distribution systems using GAN-DRL-based approach

Active distribution grids can experience voltage fluctuations and violations due to the high penetration of variable distributed energy resources (DERs). These problems might occur because of the uncertain and variable generation natures of these resources, especially solar photovoltaic resources, during panel shadowing scenarios. Volt-VAR control (VVC) is an efficient method that controls the reactive power set-points of the inverters to regulate the voltage of distribution grids. Although several VVC approaches have been proposed recently, the performance of these approaches degrades significantly if behind-the-meter solar generation data are unobservable/missing. Therefore, it is necessary to impute missing/unobservable PV data accurately to be utilized in VVC approaches. Further, this paper proposes a model-free, data-driven, centrally trained, and decentrally executed multi-agent deep reinforcement learning-based VVC architecture to regulate the voltage of distribution networks. A generative adversarial network (GAN) is incorporated to impute the unobservable PV data accurately, which improves the performance of the proposed control architecture. The proposed multi-agent-soft-actor–critic algorithm (MASAC)-based VVC technique utilizes the actual PV dataset as well as the imputed dataset from the GAN framework to learn the optimal coordinated control policy for controlling the optimal reactive power set-points of PV inverters. The effectiveness of the proposed approach is analyzed on a modified IEEE 34-bus test case with added PV inverters. The results are compared and analyzed with a base case model with no VVC and VVC with a local droop control approach, genetic algorithm optimization, and a centralized soft actor–critic-based approach. Moreover, the performance of the proposed approach is compared with that of a multi-agent VVC framework without using the PV generation data and load information as the system state. The results illustrate that the proposed method with more state input improves the voltage profile and reduces the power loss of the network across various loading and PV generation scenarios.

14 SOLAR ENERGY↗

Distributional Deep Reinforcement Learning-Based Emergency Frequency Control

Emergency frequency control is one of the most critical approaches to maintain power system stability after major disturbances. With the increasing number of grid-connected renewable energy sources, existing model-based methods of frequency control are facing up with challenges of computational speed and scalability for large-scale systems. In this paper, the emergency frequency control problem is formulated as a Markov Decision Process (MDP) and solved through a novel Distributional Deep Reinforcement Learning (DDRL) method, namely the distributional soft actor critic (DSAC) method. Compared with other RL methods that only estimate the mean value, the proposed DSAC model estimates the distribution of value function over returns. This advancement can lead to more insights and knowledge for the agent, with the benefit of a much faster and more stable learning process, and the improved frequency control performance. Here, the simulation results on IEEE 39-bus and IEEE 118-bus systems demonstrate the effectiveness and robustness of proposed models, as well as the advantage compared to other state-of-the-art DRL algorithms.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Topology-Aware Reinforcement Learning for Voltage Control: Centralized and Decentralized Strategies

Volt-VAR control (VVC) methods based on deep reinforcement learning (DRL) can effectively control distribution grid voltage and minimize power loss by implementing corrective and preventive control measures on the reactive power output of inverter-based distributed energy resources (DERs). However, model-free DRL-based VVC approaches usually cannot capture the important topological feature of the power system since they use a fully-connected network (FCN) to deliver the action. Therefore, this paper proposes a graph convolutional network (GCN)-based DRL approach that can employ the topological information of the network to take better control action for regulating the voltage. Our implementation allows for both centralized and decentralized configurations, utilizing a single agent and multiple agents respectively. Although the centralized GCN-based DRL approach has its advantages of minimizing voltage fluctuation and power loss, it is not suitable for large scale power systems due to its challenges in terms of scalability, computation speed and potential single points of failure. Therefore, these problems can be resolved using the decentralized GCN-based DRL approach. Moreover, to ensure the safe operation of the model, our proposed approach incorporates an exponential barrier function while formulating the reward function for each agent. To validate performance of the proposed approaches, the proposed model is tested on modified IEEE test systems and the performances are measured in terms on voltage fluctuation reduction, minimization of power loss and computational speed. Finally, the results show that the proposed topology-aware approach outperforms the FCN-based DRL approach in terms of reducing voltage fluctuation and minimizing power loss of the network. Moreover, it is shown that the decentralized GCN-based DRL has faster computational speed than other approaches.

42 ENGINEERING↗

Deep Reinforcement Learning for Distribution System Restoration Using Distributed Energy Resources and Tie-Switches

Distributed energy resources (DERs), such as solar PVs and energy storage, can be used to restore distribution system critical loads after the extreme weather events to increase grid resilience. However, coordinating multiple DERs together with tie-switches for multi-step restoration process under renewable uncertainty is challenging. This paper proposes a deep reinforcement learning to control discrete actions of switching on/off tie switches and DERs for critical load restoration. The restoration problem is first cast into the Markov decision process suitable for DRL. Then, the original soft actor critic (SAC) method for continuous actions has been extended to handle discrete and continuous actions. Numerical comparison results with other stochastic optimization-based approaches on the modified IEEE 33-bus system show that the proposed method can achieve fast critical load restoration in the presence of substation power outage while maintaining system voltage limit throughout the restoration process.

active distribution systems↗

Online transfer learning strategy for enhancing the scalability and deployment of deep reinforcement learning control in smart buildings

In recent years, advanced control strategies based on Deep Reinforcement Learning (DRL) proved to be effective in optimizing the management of integrated energy systems in buildings, reducing energy costs and improving indoor comfort conditions when compared to traditional reactive controllers. However, the scalability and implementation of DRL controllers are still limited since they require a considerable amount of time before converging to a near-optimal solution. This issue is currently addressed in literature through the offline pre-training of the DRL agent. However this solution results in two main critical issues: (1) the need to develop a building surrogate model to perform the training task, and (2) the need to perform a fine-tuning process over several training episodes to obtain a near-optimal control policy. In this context, this paper introduces an Online Transfer Learning (OTL) strategy that exploits two knowledge-sharing techniques, weight-initialization and imitation learning, to transfer a DRL control policy from a source office building to various target buildings in a simulation environment coupling EnergyPlus and Python. A DRL controller based on discrete Soft Actor–Critic (SAC) is trained on the source building to manage the operation of a cooling system consisting of a chiller and a thermal storage. Several target buildings are defined to benchmark the performance of the OTL strategy with that of a Rule-Based Controller (RBC) and two DRL-based control strategies, deployed in offline and online fashion. The strategy adopted for OTL emulates the real world implementation with a simulation process by implementing the transferred DRL agent for a single episode in the target buildings. Target buildings have the same geometrical features and are served by the same energy system as the source building, but differ in terms of weather conditions, electricity price schedules, occupancy patterns, and building envelope efficiency levels. The results show that the OTL strategy can reduce the cumulated sum of temperature violations on average by 50% and 80% respectively when compared to RBC and online DRL while enhancing the energy system operation with electricity cost savings ranging between 20% and 40%. Furthermore, the OTL agent performs slightly worse than the offline DRL controller but it does not require any modeling effort and can be implemented directly on target buildings emulating a real-world implementation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A reinforcement learning approach to long-horizon operations, health, and maintenance supervisory control of advanced energy systems

In this work, we develop a Reinforcement Learning (RL) approach to the supervisory control problem for advanced energy systems, such as novel nuclear reactors and other demand-driven, mission-critical, and component-health-sensitive energy plants. The inclusive problem landscape considered captures the stochastic confluence of plant performance, component health evolution, power demand from the grid, diverse maintenance actions, and operator-defined goals and constraints, all considered over meaningfully long-enough reasoning horizons. Key aspects of the proposed approach are a receding horizon control-inspired technique dictating time- or event-triggered supervisory policy (re-)constructions, as well as additional capability-enabling contributions such as timescale compression, to handle long reasoning horizons and uncertainty in parts of the problem, and practical yet demonstrably-effective handling of hybrid action spaces with continuous and discrete decision variables. The resulting algorithm consists of a simulation-based RL agent constructing stochastic supervisory control policies over nontrivial action spaces and for long horizons, applying the learned policy to the system for a much shorter interval, and perpetually repeating, to construct the next long-horizon policy. That next policy will only be applied, again, for a short interval, yet originally far-in-time events move progressively closer, their associated uncertainty decreases, and new events and aspects enter the reasoning horizon. The proposed methodology bridges fundamental receding horizon concepts with the unequivocally stronger and more scalable reasoning of contemporary RL. Numerical examples using Soft Actor–Critic Deep RL illustrate the operation and efficacy of the proposed technique for a power plant tasked with health-aware load following missions in a dynamic electricity market landscape.

97 MATHEMATICS AND COMPUTING↗

Deep reinforcement learning based optimization for a tightly coupled nuclear renewable integrated energy system

New ways to integrate energy systems to maximize efficiency are being sought to meet carbon emissions goals. Nuclear-renewable integrated energy system (NR-IES) concepts are a leading solution that couples a nuclear power plant with renewable energy, hydrogen generation plants, and energy storage systems, such that thermal and electrical power are dispatchable to fulfill grid-flexibility requirements while also producing hydrogen and maximizing revenue. Here, this paper introduces a deep reinforcement learning (DRL)-based framework to address the complex decision-making tasks for NR-IES. The objective is to maximize revenue by generating and selling hydrogen and electricity simultaneously according to their time-varying prices while keeping the energy flow in the subsystems in balance. A Python-based simulator for a NR-IES concept has been developed to integrate with OpenAI Gym and Ray/RLlib to enable an efficient and flexible computational framework for DRL research and development. Three state-of-the-art DRL algorithms have been investigated, including two-delayed deep deterministic policy gradient (TD3), soft-actor critic (SAC), proximal policy optimization (PPO), to illustrate DRL’s superiority for controlling NR-IES by comparing it with a conventional control approach, particle swarm optimization (PSO). In this effort, PPO has shown more-stable performance and also better generalization capability than SAC and TD3. Comparisons with PSO have demonstrated that, on average, PPO can achieve 13.9% more mean episode returns from the training process and 29.4% more mean episode returns from the testing process when different hydrogen-production targets are applied.

08 HYDROGEN↗

Demonstration of reconstruction-free static magnetic control of DIII-D plasma with deep reinforcement learning

This paper presents the development and experimental validation of a reinforcement learning (RL)-based magnetic controller on the DIII-D tokamak. The controller directly maps raw magnetic diagnostic signals to actuator commands, replacing the traditional isoflux control algorithm based on equilibrium reconstruction. Four RL controllers are trained using the Soft Actor–Critic algorithm with an asymmetric Actor–Critic architecture in the NSFsim simulator. All controllers are deployed in the DIII-D Plasma Control System and operated with a 4 kHz feedback loop. Two randomization strategies are evaluated during training: evolving kinetic profiles and fixed kinetic profiles within each episode. The latter approach is found to better capture experimental deviations in the current density profile and to provide overall improved control performance. Robust operation is demonstrated across heating power scans in both L- and H-mode plasmas, as well as during transient events such as L–H transitions and pellet injections. Control errors in plasma shape and radial position remained within 1.5–2.0 cm and 1 cm, respectively. A notable discrepancy was observed in the vertical X-point position, with errors of up to approximately 4 cm, attributed to the current density distribution mismatches between simulations and experiments.

DIII-D↗

Deep Reinforcement Learning for Microgrid Cost Optimization Considering Load Flexibility

This paper proposes a novel Soft-Actor-Critic (SAC) based Deep Reinforcement Learning (DRL) method for optimizing the cost of microgrid operation by leveraging load flexibility. The proposed SAC-DRL method is designed to coordinate the control of distributed energy resources (DERs) and flexible load, addressing practical energy billing formation by power distribution utilities. Key contributions include an innovative reward function to mitigate sparse reward challenges and a mixed control strategy for discrete and continuous variables, ensuring radial network topology and minimizing power loss. We evaluate the proposed method on the model of a real microgrid located in Southern California, U.S.. The SAC-DRL model is tested to demonstrate its efficacy in reducing grid dependence, optimizing resource use, and minimizing costs. The results highlight the potential of DRL in modern energy systems, offering a sustainable and economically efficient solution for energy management in microgrids.

deep reinforcement learning↗

Design, Selection, and Evaluation of Reinforcement Learning Single Agents for Ground Target Tracking

Previous approaches for small fixed-wing unmanned air systems that carry strapdown rather than gimbaled cameras achieved satisfactory ground target tracking performance using both standard and deep reinforcement learning algorithms. However, these approaches have significant restrictions and abstractions to the dynamics of the vehicle, such as constant airspeed and constant altitude, because the number of states and actions was necessarily limited. Thus, extensive tuning was required to obtain good tracking performance. The expansion from 4 state–action degrees of freedom to 15 enabled the agent to exploit previous reward functions that produced novel yet undesirable emergent behavior. This paper investigates the causes of and various potential solutions to undesirable emergent behavior in the ground target tracking problem. A combination of changes to the environment, reward structure, action space simplification, command rate, and controller implementation provides insight into obtaining stable tracking results. Consideration is given to reward structure selection and refinement to mitigate undesirable emergent behavior. Results presented in the paper for a simulated environment of a single unmanned air system tracking a randomly moving single ground target show that a soft actor–critic algorithm can produce feasible tracking trajectories without limiting the state space and action space, provided that the environment is properly posed.

Engineering↗

Enhancing Cyber Resilience of Networked Microgrids using Vertical Federated Reinforcement Learning

This paper presents a novel federated reinforcement learning (Fed-RL) methodology to inject sufficient resiliency into the operations of the network of microgrids. We consider adversarial actions to the voltage and power control loop reference signals at the grid forming (GFM) inverters in the microgrids which are essential to integrate renewable resources. Therefore, we formulate a resilient reinforcement learning training setup that uses these adversarial injections to generate episodic trajectories and train the RL agents to alleviate their impact on performance. To circumvent the concerns about data-sharing and privacy for different owners of the microgrids in the networked setting, we bring in the aspects of the federated operation to propose novel Fed-RL algorithms. As the dynamics of each microgrid are coupled due to electrical interlinks, the conventional federated RL approaches using decoupled independent environments are not applicable, which leads us to propose a multi-agent vertically federated variation of actor-critic algorithms, namely federated soft actor-critic (FedSAC). We have performed numerical simulations on an IEEE 123-bus benchmark test feeder with three microgrids by creating a customized simulation setup by encapsulating the microgrid dynamic simulations in GridLAB-D/HELICS co-simulation platform with the OpenAI Gym environment and validated the proposed resilient and secured learning methodology.

Artificial Intelligence (AI), reinforcement learni↗

Off-policy deep reinforcement learning with automatic entropy adjustment for adaptive online grid emergency control

Electric overloading conditions and contingencies put modern power systems at risk of voltage collapse and blackouts. Load shedding is crucial to maintain voltage stability for grid emergency control. However, the rule- or model-based schemes rely on accurate dynamic system models and face considerable challenges in adapting to various operating conditions and uncertain event occurrences. Here, to address these issues, this paper proposes a novel deep reinforcement learning (DRL)-based voltage stability control algorithm with automatic entropy adjustment (AEA) for grid emergency control. Various dynamic network components for complex system operations are modeled to construct the DRL environment. An off-policy soft actor-critic architecture is developed to maximize the expected reward and policy entropy simultaneously. The AEA mechanism is proposed to facilitate the policy maximum entropy procedure, and the proposed method can automatically provide effective discrete and continuous actions against various fault scenarios. Our approach accomplishes high sampling efficiency, scalability, and auto-adaptivity of the control policies under high uncertainties. Comparative studies with the existing DRL-based control methods in IEEE benchmarks indicate salient performance improvement of the proposed method for dynamic system emergency control.

24 POWER TRANSMISSION AND DISTRIBUTION↗