Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “actor critic”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Variational actor-critic algorithms,

We introduce a class of variational actor-critic algorithms based on a variational formulation over both the value function and the policy. The objective function of the variational formulation consists of two parts: one for maximizing the value function and the other for minimizing the Bellman residual. Besides the vanilla gradient descent with both the value function and the policy updates, we propose two variants, the clipping method and the flipping method, in order to speed up the convergence. We also prove that, when the prefactor of the Bellman residual is sufficiently large, the fixed point of the algorithm is close to the optimal policy.

97 MATHEMATICS AND COMPUTING↗

Power System Frequency Dynamics Modeling, State Estimation, and Control using Neural Ordinary Differential Equations (NODEs) and Soft Actor-Critic (SAC) Machine Learning Approaches

With the global energy transition of the electric power system, grid control, supervision, and protection is becoming more challenging. With the increasing integration of renewable energy sources (RES), the system dynamics are changing, causing traditional power system dynamic modeling with swing equation-based modeling approaches to fail. Additionally, the converter-dominated power grid is decreasing the system inertia, making the power system more fragile to the frequency swings. This paper first investigates and compares the application of a model-based Kalman filter state estimation approach with (i) a model-free machine learning approach --- neural ordinary differential equations (NODEs) --- and (ii) a data-driven system identification (SysId) approach to model and infer critical state values of the power system frequency dynamics. Then a model predictive control (MPC) framework is compared to a model-free Soft Actor-Critic (SAC) reinforcement learning (RL) control algorithm in providing efficient fast frequency response (FFR) to the power system frequency dynamics. The approaches are compared in terms of their performance goals as well as their per-timestep computational efficiency. Furthermore, the comparative study for state estimation shows that for the model-free requirement, both NODEs and SysId can provide accurate state estimates; however, with increasing model complexity, NODEs can be a better choice for model identification. Similarly, the results from the FFR comparative study show that the SAC RL-based FFR, once trained, outperforms MPC with better control signals and faster computation time, making the SAC RL-based FFR better option for providing FFR to the power system.

97 MATHEMATICS AND COMPUTING↗

Soft Actor Critic Based Volt-VAR Co-optimization in Active Distribution Grids

Modern distribution networks are undergoing several technical challenges, such as voltage fluctuations, because of high penetration of distributed energy resources (DERs). This paper proposes a deep reinforcement learning (DRL)-based Volt VAR co-optimization technique for reducing voltage fluctuations as well as power loss under high penetration of DERs. In addition, the proposed approach minimizes the operational cost of the grid. A stochastic policy optimization based soft actor critic (SAC) agent is proposed to configure the optimal set-points of the reactive power outputs of the inverters. The performance of the proposed model is verified on the modified IEEE 34- and 123-bus systems and compared with a base case scenario with no reactive supply by inverters, and a local droop control approach. The results demonstrate that the proposed framework outperforms the conventional droop control method in improving the voltage profile, minimizing the network power loss, and reducing grid operational cost.

—Distribution grids, deep reinforcement learning, ↗

MARS: Malleable Actor-Critic Reinforcement Learning Scheduler

In this paper, we introduce MARS, a new scheduling system for HPC-cloud infrastructures based on a cost-aware, flexible reinforcement learning approach, which serves as an intermediate layer for next generation HPC-cloud resource manager. MARS ensembles the pre-trained models from heuristic workloads and decides on the most cost-effective strategy for optimization. A whole workflow application would be split into several optimizable dependent sub-tasks, then based on the pre- defined resource management plan, a reward will be generated after executing a scheduled task. Lastly, MARS updates the Deep Neural Network (DNN) model based on the reward. MARS is designed to optimize the existing models through reinforcement mechanisms. MARS adapts to the dynamics of workflow applications, selects the most cost-effective scheduling solution among pre-built scheduling strategies (backfilling, SJF, etc.) and self- learning deep neural network model at run-time. We evaluate MARS with different real-world workflow traces. MARS can achieve 5%-60% increased performance compare to state-of-the- art approaches.

Baheri, Betis↗

Demonstration of reconstruction-free static magnetic control of DIII-D plasma with deep reinforcement learning

This paper presents the development and experimental validation of a reinforcement learning (RL)-based magnetic controller on the DIII-D tokamak. The controller directly maps raw magnetic diagnostic signals to actuator commands, replacing the traditional isoflux control algorithm based on equilibrium reconstruction. Four RL controllers are trained using the Soft Actor–Critic algorithm with an asymmetric Actor–Critic architecture in the NSFsim simulator. All controllers are deployed in the DIII-D Plasma Control System and operated with a 4 kHz feedback loop. Two randomization strategies are evaluated during training: evolving kinetic profiles and fixed kinetic profiles within each episode. The latter approach is found to better capture experimental deviations in the current density profile and to provide overall improved control performance. Robust operation is demonstrated across heating power scans in both L- and H-mode plasmas, as well as during transient events such as L–H transitions and pellet injections. Control errors in plasma shape and radial position remained within 1.5–2.0 cm and 1 cm, respectively. A notable discrepancy was observed in the vertical X-point position, with errors of up to approximately 4 cm, attributed to the current density distribution mismatches between simulations and experiments.

DIII-D↗

Enhancing Cyber Resilience of Networked Microgrids using Vertical Federated Reinforcement Learning

This paper presents a novel federated reinforcement learning (Fed-RL) methodology to inject sufficient resiliency into the operations of the network of microgrids. We consider adversarial actions to the voltage and power control loop reference signals at the grid forming (GFM) inverters in the microgrids which are essential to integrate renewable resources. Therefore, we formulate a resilient reinforcement learning training setup that uses these adversarial injections to generate episodic trajectories and train the RL agents to alleviate their impact on performance. To circumvent the concerns about data-sharing and privacy for different owners of the microgrids in the networked setting, we bring in the aspects of the federated operation to propose novel Fed-RL algorithms. As the dynamics of each microgrid are coupled due to electrical interlinks, the conventional federated RL approaches using decoupled independent environments are not applicable, which leads us to propose a multi-agent vertically federated variation of actor-critic algorithms, namely federated soft actor-critic (FedSAC). We have performed numerical simulations on an IEEE 123-bus benchmark test feeder with three microgrids by creating a customized simulation setup by encapsulating the microgrid dynamic simulations in GridLAB-D/HELICS co-simulation platform with the OpenAI Gym environment and validated the proposed resilient and secured learning methodology.

Artificial Intelligence (AI), reinforcement learni↗

Dynamic, resilient virtual sensing system and shadow controller for cyber-attack neutralization

An industrial asset may have monitoring nodes (e.g., sensor or actuator nodes) that generate current monitoring node values. An abnormality detection and localization computer may receive the series of current monitoring node values and output an indication of at least one abnormal monitoring node that is currently being attacked or experiencing a fault. An actor-critic platform may tune a dynamic, resilient state estimator for a sensor node and output tuning parameters for a controller that improve operation of the industrial asset during the current attack or fault. The actor-critic platform may include, for example, a dynamic, resilient state estimator, an actor model, and a critic model. According to some embodiments, a value function of the critic model is updated for each action of the actor model and each action of the actor model is evaluated by the critic model to update a policy of the actor-critic platform.

Roychowdhury, Subhrajit↗

Extreme Risk Mitigation in Reinforcement Learning using Extreme Value Theory

Risk-sensitive reinforcement learning (RL) has garnered significant attention in recent years due to the growing interest in deploying RL agents in real-world scenarios. A critical aspect of risk awareness involves modelling highly rare risk events (rewards) that could potentially lead to catastrophic outcomes. These infrequent occurrences present a formidable challenge for data-driven methods aiming to capture such risky events accurately. While risk-aware RL techniques do exist, they suffer from high variance estimation due to the inherent data scarcity. Our work proposes to enhance the resilience of RL agents when faced with very rare and risky events by focusing on refining the predictions of the extreme values predicted by the state-action value distribution. To achieve this, we formulate the extreme values of the state-action value function distribution as parameterized distributions, drawing inspiration from the principles of extreme value theory (EVT). We propose an extreme value theory based actor-critic approach, namely, Extreme Valued Actor-Critic (EVAC) which effectively addresses the issue of infrequent occurrence by leveraging EVT-based parameterization. Importantly, we theoretically demonstrate the advantages of employing these parameterized distributions in contrast to other risk-averse algorithms. Our evaluations show that the proposed method outperforms other risk averse RL algorithms on a diverse range of benchmark tasks, each encompassing distinct risk scenarios.

Wang, Yu↗

Multi-agent voltage control in distribution systems using GAN-DRL-based approach

Active distribution grids can experience voltage fluctuations and violations due to the high penetration of variable distributed energy resources (DERs). These problems might occur because of the uncertain and variable generation natures of these resources, especially solar photovoltaic resources, during panel shadowing scenarios. Volt-VAR control (VVC) is an efficient method that controls the reactive power set-points of the inverters to regulate the voltage of distribution grids. Although several VVC approaches have been proposed recently, the performance of these approaches degrades significantly if behind-the-meter solar generation data are unobservable/missing. Therefore, it is necessary to impute missing/unobservable PV data accurately to be utilized in VVC approaches. Further, this paper proposes a model-free, data-driven, centrally trained, and decentrally executed multi-agent deep reinforcement learning-based VVC architecture to regulate the voltage of distribution networks. A generative adversarial network (GAN) is incorporated to impute the unobservable PV data accurately, which improves the performance of the proposed control architecture. The proposed multi-agent-soft-actor–critic algorithm (MASAC)-based VVC technique utilizes the actual PV dataset as well as the imputed dataset from the GAN framework to learn the optimal coordinated control policy for controlling the optimal reactive power set-points of PV inverters. The effectiveness of the proposed approach is analyzed on a modified IEEE 34-bus test case with added PV inverters. The results are compared and analyzed with a base case model with no VVC and VVC with a local droop control approach, genetic algorithm optimization, and a centralized soft actor–critic-based approach. Moreover, the performance of the proposed approach is compared with that of a multi-agent VVC framework without using the PV generation data and load information as the system state. The results illustrate that the proposed method with more state input improves the voltage profile and reduces the power loss of the network across various loading and PV generation scenarios.

14 SOLAR ENERGY↗

Off-policy deep reinforcement learning with automatic entropy adjustment for adaptive online grid emergency control

Electric overloading conditions and contingencies put modern power systems at risk of voltage collapse and blackouts. Load shedding is crucial to maintain voltage stability for grid emergency control. However, the rule- or model-based schemes rely on accurate dynamic system models and face considerable challenges in adapting to various operating conditions and uncertain event occurrences. Here, to address these issues, this paper proposes a novel deep reinforcement learning (DRL)-based voltage stability control algorithm with automatic entropy adjustment (AEA) for grid emergency control. Various dynamic network components for complex system operations are modeled to construct the DRL environment. An off-policy soft actor-critic architecture is developed to maximize the expected reward and policy entropy simultaneously. The AEA mechanism is proposed to facilitate the policy maximum entropy procedure, and the proposed method can automatically provide effective discrete and continuous actions against various fault scenarios. Our approach accomplishes high sampling efficiency, scalability, and auto-adaptivity of the control policies under high uncertainties. Comparative studies with the existing DRL-based control methods in IEEE benchmarks indicate salient performance improvement of the proposed method for dynamic system emergency control.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Neural network approaches for parameterized optimal control

Here, we consider numerical approaches for deterministic, finite-dimensional optimal control problems whose dynamics depend on unknown or uncertain parameters. We seek to amortize the solution over a set of relevant parameters in an offline stage to enable rapid decision-making and be able to react to changes in the parameter in the online stage. To tackle the curse of dimensionality arising when the state and/or parameter are high-dimensional, we represent the policy using neural networks. We compare two training paradigms: First, our model-based approach leverages the dynamics and definition of the objective function to learn the value function of the parameterized optimal control problem and obtain the policy using a feedback form. Second, we use actor-critic reinforcement learning to approximate the policy in a data-driven way. Using an example involving a two-dimensional convection-diffusion equation, which features high-dimensional state and parameter spaces, we investigate the accuracy and efficiency of both training paradigms. While both paradigms lead to a reasonable approximation of the policy, the model-based approach is more accurate and considerably reduces the number of PDE solves.

97 MATHEMATICS AND COMPUTING↗

Network Reconfiguration for Enhanced Operational Resilience Using Reinforcement Learning

This paper proposes a reinforcement learning-based approach for distribution network reconfiguration(DNR) to enhance the resilience of the electric power supply. Resilience enhancements usually require solving large-scale stochastic optimization problems that are computationally expensive and sometimes infeasible. The exceptional performance of reinforcement learning techniques has encouraged their adoption in various power system control studies, specifically resilience-based real-time applications. In this paper, a single agent framework is developed using an Actor-Critic algorithm (ACA) to determine statuses of tie-switches in a distribution feeder impacted by an extreme weather event. The proposed approach provides a fast-acting control algorithm that reconfigures the feeder topology to reduce or even avoid load shedding. The problem is formulated as a discrete Markov decision process in such a way that a system state captures the system topology and its operational characteristics. An action is made to open or close a specific set of tie-switches after which a reward is calculated to evaluate the practicality and advantage of that action. The iterative Markov process is used to train the proposed ACA under diverse failure scenarios and is demonstrated on the 33-node distribution feeder system. Results show the capability of the proposed ACA to determine proper switching action of tie-switches with accuracy exceeding 93%.

actor critic↗

Deep Reinforcement Learning for Distribution System Restoration Using Distributed Energy Resources and Tie-Switches

Distributed energy resources (DERs), such as solar PVs and energy storage, can be used to restore distribution system critical loads after the extreme weather events to increase grid resilience. However, coordinating multiple DERs together with tie-switches for multi-step restoration process under renewable uncertainty is challenging. This paper proposes a deep reinforcement learning to control discrete actions of switching on/off tie switches and DERs for critical load restoration. The restoration problem is first cast into the Markov decision process suitable for DRL. Then, the original soft actor critic (SAC) method for continuous actions has been extended to handle discrete and continuous actions. Numerical comparison results with other stochastic optimization-based approaches on the modified IEEE 33-bus system show that the proposed method can achieve fast critical load restoration in the presence of substation power outage while maintaining system voltage limit throughout the restoration process.

active distribution systems↗

Optimization of the FRIB beam dump: a hybrid genetic algorithm and reinforcement learning approach

The operational envelope of high-power-density systems, such as particle accelerators and advanced nuclear energy systems, is critically constrained by the need to manage extreme thermal loads. To address this, we present a novel hybrid optimization framework combining a genetic algorithm (GA) with a soft actor-critic (SAC) deep reinforcement learning agent. This framework was applied to a practical high-heat-flux problem: redesigning the beam dump at the Facility for Rare Isotope Beams (FRIB) for a power upgrade from 20 kW to 50 kW. The resulting design, validated by three-dimensional conjugate heat transfer simulations, suppresses hazardous hot spots and yields a markedly more uniform temperature distribution. This provides a robust operating margin, increasing the average power-handling capability by 72% relative to the current design, demonstrating the framework’s potential to solve complex thermal management challenges in both accelerator technology and advanced nuclear systems.

Accelerator↗