Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

PowerNet: Multi-agent Deep Reinforcement Learning for Scalable Powergrid Control

This paper develops an efficient multi-agent deep reinforcement learning algorithm for cooperative controls in powergrids. Specifically, we consider the decentralized inverter-based secondary voltage control problem in distributed generators (DGs), which is first formulated as a cooperative multi-agent reinforcement learning (MARL) problem. We then propose a novel on-policy MARL algorithm, PowerNet, in which each agent (DG) learns a control policy based on (sub-)global reward but local states and encoded communication messages from its neighbors. Motivated by the fact that a local control from one agent has limited impact on agents distant from it, we exploit a novel spatial discount factor to reduce the effect from remote agents, to expedite the training process and improve scalability. Furthermore, a differentiable, learning-based communication protocol is employed to foster the collaborations among neighboring agents. In addition, to mitigate the effects of system uncertainty and random noise introduced during on-policy learning, we utilize an action smoothing factor to stabilize the policy execution. To facilitate training and evaluation, we develop PGSim, an efficient, high-fidelity powergrid simulation platform. Here, experimental results in two microgrid setups show that the developed PowerNet outperforms the conventional model-based control method, as well as several state-of-the-art MARL algorithms. The decentralized learning scheme and high sample efficiency also make it viable to large-scale power grids.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Explainable and Differentiable Reinforcement Learning for Multi-objective Optimization in Particle Accelerators

Operating particle accelerators involves optimizing multiple goals simultaneously, which can be challenging due to trade-offs among objectives. While evolutionary algorithms like the genetic algorithm (GA) have been used for various Multi-Objective Optimization (MOO) tasks, they are not inherently suited for complex control problems. This talk highlights two variations of Reinforcement Learning (RL) for concurrently optimizing heat load and trip rates at the Continuous Electron Beam Accelerator Facility (CEBAF). The problem involves strict constraints on individual states, actions, and overall energy requirements of the beam. First, this talk highlights how differentiability can be harnessed through a Deep Differentiable Reinforcement Learning (DDRL) approach to address MOO issues within particle accelerators. We examine the DDRL method alongside Model Free Reinforcement Learning (MFRL), GA, and Bayesian Optimization (BO). The performance of these methods is assessed by generating a Pareto-front for two objectives. Our findings indicate that DDRL excels in handling high-dimensional problems more effectively than MFRL, BO, and GA. Next, we will show integration of explainable physics-based constraints into RL algorithms to enhance trans- parency and trust in decision-making processes by enabling users to verify that agents adhere to established physical principles. This surrogate function can be modeled using neural networks or sparse dictionary mod- els. By examining the mathematical form of the learned constraint function, we are able to confirm the agent has learned to use the established physics of each environment provided but the surrogate model. In addi- tion, we find that the introduction of a mathematical functional dictionary based surrogate model enables our reinforcement learning algorithms to reliably converge for difficult high-dimensional accelerator controls environments.

Rajput, Kishansingh [Thomas Jefferson National Acc↗

Variational quantum reinforcement learning via evolutionary optimization

Abstract Recent advances in classical reinforcement learning (RL) and quantum computation point to a promising direction for performing RL on a quantum computer. However, potential applications in quantum RL are limited by the number of qubits available in modern quantum devices. Here, we present two frameworks for deep quantum RL tasks using gradient-free evolutionary optimization. First, we apply the amplitude encoding scheme to the Cart-Pole problem, where we demonstrate the quantum advantage of parameter saving using amplitude encoding. Second, we propose a hybrid framework where the quantum RL agents are equipped with a hybrid tensor network-variational quantum circuit (TN-VQC) architecture to handle inputs of dimensions exceeding the number of qubits. This allows us to perform quantum RL in the MiniGrid environment with 147-dimensional inputs. The hybrid TN-VQC architecture provides a natural way to perform efficient compression of the input dimension, enabling further quantum RL applications on noisy intermediate-scale quantum devices.

97 MATHEMATICS AND COMPUTING↗

Reinforcement learning applied to dilute combustion control for increased fuel efficiency

To reduce the modeling burden for control of spark-ignition engines, reinforcement learning (RL) has been applied to solve the dilute combustion limit problem. Q-learning was used to identify an optimal control policy to adjust the fuel injection quantity in each combustion cycle. A physics-based model was used to determine the relevant states of the system used for training the control policy in a data-efficient manner. The cost function was chosen such that high cycle-to-cycle variability (CCV) at the dilute limit was minimized while maintaining stoichiometric combustion as much as possible. Experimental results demonstrated a reduction of CCV after the training period with slightly lean combustion, contributing to a net increase in fuel conversion efficiency of 1.33%. To ensure stoichiometric combustion for three-way catalyst compatibility, a second feedback loop based on an exhaust oxygen sensor was incorporated into the fuel quantity controller using a slow proportional-integral (PI) controller. The closed-loop experiments showed that both feedback loops can cooperate effectively, maintaining stoichiometric combustion while reducing combustion CCV and increasing fuel conversion efficiency by 1.09%. Finally, a modified cost function was proposed to ensure stoichiometric combustion with a single controller. In addition, the learning period was shortened by half to evaluate the RL algorithm performance on limited training time. Experimental results showed that the modified cost function could achieve the desired CCV targets, however, the learning time was reduced by half and the fuel conversion efficiency increased only by 0.30%.

33 ADVANCED PROPULSION SYSTEMS↗

Solving the $k$-Sparse Eigenvalue Problem with Reinforcement Learning

We examine the possibility of using a reinforcement learning (RL) algorithm to solve large-scale eigenvalue problems in which the desired the eigenvector can be approximated by a sparse vector with at most k nonzero elements, where k is relatively small compare to the dimension of the matrix to be partially diagonalized. Here, this type of problem arises in applications in which the desired eigenvector exhibits localization properties and in large-scale eigenvalue computations in which the amount of computational resource is limited. When the positions of these nonzero elements can be determined, we can obtain the k-sparse approximation to the original problem by computing eigenvalues of a k × k submatrix extracted from k rows and columns of the original matrix. We review a previously developed greedy algorithm for incrementally probing the positions of the nonzero elements in a k-sparse approximate eigenvector and show that the greedy algorithm can be improved by using an RL method to refine the selection of k rows and columns of the original matrix. We describe how to represent states, actions, rewards and policies in an RL algorithm designed to solve the k-sparse eigenvalue problem and demonstrate the effectiveness of the RL algorithm on two examples originating from quantum many-body physics.

97 MATHEMATICS AND COMPUTING↗

The Effect of Antagonistic Behavior in Reinforcement Learning

The significant achievements of deep reinforcement learning (RL) have motivated researchers to also investigate its shortcomings. Such work has shown that typical methods in deep RL tend to produce brittle policies that overfit to the training environment. In this paper, we introduce the notion of purely antagonistic behavior in value-based agents, where the objective is not to maximize reward but to minimize the victim’s value over time. This notion is motivated by the scenario in which an antagonistic human architect, without access to the environment’s reward function, wants to build an RL agent that can impede another well-trained RL victim agent. First, we formalize a notion of antagonistic behavior in RL. Then, we provide experiments that show how a purely antagonistic agent performs compared to a well-trained victim that learns directly from the game’s rewards. Our results suggest that if one’s goal is to find vulnerabilities in well-trained agents, direct access to the environment’s rewards is not necessary, and antagonistic behavior can be measured independently from environment wins and losses.

Fujimoto, Ted C.↗

Proof-of-concept of a reinforcement learning framework for wind farm energy capture maximization in time-varying wind

Here, we present a proof-of-concept distributed reinforcement learning framework for wind farm energy capture maximization. The algorithm we propose uses Q-Learning in a wake-delayed wind farm environment and considers time-varying, though not yet fully turbulent, wind inflow conditions. These algorithm modifications are used to create the Gradient Approximation with Reinforcement Learning and Incremental Comparison (GARLIC) framework for optimizing wind farm energy capture in time-varying conditions, which is then compared to the FLOw Redirection and Induction in Steady State (FLORIS) static lookup table wind farm controller baseline.

17 WIND ENERGY↗

Dyn$\mathrm{AMO}$: Multi-agent reinforcement learning for dynamic anticipatory mesh optimization with applications to hyperbolic conservation laws

Here we introduce DynAMO, a reinforcement learning paradigm for Dynamic Anticipatory Mesh Optimization. Adaptive mesh refinement is an effective tool for optimizing computational cost and solution accuracy in numerical methods for partial differential equations. However, traditional adaptive mesh refinement approaches for time-dependent problems typically rely only on instantaneous error indicators to guide adaptivity. As a result, standard strategies often require frequent remeshing to maintain accuracy. In the DynAMO approach, multi-agent reinforcement learning is used to discover new local refinement policies that can anticipate and respond to future solution states by producing meshes that deliver more accurate solutions for longer time intervals. By applying DynAMO to discontinuous Galerkin methods for the linear advection and compressible Euler equations in two dimensions, we demonstrate that this new mesh refinement paradigm can outperform conventional threshold-based strategies while also generalizing to different mesh sizes, remeshing and simulation times, and initial conditions.

97 MATHEMATICS AND COMPUTING↗

Deep reinforcement learning to discover multi-fuel injection strategies for compression ignition engines

Over the past several decades, regulation of compression ignition engine emissions has become increasingly stringent as concern about the environmental and health implications of these emissions has grown. These changing constraints have led to a series of new, alternative fuel injection strategies that aim to maintain power output while reducing in-cylinder generated emissions by operating in the low-temperature combustion (LTC) regime. These advanced injection strategies are created and retuned for individual combinations of engine geometry, fuel, and emissions constraints. Deep reinforcement learning has been shown to be an effective alternative to traditional optimization approaches for highly combinatorial control problems, such as discovering the optimal injection schedules for compression ignition engines. In this study, we deploy a previously presented deep reinforcement learning framework to iteratively optimize a series of engine geometries over a range of increasingly strict NO x emissions regulations. We then examine the resulting injection schedules. We discuss the potential for using this deep reinforcement learning framework for fuel selection screening and for discovering unique injection strategies for different engine geometries and future emissions standards.

33 ADVANCED PROPULSION SYSTEMS↗

Deep reinforcement learning with online data augmentation to improve sample efficiency for intelligent HVAC control

Deep Reinforcement Learning (DRL) has started showing success in real-world applications such as building energy optimization. Much of the research in this space utilized simulated environments to train RL-agent in an offline mode. Very few research have used DRL-based control in real-world systems due to two main reasons: 1) sample efficiency challenge---DRL approaches need to perform a lot of interactions with the environment to collect sufficient experiences to learn from, which is difficult in real systems, and 2) comfort or safety related constraints---user's comfort must never or at least rarely be violated. In this work, we propose a novel deep Reinforcement Learning framework with online Data Augmentation (RLDA) to address the sample efficiency challenge of real-world RL. We used a time series Generative Adversarial Network (TimeGAN) architecture as a data generator. We further evaluated the proposed RLDA framework using a case study of an intelligent HVAC control. With a ≈28% improvement in the sample efficiency, RLDA framework lays the way towards increased adoption of DRL-based intelligent control in real-world building energy management systems.

Kurte, Kuldeep↗

A Reinforcement Learning Approach to Augment Conventional PID Control in Nuclear Power Plant Transient Operation

The ability of nuclear reactors to operate their power conversion cycles more flexibly will enhance their value to energy grids with variable pricing. Current nuclear control systems are typically classical controllers that are often based on proportional-integral-derivative (PID) control. This paper presents a method of augmenting the existing PID control for difficult transient operations in nuclear power plants using a reinforcement learning–derived feedforward signal applied in real time. The agents, which are trained on a test thermal load-following problem, are designed to improve steam generator outlet temperature control for a range of fast load-following scenarios covering ramp rates from 9%/min to 15%/min. Several reinforcement learning algorithms were initially investigated for the training of the feedforward agents with deep Q-learning (DQN) and proximal policy optimization (PPO) networks, which were found to be the most promising. The DQN controllers utilize discrete actions, giving them a better disturbance rejection at steady state but inconsistent response to initial temperature deviations. In contrast, PPO-trained agents, which take continuous actions except for a dead zone around zero, were shown to have the best combination of high disturbance rejection at steady state and good tracking of the desired temperature value. The ability of the PPO agent was also examined, with the average time of decision making found to be on the order of 1 ms. The fault properties of the controller under the loss of the reinforcement learning agent feedforward signal were also examined. The controller showed strong performance in situations of “no-signal” faults. but was less good at handling “stuck-at” faults, where the feedforward signal remains at a set value. In both cases, however, the PID was able to successfully maintain stability, eventually returning the system to a steady state. It is hoped that this work will allow for the proposed control architecture to be examined for more difficult control problems such that it may eventually be used to adapt existing nuclear plants for more aggressive load-following on grids of the future.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Reinforcement Learning Configuration Interaction

Selected configuration interaction (sCI) methods exploit the sparsity of the full configuration interaction (FCI) wave function, yielding significant computational savings and wave function compression without sacrificing the accuracy. Despite recent advances in sCI methods, the selection of important determinants remains an open problem. Furthermore, we explore the possibility of utilizing reinforcement learning approaches to solve the sCI problem. By mapping the configuration interaction problem onto a sequential decision-making process, the agent learns on-the-fly which determinants to include and which to ignore, yielding a compressed wave function at near-FCI accuracy. This method, which we call reinforcement-learned configuration interaction, adds another weapon to the sCI arsenal and highlights how reinforcement learning approaches can potentially help solve challenging problems in electronic structure theory.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data-driven Optimal Control Strategy for Virtual Synchronous Generator via Deep Reinforcement Learning Approach

This paper aims at developing a data-driven optimal control strategy for virtual synchronous generator (VSG) in the scenario where no expert knowledge or requirement for system model is available. Firstly, the optimal and adaptive control problem for VSG is transformed into a reinforcement learning task. Specifically, the control variables, i.e., virtual inertia and damping factor, are defined as the actions. Meanwhile, the active power output, angular frequency and its derivative are considered as the observations. Moreover, the reward mechanism is designed based on three preset characteristic functions to quantify the control targets: (1) maintaining the deviation of angular frequency within special limits; (2) preserving well-damped oscillations for both the angular frequency and active power output; (3) obtaining slow frequency drop in the transient process. Next, to maximize the cumulative rewards, a decentralized deep policy gradient algorithm, which features model-free and faster convergence, is developed and employed to find the optimal control policy. With this effort, a data-driven adaptive VSG controller can be obtained. By using the proposed controller, the inverter-based distributed generator can adaptively adjust its control variables based on current observations to fulfill the expected targets in model-free fashion. Finally, simulation results validate the feasibility and effectiveness of the proposed approach.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Gaming the beamlines—employing reinforcement learning to maximize scientific outcomes at large-scale user facilities

Abstract Beamline experiments at central facilities are increasingly demanding of remote, high-throughput, and adaptive operation conditions. To accommodate such needs, new approaches must be developed that enable on-the-fly decision making for data intensive challenges. Reinforcement learning (RL) is a domain of AI that holds the potential to enable autonomous operations in a feedback loop between beamline experiments and trained agents. Here, we outline the advanced data acquisition and control software of the Bluesky suite, and demonstrate its functionality with a canonical RL problem: cartpole. We then extend these methods to efficient use of beamline resources by using RL to develop an optimal measurement strategy for samples with different scattering characteristics. The RL agents converge on the empirically optimal policy when under-constrained with time. When resource limited, the agents outperform a naive or sequential measurement strategy, often by a factor of 100%. We interface these methods directly with the data storage and provenance technologies at the National Synchrotron Light Source II, thus demonstrating the potential for RL to increase the scientific output of beamlines, and layout the framework for how to achieve this impact.

36 MATERIALS SCIENCE↗

A Modified Maximum Entropy Inverse Reinforcement Learning Approach for Microgrid Energy Scheduling

Increasing popularity of integrating distributed energy resources (DERs) into the power system brings a challenge to optimize the microgrid dispatch policy. The reinforcement learning methods suffer from a long-time problem with the theoretical assumption of the objective/reward function for the microgrid system. Although the traditional inverse reinforcement learning (IRL) approaches can solve this problem to some extent, they encounter a limitation of complex computations for state visitation frequency in the large and continuous state space. To alleviate this limitation, we propose a modified maximum entropy IRL (MMIRL) method to extract the reward function from the expert demonstrations for solving the microgrid energy scheduling problem. The proposed MMIRL algorithm is promising in recovering the reward function and learning the dispatch policy compared to conventional approaches. Case studies are performed in an energy arbitrage problem and a microgrid system with DERs. Results substantiate that the proposed MMIRL approach can learn the dispatch policy with more than 99% efficiency and outperforms other comparative methods.

artificial intelligence, reinforcement learning, m↗