Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Distributed Power Allocation for 6-GHz Unlicensed Spectrum Sharing via Multi-agent Deep Reinforcement Learning

We consider the problem of power allocation over the 6 GHz Unlicensed National Information Infrastructure (UNII)- 5 spectrum. We propose a novel deep Reinforcement Learning (DRL)-based distributed power allocation scheme which utilizes the multi-agent Deep Deterministic Policy Gradient (MADDPG) algorithm. In particular, we model the base stations (BSs) as DRL agents that simultaneously determine the transmit powers to their scheduled user equipment (UE) in a synchronized manner. The power decision of each BS is based on its own observation of the radio environment, which consists of several local interference measurements and a limited amount of information obtained from other BSs. One advantage of the proposed scheme is that it addresses the single-agent non-stationarity problem of RL in the multi-agent scenario by incorporating the actions and observations of other BSs into each BS’s own critic which helps it to gain a more accurate perception of the overall radio environment. A centralized-training-distributed execution framework is used to train the policies where the critics are trained over the joint actions and observations of all BSs while the actor of each BS only takes the local observation as input in order to produce the transmit power. Simulation shows that the proposed power allocation scheme can achieve better throughput performance than several state-of-the-art approaches.

99 GENERAL AND MISCELLANEOUS↗

Optimal carbon storage reservoir management through deep reinforcement learning

Model-based optimization plays a central role in energy system design and management. The complexity and high-dimensionality of many process-level models, especially those used for geosystem energy exploration and utilization, often lead to formidable computational costs when the dimension of decision space is also large. This work adopts elements of recently advanced deep learning techniques to solve a sequential decision-making problem in applied geosystem management. Specifically, a deep reinforcement learning framework was formed for optimal multiperiod planning, in which a deep Q-learning network (DQN) agent was trained to maximize rewards by learning from high-dimensional inputs and from exploitation of its past experiences. To expedite computation, deep multitask learning was used to approximate high-dimensional, multistate transition functions. Both DQN and deep multitask learning are pattern based. As a demonstration, the framework was applied to optimal carbon sequestration reservoir planning using two different types of management strategies: monitoring only and brine extraction. Both strategies are designed to mitigate potential risks due to pressure buildup. Results show that the DQN agent can identify the optimal policies to maximize the reward for given risk and cost constraints. Finally, experiments also show that knowledge the agent gained from interacting with one environment is largely preserved when deploying the same agent in other similar environments.

15 GEOTHERMAL ENERGY↗

A Reinforcement Learning Approach to Parameter Selection for Distributed Optimal Power Flow

With the increasing penetration of distributed energy resources, distributed optimization algorithms have attracted significant attention for power systems applications due to their potential for superior scalability, privacy, and robustness to a single point-of-failure. The Alternating Direction Method of Multipliers (ADMM) is a popular distributed optimization algorithm; however, its convergence performance is highly dependent on the selection of penalty parameters, which are usually chosen heuristically. In this work, we use reinforcement learning (RL) to develop an adaptive penalty parameter selection policy for alternating current optimal power flow (ACOPF) problem solved via ADMM with the goal of minimizing the number of iterations until convergence. We train our RL policy using deep Q-learning and show that this policy can result in significantly accelerated convergence (up to a 59% reduction in the number of iterations compared to existing, curvatureinformed penalty parameter selection methods). Furthermore, we show that our RL policy demonstrates promise for generalizability, performing well under unseen loading schemes as well as under unseen losses of lines and generators (up to a 50% reduction in iterations). This work thus provides a proof-of-concept for using RL for parameter selection in ADMM for power systems applications.

alternating current optimal power flow↗

Grid-Interactive Multi-Zone Building Control Using Reinforcement Learning with Global-Local Policy Search

In this paper, we develop a grid-interactive multi-zone building controller based on a deep reinforcement learning (RL) approach. The controller is designed to facilitate building operation during normal conditions and demand response events, while ensuring occupants comfort and energy efficiency. We leverage a continuous action space RL formulation, and devise a two-stage global-local RL training framework. In the first stage, a global fast policy search is performed using a gradient-free RL algorithm. In the second stage, a local fine-tuning is conducted using a policy gradient method. In contrast to the state-of-the-art model predictive control (MPC) approach, the proposed RL controller does not require complex computation during real-time operation and can adapt to nonlinear building models. We illustrate the controller performance numerically using a five-zone commercial building.

30 DIRECT ENERGY CONVERSION↗

Grid-Interactive Multi-Zone Building Control Using Reinforcement Learning with Global-Local Policy Search: Preprint

In this paper, we develop a grid-interactive multi-zone building controller based on a deep reinforcement learning (RL) approach. The controller is designed to facilitate building operation during normal conditions and demand response events, while ensuring occupants comfort and energy efficiency. We leverage a continuous action space RL formulation, and devise a two-stage global-local RL training framework. In the first stage, a global fast policy search is performed using a gradient-free RL algorithm. In the second stage, a local fine-tuning is conducted using a policy gradient method. In contrast to the state-of-the-art model predictive control (MPC) approach, the proposed RL controller does not require complex computation during real-time operation and can adapt to non-linear building models. We illustrate the controller performance numerically using a five-zone commercial building.

30 DIRECT ENERGY CONVERSION↗

Challenges and Opportunities in Deep Reinforcement Learning With Graph Neural Networks: A Comprehensive Review of Algorithms and Applications

Deep reinforcement learning (DRL) has empowered a variety of artificial intelligence fields, including pattern recognition, robotics, recommendation-systems, and gaming. Similarly, graph neural networks (GNN) have also demonstrated their superior performance in supervised learning for graph-structured data. In recent times, the fusion of GNN with DRL for graph-structured environments has attracted a lot of attention. Here, this paper provides a comprehensive review of these hybrid works. These works can be classified into two categories: (1) algorithmic enhancement, where DRL and GNN complement each other for better utility; (2) application-specific enhancement, where DRL and GNN support each other. This fusion effectively addresses various complex problems in engineering and life sciences. Based on the review, we further analyze the applicability and benefits of fusing these two domains, especially in terms of increasing generalizability and reducing computational complexity. Finally, the key challenges in integrating DRL and GNN, and potential future research directions are highlighted, which will be of interest to the broader machine learning community.

97 MATHEMATICS AND COMPUTING↗

Exploring the Low-Thrust Transfer Design Space in an Ephemeris Model via Multi-Objective Reinforcement Learning

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective reinforcement learning algorithm used to train multiple policies to uncover solutions within a multi-objective solution space. MRPPO is used in this paper to train policies to construct low-thrust transfers for a SmallSat from the vicinity of !2 to an !5 short period orbit in the Sun-Earth-Moon system. First, the policies are trained in this scenario in the circular restricted three-body problem. This information is used to initialize the policies before training in a higher-fidelity ephemeris model; a process known as transfer learning. The recovered segments of the solution space will be compared to fundamental dynamical structures to both examine the results of MRPPO in this complex design scenario and explore the effectiveness of transfer learning.

Christopher J Sullivan↗

Exploring the Low-Thrust Transfer Design Space in an Ephemeris Model via Multi-Objective Reinforcement Learning

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective reinforcement learning algorithm used to train multiple policies to uncover solutions within a multi-objective solution space. MRPPO is used in this paper to train policies to construct low-thrust transfers for a SmallSat from the vicinity of L2 to an L5 short period orbit in the Sun-Earth-Moon system. First, the policies are trained in this scenario in the circular restricted three-body problem. This information is used to initialize the policies before training in a higher-fidelity ephemeris model; a process known as transfer learning. The recovered segments of the solution space will be compared to fundamental dynamical structures to both examine the results of MRPPO in this complex design scenario and explore the effectiveness of transfer learning.

Christopher J. Sullivan↗

Exploring the Low-Thrust Transfer Design Space in an Ephemeris Model via Multi-Objective Reinforcement Learning

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective reinforcement learning algorithm used to train multiple policies to uncover solutions within a multi-objective solution space. MRPPO is used in this paper to train policies to construct low-thrust transfers for a SmallSat from the vicinity of 𝐿2 to an 𝐿5 short period orbit in the Sun-Earth-Moon system. First, the policies are trained in this scenario in the circular restricted three-body problem. This information is used to initialize the policies before training in a higher-fidelity ephemeris model; a process known as transfer learning. The recovered segments of the solution space will be compared to fundamental dynamical structures to both examine the results of MRPPO in this complex design scenario and explore the effectiveness of transfer learning.

Mashiku, Alinda K.↗

Deep Reinforcement Learning Based Volt-VAR Optimization in Smart Distribution Systems

This paper develops a model-free volt-VAR optimization (VVO) algorithm via multi-agent deep reinforcement learning (DRL) in unbalanced distribution systems. This method is novel since we cast the VVO problem in distribution networks to an intelligent deep Q-network (DQN) framework, which avoids solving a specific optimization model directly when facing time-varying operating conditions in the systems. We consider statuses/ratios of switchable capacitors, voltage regulators, and smart inverters installed at distributed generators as the action variables of the agents. A delicately designed reward function guides these agents to interact with the distribution system, in the direction of reinforcing voltage regulation and power loss reduction simultaneously. The forward-backward sweep method for radial three-phase distribution systems provides accurate power flow results within a few iterations to the DRL environment. The proposed method realizes the dual goals for VVO. We test this algorithm on the unbalanced IEEE 13-bus and 123-bus systems. Numerical simulations validate the excellent performance of this method in voltage regulation and power loss reduction.

41 EE - Solar Energy Technologies Office (EE-4S)↗

Inverse Reinforcement Learning based Bayesian Goal Inference Method for Early Nuclear Proliferation Detection

Traditional methods for detection of nuclear proliferation indicators are usually applied after nuclear proliferation has already occurred. There is a need to advance these methods to perform early detection of nuclear proliferation indicators. In this project, we formulated an early detection problem as a sequential, decision-making, goal inference problem based on research publications of authors, to determine whether it is possible to infer whether an author will publish on a research activity before it has occurred. To develop and test our approach, we selected a civil nuclear activity for our case study. We constructed a state-action-state transition graph from publications of authors associated with the activity and the co-authors of their publications, using titles, abstracts, and author publication sequences. We then used inverse reinforcement learning to model the goal-directed behavior of authors in trajectories that terminate at selected goal states. Using a Bayesian formulation, we computed the probability that authors would reach each selected state from partially observed trajectories of their state transitions in their research topic space. The state with the highest probability was selected as the most probable goal state. Based on our results, we found that 60% of the times we can infer the correct goal state early; sometimes the inference is either delayed, or multiple states could be inferred as goal states. Overall, our results show that it is possible to perform early detection of research activities of authors in a nuclear technology area. Further research is necessary to establish a more accurate understanding of how topic modeling, topic space grid discretization, and the extent of overlap among trajectories of different goal states, affect the goal inference results. The methods developed in this work may be used to enhance data-driven methods for early detection of nuclear proliferation indicators.

97 MATHEMATICS AND COMPUTING↗

Enhancing Autonomous Control of Microreactors Using Multi-Agent Reinforcement Learning

In order for microreactors to be economically competitive, operation costs will need to be minimized through some degree of autonomous control. Previous work has demonstrated the effectiveness of reinforcement learning (RL) for load-following control in a drum-controlled microreactor. This study extends that work by exploring the potential of RL to independently control each of the reactor’s drums. We compare a single-agent RL approach with a multi-agent RL (MARL) framework, testing them for generalization across different load-following power profiles and control timescales, and for robustness in cases of randomly disabled control drums. Since the point kinetics simulation environment used in this study cannot resolve spatial effects, we assume that in the absence of spatially localized disturbances, optimal drum movements should be symmetrical. We demonstrate that single-agent RL is able to achieve accurate performance only when symmetric actions are ignored; otherwise, it fails to train a useful controller. Meanwhile, the MARL framework performs symmetric actions by design and trains a robust, accurate agent, as evidenced by mean absolute errors in power matching of 0.41% for the training power profile, 0.68% for a profile with half the drums disabled, and 0.21% for a profile on a realistic load-following time horizon.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Data-Driven Multi-agent Deep Reinforcement Learning for Distribution System Decentralized Voltage Control with High Penetration of PVs

This paper proposes a novel model-free/data-driven centralized training and decentralized execution multi-agent deep reinforcement learning (MADRL) framework for distribution system voltage control with high penetration of PVs. The proposed MADRL can coordinate both the real and reactive power control of PVs with existing static var compensators and battery storage systems. Unlike the existing DRL-based voltage control methods, our proposed method does not rely on a system model during both the training and execution stages. This is achieved by developing a new interaction scheme between the surrogate modeling of the original system and the multi-agent soft actor critic (MASAC) MADRL algorithm. In particular, the sparse pseudo-Gaussian process with a few-shots of measurements is utilized to construct the surrogate model of the original environment, i.e., power flow model. This is a data-driven process and no model parameters are needed. Furthermore, the MASAC enabled MADRL allows to achieve better scalability by dividing the original system into different voltage control regions with the aid of real and reactive power sensitivities to voltage, where each region is treated as an agent. This also serves as the foundation for the centralized training and decentralized execution, thus significantly reducing the communication requirements as only local measurements are required for control. Comparative results with other alternatives on the IEEE 123-nodes and 342-nodes systems demonstrate the superiority of the proposed method.

14 SOLAR ENERGY↗

Predicting Pilot Behavior in Medium Scale Scenarios Using Game Theory and Reinforcement Learning

Effective automation is critical in achieving the capacity and safety goals of the Next Generation Air Traffic System. Unfortunately creating integration and validation tools for such automation is difficult as the interactions between automation and their human counterparts is complex and unpredictable. This validation becomes even more difficult as we integrate wide-reaching technologies that affect the behavior of different decision makers in the system such as pilots, controllers and airlines. While overt short-term behavior changes can be explicitly modeled with traditional agent modeling systems, subtle behavior changes caused by the integration of new technologies may snowball into larger problems and be very hard to detect. To overcome these obstacles, we show how integration of new technologies can be validated by learning behavior models based on goals. In this framework, human participants are not modeled explicitly. Instead, their goals are modeled and through reinforcement learning their actions are predicted. The main advantage to this approach is that modeling is done within the context of the entire system allowing for accurate modeling of all participants as they interact as a whole. In addition such an approach allows for efficient trade studies and feasibility testing on a wide range of automation scenarios. The goal of this paper is to test that such an approach is feasible. To do this we implement this approach using a simple discrete-state learning system on a scenario where 50 aircraft need to self-navigate using Automatic Dependent Surveillance-Broadcast (ADS-B) information. In this scenario, we show how the approach can be used to predict the ability of pilots to adequately balance aircraft separation and fly efficient paths. We present results with several levels of complexity and airspace congestion.

Game Theory↗

Flexible Reinforcement Learning Framework for Building Control using EnergyPlus-Modelica Energy Models

In recent years, reinforcement learning (RL) methods have been greatly enhanced by leveraging deep learning approaches. RL methods applied to building control have shown potential in many applications due to their ability to complement or replace conventional methods such as model-based or rule-based controls. However, RL-based building control software is likely tailored either to one target building system or to a specific RL method so that significant additional effort would be required to customize the RL-based controller for use in other building systems or with other RL approaches. Also, RL-based building controls usually depend on building energy simulations to train controllers, so emulating building dynamics (i.e., thermal dynamics and control dynamics) and capturing sub-hourly dynamic profiles are crucial to further the development of effective RL-based building control methods. To address these challenges, we present an open source RL-based control software employing a high-fidelity hybrid EnergyPlus-Modelica building energy model which emulates building dynamics at 1-minute resolution. This software consists of decoupled components (environment, building emulator, control agent, and RL algorithm), which allows for quick prototyping and benchmarking of standard RL algorithms in different systems; for example, a single component can be replaced without revising all of the software. To demonstrate this software framework, we conducted a benchmark study using an EnergyPlus-Modelica building energy model for a Chicago office building with an RL-based controller to dynamically control the chilled water temperature setpoint and the air handling unit supply air temperature setpoint on selected floors.

Lee, Joon-Yong↗

Deep Reinforcement Learning based Model-free On-line Dynamic Multi-Microgrid Formation to Enhance Resilience

Multi-microgrid formation (MMGF) is a promising solution for enhancing power system resilience. This paper proposes a new deep reinforcement learning (RL) based model-free on-line dynamic MMGF scheme. Additionally, the dynamic MMGF problem is formulated as a Markov decision process, and a complete deep RL framework is specially designed for the topologytransformable micro-grids. In order to reduce the large action space caused by flexible switch operations, a topology transformation method is proposed and an action-decoupling Q-value is applied. Then, a convolutional neural network (CNN) based multi-buffer double deep Q-network (CM-DDQN) is developed to further improve the learning ability of the original DQN method. The proposed deep RL method provides real-time computing to support the on-line dynamic MMGF scheme, and the scheme handles a long-term resilience enhancement problem using an adaptive on-line MMGF to defend changeable conditions. The effectiveness of the proposed method is validated using a 7-bus system and the IEEE 123-bus system. The results show strong learning ability, timely response for varying system conditions and convincing resilience enhancement.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Transferable Reinforcement Learning for Smart Homes: Preprint

To harness the great amount of untapped resources at the demand side, smart home technology plays a vital role in solving the "last mile" problem in smart grid. Reinforcement learning (RL), which has demonstrated an outstanding performance in solving many sequential decision-making problems, can be a great candidate to be used in smart home control. For instance, many studies have started investigating the load scheduling problem under dynamic pricing scheme. Based on those, this study aims at providing an affordable solution to encourage a higher smart home adoption rate. Specifically, we investigate combining transfer learning (TL) with RL to reduce the training cost of an optimal RL control policy. Given an optimal policy for a benchmark home, TL can jump-start the RL training of a policy for a new home, which has different appliances and user preferences. Simulation results show that by leveraging TL, RL training converges faster and requires much less computing time for new homes that are similar to the benchmark home. In all, this study proposes a cost-effective approach for training RL control policies for homes at scale, which ultimately reduces the controller's implementation costs, increases the adoption rate of RL controllers, and makes more homes grid-interactive.

ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATION,↗