Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Off-policy deep reinforcement learning with automatic entropy adjustment for adaptive online grid emergency control

Electric overloading conditions and contingencies put modern power systems at risk of voltage collapse and blackouts. Load shedding is crucial to maintain voltage stability for grid emergency control. However, the rule- or model-based schemes rely on accurate dynamic system models and face considerable challenges in adapting to various operating conditions and uncertain event occurrences. Here, to address these issues, this paper proposes a novel deep reinforcement learning (DRL)-based voltage stability control algorithm with automatic entropy adjustment (AEA) for grid emergency control. Various dynamic network components for complex system operations are modeled to construct the DRL environment. An off-policy soft actor-critic architecture is developed to maximize the expected reward and policy entropy simultaneously. The AEA mechanism is proposed to facilitate the policy maximum entropy procedure, and the proposed method can automatically provide effective discrete and continuous actions against various fault scenarios. Our approach accomplishes high sampling efficiency, scalability, and auto-adaptivity of the control policies under high uncertainties. Comparative studies with the existing DRL-based control methods in IEEE benchmarks indicate salient performance improvement of the proposed method for dynamic system emergency control.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Towards Autonomous Lunar Resource Excavation via Deep Reinforcement Learning

To support sustainable infrastructure on the Moon, NASA needs to leverage lunar resources for in-situ processing and construction. NASA’s Regolith Advanced Surface Systems Operations Robot (RASSOR) is principally designed to mine and deliver regolith for these tasks. To reliably perform these operations on the lunar surface, RASSOR's sensors and control systems need to be robust and maximize information extracted from a reduced sensor payload. Herein, we present our findings from the Intelligent Capabilities Enhanced RASSOR project. We created reduced-order simulation environments in which we applied reinforcement learning algorithms to learn autonomous trenching controllers and produced state estimation architectures. We developed two simulations: a 2D excavation simulation used to facilitate parameter selection, and a 3D simulation developed using a game physics engine to simulate simplified soil interactions and incorporate robotic agents parameterized by dynamic models. Within these simulations, we learned autonomous excavation routines that exceed excavation efficiency measures as compared against RASSOR's existing control and teleoperation-based methods.

RASSOR↗

Towards Autonomous Lunar Resource Excavation via Deep Reinforcement Learning

To support sustainable infrastructure on the Moon, NASA needs to leverage lunar resources for in-situ processing and construction. NASA’s Regolith Advanced Surface Systems Operations Robot (RASSOR) is principally designed to mine and deliver regolith for these tasks. To reliably perform these operations on the lunar surface, RASSOR's sensors and control systems need to be robust and maximize information extracted from a reduced sensor payload. Herein, we present our findings from the Intelligent Capabilities Enhanced RASSOR project. We created reduced-order simulation environments in which we applied reinforcement learning algorithms to learn autonomous trenching controllers and produced state estimation architectures. We developed two simulations: a 2D excavation simulation used to facilitate parameter selection, and a 3D simulation developed using a game physics engine to simulate simplified soil interactions and incorporate robotic agents parameterized by dynamic models. Within these simulations, we learned autonomous excavation routines that exceed excavation efficiency measures as compared against RASSOR's existing control and teleoperation-based methods.

RASSOR↗

Adaptive Power System Emergency Control using Deep Reinforcement Learning

Power system emergency control is generally regarded as the final safety net for grid security and resiliency. Existing emergency control schemes are usually designed off-line based on either the conceived “worst” case scenarios or a few typical operation scenarios. These schemes are facing significant adaptiveness and robustness issues as increasing uncertainties and variations occur in modern electrical grids. To address these challenges, for the first time, this paper proposes a novel adaptive emergency control scheme using deep reinforcement learning (DRL), by leveraging the high-dimensional feature extraction and non-linear generalization capabilities DRL has for complex systems with high-dimensional variations. Furthermore, an open-source platform named DeepGrid has been designed for the first time to assist the DRL development and benchmarking processes in power system emergency control. Details of the platform, DRL, and emergency control schemes that use dynamic braking or under-voltage load shedding are presented. Extensive case studies performed in both two-area four-machine system and IEEE 39-Bus system have demonstrated the excellent performance and robustness of the proposed schemes.

Deep reinforcement learning, emergency control, lo↗

An Edge-Cloud Integrated Solution for Buildings Demand Response Using Reinforcement Learning

Buildings, as major energy consumers, can provide great untapped demand response (DR) resources for grid services. However, their participation remains low in real-life. One major impediment for popularizing DR in buildings is the lack of cost-effective automation systems that can be widely adopted. Existing optimization-based smart building control algorithms suffer from high costs on both building-specific modeling and on-demand computing resources. To tackle these issues, this paper proposes a cost-effective edge-cloud integrated solution using reinforcement learning (RL). Beside RL’s ability to solve sequential optimal decision-making problems, its adaptability to easy-to-obtain building models and the off-line learning feature are likely to reduce the controller’s implementation cost. Using a surrogate building model learned automatically from building operation data, an RL agent learns an optimal control policy on cloud infrastructure, and the policy is then distributed to edge devices for execution. Simulation results demonstrate the control efficacy and the learning efficiency in buildings of different sizes. A preliminary cost analysis on a 4-zone commercial building shows the annual cost for optimal policy training is only 2.25% of the DR incentive received. Results of this study show a possible approach with higher return on investment for buildings to participate in DR programs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Deep Reinforcement Learning for Distribution System Operations: A Tutorial and Survey

Here, the rapid evolution of modern electric power distribution systems into complex networks of interconnected active devices, distributed generation (DG), and storage poses increasing difficulties for system operators. The large-scale integration of distributed energy resources (DERs) and the rapid exchange of measurement data via communication networks present major opportunities for advancing grid operations but also introduce greater uncertainty, higher data dimensionality, more complex network and device models, and challenging control and optimization problems. Deep reinforcement learning (DRL) algorithms are promising in addressing these challenges. However, they have not been effectively adapted for power systems applications, requiring extensive customization for implementation and evaluation. This has resulted in reproducibility challenges and a steep learning curve for researchers new to applying DRL algorithms to the power systems domain. To bridge these gaps, this tutorial aims to serve as a valuable resource for researchers interested in exploring learning-based algorithms to operate active power distribution networks. Specifically, this work presents a generalized process for translating sequential decision-making problems in power distribution systems into Markov decision process (MDP) formulations, illustrated through concrete grid service examples. Additionally, we introduce a simple environment design strategy to develop and evaluate example DRL algorithms for distribution system applications, complete with an included code repository to guide users through environment construction.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Deep reinforcement learning for dynamic control of fuel injection timing in multi-pulse compression ignition engines

Conventional compression-ignition (CI) engines have long offered high thermal efficiencies and torque across a wide range of loads, but often require extensive exhaust gas treatment that decreases efficiency to meet ever-increasing emissions regulations. One strategy to decrease emissions is to split the fuel injection into a series of smaller injections. In this paper, we explore a new way of discovering optimal control strategies for the next generation of CI engines using deep reinforcement learning (DRL). We outline a DRL procedure to maximize the weighted reward of engine work while minimizing end-of-cycle NO x emissions. Through the procedure outlined in this paper, we show that the DRL agent is able to reduce NO x emissions threefold while only decreasing network by 2%. We demonstrate the use of transfer learning (TL) across hierarchies of physical models to accelerate the learning process, making this approach feasible for a range of control problems within this space. This paper presents a framework and demonstration for using DRL to design control systems in technology areas such as multi-pulse engine control where a hierarchy of models combined with multi-objective rewards are used for optimal operation.

33 ADVANCED PROPULSION SYSTEMS↗

Safe Reinforcement Learning for Emergency Load Shedding of Power Systems

The paradigm shift in the electric power grid necessitates a revisit of existing control methods to ensure the grid’s security and resilience. In particular, the increased uncertainties and rapidly changing operational conditions in power systems have revealed outstanding issues in terms of either speed, adaptiveness, or scalability of the existing control methods for power systems. On the other hand, the availability of massive real-time data can provide a clearer picture of what is happening in the grid. Recently, deep reinforcement learning (RL) has been regarded and adopted as a promising approach leveraging massive data for fast and adaptive grid control. However, like most existing machine learning (ML)- based control techniques, RL control usually cannot guarantee the safety of the power systems. In this paper, we introduce a novel method for safe RL-based load shedding of power systems that can enhance the safe voltage recovery of the electric power grid after experiencing faults. Numerical simulation on the IEEE 39-bus testcase is performed to demonstrate the effectiveness of the proposed safe RL emergency control, as well as its adaptive capability to faults not seen in the training.

reinforcement learning, emergency control, power g↗

A Bayesian Approach for Quantifying Data Scarcity when Modeling Human Behavior via Inverse Reinforcement Learning

Computational models that formalize complex human behaviors enable study and understanding of such behaviors. However, collecting behavior data required to estimate the parameters of such models is often tedious and resource intensive. Thus, estimating dataset size as part of data collection planning (also known as Sample Size Determination) is important to reduce the time and effort of behavior data collection while maintaining an accurate estimate of model parameters. In this paper, we present a sample size determination method based on Uncertainty Quantification (UQ) for a specific Inverse Reinforcement Learning (IRL) model of human behavior, in two cases: 1) pre-hoc experiment design—conducted in the planning stage before any data is collected, to guide the estimation of how many samples to collect; and 2) post-hoc dataset analysis—performed after data is collected, to decide if the existing dataset has sufficient samples and whether more data is needed. Here, we validate our approach in experiments with a realistic model of behaviors of people with Multiple Sclerosis (MS) and illustrate how to pick a reasonable sample size target. Our work enables model designers to perform a deeper, principled investigation of effects of dataset size on IRL.

97 MATHEMATICS AND COMPUTING↗

Deep reinforcement learning based optimization for a tightly coupled nuclear renewable integrated energy system

New ways to integrate energy systems to maximize efficiency are being sought to meet carbon emissions goals. Nuclear-renewable integrated energy system (NR-IES) concepts are a leading solution that couples a nuclear power plant with renewable energy, hydrogen generation plants, and energy storage systems, such that thermal and electrical power are dispatchable to fulfill grid-flexibility requirements while also producing hydrogen and maximizing revenue. Here, this paper introduces a deep reinforcement learning (DRL)-based framework to address the complex decision-making tasks for NR-IES. The objective is to maximize revenue by generating and selling hydrogen and electricity simultaneously according to their time-varying prices while keeping the energy flow in the subsystems in balance. A Python-based simulator for a NR-IES concept has been developed to integrate with OpenAI Gym and Ray/RLlib to enable an efficient and flexible computational framework for DRL research and development. Three state-of-the-art DRL algorithms have been investigated, including two-delayed deep deterministic policy gradient (TD3), soft-actor critic (SAC), proximal policy optimization (PPO), to illustrate DRL’s superiority for controlling NR-IES by comparing it with a conventional control approach, particle swarm optimization (PSO). In this effort, PPO has shown more-stable performance and also better generalization capability than SAC and TD3. Comparisons with PSO have demonstrated that, on average, PPO can achieve 13.9% more mean episode returns from the training process and 29.4% more mean episode returns from the testing process when different hydrogen-production targets are applied.

08 HYDROGEN↗

A Cyber-Physical System for Freeway Ramp Meter Signal Control Using Deep Reinforcement Learning in a Connected Environment

Freeway bottlenecks such as on-ramp merging areas account for about 40% of recurring freeway congestion. It is generally agreed that building more roads and adding more lanes to existing infrastructure does not solve the congestion problem, and so dynamic traffic control measures offer a more cost-effective alternative. Ramp meters, traffic signal devices that regulate traffic flow entering freeways, are among the most effective measures to mitigate congestion at on-ramp merging areas on freeways. The confluence of deep reinforcement learning (RL) and connectivity provides a possible solution to advance ramp meter signal control. Deep RL is a group of machine-learning methods that enables an agent learning from the environment to improve its performance. In this study, three deep RL methods-proximal policy optimization (PPO), Ape-X deep Q-network (DQN), and asynchronous advantage actor-critic agents (A3C)-are explored for ramp meter signal control to maximize vehicle speed and traffic throughput, as well as to minimize energy consumption and emissions at freeway on-ramp merging areas in a connected environment. The low computational requirement and scalability of deep RL for deployment make it a powerful optimization tool for time-sensitive applications such as ramp meter signal control. The results of this study show that deep RL methods yield superior performance to both a fixed-time controller and ALINE A, a state-of-the-art feedback controller.

ADVANCED PROPULSION SYSTEMS,MATHEMATICS AND COMPUT↗

A Barrier-Certificated Reinforcement Learning Approach for Enhancing Power System Transient Stability

Increasing integration of renewable resources brings more flexibility and poses new challenges to modern power systems, leading to highly nonlinear and complex dynamics. Here, this paper aims to provide a general solution framework to traditional control problems, such as frequency control and voltage control, which attempt to maintain the stability of either synchronous generators-governed or inverter-governed systems when subjected to a disturbance and simultaneously guarantee operational constraints, providing a complete complement to existing works on control design. Building on reinforcement learning (RL) and control barrier functions, the framework includes two subsystems, i.e., a model-free controller and a barrier-certification system, which discover RL-based control actions and sequentially filter them using a barrier certificate to satisfy operational constraints. Calculating a barrier function is generally challenging for a complex power system. This is addressed by representing the barrier function using neural networks (NNs) and data-based approaches. An adaptive method is introduced to certify the neural barrier function that perseveres barrier conditions, which is more compatible with online implementation. The proposed framework synthesizes a stabilizing controller that satisfies predefined safety regions. The effectiveness of the proposed framework is demonstrated via several comparative case studies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Resilient Control of Networked Microgrids Using Vertical Federated Reinforcement Learning: Designs and Real-Time Test-Bed Validations

Improving system-level resiliency of networked microgrids against adversarial cyber-attacks is an important aspect in the current regime of increased inverter-based resources (IBRs). To achieve that, this paper contributes in designing a hierarchical control layer, in conjunction with the existing control layers, resilient to adversarial attack signals. Considering model complexities, unknown dynamical behaviors of IBRs, and privacy issues regarding data sharing in multi-party-owned microgrids, designing such a control layer is non-trivial. Here, to tackle these issues, a novel federated reinforcement learning (Fed-RL) method is proposed. To grasp the interconnected dynamics of networked microgrids, the paper develops Federated Soft Actor-Critic (FedSAC) algorithm following the vertical structure of implementing Fed-RL. Next, utilizing the OpenAI Gym interface, we built a custom set-up in GridLAB-D/HELICS co-simulation platform, named Resilient RL Co-simulation (ResRLCoSIM), to train the RL agents with IEEE 123-bus benchmark comprising 3 interconnected microgrids. Finally, the learned policies in the simulation are transferred to the real-time hardware-in-the-loop (HIL) test-bed developed using the high-fidelity Hypersim platform. Finally, experiments show that the simulator-trained RL controllers achieve desirable performance with the test-bed platform, validating the minimization of the sim-to-real gap.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Multiple-objective Reinforcement Learning for Inverse Design and Identification

The aim of inverse chemical design is to develop new molecules with given optimized molecular properties or objectives. Recently, generative deep learning (DL) networks are considered as the state-of-the-art in inverse chemical design and have achieved early success in generating molecular structures with desired properties in the pharmaceutical and material chemistry fields. However, satisfying a large number (> 10 objectives) of molecular objectives is a limitation of current generative models. To improve the model’s ability to handle a large number of molecule design objectives, we developed a Reinforcement Learning (RL) based generative framework to optimize chemical molecule generation. Our use of Curriculum Learning (CL) to fine-tune the pre-trained generative network allowed the model to satisfy up to 21 objectives and increase the generative network’s robustness. The experiments show that the proposed multiple-objective RL-based generative model can correctly identify unknown molecules with an 83% to 100% success rate, compared to the baseline approach of 0%. Additionally, this proposed generative model is not limited to just chemistry research challenges; we anticipate that problems that utilize RL with multiple objectives will benefit from this framework.

reinforcement learning, generative model, inverse ↗

Application of fuzzy logic-neural network based reinforcement learning to proximity and docking operations

As part of the Research Institute for Computing and Information Systems (RICIS) activity, the reinforcement learning techniques developed at Ames Research Center are being applied to proximity and docking operations using the Shuttle and Solar Max satellite simulation. This activity is carried out in the software technology laboratory utilizing the Orbital Operations Simulator (OOS). This interim report provides the status of the project and outlines the future plans.

Jani, Yashvant↗