Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Environment Adversarial Reinforcement Learning

This paper presents a training method for increasing performance of reinforcement learning agents. The method is named Environment Adversarial Reinforcement Learning. The method requires the reinforcement learning environment to be parameterizeable. Over the course of training, environment parameters are updated in a direction of increasing difficulty for the agent. The direction for these updates is found using a performance prediction network trained on data from tests of the agent under varying environment parameters. The method was tested on a CartPole environment. A 28-58\% improvement in mean return was found when comparing performance to a baseline reinforcement learning algorithm on both easy and hard versions of the task.

machine learning↗

Inverse reinforcement learning control for building energy management

Reinforcement learning (RL) based control is widely considered a promising approach in building automation and control as it has demonstrated the potential to deal with complex objectives in adjacent domains like robotics, autonomous vehicles, gaming applications, and advertisement recommendations. When applied to any environment, model-free RL learns to improve its control performance over time without requiring a control model, by receiving and then analyzing feedback from the building environment after each control action. Operational objectives are becoming increasingly complex through the simultaneous consideration of thermal comfort, carbon emissions, grid services, and indoor air quality. In this context, conventional rule-based control approaches are proving sub-optimal, mostly heuristic, and inadequate. The model-free and self learning nature of RL appears promising and attractive as it may address the scalability issues associated with advanced control approaches. However, it suffers from long training times and unstable control behavior during the early stages of its learning process, which makes it unsuitable to be applied directly to buildings. This paper addresses these issues using an inverse reinforcement learning approach (IRL), a technique utilized to learn the objective of a controller agent which is considered an expert in its respective domain. Here, we consider a rule-based control as the expert demonstrator. IRL is different from a direct imitation (i.e., direct mapping of states to actions) of control actions as it tries to find the underlying intent of an expert's policy, providing the controller with a better-generalized policy for unseen states or environments with slightly different dynamics. This approach propels the RL controller's policy to levels similar to or better than that of a rule-based policy before it starts learning by interacting with the building. This makes RL for building energy management applications more practical as it prevents the erratic and exploratory behavior in the initial training period, simultaneously speeding up the learning process when compared to applying an untrained RL agent directly to a building environment.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Non-Stationary Policy Learning for Multi-Timescale Multi-Agent Reinforcement Learning

In multi-timescale multi-agent reinforcement learning (MARL), agents interact across different timescales. In general, policies for time-dependent behaviors, such as those induced by multiple timescales, are non-stationary. Learning non-stationary policies is challenging and typically requires sophisticated or inefficient algorithms. Motivated by the prevalence of this control problem in real-world complex systems, we introduce a simple framework for learning non-stationary policies for multi-timescale MARL. Our approach uses available information about agent timescales to define and learn periodic multi-agent policies. In detail, we theoretically demonstrate that the effects of non-stationarity introduced by multiple timescales can be learned by a periodic multi-agent policy. To learn such policies, we propose a policy gradient algorithm that parameterizes the actor and critic with phase-functioned neural networks, which provide an inductive bias for periodicity. The framework's ability to effectively learn multi-timescale policies is validated on a gridworld and building energy management environment.

control↗

Energy-Optimized Path Planning for Uas in Varying Winds Via Reinforcement Learning

In this paper we propose a reinforcement learning (RL) algorithm for path planning of Unmanned Aviation Vehicles (UAVs) under varying wind conditions. Solutions to UAV path planning problems are becoming increasingly necessary as autonomous UAVs continue to enter commercial and government spaces. Path-planning is inherently challenging, as UAVs need to account for dynamically changing flying conditions such as weather, obstacle or no-fly zones, degraded vehicle health, and off-nominal battery power consumption. Machine learning methods such as reinforcement learning (RL) have the potential to revolutionize how vehicles navigate in such uncertain environments. In this study, we compute UAV trajectories from a pre-determined starting position to a target cell within a 7X7 grid environment by optimizing parameters for mission assurance and safety limits in addition to the energy consumption and operation time. The UAV navigates the grid by taking actions to move in any of the eight cardinal and inter-cardinal directions, under constant thrust profile. The resultant UAV state is sampled from a probability distribution which accounts for the UAV’s action, local wind velocity, and the presence of obstacles or boundaries. As the unmanned airspace gets more complex due to multiple vehicles and environmental uncertainties, trade-offs between energy consumption, operation time, risk tolerance, and mission assurance need to be made. Our Markov Decision Process (MDP) environment model can capture any combination of these in the optimization objective, making it novel compared to other work in the field.

trajectory planning↗

A Study on Efficient Reinforcement Learning Through Knowledge Transfer

Although Reinforcement Learning (RL) algorithms have made impressive progress in learning complex tasks over the past years, there are still prevailing short-comings and challenges. Specifically, the sample-inefficiency and limited adaptation across tasks often make classic RL techniques impractical for real-world applications despite the gained representational power when combining deep neural networks with RL, known as Deep Reinforcement Learning (DRL). Recently, a number of approaches to address those issues have emerged. Many of those solutions are based on smart DRL architectures that enhance single task algorithms with the capability to share knowledge between agents and across tasks by introducing Transfer Learning (TL) capabilities. Here this survey addresses strategies of knowledge transfer from simple parameter sharing to privacy preserving federated learning and aims at providing a general overview of the field of TL in the DRL domain, establishes a classification framework, and briefly describes representative works in the area.

97 MATHEMATICS AND COMPUTING↗

Integration of Decentralized Graph-Based Multi-Agent Reinforcement Learning with Digital Twin for Traffic Signal Optimization

Machine learning (ML) methods, particularly Reinforcement Learning (RL), have gained widespread attention for optimizing traffic signal control in intelligent transportation systems. However, existing ML approaches often exhibit limitations in scalability and adaptability, particularly within large traffic networks. This paper introduces an innovative solution by integrating decentralized graph-based multi-agent reinforcement learning (DGMARL) with a Digital Twin to enhance traffic signal optimization, targeting the reduction of traffic congestion and network-wide fuel consumption associated with vehicle stops and stop delays. In this approach, DGMARL agents are employed to learn traffic state patterns and make informed decisions regarding traffic signal control. The integration with a Digital Twin module further facilitates this process by simulating and replicating the real-time asymmetric traffic behaviors of a complex traffic network. The evaluation of this proposed methodology utilized PTV-Vissim, a traffic simulation software, which also serves as the simulation engine for the Digital Twin. The study focused on the Martin Luther King (MLK) Smart Corridor in Chattanooga, Tennessee, USA, by considering symmetric and asymmetric road layouts and traffic conditions. Comparative analysis against an actuated signal control baseline approach revealed significant improvements. Experiment results demonstrate a remarkable 55.38% reduction in Eco_PI, a developed performance measure capturing the cumulative impact of stops and penalized stop delays on fuel consumption, over a 24 h scenario. In a PM-peak-hour scenario, the average reduction in Eco_PI reached 38.94%, indicating the substantial improvement achieved in optimizing traffic flow and reducing fuel consumption during high-demand periods. These findings underscore the effectiveness of the integrated DGMARL and Digital Twin approach in optimizing traffic signals, contributing to a more sustainable and efficient traffic management system.

actuated signal control↗

Integration of Decentralized Graph-Based Multi-Agent Reinforcement Learning with Digital Twin for Traffic Signal Optimization

Machine learning (ML) methods, particularly Reinforcement Learning (RL), have gained widespread attention for optimizing traffic signal control in intelligent transportation systems. However, existing ML approaches often exhibit limitations in scalability and adaptability, particularly within large traffic networks. This paper introduces an innovative solution by integrating decentralized graph-based multi-agent reinforcement learning (DGMARL) with a Digital Twin to enhance traffic signal optimization, targeting the reduction of traffic congestion and network-wide fuel consumption associated with vehicle stops and stop delays. In this approach, DGMARL agents are employed to learn traffic state patterns and make informed decisions regarding traffic signal control. The integration with a Digital Twin module further facilitates this process by simulating and replicating the real-time asymmetric traffic behaviors of a complex traffic network. The evaluation of this proposed methodology utilized PTV-Vissim, a traffic simulation software, which also serves as the simulation engine for the Digital Twin. The study focused on the Martin Luther King (MLK) Smart Corridor in Chattanooga, Tennessee, USA, by considering symmetric and asymmetric road layouts and traffic conditions. Comparative analysis against an actuated signal control baseline approach revealed significant improvements. Experiment results demonstrate a remarkable 55.38% reduction in Eco_PI, a developed performance measure capturing the cumulative impact of stops and penalized stop delays on fuel consumption, over a 24 h scenario. In a PM-peak-hour scenario, the average reduction in Eco_PI reached 38.94%, indicating the substantial improvement achieved in optimizing traffic flow and reducing fuel consumption during high-demand periods. These findings underscore the effectiveness of the integrated DGMARL and Digital Twin approach in optimizing traffic signals, contributing to a more sustainable and efficient traffic management system.

42 ENGINEERING↗

Learning and Fast Adaptation for Grid Emergency Control via Deep Meta Reinforcement Learning

As power systems are undergoing a significant transformation with more uncertainties, less inertia and closer to operation limits, there is increasing risk of large outages. Thus, there is an imperative need to enhance grid emergency control to maintain system reliability and security. Towards this end, great progress has been made in developing deep reinforcement learning (DRL) based grid control solutions in recent years. However, existing DRL-based solutions have two main limitations: 1) they cannot handle well with a wide range of grid operation conditions, system parameters, and contingencies; 2) they generally lack the ability to fast adapt to new grid operation conditions, system parameters, and contingencies, limiting their applicability for real-world applications. Here, in this paper, we mitigate these limitations by developing a novel deep meta-reinforcement learning (DMRL) algorithm. The DMRL combines the meta strategy optimization together with DRL, and trains policies modulated by a latent space that can quickly adapt to new scenarios. We test the developed DMRL algorithm on the IEEE 300-bus system. We demonstrate fast adaptation of the meta-trained DRL polices with latent variables to new operating conditions and scenarios using the proposed method, which achieves superior performance compared to the state-of-the-art DRL and model predictive control (MPC) methods.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Demonstration of Intelligent HVAC Load Management With Deep Reinforcement Learning: Real-World Experience of Machine Learning in Demand Control

We report that buildings account for 40% of total primary energy consumption and 30% of all CO 2 emissions worldwide. A large portion of building energy consumption is due to heating, ventilation, and air-conditioning (HVAC) systems. In the summer, for example, more than 50% of a building’s electricity consumption is used for cooling. With proper energy management, buildings can provide load shifting, peak shaving, frequency regulation, and many other demand response services.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Exploring the holographic entropy cone via reinforcement learning

We develop a reinforcement learning algorithm to study the holographic entropy cone. Given a target entropy vector, our algorithm searches for a graph realization whose min-cut entropies match the target vector. If the target vector does not admit such a graph realization, it must lie outside the cone, in which case the algorithm finds a graph whose corresponding entropy vector most nearly approximates the target and allows us to probe the location of the facets. For the N = 3 cone, we confirm that our algorithm successfully rediscovers monogamy of mutual information beginning with a target vector outside the holographic entropy cone. We then apply the algorithm to the N = 6 cone, analyzing the 6 mystery extreme rays of the subadditivity cone from [1] that satisfy all known holographic entropy inequalities yet lacked graph realizations. We found realizations for 3 of them, proving they are genuine extreme rays of the holographic entropy cone, while providing evidence that the remaining 3 are not realizable, implying unknown holographic inequalities exist for N = 6.

AdS-CFT correspondence↗

Enhancing Cyber Resilience of Networked Microgrids using Vertical Federated Reinforcement Learning

This paper presents a novel federated reinforcement learning (Fed-RL) methodology to inject sufficient resiliency into the operations of the network of microgrids. We consider adversarial actions to the voltage and power control loop reference signals at the grid forming (GFM) inverters in the microgrids which are essential to integrate renewable resources. Therefore, we formulate a resilient reinforcement learning training setup that uses these adversarial injections to generate episodic trajectories and train the RL agents to alleviate their impact on performance. To circumvent the concerns about data-sharing and privacy for different owners of the microgrids in the networked setting, we bring in the aspects of the federated operation to propose novel Fed-RL algorithms. As the dynamics of each microgrid are coupled due to electrical interlinks, the conventional federated RL approaches using decoupled independent environments are not applicable, which leads us to propose a multi-agent vertically federated variation of actor-critic algorithms, namely federated soft actor-critic (FedSAC). We have performed numerical simulations on an IEEE 123-bus benchmark test feeder with three microgrids by creating a customized simulation setup by encapsulating the microgrid dynamic simulations in GridLAB-D/HELICS co-simulation platform with the OpenAI Gym environment and validated the proposed resilient and secured learning methodology.

Artificial Intelligence (AI), reinforcement learni↗

A Transfer Learning Strategy for Improving the Data Efficiency of Deep Reinforcement Learning Control in Smart Buildings

Reinforcement learning (RL) is a powerful tool that has shown promising results in many domains such as robotics and game-playing. Because RL algorithms learn optimal control policies by continuously interacting with their environments, these algorithms require a lot of data to learn, which limits their application to a wide range of domains. For this reason, there is an immense need for improving the training and data efficiency of RL. Towards addressing this research gap, this paper proposes a transfer learning (TL) approach to improve the efficiency of the RL algorithms by reducing data need and, thus, reducing training time. To demonstrate the proposed approach, a knowledge transfer from a set of buildings to another building was conducted. The results show that the proposed TL approach is a promising method that can efficiently harness the information from similar RL tasks and reduce the data needs of RL algorithms.

Amasyali, Kadir↗

Graph reinforcement learning for exploring model spaces beyond the standard model

We present a methodology for performing scans of beyond the standard model (BSM) parameter spaces with reinforcement learning. We identify a novel procedure using graph neural networks that is capable of exploring spaces of models without the user specifying a fixed particle content, allowing broad classes of BSM models to be explored—in theory, the technique is applicable to nearly any model space with a prespecified gauge group. We provide a generic procedure by which a suitable graph grammar can be developed for any BSM model that features user-specified symmetry groups and a finite number of different possible particle species, the use of which is applicable to a variety of machine learning tasks over the actions of BSM theories beyond our particular reinforcement learning use case. As a proof of concept, we construct the graph grammar for theories with vectorlike leptons that may or may not be charged under a dark U ( 1 ) group, inspired by portal matter extensions of the sub-GeV vector portal/kinetic mixing simplified dark matter models. We then use this graph grammar to create a reinforcement learning environment tasked with creating models with these vectorlike leptons that are consistent with a list of a variety of precision observables. The reinforcement learning agent succeeds in developing models that can address the observed muon anomalous magnetic moment discrepancy while remaining consistent with flavor violation and electroweak precision observables, including both constructions that have previously been studied as well as new models that have not, to our knowledge, previously been identified. By inspecting the resulting ensembles of models that the agent produces and experimenting with different configurations for our reinforcement learning environment and graph grammar, we also infer various lessons about the development of these environments that can be transferable to reinforcement learning scans of more complicated model spaces and comment on future directions for the development of this technique into a more mature tool. Published by the American Physical Society 2025

Wojcik, George N.↗

Optimizing thermodynamic trajectories using evolutionary and gradient-based reinforcement learning

Here using a model heat engine, we show that neural-network-based reinforcement learning can identify thermodynamic trajectories of maximal efficiency. We consider both gradient and gradient-free reinforcement learning. We use an evolutionary learning algorithm to evolve a population of neural networks, subject to a directive to maximize the efficiency of a trajectory composed of a set of elementary thermodynamic processes; the resulting networks learn to carry out the maximally efficient Carnot, Stirling, or Otto cycles. When given an additional irreversible process, this evolutionary scheme learns a previously unknown thermodynamic cycle. Gradient-based reinforcement learning is able to learn the Stirling cycle, whereas an evolutionary approach achieves the optimal Carnot cycle. Our results show how the reinforcement learning strategies developed for game playing can be applied to solve physical problems conditioned upon path-extensive order parameters.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Interpreting Primal-Dual Algorithms for Constrained Multiagent Reinforcement Learning: Preprint

We study multiagent reinforcement learning (MARL) with constraints. This setting is gaining importance as MARL algorithms find new applications in real-world systems ranging from power grids to drone swarms. Most constrained MARL (C-MARL) algorithms use a primal-dual approach to enforce constraints through a penalty function added to the reward. In this paper, we study the structural effects of the primal-dual approach on the constraints and value function. First, we show that using the constraint evaluation as the penalty leads to a weak notion of safety, but by making simple modifications to the penalty function, we can enforce meaningful probabilistic safety constraints. Second, we show that the penalty term changes the value function in a way that is easy to model, and demonstrate the consequences of not doing so. We conclude with simulations in a simple constrained multiagent environment to back up the theoretical results.

data-driven control↗

A Generation-Storage Coordination Dispatch Strategy for Power System Based on Causal Reinforcement Learning

In the backdrop of global energy transformation, power systems integrating high proportions of renewable energy sources are facing unprecedented challenges in operational stability and dispatch efficiency. To address these challenges, this study introduces a generation-storage coordination real-time dispatch strategy based on Causal Power System Dynamic Reinforcement Learning (CPSDRL). Diverging from traditional reinforcement learning approaches, CPSDRL innovatively incorporates causal inference within the state prediction model - the crux of model-based reinforcement learning - thereby establishing the Power Causal Dynamic Model (PCDM). Assisted by the prior knowledge of power systems, the model significantly enhances prediction accuracy and reliability through a two-stage training process. Utilizing PCDM, this study further applies a direct policy search algorithm to optimize the real-time dispatch strategy. Experimental results indicate that the proposed method improves the stability of generation-storage coordination real-time dispatch and exhibits competitive advantages in sample efficiency and computational speed, compared to traditional model-based and model-free reinforcement learning algorithms. This method is expected to enhance the practicality and adaptability of causal reinforcement learning techniques in power system scheduling and control.

causal reinforcement learning↗

Predicting Catalyst Surface Stability Under Reaction Conditions Using Deep Reinforcement Learning and Machine Learning Potentials

Catalysts are critical for most large-scale energy intensive chemical transformation processes, such as energy storage, liquid fuel production, and the formation of chemical building blocks. The catalyst composition, structure, and morphology impact the performance under reaction conditions and influences properties like activity and selectivity. Many industrial catalysts offer imperfect activity/selectivity, or contain expensive metals. Further, catalysts can deactivate over time as the morphology changes or harsh reactive environments alter the surface structure. Methods to automatically model and predict the kinetics of how catalyst surfaces will restructure would enable engineers to design around these challenges, improve performance, increase longevity. This project investigated a specific type of machine learning model, deep reinforcement learning, machine learning models to act as surrogates for the physical system, and compared the results of these approaches with a specialized high-throughput experimental synthesis and measurement platform.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Use of Inverse Reinforcement Learning for Identity Prediction

We adopt Markov Decision Processes (MDP) to model sequential decision problems, which have the characteristic that the current decision made by a human decision maker has an uncertain impact on future opportunity. We hypothesize that the individuality of decision makers can be modeled as differences in the reward function under a common MDP model. A machine learning technique, Inverse Reinforcement Learning (IRL), was used to learn an individual's reward function based on limited observation of his or her decision choices. This work serves as an initial investigation for using IRL to analyze decision making, conducted through a human experiment in a cyber shopping environment. Specifically, the ability to determine the demographic identity of users is conducted through prediction analysis and supervised learning. The results show that IRL can be used to correctly identify participants, at a rate of 68% for gender and 66% for one of three college major categories.

Hayes, Roy↗