Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Deep reinforcement learning control of hydraulic fracturing

Hydraulic fracturing is a technique to extract oil and gas from shale formations, and obtaining a uniform proppant concentration along the fracture is key to its productivity. Recently, various model predictive control schemes have been proposed to achieve this objective. But such controllers require an accurate and computationally efficient model which is difficult to obtain given the complexity of the process and uncertainties in the rock formation properties. In this article, we design a model-free data-based reinforcement learning controller which learns an optimal control policy through interactions with the process. Deep reinforcement learning (DRL) controller is based on the Deep Deterministic Policy Gradient algorithm that combines Deep-Q-network with actor-critic framework. In addition, we utilize dimensionality reduction and transfer learning to quicken the learning process. We show that the controller learns an optimal policy to obtain uniform proppant concentration despite the complex nature of the process while satisfying various input constraints.

42 ENGINEERING↗

Controlling colloidal crystals via morphing energy landscapes and reinforcement learning

We report a feedback control method to remove grain boundaries and produce circular shaped colloidal crystals using morphing energy landscapes and reinforcement learning–based policies. We demonstrate this approach in optical microscopy and computer simulation experiments for colloidal particles in ac electric fields. First, we discover how tunable energy landscape shapes and orientations enhance grain boundary motion and crystal morphology relaxation. Next, reinforcement learning is used to develop an optimized control policy to actuate morphing energy landscapes to produce defect-free crystals orders of magnitude faster than natural relaxation times. Morphing energy landscapes mechanistically enable rapid crystal repair via anisotropic stresses to control defect and shape relaxation without melting. This method is scalable for up to at least N = 10 3 particles with mean process times scaling as N 0.5 . Further scalability is possible by controlling parallel local energy landscapes (e.g., periodic landscapes) to generate large-scale global defect-free hierarchical structures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Graph Partitioning and Sparse Matrix Ordering using Reinforcement Learning and Graph Neural Networks

We present a novel method for graph partitioning, based on reinforcement learning and graph convolutional neural networks. Our approach is to recursively partition coarser representations of a given graph. The neural network is implemented using SAGE graph convolution layers, and trained using an advantage actor critic (A2C) agent. We present two variants, one for finding an edge separator that minimizes the normalized cut or quotient cut, and one that finds a small vertex separator. The vertex separators are then used to construct a nested dissection ordering to permute a sparse matrix so that its triangular factorization will incur less fill-in. The partitioning quality is compared with partitions obtained using METIS and SCOTCH, and the nested dissection ordering is evaluated in the sparse solver SuperLU. Our results show that the proposed method achieves similar partitioning quality as METIS and SCOTCH. Furthermore, the method generalizes across different classes of graphs, and works well on a variety of graphs from the SuiteSparse sparse matrix collection.

97 MATHEMATICS AND COMPUTING↗

Development of algorithms for augmenting and replacing conventional process control using reinforcement learning

Here, this work seeks to allow for the online operation and training of model-free reinforcement learning (RL) agents but limit the risk to system equipment and personnel. The parallel implementation of RL alongside more conventional process control (CPC) allows for the RL algorithm to learn from CPC. The past performance of both methods are assessed on a continuous basis allowing for a transition from CPC to RL and, if needed, transitioning back to CPC from RL. This allows for the RL algorithm to slowly and safely assume control of the process without significant degradation in control performance. It is shown that the RL can derive a near optimal policy even when coupled with a suboptimal CPC. It is also demonstrated that the coupled RL-CPC algorithm learns at a faster rate than traditional RL methods of exploration while the algorithm’s performance does not deteriorate below CPC, even when exposed to an unknown operating condition.

30 DIRECT ENERGY CONVERSION↗

Surfactant-Specific AI-Driven Molecular Design: Integrating Generative Models, Predictive Modeling, and Reinforcement Learning for Tailored Surfactant Synthesis

Molecular design is a critical aspect of various scientific and industrial fields, where the properties of molecules hold significant importance. In this study, a 3-fold methodology design is presented that leverages the power of generative artificial intelligence (AI), predictive modeling, and reinforcement learning to create tailored molecules with desired properties. This model synergistically combines deep learning techniques with Self-Referencing Embedded Strings (SELFIES) molecular representation to build a generative model that generates valid molecules and a graphical neural network model that accurately forecasts molecular properties. The Variational Autoencoder (VAE) coupled with reinforcement learning helps refine molecule generation based on targeted attributes. Data from an experimental study involving surfactants were used to test the framework. A validation of the structural integrity of the molecules generated was conducted, and Tanimoto similarities were used to quantify the similarity and diversity between the original and generated molecular structures. Also, saliency maps for the generated surfactants were produced to identify the features explaining the property values. Lastly, molecular dynamics simulations were used to validate the stability of the generated molecules. The results showed that the proposed framework can effectively produce valid molecules within the set property threshold value.

36 MATERIALS SCIENCE↗

Optimal Coordination of Electric Vehicles for Grid Services using Deep Reinforcement Learning

Recent research has shown the effectiveness of reinforcement learning (RL) in coordinating electric vehicles (EVs) with vehicle-to-grid capabilities for grid services. However, many of these studies rely on lookup table and deep Q-network techniques, which can be impractical when dealing with continuous states and actions. In addition, existing RL designs inadequately account for battery aging effects, EV user satisfaction, uncertain departure and arrival time, and trip distance, which may compromise effective coordination. This paper aims to bridge these gaps by developing an innovative deep deterministic policy gradient-based RL framework for optimal coordination of EVs. Case studies were carried out using a test system with 100 EVs, and numerical analysis results showed that the proposed RL framework can effectively coordinate EVs to maximize economic benefits and user satisfaction while ensuring the expected battery lifespan.

Das, Avijit↗

MPRL (Multi-Pulse Reinforcement Learning)

MPRL contains code for defining and deploying reinforcement learning control agents for multi-pulse, multi-fuel advanced compression ignition engines.

Wimer, Nicholas↗

Extreme Risk Mitigation in Reinforcement Learning using Extreme Value Theory

Risk-sensitive reinforcement learning (RL) has garnered significant attention in recent years due to the growing interest in deploying RL agents in real-world scenarios. A critical aspect of risk awareness involves modelling highly rare risk events (rewards) that could potentially lead to catastrophic outcomes. These infrequent occurrences present a formidable challenge for data-driven methods aiming to capture such risky events accurately. While risk-aware RL techniques do exist, they suffer from high variance estimation due to the inherent data scarcity. Our work proposes to enhance the resilience of RL agents when faced with very rare and risky events by focusing on refining the predictions of the extreme values predicted by the state-action value distribution. To achieve this, we formulate the extreme values of the state-action value function distribution as parameterized distributions, drawing inspiration from the principles of extreme value theory (EVT). We propose an extreme value theory based actor-critic approach, namely, Extreme Valued Actor-Critic (EVAC) which effectively addresses the issue of infrequent occurrence by leveraging EVT-based parameterization. Importantly, we theoretically demonstrate the advantages of employing these parameterized distributions in contrast to other risk-averse algorithms. Our evaluations show that the proposed method outperforms other risk averse RL algorithms on a diverse range of benchmark tasks, each encompassing distinct risk scenarios.

Wang, Yu↗

Reinforcement Learning for Load-balanced Parallel Particle Tracing

We explore an online reinforcement learning (RL) paradigm to dynamically optimize parallel particle tracing performance in distributed-memory systems. Our method combines three novel components: (1) a work donation algorithm, (2) a high-order workload estimation model, and (3) a communication cost model. First, we design an RL-based work donation algorithm. Our algorithm monitors workloads of processes and creates RL agents to donate data blocks and particles from high-workload processes to low-workload processes to minimize program execution time. The agents learn the donation strategy on the fly based on reward and cost functions designed to consider processes' workload changes and data transfer costs of donation actions. Second, we propose a workload estimation model, helping RL agents estimate the workload distribution of processes in future computations. Third, we design a communication cost model that considers both block and particle data exchange costs, helping RL agents make effective decisions with minimized communication costs. We demonstrate that our algorithm adapts to different flow behaviors in large-scale fluid dynamics, ocean, and weather simulation data. Our algorithm improves parallel particle tracing performance in terms of parallel efficiency, load balance, and costs of I/O and communication for evaluations with up to 16,384 processors.

Distributed and parallel particle tracing↗

Reinforcement learning for bluff body active flow control in experiments and simulations

Significance Reinforcement learning (RL) has been applied effectively in games and robotic manipulation. We demonstrate the effectiveness of RL in experimental fluid mechanics by applying it to reduce the drag of circular cylinders in turbulent flow, a canonical fluid–structure interaction problem. Although physics agnostic, RL managed to reduce the drag by 30 % or reach another specified optimum point very quickly. Following this discovery, we used high-fidelity simulations to probe the underlying physical mechanisms so that the discovered control techniques can be generalized to other similar flow problems. More broadly, RL-guided active control can lead to efficient exploration of additional flow-control strategies in experimental fluid mechanics, potentially paving the way for accelerating scientific discovery and different designs in flow-related engineering problems.

42 ENGINEERING↗

Deep reinforcement learning for complex evaluation of one-loop diagrams in quantum field theory

In this paper we present a technique based on deep reinforcement learning that allows for numerical analytic continuation of integrals that are often encountered in one-loop diagrams in quantum field theory. To extract certain quantities of two-point functions, such as spectral densities, mass poles or multiparticle thresholds, it is necessary to perform an analytic continuation of the correlator in question. At one-loop level in Euclidean space, this results in the necessity to deform the integration contour of the loop integral in the complex plane of the square of the loop momentum, to avoid nonanalyticities in the integration plane. Using a toy model for which an exact solution is known, we train a reinforcement learning agent to perform the required contour deformations. We report our study shows great promise for an agent to be deployed in iterative numerical approaches used to compute nonperturbative two-point functions, such as the quark propagator Dyson-Schwinger equation, or more generally, Fredholm equations of the second kind, in the complex domain.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Microgrid energy scheduling under uncertain extreme weather: Adaptation from parallelized reinforcement learning agents

Microgrids are useful solutions for integrating renewable energy resources and providing seamless green electricity to minimize carbon footprint. In recent years, extreme weather events happened often worldwide and caused significant economic and societal losses. Such events bring uncertainties to the microgrid energy scheduling problems and increase the challenges of microgrid operation. Traditional optimization approaches suffer from the inaccuracy of the uncertain microgrid model and the unseen events. Existing reinforcement learning (RL) - based approaches are also hampered by the limited generalization and the increasing computational burden when stochastic formulations are required to accommodate the uncertainties. This paper proposes a new parallelized reinforcement learning (PRL) method based on the probabilistic events to handle the microgrid energy uncertainties. Specifically, several local learning agents are employed to interact with pertinent microgrid environments in a distributed manner and report outcomes to the global agent, which will optimize microgrid energy resources online during extreme events. The stochastic microgrid energy optimization problem is reformulated to include all possible scenarios with probabilities. The advantage estimate functions of learning agents are designed with a backward sweep to transfer the outcomes to the value function updating process. Two simulation studies, stochastic optimization and online testing, are performed to compare with several existing RL approaches. Results substantiate that the proposed PRL method can achieve up to 20% optimization performance improvement with 4 and 28 times less computation cost than Q-learning with experience replay and multi-agent Q-learning approaches, respectively.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The influence of exposure to early-life adversity on agency-modulated reinforcement learning

Agency beliefs influence how humans learn from different contexts and outcomes. Research demonstrates that stressors, such as exposure to early-life adversity (ELA), are associated with both agency beliefs and learning, but how these processes interact remains unclear. The current study investigated whether exposure to ELA influences agency and interacts with reinforcement learning in adults. Replicating prior behavioral and computational work, ELA resulted in decreased learning, while increased adversity severity was associated with decreased latent agency beliefs. These findings suggest that exposure to adversity in childhood has a nuanced impact on reinforcement learning and agency beliefs in adulthood.

Neurosciences & Neurology↗

Discovery of False Data Injection Attacks on Power Grid Frequency Controllers with Reinforcement Learning [Poster]

While inverter-based DER (distributed energy resources) are instrumental to integrating renewable energy into the power grid, they reduce the grid's mechanical inertia, thereby increasing the risk of frequency instabilities. To compensate for frequency instability risks, the grid must also undergo a transformation to include digital technologies that allow for two-way communication between the utility and customers. The current and future state of the power grid allows for building a cleaner energy landscape. However, the grid may also become vulnerable to novel cyber threats. To preemptively protect the power grid against elaborate cyber-attacks, we propose to discover potential threats via reinforcement learning. In this work, the focus is on studying false data injection attacks that target the control logic of frequency controllers. We show that a reinforcement learning agent can successfully discover how to best inject false data into linear droop controllers.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Quantum Reinforcement Learning for Volt-VAR Control in Power Distribution Systems

Volt-VAR control (VVC) is crucial in active distribution networks for optimizing voltage profiles and minimizing network losses. While traditional deep reinforcement learning (DRL) algorithms exhibit promise for VVC, they often require extensive computational resources to handle such a high-dimensional problem. As a potential solution, quantum reinforcement learning (QRL) algorithms integrate the computational capabilities of quantum computing into the DRL framework. However, existing QRL algorithms struggle with complex VVC problems due to the limitations of current quantum hardware. To bridge this gap, this paper proposes an innovative QRL algorithm featuring an end-to-end architecture that integrates a classical autoencoder, variational quantum circuits (VQCs), and classical post-processing layers. This design efficiently compresses high-dimensional grid states, enabling VQCs to leverage quantum advantages while producing multiple control device outputs tailored for VVC tasks. Numerical studies on three representative distribution systems verify the effectiveness and scalability of the proposed QRL algorithm, and demonstrate its enhanced performance over classical approaches with only approximately 1% of the parameters. Additionally, the robustness of our developed algorithm is validated through noisy quantum environments.

97 MATHEMATICS AND COMPUTING↗

Exploring electron beam induced atomic assembly via reinforcement learning in a molecular dynamics environment

We report atom-by-atom assembly of functional materials and devices is perceived as one of the ultimate targets of nanotechnology. Recently it has been shown that the beam of a scanning transmission electron microscope can be used for targeted manipulation of individual atoms. However, the process is highly dynamic in nature rendering control difficult. One possible solution is to instead train artificial agents to perform the atomic manipulation in an automated manner without need for human intervention. As a first step to realizing this goal, we explore how artificial agents can be trained for atomic manipulation in a simplified molecular dynamics environment of graphene with Si dopants, using reinforcement learning. We find that it is possible to engineer the reward function of the agent in such a way as to encourage formation of local clusters of dopants under different constraints. This study shows the potential for reinforcement learning in nanoscale fabrication, and crucially, that the dynamics learned by agents encode specific elements of important physics that can be learned.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Assessing Adaptive Irrigation Impacts on Water Scarcity in Nonstationary Environments—A Multi‐Agent Reinforcement Learning Approach

Abstract One major challenge in water resource management is to balance the uncertain and nonstationary water demands and supplies caused by the changing anthropogenic and hydroclimate conditions. To address this issue, we developed a reinforcement learning agent‐based modeling (RL‐ABM) framework where agents (agriculture water users) are able to learn and adjust water demands based on their interactions with the water systems. The intelligent agents are created by a reinforcement learning algorithm adapted from the Q‐learning algorithm. We illustrated this framework in a case study where the RL‐ABM is two‐way coupled with the Colorado River Simulation System (CRSS), a long‐term planning model used for the administration of the Colorado River Basin, for assessing agriculture water uses impacts on water scarcity. Seventy‐eight intelligent agents are simulated, which can be grouped into three categories based on their parameter values: the “aggressive” (swift actions; low regrets), the “forward‐looking conservative” (mild actions; high regrets; fast learning), and the “myopic conservative” (mild actions; median regrets; slow learning). The ABM‐CRSS results showed that the major reservoirs in the Upper Colorado Basin might experience more frequent water shortages due to the increasing water uses compared to the original CRSS results. If the drought continues, the case study also demonstrates that agents can learn and adjust their demands.

Hung, Fengwei↗

Buyers Collusion in Incentivized Forwarding Networks: A Multi-Agent Reinforcement Learning Study

We present the issue of monetarily incentivized forwarding in a multi-hop mesh network architecture from an economic perspective. It is anticipated that credit-incentivized forwarding and relaying will be a simple method of exchanging transmission power and spectrum for connectivity. However, gateways and forwarding nodes, like any other free market, may create an oligopolistic market for the users they serve. In this study, a coalition scheme between buyers aims to address price control by gateways or nodes closer to gateways. In a Stackelberg competition game, buyer agents (users) and sellers (gateways) make decisions using reinforcement learning (RL), with decentralized Deep Q-Networks to buy and sell forwarding resources. We allow communication links between the buyers with a limited messaging space, without defining a collusion mechanism. The idea is to demonstrate that through messaging, and RL tacit collusion can emerge between agents in a decentralized setup. The multi-agent reinforcement learning (MARL) system is presented and analyzed from a machine-learning perspective. Moreover, MARL dynamics are discussed via mean field analysis to better understand divergence causes and make implementation recommendations for such systems. Finally, the simulation results show the results of coordination among the users.

42 ENGINEERING↗