Engineering PapersSearch

SEARCH · Engineering Papers

Results for “RL”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Artificial-intelligence-driven shot reduction in quantum measurement

Variational Quantum Eigensolver (VQE) provides a powerful solution for approximating molecular ground state energies by combining quantum circuits and classical computers. However, estimating probabilistic outcomes on quantum hardware requires repeated measurements (shots), incurring significant costs as accuracy increases. Optimizing shot allocation is thus critical for improving the efficiency of VQE. Current strategies rely heavily on hand-crafted heuristics requiring extensive expert knowledge. This paper proposes a reinforcement learning (RL)-based approach that automatically learns shot assignment policies to minimize total measurement shots while achieving convergence to the minimum of the energy expectation in VQE. The RL agent assigns measurement shots across VQE optimization iterations based on the progress of the optimization. This approach reduces VQE's dependence on static heuristics and human expertise. When the RL-enabled VQE is applied to a small molecule, a shot reduction policy is learned. The policy demonstrates transferability across systems and compatibility with other wavefunction Ansätze. In addition to these specific findings, this work highlights the potential of RL for automatically discovering efficient and scalable quantum optimization strategies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Reinforcement Learning for Anomaly Detection in Nuclear Power Plant Operation and Maintenance

In nuclear power plants (NPPs), timely identification of sensor and human errors is critical to ensure safe and efficient plant operations. Anomaly detection models can be employed for this task. However, traditional anomaly detection approaches may have high dependency on labeled datasets and struggle with adaptability in complex, dynamic environments. Reinforcement learning (RL) has demonstrated significant potential in fault diagnosis and anomaly detection; however, its application to anomaly detection in NPPs remains a relatively underexplored research direction. Hence, to address this gap, in this study, we present a novel physics-informed reinforcement learning model, PIRL-AD: Physics-Informed Reinforcement Learning for Anomaly Detection, that integrates domain knowledge from calorimetric equations into the RL framework for enhanced sensor and human error anomaly detection. We evaluate the performance of PIRL-AD against a non-physics informed RL benchmark and a support vector machine (SVM) on data collected from a forced flow loop testbed. Experimental results suggest that PIRL-AD outperforms other baselines on a range of anomalous datasets that include both sensor and human-induced anomalies across key performance metrics, statistically outperforming the RL and SVM benchmarks with respect to geometric mean (respectively, 92.96% vs. 91.06% vs. 83.01%) and F1-score (respectively, 89.23% vs. 86.98% vs. 77.01%). Furthermore, the findings suggest the potential of physics-integrated reinforcement learning models for enhanced anomaly detection performance in NPPs.

Reinforcement learning

Source function from two-particle correlation function through entropy-regularized Richardson-Lucy deblurring

Source functions are obtained from p – p and d – α correlation functions by applying the Richardson-Lucy (RL) deblurring to the Koonin-Pratt (KP) equation. To prevent fitting of noise in the correlation function, total-variation (TV) regularization is employed that has been effective in ordinary image restoration. TV alone cannot ensure normalization of the source functions. To ensure the latter, we propose a maximum-entropy regularized RL algorithm (MEM-RL). We outline the MEM-RL formalism and optimization strategy for the KP equation, demonstrating its effectiveness on both simulated and experimental data, including the p – p and d – α correlation functions.

62 RADIOLOGY AND NUCLEAR MEDICINE

Optimal Management of Grid-Interactive Efficient Buildings via Safe Reinforcement Learning

Reinforcement learning (RL)-based methods have achieved significant success in managing grid-interactive efficient buildings (GEBs). However, RL does not carry intrinsic guarantees of constraint satisfaction, which may lead to severe safety consequences. Besides, in GEB control applications, most existing safe RL approaches rely only on the regularisation parameters in neural networks or penalty of rewards, which often encounter challenges with parameter tuning and lead to catastrophic constraint violations. To provide enforced safety guarantees in controlling GEBs, this paper designs a physics-inspired safe RL method whose decision-making is enhanced through safe interaction with the environment. Different energy resources in GEBs are optimally managed to minimize energy costs and maximize customer comfort. The proposed approach can achieve strict constraint guarantees based on prior knowledge of a set of developed hard steady-state rules. Simulations on the optimal management of GEBs, including heating, ventilation, and air conditioning (HVAC), solar photovoltaics, and energy storage systems, demonstrate the effectiveness of the proposed approach.

Huo, Xiang

Optimal Coordination of Electric Vehicles for Grid Services using Deep Reinforcement Learning

Recent research has shown the effectiveness of reinforcement learning (RL) in coordinating electric vehicles (EVs) with vehicle-to-grid capabilities for grid services. However, many of these studies rely on lookup table and deep Q-network techniques, which can be impractical when dealing with continuous states and actions. In addition, existing RL designs inadequately account for battery aging effects, EV user satisfaction, uncertain departure and arrival time, and trip distance, which may compromise effective coordination. This paper aims to bridge these gaps by developing an innovative deep deterministic policy gradient-based RL framework for optimal coordination of EVs. Case studies were carried out using a test system with 100 EVs, and numerical analysis results showed that the proposed RL framework can effectively coordinate EVs to maximize economic benefits and user satisfaction while ensuring the expected battery lifespan.

Das, Avijit

Nonlinear evolution of high frequency R-mode waves excited by water group ions near comets - Computer experiments

An ion beam resonates with R-mode waves at a high-frequency RH mode and a low-frequency RL mode. The nonlinear evolution of ion beam-generated RH waves is studied here by one-dimensional hybrid computer experiments. Both wave-particle and subsequent wave-wave interactions are examined. The competing process among coexisting RH and RL mode beam instabilities and repeated decay instabilities triggered by the beam-excited RH mode waves is clarified. It is found that the quenching of the RH instability is not caused by a thermal spreading of the ion beam, but by the nonlinear wave-wave coupling process. The growing RH waves become unstable against the decay instability. This instability involves a backward-traveling RH electromagnetic wave and a forward-traveling longitudinal sound wave. The inverse cascading process is found to occur faster than the growth of the RL mode. Wave spectra decaying from the RH waves weaken as time elapses and the RL mode waves become dominant at the end of the computer experiment.

Kojima, H.

Dynamic Altitude Simulation System Performance Modeling

The "Dynamic Altitude Simulation Prediction Program," DASSPP, is a program to predict the transient response of an engine, test cell, ejector system under engine shutdown conditions. These transients are important to know so that corrective modifications can be adapted to prevent any damage to the engine or test cell. The "Dynamic Altitude System Simulation Prediction Program," DASSPP, is a major rewrite of the existing program "RL-1000" written in BASIC. The RL-1000 program was written to analyze the transients of the RL-10 system only. The new program is written to run in Excell 97 and utilizes the Visual BASIC language in Excell. The program has many added features not included in the original "Rl-1000" program. The program utilizes the ejector models developed during the summer of 1997. The new program is very user friendly and utilizes a dialog box for data input.

LaFrance, Leo J.

Safety Analysis of FMS/CTAS Interactions During Aircraft Arrivals

This grant funded research on human-computer interaction design and analysis techniques, using future ATC environments as a testbed. The basic approach was to model the nominal behavior of both the automated and human procedures and then to apply safety analysis techniques to these models. Our previous modeling language, RSML, had been used to specify the system requirements for TCAS II for the FAA. Using the lessons learned from this experience, we designed a new modeling language that (among other things) incorporates features to assist in designing less error-prone human-computer interactions and interfaces and in detecting potential HCI problems, such as mode confusion. The new language, SpecTRM-RL, uses "intent" abstractions, based on Rasmussen's abstraction hierarchy, and includes both informal (English and graphical) specifications and formal, executable models for specifying various aspects of the system. One of the goals for our language was to highlight the system modes and mode changes to assist in identifying the potential for mode confusion. Three published papers resulted from this research. The first builds on the work of Degani on mode confusion to identify aspects of the system design that could lead to potential hazards. We defined and modeled modes differently than Degani and also defined design criteria for SpecTRM-RL models. Our design criteria include the Degani criteria but extend them to include more potential problems. In a second paper, Leveson and Palmer showed how the criteria for indirect mode transitions could be applied to a mode confusion problem found in several ASRS reports for the MD-88. In addition, we defined a visual task modeling language that can be used by system designers to model human-computer interaction. The visual models can be translated into SpecTRM-RL models, and then the SpecTRM-RL suite of analysis tools can be used to perform formal and informal safety analyses on the task model in isolation or integrated with the rest of the modeled system. We had hoped to be able to apply these modeling languages and analysis tools to a TAP air/ground trajectory negotiation scenario, but the development of the tools took more time than we anticipated.

Nancy G. Leveson

A Survey of Collective Intelligence

This chapter presents the science of "COllective INtelligence" (COIN). A COIN is a large multi-agent systems where: i) the agents each run reinforcement learning (RL) algorithms; ii) there is little to no centralized communication or control; iii) there is a provided world utility function that, rates the possible histories of tile full system. Tile conventional approach to designing large distributed systems to optimize a world utility does not use agents running RL algorithms. Rather that approach begins with explicit modeling of the overall system's dynamics, followed by detailed hand-tuning of the interactions between the components to ensure that they "cooperate" as far as the world utility is concerned. This approach is labor-intensive, often results in highly non-robust systems, and usually results in design techniques that, have limited applicability. In contrast, with COINs we wish to solve the system design problems implicitly, via the 'adaptive' character of the RL algorithms of each of the agents. This COIN approach introduces an entirely new, profound design problem: Assuming the RL algorithms are able to achieve high rewards, what reward functions for the individual agents will, when pursued by those agents, result in high world utility? In other words, what reward functions will best ensure that we do not have phenomena like the tragedy of the commons, or Braess's paradox? Although still very young, the science of COINs has already resulted in successes in artificial domains, in particular in packet-routing, the leader-follower problem, and in variants of Arthur's "El Farol bar problem". It is expected that as it matures not only will COIN science expand greatly the range of tasks addressable by human engineers, but it will also provide much insight into already established scientific fields, such as economics, game theory, or population biology.

Wolpert, David H.

Shape deformation of the organ of Corti associated with length changes of outer hair cell

Cochlear outer hair cells (OHC) are commonly assumed to function as mechanical effectors as well as sensory receptors in the organ of Corti (OC) of the inner ear. OHC in vitro and in organ explants exhibit mechanical responses to electrical, chemical or mechanical stimulation which may represent an aspect of their effector process that is expected in vivo. A detailed description, however, of an OHC effector operation in situ is still missing. Specifically, little is known as to how OHC movements influence the geometry of the OC in situ. Previous work has demonstrated that the motility of isolated OHCs in response to electrical stimulation and to K(+)-gluconate is probably under voltage control and causes depolarisation (shortening) and hyperpolarization (elongation). This work was undertaken to investigate if the movements that were observed in isolated OHC, and which are induced by ionic stimulation, could change the geometry of the OC. A synchronized depolarization of OHC was induced in guinea pig cochleae by exposing the entire OC to artificial endolymph (K+). Subsequent morphometry of mid-modiolar sections from these cochleae revealed that the distance between the basilar membrane (BM) and the reticular lamina (RL) had decreased considerably. Furthermore, in the three upper turns OHC had significantly shortened in all rows. The results suggest that OHC can change their length in the organ of Corti (OC) thus deforming the geometry of the OC. The experiments reveal a tonic force generation within the OC that may change the position of RL and/or BM, contribute to damping, modulate the BM-RL-distance and control the operating points of RL and sensory hair bundles. Thus, the results suggest active self-adjustments of cochlear mechanics by slow OHC length changes. Such mechanical adjustments have recently been postulated to correspond to timing elements of animal communication, speech or music.

Non-NASA Center

Broadband and Tunable Microwave Absorption Properties from Large Magnetic Loss in Ni–Zn Ferrite

Highly effective electromagnetic (EM) wave absorber materials with strong reflection loss (RL) and a wide absorption bandwidth (EBW) in gigahertz (GHz) frequencies are crucial for advanced wireless applications and portable electronics. Traditional microwave absorbers lack magnetic loss and struggle with impedance matching, while ferrites are stable, exhibit excellent magnetic and dielectric losses, and offer better impedance matching. However, achieving the desired EBW in ferrites remains a challenge, necessitating further composition design. In this study, impedance matching is successfully enhanced and EBW in Ni–Zn ferrite is broadened by successive doping with Mn and Co , without incorporation of any polymer filler. It is found that Ni 0.4 Co 0.1 Zn 0.5 Fe 1.9 Mn 0.1 O 4 material exhibits exceptional EM wave absorption, with a maximum RL of −48.7 dB. It also featured a significant EBW of 10.8 GHz, maintaining a 90% absorption rate (RL < −10 dB) for a thickness of 4.5 mm. These outstanding properties result from substantial magnetic losses and favorable impedance matching. These findings represent a significant step forward in the development of microwave absorber materials, addressing EM wave pollution concerns within GHz frequencies, including the frequency band used in popular 5G technology.

36 MATERIALS SCIENCE

A safe reinforcement learning algorithm for supervisory control of power plants

Traditional control theory-based methods require tailored engineering for each system and constant fine-tuning. In power plant control, one often needs to obtain a precise representation of the system dynamics and carefully design the control scheme accordingly. Model-free Reinforcement learning (RL) has emerged as a promising solution for control tasks due to its ability to learn from trial-and-error interactions with the environment. It eliminates the need for explicitly modeling the environment’s dynamics, which is potentially inaccurate. However, the direct imposition of state constraints in power plant control raises challenges for standard RL methods. To address this, we propose a chance-constrained RL algorithm based on Proximal Policy Optimization for supervisory control. Our method employs Lagrangian relaxation to convert the constrained optimization problem into an unconstrained objective, where trainable Lagrange multipliers enforce the state constraints. In conclusion, our approach achieves the smallest distance of violation and violation rate in a load-follow maneuver for an advanced Nuclear Power Plant design.

constrained optimization

Hierarchical Reinforcement Learning of a Short-Range Bond-Order Potential for Silica: Analytic Embedding of Coordination with Classical Efficiency

Reinforcement learning (RL) has recently emerged as a data-efficient strategy to parametrize short-range interatomic potentials. Building on our past RL optimization of pairwise silica models, we extend the framework to a bond-order (Tersoff-type) potential that provides an analytic embedding of local coordination through a three-body term. A hierarchical RL workflow combining continuous-action Monte Carlo Tree Search and property-based rewards efficiently explores the 26-dimensional parameter space, sequentially optimizing lattice parameters, densities, angles, and cohesive energies of 21 silica polymorphs. The resulting models, Q-Tersoff and ML-Tersoff, reproduce the energetic ordering of low-energy phases and capture the angular correlations and amorphous structure factors of silica with improved fidelity over pairwise force fields, while remaining orders of magnitude faster than high-dimensional machine-learned potentials. Both models underperform for elastic constants and high-energy frameworks, delineating the limits of the current analytic form. The approach establishes a general and interpretable route to angle-aware, short-range potentials that bridge physics-based and machine-learned descriptions of silicate materials.

36 MATERIALS SCIENCE

Effects of Quantum Dot Loading on the Radioluminescence Efficiency in Quantum-Dot-Embedded Composites

Nanoparticle-embedded plastic scintillators are an emerging technology for fast, large-area, high-resolution radiation detection and imaging. Here, this study investigates the properties of such composites, focusing on the effects of the quantum dot (QD) concentration on the radioluminescence (RL) intensity, spectra, and dynamics. Experiments using CdSe/CdS QDs in a polymer reveal a superlinear increase in RL with the QD concentration despite optical losses from inner filtering and interparticle interactions. When corrected for inner filtering, RL shows a quadratic concentration dependence, consistent with simple analytical models of improving the secondary electron capture. Practically, the benefits of high QD concentrations are muted by optical losses, but the findings apply to other systems with insulating hosts. In addition to manipulating emission for large effective Stokes shifts, future improvements may come from hosts with higher stopping power and better charge transport, which enable more effective funneling of excitations but without concomitant optical losses associated with high nanoparticle concentrations.

Auger recombination

Automated Construction of Artificial Lattice Structures with Designer Electronic States

Manipulating matter with a scanning tunneling microscope (STM) enables the creation of atomically defined artificial structures that host designer quantum states. However, the time-consuming nature of the manipulation process, coupled with the sensitivity of the STM tip, constrains the exploration of diverse configurations and limits the size of the designed features. In this study, we present a reinforcement learning (RL)-based framework for creating artificial structures by spatially manipulating carbon monoxide (CO) molecules on a copper substrate by using the STM tip. The automated workflow combines molecule detection and manipulation, employing deep-learning-based object detection to locate CO molecules and linear assignment algorithms to allocate these molecules to designated target sites. We initially perform molecule maneuvering based on randomized parameter sampling for sample bias, tunneling current set point, and manipulation speed. This data set is then structured into an action trajectory used to train an RL agent. The model is subsequently deployed on the STM for real-time fine-tuning of the manipulation parameters during structure construction. Our approach incorporates path-planning protocols coupled with active drift compensation to enable atomically precise fabrication of structures with significantly reduced human input while realizing larger-scale artificial lattices with the desired electronic properties. Furthermore, using our approach, we demonstrate the automated construction of an extended artificial graphene lattice and confirm the existence of a characteristic Dirac point in its electronic structure. Further challenges regarding the RL-based structural assembly scalability are discussed.

Algorithms

The effects of photosynthetic rate on respiration in light, starch/sucrose partitioning, and other metabolic fluxes within photosynthesis

In the future, plants may encounter increased light and elevated CO 2 levels. How consequent alterations in photosynthetic rates will impact fluxes in photosynthetic carbon metabolism remains uncertain. Respiration in light ( R L ) is pivotal in plant carbon balance and a key parameter in photosynthesis models. Understanding the dynamics of photosynthetic metabolism and R L under varying environmental conditions is essential for optimizing plant growth and agricultural productivity. However, measuring R L under high light and high CO 2 (HLHC) conditions poses challenges using traditional gas exchange methods. In this study, we employed isotopically nonstationary metabolic flux analysis (INST-MFA) to estimate RL and investigate photosynthetic carbon flux, unveiling nuanced adjustments in Camelina sativa under HLHC. Despite numerous flux alterations in HLHC, RL remained stable. HLHC affects several factors influencing RL, such as starch and sucrose partitioning, v o /v c ratio, triose phosphate partitioning, and hexose kinase activity. Analysis of A/C i curve operational points reveals that HLHC’s major changes primarily stem from CO 2 suppressing photorespiration. Integration of these fluxes into a simplified model predicts changes in CBC labeling under HLHC. This study extends our prior discovery that incomplete CBC labeling is due to unlabeled carbon reimported during R L , offering insights into manipulating labeling through adjustments in photosynthetic rates.

Elevated CO2

Deep Multi-Agent Reinforcement Learning for Real-World Signalized Traffic Corridor Control

Signalized traffic control problem has been addressed recently with deep Reinforcement Learning (RL) approaches involving diverse state, action, and reward structures. While significant progress has been noted in the literature, open challenges still remain in the areas of adaptive signal phase timing, coordination in a multi-intersection corridor setting, and consideration of real-world traffic conditions. In the context of deep RL-based problem framing, extensions are needed that enable adaptive signal phase timings in an intersection agent's action space, computationally efficient information sharing among neighboring signalized intersection agents along a corridor, and experimentation in realistic simulation environments. In this paper, we develop a deep Advantage Actor Critic (A2C) multi-agent RL (MARL) approach capturing the research extensions above and apply it within a real-world calibrated Aimsun Next traffic corridor simulation model based on traffic data from the City of Coral Gables, Florida. For a multi-intersection corridor control setting, our numerical simulation experiments with a decentralized A2C MARL algorithm applied at different time periods led to a total average corridor travel delay reduction (expressed in seconds/mile averaged over vehicles) from 4.9% to 19.9% compared to state-of-the-art actuated control.

Shuvo, Salman S. [BATTELLE (PACIFIC NW LAB)]

Reinforcement Learning-Based Approach for EMT Automation of Large-Scale PV Plants

In the pursuit of efficient and precise modeling of large-scale power systems, particularly utility-scale photovoltaic (PV) plants, Electromagnetic Transient (EMT) simulations play a crucial role. As utility-scale PV plants increase in size and complexity, traditional computational methods become inadequate, necessitating more advanced techniques. This paper highlights the progressive efforts made to accelerate EMT simulations. A novel continuous reinforcement learning (RL) strategy is explored to automate the differentiation and categorization of stiff and non-stiff differential algebraic equations (DAEs). The use of stiff and non-stiff integration methods applied to relevant parts of the DAEs assists with the speed-up of the simulations. The paper details the data acquisition, development and offline training of the RL model, leading to its validation that demonstrates a high precision in optimizing simulation methods. The proposed RL promises to significantly enhance the efficacy of EMT simulations, offering a robust framework for the future of power system analysis.

Xia, Qianxue