Engineering PapersSearch

SEARCH · Engineering Papers

Results for “RL”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Enhancing Autonomous Control of Microreactors Using Multi-Agent Reinforcement Learning

In order for microreactors to be economically competitive, operation costs will need to be minimized through some degree of autonomous control. Previous work has demonstrated the effectiveness of reinforcement learning (RL) for load-following control in a drum-controlled microreactor. This study extends that work by exploring the potential of RL to independently control each of the reactor’s drums. We compare a single-agent RL approach with a multi-agent RL (MARL) framework, testing them for generalization across different load-following power profiles and control timescales, and for robustness in cases of randomly disabled control drums. Since the point kinetics simulation environment used in this study cannot resolve spatial effects, we assume that in the absence of spatially localized disturbances, optimal drum movements should be symmetrical. We demonstrate that single-agent RL is able to achieve accurate performance only when symmetric actions are ignored; otherwise, it fails to train a useful controller. Meanwhile, the MARL framework performs symmetric actions by design and trains a robust, accurate agent, as evidenced by mean absolute errors in power matching of 0.41% for the training power profile, 0.68% for a profile with half the drums disabled, and 0.21% for a profile on a realistic load-following time horizon.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Designing reinforcement learning algorithms for building HVAC control: From experimental observation to simulation comparisons

Advanced supervisory-level control with reinforcement learning (RL) is regarded as a promising solution for HVAC systems to minimize energy consumption while maintaining thermal comfort and indoor air quality. However, most RL applications were conducted in the simulation environment rather than real-world HVAC systems. This paper developed a value-based RL controller termed Deep Q-Network (DQN) for a typical central HVAC system and evaluated its performance in a building test facility. By comparing DQN with a rule-based controller, the study not only demonstrated the cases where DQN could properly maintain indoor comfort but also discussed possible reasons why DQN failed in some other situations. Recognizing the limitations of value-based RL algorithms from the experimental tests, a simulation study was conducted to compare DQN with an alternative RL approach, an actor–critic algorithm termed Deep Deterministic Policy Gradient (DDPG). In scenarios with a relatively large action space, DDPG outperformed DQN by requiring fewer computational resources and achieving better thermal comfort, lower energy consumption, and more stable control actions. The findings suggest that the ability of DDPG to handle continuous control variables more effectively allows for faster convergence in training and more precise control in practice, which enhances the overall efficiency and reliability of the HVAC system.

Guo, Fangzhou

Energy performance evaluation of the ASHRAE Guideline 36 control and reinforcement learning–based control using field measurements

This study evaluates the energy performance of ASHRAE Guideline 36–compliant control (ASHRAE 36 control) and reinforcement learning (RL)–based control through experimental field tests and a simulation study. Three field tests were conducted at Oak Ridge National Laboratory’s commercial building test facility in Oak Ridge, Tennessee: a baseline with a baseline conventional control, a test with ASHRAE 36 control, and a test with RL-based control. The selected ASHRAE 36 controls were trim and respond control, as well as variable air volume (VAV) box control. We compared the measured supply air temperature of the rooftop unit, VAV box supply air temperature, and VAV box supply airflow rate across the three test cases. The field data indicated that ASHRAE 36 controls operated as specified by ASHRAE Guideline 36. Based on these data, ASHRAE 36 control achieved a 45 % reduction in hourly averaged HVAC energy consumption compared with the baseline, and RL-based control achieved a 66 % reduction. These potential annual energy savings were confirmed using a calibrated whole-building energy model. Compared with the baseline, ASHRAE 36 control reduced HVAC energy consumption by 42 %, and RL-based control achieved a 54 % reduction. Furthermore, RL-based control reduced total HVAC energy consumption by 21 % more than ASHRAE 36 control.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Reinforcement Learning‐Based Adaptation of Grid Following Inverter's Internal Controller to Networked Microgrids' Strengths

The varying topological configurations, generator commitments and dispatches, and dynamic load demand lead to changing system's strengths during the operations of networked microgrids. When the system's strengths significantly change, the fixed control gains at large devices may result in unsatisfactory system performance; this necessitates the tuning of the control gains at large devices to adapt to the changing system's strengths. In this paper, observer-based reinforcement learning (RL) is utilised to automatically tune the proportional-integral (PI) gains of phase lock loop (PLL) controller of grid-following (GFL) inverters to adapt to the changing strengths of microgrids and networked microgrids. The RL agent in this framework augments an observer predicting system's strengths, from which the RL control policy will adjust accordingly to tune the PLL controller's gains towards the system's strengths. Also, to enhance the control performance, the recently introduced Barrier function-based RL framework is leveraged for the design of reward function to prevent the high frequency nadir. An operational 26 kV electric distribution system, which is modelled as networked microgrids, is used to illustrate the need and effectiveness of the proposed RL-tuned control.

frequency response

Microsecond-latency feedback at a particle accelerator by online reinforcement learning on hardware

The commissioning and operation of future large-scale scientific experiments will challenge current tuning and control methods. Reinforcement learning (RL) algorithms are a promising solution due to their ability to dynamically adapt to changing environments and consider delayed consequences. In many real-world applications, RL policies must produce actions in real time, often within microseconds to milliseconds, imposing significant constraints on system latency and computational overhead that conventional machine learning libraries are not designed to handle. To control phenomena in real time at these timescales, RL needs to be deployed on-the-edge, namely on dedicated hardware located near the system it controls, without relying on a host CPU or cloud-based inference. In this work we present the design and deployment of an experience accumulator system in a particle accelerator. In this system, deep-RL algorithms run using hardware acceleration and act within a few microseconds, enabling the use of RL for control of phenomena like beam instabilities. The training uses the collected data offline to reduce the number of operations carried out on the acceleration hardware. The proposed architecture was tested in real experimental conditions at the Karlsruhe research accelerator, a synchrotron light source, where the system was used to control artificially induced horizontal betatron oscillations in real-time, with a control loop period of just 2.7 μs. The results showed a performance comparable to the commercial feedback system available at the accelerator, demonstrating the viability and potential of this approach. Due to the self-learning and reconfiguration capability of this implementation, a seamless application to other control problems is possible. Applications range from particle accelerators to large-scale research and industrial facilities.

FPGA

Townsend’s Ground Squirrel Conservation on the Hanford Site for Calendar Year 2023: Translocation to Support At-Risk Ground Squirrel Populations

The U.S. Department of Energy, Richland Operations Office (RL) conducts ecological monitoring on the Hanford Site to collect and track data needed to ensure compliance with an array of environmental laws, regulations, and policies governing RL activities. Ecological monitoring data provide baseline information about the plants, animals, and habitat under RL stewardship at the Hanford Site, which is required for accurate ecological impact assessment decision making under the National Environmental Policy Act and the Comprehensive Environmental Response, Compensation, and Liability Act. In addition, ecological monitoring helps ensure that RL, its contractors, and other entities conducting activities on the Hanford Site are in compliance with DOE/EIS-0222-F, Final Hanford Comprehensive Land Use Plan Environmental Impact Statement. RL places priority on monitoring those plant and animal species or habitats with specific regulatory protections or requirements; or that are rare and/or declining (federal or state listed endangered, threatened, or sensitive species) or of significant interest to federal, state, or tribal governments or the public.

54 ENVIRONMENTAL SCIENCES

Interplanetary Shocks Lacking Type 2 Radio Bursts

We report on the radio-emission characteristics of 222 interplanetary (IP) shocks detected by spacecraft at Sun-Earth L1 during solar cycle 23 (1996 to 2006, inclusive). A surprisingly large fraction of the IP shocks (approximately 34%) was radio quiet (RQ; i.e., the shocks lacked type II radio bursts). We examined the properties of coronal mass ejections (CMEs) and soft X-ray flares associated with such RQ shocks and compared them with those of the radio-loud (RL) shocks. The CMEs associated with the RQ shocks were generally slow (average speed approximately 535 km/s) and only approximately 40% of the CMEs were halos. The corresponding numbers for CMEs associated with RL shocks were 1237 km/s and 72%, respectively. Thus, the CME kinetic energy seems to be the deciding factor in the radio-emission properties of shocks. The lower kinetic energy of CMEs associated with RQ shocks is also suggested by the lower peak soft X-ray flux of the associated flares (C3.4 versus M4.7 for RL shocks). CMEs associated with RQ CMEs were generally accelerating within the coronagraph field of view (average acceleration approximately +6.8 m/s (exp 2)), while those associated with RL shocks were decelerating (average acceleration approximately 3.5 m/s (exp 2)). This suggests that many of the RQ shocks formed at large distances from the Sun, typically beyond 10 Rs, consistent with the absence of metric and decameter-hectometric (DH) type II radio bursts. A small fraction of RL shocks had type II radio emission solely in the kilometric (km) wavelength domain. Interestingly, the kinematics of the CMEs associated with the km type II bursts is similar to those of RQ shocks, except that the former are slightly more energetic. Comparison of the shock Mach numbers at 1 AU shows that the RQ shocks are mostly subcritical, suggesting that they were not efficient in accelerating electrons. The Mach number values also indicate that most of these are quasi-perpendicular shocks. The radio-quietness is predominant in the rise phase and decreases through the maximum and declining phases of solar cycle 23. About 18% of the IP shocks do not have discernible ejecta behind them. These shocks are due to CMEs moving at large angles from the Sun-Earth line and hence are not blast waves. The solar sources of the shock-driving CMEs follow the sunspot butterfly diagram, consistent with the higher-energy requirement for driving shocks.

Gopalswamy, N.

A transfer learning approach to energy-efficient control of small and medium-sized commercial buildings

Model-free reinforcement learning (RL) provides a data-driven and adaptive approach to optimize building energy use while satisfying occupant comfort. This powerful tool does not need any prior knowledge about the environment and system it is optimizing and can adapt its policy based on the changes in captures. Like any other data-driven tool, it faces high training costs due to the extensive agent-environment interactions required to capture long-term building dynamics and user comfort. Transfer learning, particularly policy distillation, offers a promising way to accelerate training by leveraging pretrained RL agents in different building and system types. Here, this study investigates online student distillation, in which the student model updates its neural network weights using outputs from teacher models. The work introduces a student distillation strategy designed for efficient knowledge transfer, along with a teacher selection method that ensures high-quality guidance. The approach is validated using a highly calibrated whole building energy model for a small/medium commercial building test facility. Results show substantial reductions in training time and data requirements while surpassing the performance of ASHRAE Guideline 36, an advanced rule-based control strategy. The distilled RL model required 45% less data and achieved 20% higher cumulative rewards than a state-of-the-art RL model, with faster convergence and lower energy consumption. These outcomes demonstrate that effective transfer learning enables a scalable and data-efficient energy management solution for commercial buildings.

ASHRAE guideline 36

RLGBS: Reinforcement Learning-Guided Beam Search for process optimization in a paper machine dryer section

Paper drying is responsible for over two-thirds of energy consumption in the U.S. pulp and paper industry, presenting significant potential for energy savings through optimization of process parameters. Current approaches often assume fixed operating conditions, neglecting dynamic ambient and process variations that limit achievable savings and real-world applicability. To this end, we develop a physics-based simulation environment for a paper machine dryer section and propose a reinforcement learning (RL) framework to minimize overall energy consumption by optimizing drying process parameters under diverse operating conditions. To mitigate overdrying and numerical instabilities caused by suboptimal local RL actions, we introduce Reinforcement Learning-Guided Beam Search (RLGBS), which explores multiple action sequences in parallel using beam search. Instead of making step-by-step decisions, RLGBS prioritizes solutions based on cumulative probability, reducing the impact of individual suboptimal actions. Experiments demonstrate that RLGBS achieves consistent energy savings under unseen operating conditions not encountered during training, outperforming conventional RL methods. While validated in drying optimization, this framework is broadly applicable to other RL-based industrial process control problems.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Artificial-intelligence-driven shot reduction in quantum measurement

Variational Quantum Eigensolver (VQE) provides a powerful solution for approximating molecular ground state energies by combining quantum circuits and classical computers. However, estimating probabilistic outcomes on quantum hardware requires repeated measurements (shots), incurring significant costs as accuracy increases. Optimizing shot allocation is thus critical for improving the efficiency of VQE. Current strategies rely heavily on hand-crafted heuristics requiring extensive expert knowledge. This paper proposes a reinforcement learning (RL)-based approach that automatically learns shot assignment policies to minimize total measurement shots while achieving convergence to the minimum of the energy expectation in VQE. The RL agent assigns measurement shots across VQE optimization iterations based on the progress of the optimization. This approach reduces VQE's dependence on static heuristics and human expertise. When the RL-enabled VQE is applied to a small molecule, a shot reduction policy is learned. The policy demonstrates transferability across systems and compatibility with other wavefunction Ansätze. In addition to these specific findings, this work highlights the potential of RL for automatically discovering efficient and scalable quantum optimization strategies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Reinforcement Learning for Anomaly Detection in Nuclear Power Plant Operation and Maintenance

In nuclear power plants (NPPs), timely identification of sensor and human errors is critical to ensure safe and efficient plant operations. Anomaly detection models can be employed for this task. However, traditional anomaly detection approaches may have high dependency on labeled datasets and struggle with adaptability in complex, dynamic environments. Reinforcement learning (RL) has demonstrated significant potential in fault diagnosis and anomaly detection; however, its application to anomaly detection in NPPs remains a relatively underexplored research direction. Hence, to address this gap, in this study, we present a novel physics-informed reinforcement learning model, PIRL-AD: Physics-Informed Reinforcement Learning for Anomaly Detection, that integrates domain knowledge from calorimetric equations into the RL framework for enhanced sensor and human error anomaly detection. We evaluate the performance of PIRL-AD against a non-physics informed RL benchmark and a support vector machine (SVM) on data collected from a forced flow loop testbed. Experimental results suggest that PIRL-AD outperforms other baselines on a range of anomalous datasets that include both sensor and human-induced anomalies across key performance metrics, statistically outperforming the RL and SVM benchmarks with respect to geometric mean (respectively, 92.96% vs. 91.06% vs. 83.01%) and F1-score (respectively, 89.23% vs. 86.98% vs. 77.01%). Furthermore, the findings suggest the potential of physics-integrated reinforcement learning models for enhanced anomaly detection performance in NPPs.

Reinforcement learning

Source function from two-particle correlation function through entropy-regularized Richardson-Lucy deblurring

Source functions are obtained from p – p and d – α correlation functions by applying the Richardson-Lucy (RL) deblurring to the Koonin-Pratt (KP) equation. To prevent fitting of noise in the correlation function, total-variation (TV) regularization is employed that has been effective in ordinary image restoration. TV alone cannot ensure normalization of the source functions. To ensure the latter, we propose a maximum-entropy regularized RL algorithm (MEM-RL). We outline the MEM-RL formalism and optimization strategy for the KP equation, demonstrating its effectiveness on both simulated and experimental data, including the p – p and d – α correlation functions.

62 RADIOLOGY AND NUCLEAR MEDICINE

Optimal Management of Grid-Interactive Efficient Buildings via Safe Reinforcement Learning

Reinforcement learning (RL)-based methods have achieved significant success in managing grid-interactive efficient buildings (GEBs). However, RL does not carry intrinsic guarantees of constraint satisfaction, which may lead to severe safety consequences. Besides, in GEB control applications, most existing safe RL approaches rely only on the regularisation parameters in neural networks or penalty of rewards, which often encounter challenges with parameter tuning and lead to catastrophic constraint violations. To provide enforced safety guarantees in controlling GEBs, this paper designs a physics-inspired safe RL method whose decision-making is enhanced through safe interaction with the environment. Different energy resources in GEBs are optimally managed to minimize energy costs and maximize customer comfort. The proposed approach can achieve strict constraint guarantees based on prior knowledge of a set of developed hard steady-state rules. Simulations on the optimal management of GEBs, including heating, ventilation, and air conditioning (HVAC), solar photovoltaics, and energy storage systems, demonstrate the effectiveness of the proposed approach.

Huo, Xiang

Nonlinear evolution of high frequency R-mode waves excited by water group ions near comets - Computer experiments

An ion beam resonates with R-mode waves at a high-frequency RH mode and a low-frequency RL mode. The nonlinear evolution of ion beam-generated RH waves is studied here by one-dimensional hybrid computer experiments. Both wave-particle and subsequent wave-wave interactions are examined. The competing process among coexisting RH and RL mode beam instabilities and repeated decay instabilities triggered by the beam-excited RH mode waves is clarified. It is found that the quenching of the RH instability is not caused by a thermal spreading of the ion beam, but by the nonlinear wave-wave coupling process. The growing RH waves become unstable against the decay instability. This instability involves a backward-traveling RH electromagnetic wave and a forward-traveling longitudinal sound wave. The inverse cascading process is found to occur faster than the growth of the RL mode. Wave spectra decaying from the RH waves weaken as time elapses and the RL mode waves become dominant at the end of the computer experiment.

Kojima, H.

Dynamic Altitude Simulation System Performance Modeling

The "Dynamic Altitude Simulation Prediction Program," DASSPP, is a program to predict the transient response of an engine, test cell, ejector system under engine shutdown conditions. These transients are important to know so that corrective modifications can be adapted to prevent any damage to the engine or test cell. The "Dynamic Altitude System Simulation Prediction Program," DASSPP, is a major rewrite of the existing program "RL-1000" written in BASIC. The RL-1000 program was written to analyze the transients of the RL-10 system only. The new program is written to run in Excell 97 and utilizes the Visual BASIC language in Excell. The program has many added features not included in the original "Rl-1000" program. The program utilizes the ejector models developed during the summer of 1997. The new program is very user friendly and utilizes a dialog box for data input.

LaFrance, Leo J.

Safety Analysis of FMS/CTAS Interactions During Aircraft Arrivals

This grant funded research on human-computer interaction design and analysis techniques, using future ATC environments as a testbed. The basic approach was to model the nominal behavior of both the automated and human procedures and then to apply safety analysis techniques to these models. Our previous modeling language, RSML, had been used to specify the system requirements for TCAS II for the FAA. Using the lessons learned from this experience, we designed a new modeling language that (among other things) incorporates features to assist in designing less error-prone human-computer interactions and interfaces and in detecting potential HCI problems, such as mode confusion. The new language, SpecTRM-RL, uses "intent" abstractions, based on Rasmussen's abstraction hierarchy, and includes both informal (English and graphical) specifications and formal, executable models for specifying various aspects of the system. One of the goals for our language was to highlight the system modes and mode changes to assist in identifying the potential for mode confusion. Three published papers resulted from this research. The first builds on the work of Degani on mode confusion to identify aspects of the system design that could lead to potential hazards. We defined and modeled modes differently than Degani and also defined design criteria for SpecTRM-RL models. Our design criteria include the Degani criteria but extend them to include more potential problems. In a second paper, Leveson and Palmer showed how the criteria for indirect mode transitions could be applied to a mode confusion problem found in several ASRS reports for the MD-88. In addition, we defined a visual task modeling language that can be used by system designers to model human-computer interaction. The visual models can be translated into SpecTRM-RL models, and then the SpecTRM-RL suite of analysis tools can be used to perform formal and informal safety analyses on the task model in isolation or integrated with the rest of the modeled system. We had hoped to be able to apply these modeling languages and analysis tools to a TAP air/ground trajectory negotiation scenario, but the development of the tools took more time than we anticipated.

Nancy G. Leveson

A Survey of Collective Intelligence

This chapter presents the science of "COllective INtelligence" (COIN). A COIN is a large multi-agent systems where: i) the agents each run reinforcement learning (RL) algorithms; ii) there is little to no centralized communication or control; iii) there is a provided world utility function that, rates the possible histories of tile full system. Tile conventional approach to designing large distributed systems to optimize a world utility does not use agents running RL algorithms. Rather that approach begins with explicit modeling of the overall system's dynamics, followed by detailed hand-tuning of the interactions between the components to ensure that they "cooperate" as far as the world utility is concerned. This approach is labor-intensive, often results in highly non-robust systems, and usually results in design techniques that, have limited applicability. In contrast, with COINs we wish to solve the system design problems implicitly, via the 'adaptive' character of the RL algorithms of each of the agents. This COIN approach introduces an entirely new, profound design problem: Assuming the RL algorithms are able to achieve high rewards, what reward functions for the individual agents will, when pursued by those agents, result in high world utility? In other words, what reward functions will best ensure that we do not have phenomena like the tragedy of the commons, or Braess's paradox? Although still very young, the science of COINs has already resulted in successes in artificial domains, in particular in packet-routing, the leader-follower problem, and in variants of Arthur's "El Farol bar problem". It is expected that as it matures not only will COIN science expand greatly the range of tasks addressable by human engineers, but it will also provide much insight into already established scientific fields, such as economics, game theory, or population biology.

Wolpert, David H.

Shape deformation of the organ of Corti associated with length changes of outer hair cell

Cochlear outer hair cells (OHC) are commonly assumed to function as mechanical effectors as well as sensory receptors in the organ of Corti (OC) of the inner ear. OHC in vitro and in organ explants exhibit mechanical responses to electrical, chemical or mechanical stimulation which may represent an aspect of their effector process that is expected in vivo. A detailed description, however, of an OHC effector operation in situ is still missing. Specifically, little is known as to how OHC movements influence the geometry of the OC in situ. Previous work has demonstrated that the motility of isolated OHCs in response to electrical stimulation and to K(+)-gluconate is probably under voltage control and causes depolarisation (shortening) and hyperpolarization (elongation). This work was undertaken to investigate if the movements that were observed in isolated OHC, and which are induced by ionic stimulation, could change the geometry of the OC. A synchronized depolarization of OHC was induced in guinea pig cochleae by exposing the entire OC to artificial endolymph (K+). Subsequent morphometry of mid-modiolar sections from these cochleae revealed that the distance between the basilar membrane (BM) and the reticular lamina (RL) had decreased considerably. Furthermore, in the three upper turns OHC had significantly shortened in all rows. The results suggest that OHC can change their length in the organ of Corti (OC) thus deforming the geometry of the OC. The experiments reveal a tonic force generation within the OC that may change the position of RL and/or BM, contribute to damping, modulate the BM-RL-distance and control the operating points of RL and sensory hair bundles. Thus, the results suggest active self-adjustments of cochlear mechanics by slow OHC length changes. Such mechanical adjustments have recently been postulated to correspond to timing elements of animal communication, speech or music.

Non-NASA Center