Engineering PapersSearch

SEARCH · Engineering Papers

Results for “RL”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Design study of RL10 derivatives. Volume 2: Engine design characteristics, appendices

Calculations, curves, and substantiating data which support the engine design characteristics of the RL-10 engines are presented. A description of the RL-10 ignition system is provided. The performance calculations of the RL-10 derivative engines and the performance results obtained are reported. The computer simulations used to establish the control system requirements and to define the engine transient characteristics are included.

Source record

Designing Specification Languages for Process Control Systems: Lessons Learned and Steps to the Future

Previously, we defined a blackbox formal system modeling language called RSML (Requirements State Machine Language). The language was developed over several years while specifying the system requirements for a collision avoidance system for commercial passenger aircraft. During the language development, we received continual feedback and evaluation by FAA employees and industry representatives, which helped us to produce a specification language that is easily learned and used by application experts. Since the completion of the PSML project, we have continued our research on specification languages. This research is part of a larger effort to investigate the more general problem of providing tools to assist in developing embedded systems. Our latest experimental toolset is called SpecTRM (Specification Tools and Requirements Methodology), and the formal specification language is SpecTRM-RL (SpecTRM Requirements Language). This paper describes what we have learned from our use of RSML and how those lessons were applied to the design of SpecTRM-RL. We discuss our goals for SpecTRM-RL and the design features that support each of these goals.

Leveson, Nancy G.

Reinforcement Learning in Distributed Domains: Beyond Team Games

Distributed search algorithms are crucial in dealing with large optimization problems, particularly when a centralized approach is not only impractical but infeasible. Many machine learning concepts have been applied to search algorithms in order to improve their effectiveness. In this article we present an algorithm that blends Reinforcement Learning (RL) and hill climbing directly, by using the RL signal to guide the exploration step of a hill climbing algorithm. We apply this algorithm to the domain of a constellations of communication satellites where the goal is to minimize the loss of importance weighted data. We introduce the concept of 'ghost' traffic, where correctly setting this traffic induces the satellites to act to optimize the world utility. Our results indicated that the bi-utility search introduced in this paper outperforms both traditional hill climbing algorithms and distributed RL approaches such as team games.

Wolpert, David H.

Adaptive, Distributed Control of Constrained Multi-Agent Systems

Product Distribution (PO) theory was recently developed as a broad framework for analyzing and optimizing distributed systems. Here we demonstrate its use for adaptive distributed control of Multi-Agent Systems (MASS), i.e., for distributed stochastic optimization using MAS s. First we review one motivation of PD theory, as the information-theoretic extension of conventional full-rationality game theory to the case of bounded rational agents. In this extension the equilibrium of the game is the optimizer of a Lagrangian of the (Probability dist&&on on the joint state of the agents. When the game in question is a team game with constraints, that equilibrium optimizes the expected value of the team game utility, subject to those constraints. One common way to find that equilibrium is to have each agent run a Reinforcement Learning (E) algorithm. PD theory reveals this to be a particular type of search algorithm for minimizing the Lagrangian. Typically that algorithm i s quite inefficient. A more principled alternative is to use a variant of Newton's method to minimize the Lagrangian. Here we compare this alternative to RL-based search in three sets of computer experiments. These are the N Queen s problem and bin-packing problem from the optimization literature, and the Bar problem from the distributed RL literature. Our results confirm that the PD-theory-based approach outperforms the RL-based scheme in all three domains.

Bieniawski, Stefan

Light-stimulated cell expansion in bean (Phaseolus vulgaris L.) leaves. II. Quantity and quality of light required

The quantity and quality of light required for light-stimulated cell expansion in leaves of Phaseolus vulgaris L. have been determined. Seedlings were grown in dim red light (RL; 4 micromoles photons m-2 s-1) until cell division in the primary leaves was completed, then excised discs were incubated in 10 mM sucrose plus 10 mM KCl in a variety of light treatments. The growth response of discs exposed to continuous white light (WL) for 16 h was saturated at 100 micromoles m-2 s-1, and did not show reciprocity. Extensive, but not continuous, illumination was needed for maximal growth. The wavelength dependence of disc expansion was determined from fluence-response curves obtained from 380 to 730 nm provided by the Okazaki Large Spectrograph. Blue (BL; 460 nm) and red light (RL; 660 nm) were most effective in promoting leaf cell growth, both in photosynthetically active and inhibited leaf discs. Far-red light (FR; 730 nm) reduced the effectiveness of RL, but not BL, indicating that phytochrome and a separate blue-light receptor mediate expansion of leaf cells.

NASA Discipline Plant Biology

Evolution of Interplanetary Shocks and their CME Drivers

Shock-driving coronal mass ejections (CMEs) constitute the most energetic phenomena in the heliosphere. The shocks can be identified in a number of ways based on remote-sensing and in situ observations. Type II radio bursts are the earliest indicators of shocks that accelerate electrons to energies up to -10 keV. Solar energetic particle (SEP) events are always accompanied by long wavelength type II bursts indicating that the same shock accelerates ions and electrons. A recent investigation involving a large number of interplanetary (IP) shocks revealed that about 35% of them do not produce type II bursts (radio quiet, RQ) or SEPs. Comparison of the RQ shocks with the radio loud (RL) ones revealed some interesting results such as: (1) the lack of evidence for blast waves,(2) energetic particle enhancement in the shock front in -20% of RQ shocks, and (3) determination of the difference between the RQ and RL shocks in terms of the different kinematic properties of the associated CMEs. On the other hand the shock properties measured at I AU are not too different for the RQ and RL cases. This can be attributed to the interaction with the IP medium, which seems to erase the difference. Implications of this evolution for the geoeffectiveness is also discussed.

Gopalswamy, Natchimuthuk

Influences of the Driver and Ambient Medium Characteristics on the Formation of Shocks in the Solar Atmosphere

Traveling interplanetary (IP) shocks were discovered in the early 1960s, but their solar origin has been controversial. Early research focused on solar flares as the source of the shocks, but when coronal mass ejections (CMEs) were discovered, it became clear that fast CMEs clearly can drive the shocks. Type II radio bursts are excellent signatures of shocks near the Sun. The close correspondence between type II radio bursts and solar energetic particles (SEPs) makes it clear that the same shock accelerates ions and electrons. A recent investigation involving a large number of IP shocks revealed that about 35% of IP shocks do not produce type II bursts or SEPs. Comparing these radio quiet (RQ) shocks with the radio loud (RL) ones revealed some interesting results: (1) there is no evidence for blast waves, in that all IP shocks can be attributed to CMEs, (2) a small fraction (20%) of RQ shocks is associated with ion enhancements at the shocks when they move past the observing spacecraft, (3) the primary difference between the RQ and RL shocks can be traced to the different kinematic properties of the associated CMEs and the variation of the characteristic speeds of the ambient medium, and (4) the shock properties measured at 1 AU are not too different for the RQ and RL cases due to the interaction of the shock driver with the IP medium that seems to erase the difference.

Nat, Gopalswamy

Shock-Driving CMEs Near the Sun, in the Interplanetary Medium, and Near Earth

The excellent correspondence between type II radio bursts and solar energetic particles (SEPs) made it clear that the same shock accelerates ions and electrons. A recent investigation involving a large number of I shocks revealed that about 35% of IP shocks do not produce type II bursts (radio quiet) or SEPs. Comparing the RQ shocks with the radio loud (RL) ones revealed some interesting results, which will be summarized in this poster. (1) There is no evidence for blast waves. (2) Even a small fraction (20%) of RQ shocks is associated with ion enhancements at the shock when the shock passes the spacecraft. (3) The primary difference between the RQ and RL shocks can be traced to the different kinematic properties of the associated CMEs, although the shock properties measured at 1 AU are not too different for the RQ and RL cases. This can be attributed to the interaction with the IP medium, which seems to erase the difference. More details can be found in Astrophysical Journal 710, 1111, 2010 (http://adsabs.harvard.edu/absZ2009arXivO9l2.4719G).

Gopalswamy, N.

On Interplanetary Shocks Driven by Coronal Mass Ejections

Traveling interplanetary (IP) shocks were first detected in the early 1960s, but their solar origin has been controversial. Early research focused on solar flares as the source of the shocks, but when CMEs were discovered, it became clear that fast CMEs are the shock drivers. Type radio II bursts are excellent signatures of shocks near the Sun (Type II radio bursts were known long before the detection of shocks and CMEs). The excellent correspondence between type II bursts and solar energetic particle (SEP) events made it clear that the same shock accelerates ions and electrons. Shocks near the Sun are also seen occasionally in white-light coronagraphic images. In the solar wind, shocks are observed as discontinuities in plasma parameters such as density and speed. Energetic storm particle events and sudden commencement of geomagnetic storm are also indicators of shocks arriving at Earth. After an overview on these shock signatures, I will summarize the results of a recent investigation of a large number of IP shocks. The study revealed that about 35% of IP shocks do not produce type II bursts (radio quiet - RQ) or SEPs. Comparing the RQ shocks with the radio loud (RL) ones revealed some interesting results: (1) There is no evidence for blast wave shocks. (2) A small fraction (20%) of RQ shocks is associated with ion enhancements at the shock when the shock passes the spacecraft. (3) The primary difference between the RQ and RL shocks can be traced to the different kinematic properties of the associated CMEs. On the other hand the shock properties measured at 1 AU are not too different for the RQ and RL cases. This can be attributed to the interaction with the IP medium, which seems to erase the difference between the shocks.

Gopalswarmy, Nat

Coronal Mass Ejection-driven Shocks and the Associated Sudden Commencements-sudden Impulses

Interplanetary (IP) shocks are mainly responsible for the sudden compression of the magnetosphere, causing storm sudden commencement (SC) and sudden impulses (SIs) which are detected by ground-based magnetometers. On the basis of the list of 222 IP shocks compiled by Gopalswamy et al., we have investigated the dependence of SC/SIs amplitudes on the speed of the coronal mass ejections (CMEs) that drive the shocks near the Sun as well as in the interplanetary medium. We find that about 91% of the IP shocks were associated with SC/SIs. The average speed of the SC/SI-associated CMEs is 1015 km/s, which is almost a factor of 2 higher than the general CME speed. When the shocks were grouped according to their ability to produce type II radio burst in the interplanetary medium, we find that the radio-loud (RL) shocks produce a much larger SC/SI amplitude (average approx. 32 nT) compared to the radio-quiet (RQ) shocks (average approx. 19 nT). Clearly, RL shocks are more effective in producing SC/SIs than the RQ shocks. We also divided the IP shocks according to the type of IP counterpart of interplanetary CMEs (ICMEs): magnetic clouds (MCs) and nonmagnetic clouds. We find that the MC-associated shock speeds are better correlated with SC/SI amplitudes than those associated with non-MC ejecta. The SC/SI amplitudes are also higher for MCs than ejecta. Our results show that RL and RQ type of shocks are important parameters in producing the SC/SI amplitude.

Coronal mass ejections

Scheduling the NASA Deep Space Network with Deep Reinforcement Learning

With three complexes spread evenly across the Earth, NASA’s Deep Space Network (DSN) is the primary means of communications as well as a significant scientific instrument for dozens of active missions around the world. A rapidly rising number of spacecraft and increasingly complex scientific instruments with higher bandwidth requirements have resulted in demand that exceeds the network’s capacity across its 12 antennae. The existing DSN scheduling process operates on a rolling weekly basis and is time-consuming; for a given week, generation of the final baseline schedule of spacecraft tracking passes takes roughly 5 months from the initial requirements submission deadline, with several weeks of peer-to-peer negotiations in between. This paper proposes a deep reinforcement learning (RL) approach to generate candidate DSN schedules from mission requests and spacecraft ephemeris data with demonstrated capability to address real-world operational constraints. A deep RL agent is developed that takes mission requests for a given week as input, and interacts with a DSN scheduling environment to allocate tracks such that its reward signal is maximized. A comparison is made between an agent trained using Proximal Policy Optimization and its random, untrained counterpart. The results represent a proof-of-concept that, given a well-shaped reward signal, a deep RL agent can learn the complex heuristics used by experts to schedule the DSN. A trained agent can potentially be used to generate candidate schedules to bootstrap the scheduling process and thus reduce the turnaround cycle for DSN scheduling.

Wilson, Brian

SatNet: A Benchmark for Satellite Scheduling Optimization

Satellites provide essential services such as networking and weather tracking, and the number of near-earth and deep space satellites are expected to grow rapidly in the coming years. Communications with terrestrial ground stations is one of the critical functionalities of any space mission. Satellite scheduling is a problem that has been scientifically investigated since the 1970s. A central aspect of this problem is the need to consider resource contention and satellite visibility constraints as they require line of sight. Due to the combinatorial nature of the problem, prior solutions such as linear programs and evolutionary algorithms require extensive compute capabilities to output a feasible schedule for each scenario. Machine learning based scheduling can provide an alternative solution by training a model with historical data and generating a schedule quickly with model inference. We present SatNet, a benchmark for satellite scheduling optimization based on historical data from the NASA Deep Space Network. We propose formulation of the satellite scheduling problem as a Markov Decision Process and use reinforcement learning (RL) policies to generate schedules. The nature of constraints imposed by SatNet differ from other combinatorial optimization problems such as vehicle routing studied in prior literature. Our initial results indicate that RL is an alternative optimization approach that can generate candidate solutions of comparable quality to existing state-of-the-practice results. However, we also find that RL policies overfit to the training dataset and do not generalize well to new data, thereby necessitating continued research on reusable and generalizable agents.

Wilson, Brian

Discovering the Most Severe K-Point Failure Based on Reinforcement Learning: Preprint

Smart devices are essential to ensure the stability of the power grid and resilience to intermittent energy production. However, smart devices can also be the target of cyber adversaries that may exploit false data injection attacks (FDIAs) to induce unstable grid conditions. A practical consideration of FDIA mitigation approaches is addressed here: given a finite available budget, for which smart device should cyber-threat mitigation be deployed first? In this work, this question is answered by identifying the so-called most-sensitive devices, i.e., the devices that, if compromised, can let an adversary induce the most serious grid instabilities. The method proposed utilizes an adversarial reinforcement learning (RL) framework to identify the k-mostsensitive smart devices (here, smart inverters). The adversarial agent can tamper with the compromised inverters' active and reactive operating power setup points, with the goal of maximizing voltage deviations. Numerical results show that the proposed RL method finds the optimal attack scenarios for 1-point failure and the near-optimal solution for the 2-point case. Additionally, the proposed RL method achieves an 8.8 speed-up ratio in running time compared to the brute force method for the 2-point case.

97 MATHEMATICS AND COMPUTING

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.

Naug, Avisek [Hewlett Packard Enterprise]

Machine Learning a Simple Interpretable Short-Range Potential for Silica

A wide array of models, spanning from computationally expensive ab initio methods to a spectrum of force-field approaches, have been developed and employed to probe silica polymorphs and understand growth processes and atomic-level dynamical transitions in silica. However, the quest for a model capable of making accurate predictions with high computational efficiency for various silica polymorphs is still ongoing. Recent developments in short-range machine-learned models, such as GAP and NNPScan, have shown promise in providing reasonable descriptions of silica, but their computational cost remains high compared to force fields such as BKS which are based on simple interpretable functional forms. Here, in this study, we build on the recent success of our reinforcement learning (RL) workflow to derive a new set of optimal parameters for a promising short-range BKS-based model proposed by Soules. We use RL to navigate the eight-dimensional parameter space of the Soules potential using an experimental training data set that includes both local and global structural features from approximately 21 experimentally realized silica polymorphs, including high density phases and porous zeolites. We compare the performance of our machine-learned ML-Soules model with other high quality models including our recent machine-learned parametrization of BKS (ML-BKS), a machine-learned potential (GAP), as well as predictions of ab initio calculations with the highly fidelity SCAN functional. The ML-Soules accurately captures the relative energetic ordering of various polymorphs as well as their structural features at a significantly reduced computational expense. The ML-Soules model also reasonably captures the structure, density, and elastic constants of quartz, as well as metastable silica polymorphs. We further discuss the limitations of the Soules functional form and propose potential enhancements, including the incorporation of additional three-body terms and/or the utilization of different short-ranged functional forms to achieve greater accuracy for both global and local features in the modeling of silica while retaining low computational cost.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Ab Initio-Based Bond Order Potential for Arsenene Polymorphs Developed via Hierarchical Reinforcement Learning

Arsenene, a less-explored two-dimensional material, holds the potential for applications in wearable electronics, memory devices, and quantum systems. This study introduces a bond-order potential model with Tersoff formalism, the ML-Tersoff, which leverages multireward hierarchical reinforcement learning (RL), trained on an ab initio data set. This data set covers a spectrum of properties for arsenene polymorphs, enhancing our understanding of its mechanical and thermal behaviors without the complexities of traditional models requiring multiple parameter sets. Our RL strategy utilizes decision trees coupled with a hierarchical reward strategy to accelerate convergence in high-dimensional continuous search spaces. Unlike the Stillinger-Weber approach, which demands separate formalisms for buckled and puckered forms, the ML-Tersoff model concurrently captures multiple properties of the two polymorphs by effectively representing the local environment, thereby avoiding the need for different atomic types. Here, we apply the ML model to understand the mechanical and thermal properties of the arsenene polymorphs and nanostructures. We observe an inverse relationship between the critical strain and temperature in arsenene. Thermal conductivity calculations in nanosheets show good agreement with ab initio data, reflecting a decrease in thermal conductivity attributable to increased anharmonic effects at higher temperatures. We also apply the model to predict the thermal behavior of arsenene nanotubes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Entanglement engineering of optomechanical systems by reinforcement learning

Entanglement is fundamental to quantum information science and technology, yet controlling and manipulating entanglement—so-called entanglement engineering—for arbitrary quantum systems remains a formidable challenge. There are two difficulties: the fragility of quantum entanglement and its experimental characterization. We develop a model-free deep reinforcement-learning (RL) approach to entanglement engineering, in which feedback control together with weak continuous measurement and partial state observation is exploited to generate and maintain desired entanglement. We employ quantum optomechanical systems with linear or nonlinear photon–phonon interactions to demonstrate the workings of our machine-learning-based entanglement engineering protocol. In particular, the RL agent sequentially interacts with one or multiple parallel quantum optomechanical environments, collects trajectories, and updates the policy to maximize the accumulated reward to create and stabilize quantum entanglement over an arbitrary amount of time. The machine-learning-based model-free control principle is applicable to the entanglement engineering of experimental quantum systems in general.

97 MATHEMATICS AND COMPUTING

Demonstration of reconstruction-free static magnetic control of DIII-D plasma with deep reinforcement learning

This paper presents the development and experimental validation of a reinforcement learning (RL)-based magnetic controller on the DIII-D tokamak. The controller directly maps raw magnetic diagnostic signals to actuator commands, replacing the traditional isoflux control algorithm based on equilibrium reconstruction. Four RL controllers are trained using the Soft Actor–Critic algorithm with an asymmetric Actor–Critic architecture in the NSFsim simulator. All controllers are deployed in the DIII-D Plasma Control System and operated with a 4 kHz feedback loop. Two randomization strategies are evaluated during training: evolving kinetic profiles and fixed kinetic profiles within each episode. The latter approach is found to better capture experimental deviations in the current density profile and to provide overall improved control performance. Robust operation is demonstrated across heating power scans in both L- and H-mode plasmas, as well as during transient events such as L–H transitions and pellet injections. Control errors in plasma shape and radial position remained within 1.5–2.0 cm and 1 cm, respectively. A notable discrepancy was observed in the vertical X-point position, with errors of up to approximately 4 cm, attributed to the current density distribution mismatches between simulations and experiments.

DIII-D