Engineering PapersSearch

SEARCH · Engineering Papers

Results for “reinforcement learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Adaptive Stress Testing of Collision Avoidance Systems for Small UASs with Deep Reinforcement Learning

The next-generation Airborne Collision Avoidance System for smaller UASs (ACAS sXu) is currently being developed and tested by the Federal Aviation Administration (FAA) to provide detect-and-avoid capability for small unmanned aircraft operating beyond line-of-sight. Due to the complexity and safety-critical nature of the system, safety validation is important not only for the certification of the final system, but also for informing changes during the iterative development process. In this paper, we analyze a prototype of ACAS sXu in simulated aircraft encounters to discover scenarios of small near mid-air collisions (sNMACs), an important safety event in which two aircraft come closer than 50 feet horizontally and 15 feet vertically. Due to the size and complexity of the system as well as rarity of sNMAC events, traditional methods such as Monte Carlo testing often require informed setup and targeting to elicit failures. However, such a dependence on domain knowledge can be incompatible with the independent verification and validation (IV&V) process, the aim of which is to discover unforeseen issues. To address these challenges, we apply an accelerated validation method called adaptive stress testing (AST) to find the most likely sNMAC scenarios without reliance on system introspection. AST uses reinforcement learning to adapt the search towards the most promising areas of the search space as it progresses. We use a state-of-the-art deep reinforcement learning algorithm, proximate policy optimization, to more efficiently search the large and continuous state space. We find that this approach significantly improves the performance of AST compared to a prior approach based on Monte Carlo tree search. We perform experiments using AST to find sNMAC events under various encounter configurations, varying parameters pertaining to dynamics and coordination. Our experiments show AST to be very effective at finding sNMAC scenarios. We summarize our findings, presenting high-level categories of discovered sNMACs and specific examples of encounters in each category.

aircraft collision avoidance

Integrated Routing and Traffic Signal Control for CAVs via Reinforcement Learning Approach

Incorporating Connected and Automated Vehicles (CAVs) into urban traffic networks presents opportunities and challenges for traffic management systems. This paper aims to develop an integrated routing and traffic signal control system designed explicitly for CAVs, utilizing a Reinforcement Learning (RL) approach. The objective is to enhance traffic flow and improve overall transportation efficiency in the controlled areas. We propose an innovative framework that employs the Deep Reinforcement Learning (DRL) algorithm, especially the Deep Q-network (DQN), to dynamically adjust the number of vehicles in the routes and the duration of traffic signals. Our simulation results demonstrate that a DQN agent successfully optimizes the number of vehicles in the routes and traffic signal timings of traffic signal controllers, eventually reducing total travel time. The study illustrates the potential usage of RL-based systems in managing routing and traffic signals for CAVs, offering a promising opportunity for future urban traffic management strategies.

Park, Jiho [New York University]

Reinforcement Learning with Autonomous Small Unmanned Aerial Vehicles in Cluttered Environments

We present ongoing work in the Autonomy Incubator at NASA Langley Research Center (LaRC) exploring the efficacy of a data set aggregation approach to reinforcement learning for small unmanned aerial vehicle (sUAV) flight in dense and cluttered environments with reactive obstacle avoidance. The goal is to learn an autonomous flight model using training experiences from a human piloting a sUAV around static obstacles. The training approach uses video data from a forward-facing camera that records the human pilot's flight. Various computer vision based features are extracted from the video relating to edge and gradient information. The recorded human-controlled inputs are used to train an autonomous control model that correlates the extracted feature vector to a yaw command. As part of the reinforcement learning approach, the autonomous control model is iteratively updated with feedback from a human agent who corrects undesired model output. This data driven approach to autonomous obstacle avoidance is explored for simulated forest environments furthering autonomous flight under the tree canopy research. This enables flight in previously inaccessible environments which are of interest to NASA researchers in Earth and Atmospheric sciences.

Tran, Loc

Ground Delay Program Analytics with Behavioral Cloning and Inverse Reinforcement Learning

We used historical data to build two types of model that predict Ground Delay Program implementation decisions and also produce insights into how and why those decisions are made. More specifically, we built behavioral cloning and inverse reinforcement learning models that predict hourly Ground Delay Program implementation at Newark Liberty International and San Francisco International airports. Data available to the models include actual and scheduled air traffic metrics and observed and forecasted weather conditions. We found that the random forest behavioral cloning models we developed are substantially better at predicting hourly Ground Delay Program implementation for these airports than the inverse reinforcement learning models we developed. However, all of the models struggle to predict the initialization and cancellation of Ground Delay Programs. We also investigated the structure of the models in order to gain insights into Ground Delay Program implementation decision making. Notably, characteristics of both types of model suggest that GDP implementation decisions are more tactical than strategic: they are made primarily based on conditions now or conditions anticipated in only the next couple of hours.

Bloem, Michael

Adaptive Stress Testing: Using Reinforcement Learning to Find Failures in Safety-Critical Systems

Emerging applications in artificial intelligence, such as driverless cars and autonomous aircraft promise to be more efficient, cheaper to operate, and always available. However, ensuring the safety of these systems remains a major challenge to their certification and adoption. These autonomous systems are expected to routinely make safety-critical decisions where failures can have serious consequences including loss of life and property. Testing and validation techniques aim to identify and diagnose potential failures before the system is deployed. However, finding failure scenarios in autonomous systems can be very challenging due to high-dimensional and continuous state spaces, interaction with large environments over many time steps, and the rarity of failures. This talk presents Adaptive Stress Testing (AST), a simulation-based testing framework for finding the most likely path to a failure event of a safety-critical system. The key idea of AST is that stress testing can be formulated as a Partially Observable Markov Decision Process (POMDP), which enables reinforcement learning techniques to be used for finding failure events. Reinforcement learning algorithms can efficiently explore the search space and have been shown to scale to very large systems. We present applications of AST to find failures in various safety-critical systems including the aircraft collision avoidance systems, autonomous cars, and small unmanned aerial vehicles.

autonomous vehicles

Refining Linear Fuzzy Rules by Reinforcement Learning

Linear fuzzy rules are increasingly being used in the development of fuzzy logic systems. Radial basis functions have also been used in the antecedents of the rules for clustering in product space which can automatically generate a set of linear fuzzy rules from an input/output data set. Manual methods are usually used in refining these rules. This paper presents a method for refining the parameters of these rules using reinforcement learning which can be applied in domains where supervised input-output data is not available and reinforcements are received only after a long sequence of actions. This is shown for a generalization of radial basis functions. The formation of fuzzy rules from data and their automatic refinement is an important step in closing the gap between the application of reinforcement learning methods in the domains where only some limited input-output data is available.

Berenji, Hamid R.

A Two-Stage Quantum Reinforcement Learning Method for Multi-Objective Transmission Switching

Multi-objective transmission switching (MO-TS) problems involve the strategic reconfiguration of network topology to simultaneously optimize multiple objectives. As the system scale increases, finding feasible solutions becomes increasingly challenging due to the problem's nonlinearity and high computational complexity. To address these challenges, this paper proposes a two-stage quantum reinforcement learning method that leverages potential quantum advantages for MO-TS. In the first stage, candidate switching lines are identified using a graph-theoretical approach to reduce the problem's dimensionality. The second stage introduces a quantum-classical reinforcement learning framework, where a learnable measurement-based CNN-ResVQC architecture is developed to effectively reduce the input dimension for quantum processing, mitigate vanishing gradients, and enhance trainability while improving the quantum circuit's flexibility in modeling complex decision policies for MO-TS. Numerical studies on IEEE 14-bus, 57-bus, and 118-bus systems demonstrate that the proposed algorithm achieves superior training stability and faster convergence with approximately 1% of the network parameters required by classical algorithms, highlighting its effectiveness, efficiency, and scalability. Furthermore, the practicality is validated through its stable convergence under three common quantum noise channels.

99 GENERAL AND MISCELLANEOUS

Nuclear microreactor transient and load-following control with deep reinforcement learning

The economic feasibility of nuclear microreactors will depend on minimizing operating costs through advancements in autonomous control, especially when these microreactors are operating alongside other types of energy systems (e.g., renewable energy). This study explores the application of deep reinforcement learning (RL) for real-time drum control in microreactors, exploring performance in regard to load-following scenarios. By leveraging a point kinetics model with thermal and xenon feedback, we first establish a baseline using a single-output RL agent, then compare it against a traditional proportional–integral–derivative (PID) controller. This study demonstrates that RL controllers, including both single- and multi-agent RL (MARL) frameworks, can achieve similar or even superior load-following performance as traditional PID control across a range of load-following scenarios. In short transients, the RL agent was able to reduce the tracking error rate in comparison to PID by one half to one third. Over extended 300-minute load-following scenarios in which xenon feedback becomes a dominant factor, PID maintained better accuracy, but RL still remained within a 1% error margin despite being trained only on short-duration scenarios. This highlights RL’s strong ability to generalize and extrapolate to longer, more complex transients, affording substantial reductions in training costs and reduced overfitting. Furthermore, when control was extended to multiple drums, MARL enabled independent drum control as well as maintained reactor symmetry constraints without sacrificing performance---an objective that standard single-agent RL could not learn. We also found that, as increasing levels of Gaussian noise were added to the power measurements, the RL controllers were able to maintain lower error rates than PID, and to do so with at least 10% and upwards of 150% less control effort. These findings illustrate RL's potential for autonomous nuclear reactor control, laying the groundwork for future integration into high-fidelity simulations and experimental validation efforts.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

A Reinforcement Learning Framework for Space Missions in Unknown Environments

A land-and-traverse mission to icy worlds such as Europa and Enceladus is challenging due to lack of prior knowledge regarding the terrain conditions. Previous work [1] showed that rovers with high degrees of freedom (DoF) can achieve robust traversal by leveraging redundant modes for mobility to counter terrain uncertainty (e.g. walking, driving, or inch-worming). This paper presents a generic and scalable reinforcement learning scheme for enabling on-board decision making on rovers to automatically switch between modes of traversal based on online performance feedback. The objective is to maximize energy efficiency, minimize operator input and successfully negotiate unstructured terrain conditions without relying on exhaustive prior knowledge. The proposed methodology is well grounded in the literature on reinforcement learning and has been adapted to address conformance to validation and verification requirements and JPL flight operations history of using per-sol prescribed sequences for a space mission.

Tavallali, Peyman

A Physics-Informed Reinforcement Learning Framework for Economic-Thermal Co-Optimization of Crypto Mining Data Centers: Preprint

The rapid expansion of cryptocurrency mining has created a new class of high-density data centers characterized by extreme thermal flux and high sensitivity to volatile economic markets. Traditional thermal management strategies, typically reliant on rule-based control, maintain static setpoints that fail to account for fluctuating electricity prices and cryptocurrency values - factors critical to mining profitability. To address this, we present a physics-informed reinforcement learning (PIRL) framework for economic-thermal co-optimization in crypto mining data centers. This framework consists of a proximal policy optimization (PPO) agent, a virtual testbed powered by high-fidelity physics-based models, and an interactive frontend dashboard. The PPO agent is trained using the virtual testbed and strict hardware safety limits. This physics-informed approach allows the agent to learn a stochastic policy that dynamically balances mining revenue against operational costs by co-optimizing HVAC cooling setpoints and IT computational hashrate. The simulation results demonstrate that the integrated framework achieved an 8.62% increase in net operational profit compared to traditional baseline strategies while strictly adhering to safety-critical temperature constraints (coolant supply temperature < 32 degrees C). This work provides a scalable template for the deployment of reinforcement learning in mission critical facilities where economic volatility and physical safety must be managed simultaneously.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Scheduling Mission Reconfiguration for an Interferometry Synthetic Aperture Radar Using Deep Reinforcement Learning

This paper presents a method to intelligently adapt the baseline of a synthetic aperture radar based on Deep Rein- forcement Learning to help create plans for missions that use formation flight for Earth observation purposes. The main contribution of this paper is the initial results we have found from applying the tool to a toy mission: measuring the ver- tical structure of forests by using a synthetic aperture radar mounted on a formation of 7 satellites orbiting the Earth in a Sun Synchronous Orbit. We have found that with a reward function based on expected science return over time and fuel usage, the Deep Reinforcement Learning planner is able to create plans with positive scientific returns while minimizing fuel usage. We also find that fuel usage and collision avoid- ance planning is better done with traditional methods, as Deep Reinforcement Learning does not converge to optimal solutions.

Viros-i-Martin, Antoni

Microsecond-latency feedback at a particle accelerator by online reinforcement learning on hardware

The commissioning and operation of future large-scale scientific experiments will challenge current tuning and control methods. Reinforcement learning (RL) algorithms are a promising solution due to their ability to dynamically adapt to changing environments and consider delayed consequences. In many real-world applications, RL policies must produce actions in real time, often within microseconds to milliseconds, imposing significant constraints on system latency and computational overhead that conventional machine learning libraries are not designed to handle. To control phenomena in real time at these timescales, RL needs to be deployed on-the-edge, namely on dedicated hardware located near the system it controls, without relying on a host CPU or cloud-based inference. In this work we present the design and deployment of an experience accumulator system in a particle accelerator. In this system, deep-RL algorithms run using hardware acceleration and act within a few microseconds, enabling the use of RL for control of phenomena like beam instabilities. The training uses the collected data offline to reduce the number of operations carried out on the acceleration hardware. The proposed architecture was tested in real experimental conditions at the Karlsruhe research accelerator, a synchrotron light source, where the system was used to control artificially induced horizontal betatron oscillations in real-time, with a control loop period of just 2.7 μs. The results showed a performance comparable to the commercial feedback system available at the accelerator, demonstrating the viability and potential of this approach. Due to the self-learning and reconfiguration capability of this implementation, a seamless application to other control problems is possible. Applications range from particle accelerators to large-scale research and industrial facilities.

FPGA

Energy performance evaluation of the ASHRAE Guideline 36 control and reinforcement learning–based control using field measurements

This study evaluates the energy performance of ASHRAE Guideline 36–compliant control (ASHRAE 36 control) and reinforcement learning (RL)–based control through experimental field tests and a simulation study. Three field tests were conducted at Oak Ridge National Laboratory’s commercial building test facility in Oak Ridge, Tennessee: a baseline with a baseline conventional control, a test with ASHRAE 36 control, and a test with RL-based control. The selected ASHRAE 36 controls were trim and respond control, as well as variable air volume (VAV) box control. We compared the measured supply air temperature of the rooftop unit, VAV box supply air temperature, and VAV box supply airflow rate across the three test cases. The field data indicated that ASHRAE 36 controls operated as specified by ASHRAE Guideline 36. Based on these data, ASHRAE 36 control achieved a 45 % reduction in hourly averaged HVAC energy consumption compared with the baseline, and RL-based control achieved a 66 % reduction. These potential annual energy savings were confirmed using a calibrated whole-building energy model. Compared with the baseline, ASHRAE 36 control reduced HVAC energy consumption by 42 %, and RL-based control achieved a 54 % reduction. Furthermore, RL-based control reduced total HVAC energy consumption by 21 % more than ASHRAE 36 control.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Large language models (LLMs) are increasingly adapted to downstream tasks via reinforcement learning (RL) methods like Group Relative Policy Optimization (GRPO), which often require thousands of rollouts to learn new tasks. We argue that the interpretable nature of language often provides a much richer learning medium for LLMs, compared to policy gradients derived from sparse, scalar rewards. To test this, we introduce GEPA (Genetic-Pareto), a prompt optimizer that thoroughly incorporates natural language reflection to learn high-level rules from trial and error. Given any AI system containing one or more LLM prompts, GEPA samples trajectories (e.g., reasoning, tool calls, and tool outputs) and reflects on them in natural language to diagnose problems, propose and test prompt updates, and combine complementary lessons from the Pareto frontier of its own attempts. As a result of GEPA's design, it can often turn even just a few rollouts into a large quality gain. Across six tasks, GEPA outperforms GRPO by 6% on average and by up to 20%, while using up to 35x fewer rollouts. GEPA also outperforms the leading prompt optimizer, MIPROv2, by over 10% (e.g., +12% accuracy on AIME-2025), and demonstrates promising results as an inference-time search strategy for code optimization. We release our code at https://github.com/gepa-ai/gepa.

97 MATHEMATICS AND COMPUTING

Exploring Transfers Between Earth-Moon Halo Orbits via Multi-Objective Reinforcement Learning

Multi-Reward Proximal Policy Optimization, a multi-objective deep reinforcement learning algorithm, is used to examine the design space of low-thrust trajectories for a SmallSat transferring between two libration point orbits in the Earth-Moon system. Using Multi-Reward Proximal Policy Optimization, multiple policies are simultaneously and efficiently trained on three distinct trajectory design scenarios. Each policy is trained to create a unique control scheme based on the trajectory design scenario and assigned reward function: a unique combination of weights scaling competing objectives that guide the spacecraft to the target mission orbit, incentivize faster flight times, and penalize propellant mass usage. Then, the policies are evaluated on the same set of perturbed initial conditions in each scenario to generate the propellant mass usages, flight times, and state discontinuities from a reference trajectory for each control scheme. This solution space of low-thrust trajectories for a SmallSat is used to examine the multi-objective trade space for the trajectory design scenario. By autonomously constructing the solution space, insights into the required propellant mass, flight time, and transfer geometry are rapidly achieved.

Christopher J Sullivan

Exploring Transfers Between Earth-Moon Halo Orbits via Multi-Objective Reinforcement Learning

Multi-Reward Proximal Policy Optimization, a multi-objective deep reinforcement learning algorithm, is used to examine the design space of low-thrust trajectories for a SmallSat transferring between two libration point orbits in the Earth- Moon system. Using Multi-Reward Proximal Policy Optimiza- tion, multiple policies are simultaneously and efficiently trained on three distinct trajectory design scenarios. Each policy is trained to create a unique control scheme based on the trajectory design scenario and assigned reward function: a unique combination of weights scaling competing objectives that guide the spacecraft to the target mission orbit, incentivize faster flight times, and penalize propellant mass usage. Then, the policies are evaluated on the same set of perturbed initial conditions in each scenario to generate the propellant mass usages, flight times, and state discontinuities from a reference trajectory for each control scheme. This solution space of low-thrust trajectories for a SmallSat is used to examine the multi-objective trade space for the trajectory design scenario. By autonomously constructing the solution space, insights into the required propellant mass, flight time, and transfer geometry are rapidly achieved.

Mashiku, Alinda K.

Exploring the Low-Thrust Transfer Design Space in an Ephemeris Model via Multi-Objective Reinforcement Learning

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective reinforcement learning algorithm used to train multiple policies to uncover solutions within a multi-objective solution space. MRPPO is used in this paper to train policies to construct low-thrust transfers for a SmallSat from the vicinity of !2 to an !5 short period orbit in the Sun-Earth-Moon system. First, the policies are trained in this scenario in the circular restricted three-body problem. This information is used to initialize the policies before training in a higher-fidelity ephemeris model; a process known as transfer learning. The recovered segments of the solution space will be compared to fundamental dynamical structures to both examine the results of MRPPO in this complex design scenario and explore the effectiveness of transfer learning.

Christopher J Sullivan

Exploring the Low-Thrust Transfer Design Space in an Ephemeris Model via Multi-Objective Reinforcement Learning

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective reinforcement learning algorithm used to train multiple policies to uncover solutions within a multi-objective solution space. MRPPO is used in this paper to train policies to construct low-thrust transfers for a SmallSat from the vicinity of L2 to an L5 short period orbit in the Sun-Earth-Moon system. First, the policies are trained in this scenario in the circular restricted three-body problem. This information is used to initialize the policies before training in a higher-fidelity ephemeris model; a process known as transfer learning. The recovered segments of the solution space will be compared to fundamental dynamical structures to both examine the results of MRPPO in this complex design scenario and explore the effectiveness of transfer learning.

Christopher J. Sullivan