Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Scheduling Mission Reconfiguration for an Interferometry Synthetic Aperture Radar Using Deep Reinforcement Learning

This paper presents a method to intelligently adapt the baseline of a synthetic aperture radar based on Deep Rein- forcement Learning to help create plans for missions that use formation flight for Earth observation purposes. The main contribution of this paper is the initial results we have found from applying the tool to a toy mission: measuring the ver- tical structure of forests by using a synthetic aperture radar mounted on a formation of 7 satellites orbiting the Earth in a Sun Synchronous Orbit. We have found that with a reward function based on expected science return over time and fuel usage, the Deep Reinforcement Learning planner is able to create plans with positive scientific returns while minimizing fuel usage. We also find that fuel usage and collision avoid- ance planning is better done with traditional methods, as Deep Reinforcement Learning does not converge to optimal solutions.

Viros-i-Martin, Antoni↗

RLC4CLR (Reinforcement Learning Controller for Critical Load Restoration Problems)

RLC4CLR demonstrates using a reinforcement learning controller (RLC) to solve a critical load restoration (CLR) problem, which improves the grid resilience after a substation outage event. RLC4CLR consists of two parts. (1) RL environment: This environment encapsulates the CLR problem to be solved and provides interfacing functions to follow the standard OpenAI Gym format. A power system simulator, i.e., OpenDSS, is included to provide the power flow solution. Controller inputs and outputs (RL state and action) as well as the reward are defined in this environment as well. In summary, the RL environment is the problem formulation from which the RL agent can learn. (2) RL training script: The training script enables the RL agent to learn its control policy by interacting with the RL environment. For RL training, an open-sourced RL library, i.e., RLlib, is leveraged which is based on a distributed computing framework (Ray). The training script is designed to be able to be run on both local machine or the NREL HPC system. Other components of RLC4CLR include input data, e.g., grid model (standard IEEE test feeders), and other files used for results analysis.

Zhang, Xiangyu↗

On the Verification of Deep Reinforcement Learning Solution for Intelligent Operation of Distribution Grids

Capabilities of deep reinforcement learning (DRL) in obtaining fast decision policies in high dimensional and stochastic environments have led to its extensive use in operational research, including the operation of distribution grids with high penetration of distributed energy resources (DER). However, the feasibility and robustness of DRL solutions are not guaranteed for the system operator, and hence, those solutions may be of limited practical value. This paper proposes an analytical method to find feasibility ellipsoids that represent the range of multi-dimensional system states in which the DRL solution is guaranteed to be feasible. Empirical studies and stochastic sampling determine the ratio of the discovered to the actual feasible space as a function of the sample size. In addition, the performance of logarithmic, linear, and exponential penalization of infeasibility during the DRL training are studied and compared in order to reduce the number of infeasible solutions

Hosseini, Mohammad Mehdi↗

Deep Reinforcement Learning for Microgrid Cost Optimization Considering Load Flexibility

This paper proposes a novel Soft-Actor-Critic (SAC) based Deep Reinforcement Learning (DRL) method for optimizing the cost of microgrid operation by leveraging load flexibility. The proposed SAC-DRL method is designed to coordinate the control of distributed energy resources (DERs) and flexible load, addressing practical energy billing formation by power distribution utilities. Key contributions include an innovative reward function to mitigate sparse reward challenges and a mixed control strategy for discrete and continuous variables, ensuring radial network topology and minimizing power loss. We evaluate the proposed method on the model of a real microgrid located in Southern California, U.S.. The SAC-DRL model is tested to demonstrate its efficacy in reducing grid dependence, optimizing resource use, and minimizing costs. The results highlight the potential of DRL in modern energy systems, offering a sustainable and economically efficient solution for energy management in microgrids.

deep reinforcement learning↗

Reinforcement Learning to Enhance Optimal Operation of Resilient Community Energy Systems

This paper presents a novel model-free multi-agent Reinforcement Learning (RL) control method to enhance the resilience of community energy systems in island mode, which coordinates multiple objectives without the necessity of identifying system models that require expert knowledge. Specifically, a community-level coordinator agent is designed to allocate renewable energy resources among different buildings, and multiple building-level agents are developed to optimize load schedules based on limited energy resources and requirements of building loads and occupants’ comfort. In a two-day evaluation, our RL approach demonstrated a similar performance against MPC without requiring system models and formulation of optimization problems as required in MPC.

ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATION↗

Microsecond-latency feedback at a particle accelerator by online reinforcement learning on hardware

The commissioning and operation of future large-scale scientific experiments will challenge current tuning and control methods. Reinforcement learning (RL) algorithms are a promising solution due to their ability to dynamically adapt to changing environments and consider delayed consequences. In many real-world applications, RL policies must produce actions in real time, often within microseconds to milliseconds, imposing significant constraints on system latency and computational overhead that conventional machine learning libraries are not designed to handle. To control phenomena in real time at these timescales, RL needs to be deployed on-the-edge, namely on dedicated hardware located near the system it controls, without relying on a host CPU or cloud-based inference. In this work we present the design and deployment of an experience accumulator system in a particle accelerator. In this system, deep-RL algorithms run using hardware acceleration and act within a few microseconds, enabling the use of RL for control of phenomena like beam instabilities. The training uses the collected data offline to reduce the number of operations carried out on the acceleration hardware. The proposed architecture was tested in real experimental conditions at the Karlsruhe research accelerator, a synchrotron light source, where the system was used to control artificially induced horizontal betatron oscillations in real-time, with a control loop period of just 2.7 μs. The results showed a performance comparable to the commercial feedback system available at the accelerator, demonstrating the viability and potential of this approach. Due to the self-learning and reconfiguration capability of this implementation, a seamless application to other control problems is possible. Applications range from particle accelerators to large-scale research and industrial facilities.

FPGA↗

Reinforcement learning for online adaptation of model predictive controllers: Application to a selective catalytic reduction unit

Here we present a novel application of reinforcement learning (RL) for online dynamic tuning of model predictive controllers (MPC). Applying a state-action-reward-state-action (SARSA) algorithm for temporal difference learning with a control-specific reward function improves the error tracking performance of a standard MPC formulation. The proposed RL approach is also readily adaptable to other MPCs, or entirely different control approaches. Practical details for the implementation of the RL-MPC algorithm are also presented. The proposed algorithm is applied to a case study of controlling nitrogen oxide (NO x ) emissions in an industrial selective catalytic reduction (SCR) unit, a control problem characterized by significant nonlinearity and time delay. Along with an RL-MPC formulation for NOx control, another MPC is proposed to mitigate ammonia slip and decrease ammonia consumption in the SCR. Results showing the efficacy of the RL-MPC for NO x control through learning and implementation on the nonlinear SCR dynamic model are presented.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Efficient Reinforcement Learning for Real-Time Hardware-Based Energy System Experiments: Preprint

In the context of urgent climate challenges and the pressing need for rapid technology development, Reinforcement Learning (RL) stands as a compelling data-driven method for controlling real-world physical systems. However, RL implementation often entails time-consuming and computationally intensive data collection and training processes, rendering them inefficient for real-time applications that lack non-real-time models. To address these limitations, real-time emulation techniques have emerged as valuable tools for the lab-scale rapid prototyping of intricate energy systems. While emulated systems offer a bridge between simulation and reality, they too face constraints, hindering comprehensive characterization, testing, and development. In this research, we construct a surrogate model using limited data from simulated systems, enabling an efficient and effective training process for a Double Deep Q-Network (DDQN) agent for future deployment. Our approach is illustrated through a hydropower application, demonstrating the practical impact of our approach on climate-related technology development.

deep Q-learning↗

A reinforcement learning approach to long-horizon operations, health, and maintenance supervisory control of advanced energy systems

In this work, we develop a Reinforcement Learning (RL) approach to the supervisory control problem for advanced energy systems, such as novel nuclear reactors and other demand-driven, mission-critical, and component-health-sensitive energy plants. The inclusive problem landscape considered captures the stochastic confluence of plant performance, component health evolution, power demand from the grid, diverse maintenance actions, and operator-defined goals and constraints, all considered over meaningfully long-enough reasoning horizons. Key aspects of the proposed approach are a receding horizon control-inspired technique dictating time- or event-triggered supervisory policy (re-)constructions, as well as additional capability-enabling contributions such as timescale compression, to handle long reasoning horizons and uncertainty in parts of the problem, and practical yet demonstrably-effective handling of hybrid action spaces with continuous and discrete decision variables. The resulting algorithm consists of a simulation-based RL agent constructing stochastic supervisory control policies over nontrivial action spaces and for long horizons, applying the learned policy to the system for a much shorter interval, and perpetually repeating, to construct the next long-horizon policy. That next policy will only be applied, again, for a short interval, yet originally far-in-time events move progressively closer, their associated uncertainty decreases, and new events and aspects enter the reasoning horizon. The proposed methodology bridges fundamental receding horizon concepts with the unequivocally stronger and more scalable reasoning of contemporary RL. Numerical examples using Soft Actor–Critic Deep RL illustrate the operation and efficacy of the proposed technique for a power plant tasked with health-aware load following missions in a dynamic electricity market landscape.

97 MATHEMATICS AND COMPUTING↗

Curriculum-based Reinforcement Learning for Distribution System Critical Load Restoration

This paper focuses on the critical load restoration problem in distribution systems following major outages. To provide fast online response and optimal sequential decision-making support, a reinforcement learning (RL) based approach is proposed to optimize the restoration. Due to the complexities stemming from the large policy search space, renewable uncertainty, and nonlinearity in a complex grid control problem, directly applying RL algorithms to train a satisfactory policy requires extensive tuning to be successful. To address this challenge, this paper leverages the curriculum learning (CL) technique to design a training curriculum involving a simpler steppingstone problem that guides the RL agent to learn to solve the original hard problem in a progressive and more effective manner. We demonstrate that compared with direct learning, CL facilitates controller training to achieve better performance. To study realistic scenarios where renewable forecasts used for decision-making are in general imperfect, the experiments compare the trained RL controllers against two model predictive controllers (MPCs) using renewable forecasts with different error levels and observe how these controllers can hedge against the uncertainty. Results show that RL controllers are less susceptible to forecast errors than the baseline MPCs and can provide a more reliable restoration process.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Energy performance evaluation of the ASHRAE Guideline 36 control and reinforcement learning–based control using field measurements

This study evaluates the energy performance of ASHRAE Guideline 36–compliant control (ASHRAE 36 control) and reinforcement learning (RL)–based control through experimental field tests and a simulation study. Three field tests were conducted at Oak Ridge National Laboratory’s commercial building test facility in Oak Ridge, Tennessee: a baseline with a baseline conventional control, a test with ASHRAE 36 control, and a test with RL-based control. The selected ASHRAE 36 controls were trim and respond control, as well as variable air volume (VAV) box control. We compared the measured supply air temperature of the rooftop unit, VAV box supply air temperature, and VAV box supply airflow rate across the three test cases. The field data indicated that ASHRAE 36 controls operated as specified by ASHRAE Guideline 36. Based on these data, ASHRAE 36 control achieved a 45 % reduction in hourly averaged HVAC energy consumption compared with the baseline, and RL-based control achieved a 66 % reduction. These potential annual energy savings were confirmed using a calibrated whole-building energy model. Compared with the baseline, ASHRAE 36 control reduced HVAC energy consumption by 42 %, and RL-based control achieved a 54 % reduction. Furthermore, RL-based control reduced total HVAC energy consumption by 21 % more than ASHRAE 36 control.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Safe Exploration Reinforcement Learning for Load Restoration using Invalid Action Masking

This paper addresses the load restoration problem after a power outage event. Our primary proposed methodology uses a multi-agent reinforcement learning method to make the optimal sequential decisions on picking up critical loads. Typically, a negative reward is provided to discourage the agents from selecting decisions that violate physical constraints during the restoration process. However, the main disadvantage of this approach is its difficulty in applying it to large-scale systems due to the curse of dimensionality. This paper introduces the invalid action masking technique to overcome this limitation. The features of this technique include zero physical constraint violations, reduced training time, and stabilization of the explo- ration process. Simulation results are performed in IEEE 13-node and IEEE 123-node systems showing the better performance of the proposed algorithm in comparison to the conventional approaches both in terms of restored power and learning curve.

reinforcement learning, blackstart, artificial int↗

GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Large language models (LLMs) are increasingly adapted to downstream tasks via reinforcement learning (RL) methods like Group Relative Policy Optimization (GRPO), which often require thousands of rollouts to learn new tasks. We argue that the interpretable nature of language often provides a much richer learning medium for LLMs, compared to policy gradients derived from sparse, scalar rewards. To test this, we introduce GEPA (Genetic-Pareto), a prompt optimizer that thoroughly incorporates natural language reflection to learn high-level rules from trial and error. Given any AI system containing one or more LLM prompts, GEPA samples trajectories (e.g., reasoning, tool calls, and tool outputs) and reflects on them in natural language to diagnose problems, propose and test prompt updates, and combine complementary lessons from the Pareto frontier of its own attempts. As a result of GEPA's design, it can often turn even just a few rollouts into a large quality gain. Across six tasks, GEPA outperforms GRPO by 6% on average and by up to 20%, while using up to 35x fewer rollouts. GEPA also outperforms the leading prompt optimizer, MIPROv2, by over 10% (e.g., +12% accuracy on AIME-2025), and demonstrates promising results as an inference-time search strategy for code optimization. We release our code at https://github.com/gepa-ai/gepa.

97 MATHEMATICS AND COMPUTING↗

Exploring Transfers Between Earth-Moon Halo Orbits via Multi-Objective Reinforcement Learning

Multi-Reward Proximal Policy Optimization, a multi-objective deep reinforcement learning algorithm, is used to examine the design space of low-thrust trajectories for a SmallSat transferring between two libration point orbits in the Earth-Moon system. Using Multi-Reward Proximal Policy Optimization, multiple policies are simultaneously and efficiently trained on three distinct trajectory design scenarios. Each policy is trained to create a unique control scheme based on the trajectory design scenario and assigned reward function: a unique combination of weights scaling competing objectives that guide the spacecraft to the target mission orbit, incentivize faster flight times, and penalize propellant mass usage. Then, the policies are evaluated on the same set of perturbed initial conditions in each scenario to generate the propellant mass usages, flight times, and state discontinuities from a reference trajectory for each control scheme. This solution space of low-thrust trajectories for a SmallSat is used to examine the multi-objective trade space for the trajectory design scenario. By autonomously constructing the solution space, insights into the required propellant mass, flight time, and transfer geometry are rapidly achieved.

Christopher J Sullivan↗

Exploring Transfers Between Earth-Moon Halo Orbits via Multi-Objective Reinforcement Learning

Multi-Reward Proximal Policy Optimization, a multi-objective deep reinforcement learning algorithm, is used to examine the design space of low-thrust trajectories for a SmallSat transferring between two libration point orbits in the Earth- Moon system. Using Multi-Reward Proximal Policy Optimiza- tion, multiple policies are simultaneously and efficiently trained on three distinct trajectory design scenarios. Each policy is trained to create a unique control scheme based on the trajectory design scenario and assigned reward function: a unique combination of weights scaling competing objectives that guide the spacecraft to the target mission orbit, incentivize faster flight times, and penalize propellant mass usage. Then, the policies are evaluated on the same set of perturbed initial conditions in each scenario to generate the propellant mass usages, flight times, and state discontinuities from a reference trajectory for each control scheme. This solution space of low-thrust trajectories for a SmallSat is used to examine the multi-objective trade space for the trajectory design scenario. By autonomously constructing the solution space, insights into the required propellant mass, flight time, and transfer geometry are rapidly achieved.

Mashiku, Alinda K.↗

Unsupervised, supervised and reinforced learning via spiking computation

The present invention relates to unsupervised, supervised and reinforced learning via spiking computation. The neural network comprises a plurality of neural modules. Each neural module comprises multiple digital neurons such that each neuron in a neural module has a corresponding neuron in another neural module. An interconnection network comprising a plurality of edges interconnects the plurality of neural modules. Each edge interconnects a first neural module to a second neural module, and each edge comprises a weighted synaptic connection between every neuron in the first neural module and a corresponding neuron in the second neural module.

Modha, Dharmendra S.↗

Distributed Power Allocation for 6-GHz Unlicensed Spectrum Sharing via Multi-agent Deep Reinforcement Learning

We consider the problem of power allocation over the 6 GHz Unlicensed National Information Infrastructure (UNII)- 5 spectrum. We propose a novel deep Reinforcement Learning (DRL)-based distributed power allocation scheme which utilizes the multi-agent Deep Deterministic Policy Gradient (MADDPG) algorithm. In particular, we model the base stations (BSs) as DRL agents that simultaneously determine the transmit powers to their scheduled user equipment (UE) in a synchronized manner. The power decision of each BS is based on its own observation of the radio environment, which consists of several local interference measurements and a limited amount of information obtained from other BSs. One advantage of the proposed scheme is that it addresses the single-agent non-stationarity problem of RL in the multi-agent scenario by incorporating the actions and observations of other BSs into each BS’s own critic which helps it to gain a more accurate perception of the overall radio environment. A centralized-training-distributed execution framework is used to train the policies where the critics are trained over the joint actions and observations of all BSs while the actor of each BS only takes the local observation as input in order to produce the transmit power. Simulation shows that the proposed power allocation scheme can achieve better throughput performance than several state-of-the-art approaches.

99 GENERAL AND MISCELLANEOUS↗