Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

A safe reinforcement learning algorithm for supervisory control of power plants

Traditional control theory-based methods require tailored engineering for each system and constant fine-tuning. In power plant control, one often needs to obtain a precise representation of the system dynamics and carefully design the control scheme accordingly. Model-free Reinforcement learning (RL) has emerged as a promising solution for control tasks due to its ability to learn from trial-and-error interactions with the environment. It eliminates the need for explicitly modeling the environment’s dynamics, which is potentially inaccurate. However, the direct imposition of state constraints in power plant control raises challenges for standard RL methods. To address this, we propose a chance-constrained RL algorithm based on Proximal Policy Optimization for supervisory control. Our method employs Lagrangian relaxation to convert the constrained optimization problem into an unconstrained objective, where trainable Lagrange multipliers enforce the state constraints. In conclusion, our approach achieves the smallest distance of violation and violation rate in a load-follow maneuver for an advanced Nuclear Power Plant design.

constrained optimization↗

Multi-Agent Deep Reinforcement Learning for Realistic Distribution System Voltage Control Using PV Inverters

Over the last few decades, the deployment of distributed solar photovoltaic (PV) systems has increased consistently. High PV penetration could cause adverse effects on the grid, such as voltage violations. This paper proposes a new distributed soft actor-critic based multi-agent deep reinforcement learning (SAC-MADRL) control solution to minimize the PV real power curtailment while keeping the grid voltage in an acceptable range. New reward functions have been designed to coordinate different agents during the learning process, yielding improved convergence. Comparison results with other control methods on a real feeder in western Colorado U.S. with 80% penetration of PVs demonstrate that the proposed method has better capability of effectively regulating voltage while minimizing the PV real power curtailment.

distribution system↗

Reinforcement Learning for Volt- Var Control: A Novel Two-stage Progressive Training Strategy

This paper develops a reinforcement learning (RL) approach to solve a cooperative, multi-agent Volt-Var Control (VVC) problem for high solar penetration distribution systems. The ingenuity of our RL method lies in a novel two-stage progressive training strategy that can effectively improve training speed and convergence of the machine learning algorithm. In Stage 1 (individual training), while holding all the other agents inactive, we separately train each agent to obtain its own optimal VVC actions in the action space: fconsume, generate, do-nothingg. In Stage 2 (cooperative training), all agents are trained again coordinatively to share VVC responsibility. Rewards and costs in our RL scheme include (i) a system-level reward (for taking an action), (ii) an agent-level reward (for doing-nothing), and (iii) an agent-level action cost function. This new framework allows rewards to be dynamically allocated to each agent based on their contribution while accounting for the trade-off between control effectiveness and action cost. The proposed methodology is tested and validated in a modified IEEE 123-bus system using realistic PV and load profiles. Simulation results confirm that the proposed approach is robust and computationally efficient; and it achieves desirable volt-var control performance under a wide range of operation conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Topology-Aware Reinforcement Learning for Voltage Control: Centralized and Decentralized Strategies

Volt-VAR control (VVC) methods based on deep reinforcement learning (DRL) can effectively control distribution grid voltage and minimize power loss by implementing corrective and preventive control measures on the reactive power output of inverter-based distributed energy resources (DERs). However, model-free DRL-based VVC approaches usually cannot capture the important topological feature of the power system since they use a fully-connected network (FCN) to deliver the action. Therefore, this paper proposes a graph convolutional network (GCN)-based DRL approach that can employ the topological information of the network to take better control action for regulating the voltage. Our implementation allows for both centralized and decentralized configurations, utilizing a single agent and multiple agents respectively. Although the centralized GCN-based DRL approach has its advantages of minimizing voltage fluctuation and power loss, it is not suitable for large scale power systems due to its challenges in terms of scalability, computation speed and potential single points of failure. Therefore, these problems can be resolved using the decentralized GCN-based DRL approach. Moreover, to ensure the safe operation of the model, our proposed approach incorporates an exponential barrier function while formulating the reward function for each agent. To validate performance of the proposed approaches, the proposed model is tested on modified IEEE test systems and the performances are measured in terms on voltage fluctuation reduction, minimization of power loss and computational speed. Finally, the results show that the proposed topology-aware approach outperforms the FCN-based DRL approach in terms of reducing voltage fluctuation and minimizing power loss of the network. Moreover, it is shown that the decentralized GCN-based DRL has faster computational speed than other approaches.

42 ENGINEERING↗

MSD CoP Webinar: Modeling the Operations of Reservoir Systems with LLMs and Inverse Reinforcement Learning

Context: This panel featured three presentations centered on the common theme of applying LLMs and inverse reinforcement learning (IRL) to capture the complex human-environment interactions that are central to the operation of reservoir systems. Dr. Wyatt Arnold will kick off the webinar with a talk on how analyzing LLM chain-of-thought reasoning reveals sophisticated quantitative justification and risk awareness, showing promise as a bridge between quantitative models and value-driven water management decisions. Next, Dr. Matteo Giuliani will build on this with a discussion demonstrating that AI- and IRL-driven approaches can infer the trade-offs between flood control and water supply using historical observations. Finally, Dr. Rohan Singh Wilkho will close the webinar with a talk establishing IRL as a generalizable diagnostic tool for decoding decision-making in managed hydrologic and human-infrastructure systems. Across the three presentations, the application of LLMs and IRL opens new possibilities for the development of adaptive, transparent, and human-aware models supporting water management in an increasingly uncertain future. Presenters: Wyatt Arnold (Politecnico di Milano); Matteo Giuliani (Politecnico di Milano); Rohan Singh Wilkho (Cornell University) Moderator: Patrick M. Reed (MSD CoP Facilitation Team); Stefano Galelli (MSD CoP AI Working Group Co-Chair); David Gold (MSD CoP AI Working Group Co-Chair) This webinar was held on: June 23rd, 2026 from 12-1 PM EST.

Arnold, Wyatt [Politecnico di Milano]↗

Distributed reinforcement learning and consensus control of energy systems

Disclosed herein are methods, systems, and devices for utilizing distributed reinforcement learning and consensus control to most effectively generate and utilize energy. In some embodiments, individual turbines within a wind farm may communicate to reach a consensus as to the desired yaw angle based on the wind conditions.

King, Jennifer Rose Rose↗

Discovering mechanisms for materials microstructure optimization via reinforcement learning of a generative model

Abstract The design of materials structure for optimizing functional properties and potentially, the discovery of novel behaviors is a keystone problem in materials science. In many cases microstructural models underpinning materials functionality are available and well understood. However, optimization of average properties via microstructural engineering often leads to combinatorically intractable problems. Here, we explore the use of the reinforcement learning (RL) for microstructure optimization targeting the discovery of the physical mechanisms behind enhanced functionalities. We illustrate that RL can provide insights into the mechanisms driving properties of interest in a 2D discrete Landau ferroelectrics simulator. Intriguingly, we find that non-trivial phenomena emerge if the rewards are assigned to favor physically impossible tasks, which we illustrate through rewarding RL agents to rotate polarization vectors to energetically unfavorable positions. We further find that strategies to induce polarization curl can be non-intuitive, based on analysis of learned agent policies. This study suggests that RL is a promising machine learning method for material design optimization tasks, and for better understanding the dynamics of microstructural simulations.

36 MATERIALS SCIENCE↗

Optimization of the FRIB beam dump: a hybrid genetic algorithm and reinforcement learning approach

The operational envelope of high-power-density systems, such as particle accelerators and advanced nuclear energy systems, is critically constrained by the need to manage extreme thermal loads. To address this, we present a novel hybrid optimization framework combining a genetic algorithm (GA) with a soft actor-critic (SAC) deep reinforcement learning agent. This framework was applied to a practical high-heat-flux problem: redesigning the beam dump at the Facility for Rare Isotope Beams (FRIB) for a power upgrade from 20 kW to 50 kW. The resulting design, validated by three-dimensional conjugate heat transfer simulations, suppresses hazardous hot spots and yields a markedly more uniform temperature distribution. This provides a robust operating margin, increasing the average power-handling capability by 72% relative to the current design, demonstrating the framework’s potential to solve complex thermal management challenges in both accelerator technology and advanced nuclear systems.

Accelerator↗

Technical Report for Bayesian Optimization and Reinforcement Learning for Beam Polarization Increase in the BNL Hadron Injectors

This project developed and evaluated physics-informed Bayesian learning and machine learning (ML)-based optimization methods for improving beam polarization preservation in the BNL hadron injector chain. The work focused on uncertainty-aware digital twin modeling, Bayesian calibration of accelerator simulations using beam measurements, and data-efficient optimization strategies including Bayesian optimization and reinforcement learning. These methods were applied to injector tuning and RF control problems in realistic accelerator settings to support improved operational robustness and readiness for RHIC operations and future Electron–Ion Collider facilities. No subject inventions were disclosed under this award.

43 PARTICLE ACCELERATORS↗

Deep reinforcement learning for optimal control of induction welding process

Optimizing induction welding (IW) process parameters for the application of joining thermoplastic composites is challenging as it requires achieving complex spatiotemporal thermal characteristics along the weld-line to obtain desired weld quality. We formulate an optimal control problem which captures these requirements and seeks to optimize the IW coil speed using a fast-acting dynamic IW process model. We develop a novel Deep Reinforcement Learning (DRL) framework to solve this computationally challenging control problem and demonstrate via simulation study that the learned DRL feedback control policy results in better spatiotemporal thermal characteristics as compared to the current state-of-the-art.

36 MATERIALS SCIENCE↗

Reinforcement Learning for Spacecraft Navigation & Environment Characterization in the Planar-Restricted Two-Body Problem

As science, exploration, and commercial space missions become increasingly complex, so does the need for efficient, autonomous, and integrated spacecraft navigation and operations techniques. Key operational functions, including data collection and transmission, environment characterization, systems constraints, human factors, and navigation, often are intertwined and conflicted. Deep Reinforcement Learning (DRL) offers a framework for addressing integrated spacecraft navigation and planning in an uncertain dynamical environment. The goal of this study is to evaluate the utility of DRL for integrated spacecraft navigation and planning. This is achieved by developing a simple environmental characterization training environment in the Planar-Restricted 2-Body Problem (PR2BP), establishing benchmarks and heuristic baselines, and designing a previously unstudied Markov Decision Process (MDP) formulation. This MDP formulation enables the spacecraft DRL agents to appropriately balance navigation and actuation capabilities. The resulting DRL-derived policy exceeds a random or untrained policy and meets or exceeds the level of performance of a heuristic without actuation. In the process, valuable intuition is gained about the problem with insight into how DRL methods could scale to increasingly more realistic scenarios, including net-work design and training architectures, efficient state space representations, and methods for encouraging exploration in a parametric action space, among others.

navigation↗

Adaptive Load Shedding for Grid Emergency Control via Deep Reinforcement Learning

Emergency control, typically such as under-voltage load shedding (UVLS), is broadly used to grapple with low voltage and voltage instability issues in real-world power systems under contingencies. However, existing emergency control schemes are rule-based and cannot be adaptively applied to uncertain and floating operating conditions. Here, we propose an adaptive UVLS algorithm for emergency control via deep reinforcement learning (DRL) and expert systems. We first construct dynamic components for picturing the power system operation as the environment. The transient voltage recovery criteria, which poses time-varying requirements to UVLS, is integrated into the states and reward function to advise the learning of deep neural networks. The proposed method has no tuning issue of coefficients in reward functions, and this issue was regarded as a deficiency in the existing DRL-based algorithms. Case studies illustrate that the proposed method outperforms the traditional UVLS relay in both the timeliness and efficacy for emergency control.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Restoring Distribution System Under Renewable Uncertainty Using Reinforcement Learning

Distributed energy resources (DER) in distribution systems, including renewable generation, micro-turbine, and energy storage, can be used to restore critical loads following extreme events to increase grid resiliency. However, properly coordinating multiple DERs in the system for multi-step restoration process under renewable uncertainty and fuel availability is a complicated sequential optimal control problem. Due to its capability to handle system non-linearity and uncertainty, reinforcement learning (RL) stands out as a potentially powerful candidate in solving complex sequential control problems. Moreover, the offline training of RL provides excellent action readiness during online operation, making it suitable to problems such as load restoration, where in-time, correct and coordinated actions are needed. In this study, a distribution system prioritized load restoration based on a simplified single-bus system is studied: with imperfect renewable generation forecast, the performance of an RL controller is compared with that of a deterministic model predictive control (MPC). Our experiment results show that the RL controller is able to learn from experience, adapt to the imperfect forecast information and provide a more reliable restoration process when compared with the baseline MPC controller.

61 RADIATION PROTECTION AND DOSIMETRY↗

Restoring Distribution System Under Renewable Uncertainty Using Reinforcement Learning

Distributed energy resources (DERs) in distribution systems, including renewable generation, micro-turbine, and energy storage, can be used to restore critical loads following extreme events to increase grid resiliency. However, properly coordinating multiple DERs in the system for multi-step restoration process under renewable uncertainty and fuel availability is a complicated sequential optimal control problem. Due to its capability to handle system non-linearity and uncertainty, reinforcement learning (RL) stands out as a potentially powerful candidate in solving complex sequential control problems. Moreover, the offline training of RL provides excellent action readiness during online operation, making it suitable to problems such as load restoration, where in-time, correct and coordinated actions are needed. In this study, a distribution system prioritized load restoration based on a simplified single-bus system is studied: with imperfect renewable generation forecast, the performance of an RL controller is compared with that of a deterministic model predictive control (MPC). Our experiment results show that the RL controller is able to learn from experience, adapt to the imperfect forecast information and provide a more reliable restoration process when compared with the baseline controller.

61 RADIATION PROTECTION AND DOSIMETRY↗

Restoring Distribution System Under Renewable Uncertainty Using Reinforcement Learning: Preprint

Distributed energy resources (DER) in distribution systems, including renewable generation, micro-turbine, and energy storage, can be used to restore critical loads following extreme events to increase grid resiliency. However, properly coordinating multiple DERs in the system for multi-step restoration process under renewable uncertainty and fuel availability is a complicated sequential optimal control problem. Due to its capability to handle system non-linearity and uncertainty, reinforcement learning (RL) stands out as a potentially powerful candidate in solving complex sequential control problems. Moreover, the offline training of RL provides excellent action readiness during online operation, making it suitable to problems such as load restoration, where in-time, correct and coordinated actions are needed. In this study, a distribution system prioritized load restoration based on a simplified single-bus system is studied: with imperfect renewable generation forecast, the performance of an RL controller is compared with that of a deterministic model predictive control (MPC). Our experiment results show that the RL controller is able to learn from experience, adapt to the imperfect forecast information and provide a more reliable restoration process when compared with the baseline MPC controller.

61 RADIATION PROTECTION AND DOSIMETRY↗

Designing reinforcement learning algorithms for building HVAC control: From experimental observation to simulation comparisons

Advanced supervisory-level control with reinforcement learning (RL) is regarded as a promising solution for HVAC systems to minimize energy consumption while maintaining thermal comfort and indoor air quality. However, most RL applications were conducted in the simulation environment rather than real-world HVAC systems. This paper developed a value-based RL controller termed Deep Q-Network (DQN) for a typical central HVAC system and evaluated its performance in a building test facility. By comparing DQN with a rule-based controller, the study not only demonstrated the cases where DQN could properly maintain indoor comfort but also discussed possible reasons why DQN failed in some other situations. Recognizing the limitations of value-based RL algorithms from the experimental tests, a simulation study was conducted to compare DQN with an alternative RL approach, an actor–critic algorithm termed Deep Deterministic Policy Gradient (DDPG). In scenarios with a relatively large action space, DDPG outperformed DQN by requiring fewer computational resources and achieving better thermal comfort, lower energy consumption, and more stable control actions. The findings suggest that the ability of DDPG to handle continuous control variables more effectively allows for faster convergence in training and more precise control in practice, which enhances the overall efficiency and reliability of the HVAC system.

Guo, Fangzhou↗

Harnessing deep reinforcement learning to construct time-dependent optimal fields for quantum control dynamics

Here, we present an efficient deep reinforcement learning (DRL) approach to automatically construct time-dependent optimal control fields that enable desired transitions in dynamical chemical systems. Our DRL approach gives impressive performance in constructing optimal control fields, even for cases that are difficult to converge with existing gradient-based approaches. We provide a detailed description of the algorithms and hyperparameters as well as performance metrics for our DRL-based approach. Our results demonstrate that DRL can be employed as an effective artificial intelligence approach to efficiently and autonomously design control fields in quantum dynamical chemical systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Adaptive Reinforcement Learning (ARL) Control of a Multi-port Resonant Converter in UAV Systems

This study presents an adaptive reinforcement learning (ARL) control framework for a multi-port resonant converter used in hybrid unmanned aerial vehicle (UAV) power systems. The converter integrates high-frequency half-bridge input ports connected to a rectified engine–generator set and a battery energy storage system, along with a semi-bridgeless active rectifier supplying the propulsion load. A deep RL agent is trained to dynamically regulate inter-port phase-shift commands in real time based on flight conditions and load power demand. The ARL controller autonomously identifies phase-shift combinations that maximize conversion efficiency while maintaining stable and coordinated power flow, even under rapidly varying operating scenarios. This data-driven approach eliminates the need for explicit system modeling or extensive manual tuning and enables coordinated control among multiple power ports without inter-port communication. Experimental results validate that the ARL based strategy achieves reliable power sharing and consistently high-efficiency operation across diverse UAV operating conditions.

Asa, Erdem [ORNL] (ORCID:0000000190884812)↗