Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

GraMeR: Gra ph Me ta R einforcement learning for multi-objective influence maximization

Influence maximization (IM) is a combinatorial problem of identifying a subset of seed nodes in a network (graph), which when activated, provide a maximal spread of influence in the network for a given diffusion model and a budget for seed set size. IM has numerous applications such as viral marketing, epidemic control, sensor placement and other network-related tasks. However, its practical uses are limited due to the computational complexity of current algorithms. Recently, deep reinforcement learning has been leveraged to solve IM in order to ease the computational burden. However, there are serious limitations in current approaches, including narrow IM formulation that only consider influence via spread and ignore self-activation, low scalability to large graphs, and lack of generalizability across graph families leading to a large running time for every test network. In this work, we address these limitations through a unique approach that involves: (1) Formulating a generic IM problem as a Markov decision process that handles both intrinsic and influence activations; (2)incorporating generalizability via meta-learning across graph families. There are previous works that combine deep reinforcement learning with graph neural network, but this work solves a more realistic IM problem and incorporates generalizability across graphs via meta reinforcement learning. Extensive experiments are carried out in various standard networks to validate performance of the proposed Graph Meta Reinforcement learning (GraMeR) framework. Finally, the results indicate that GraMeR is multiple orders faster and generic than conventional approaches when applied on small to medium scale graphs.

97 MATHEMATICS AND COMPUTING↗

Decentralized Distributed Proximal Policy Optimization (DD-PPO) for High Performance Computing Scheduling on Multi-User Systems

Resource allocation in High Performance Computing (HPC) environments presents a complex and multifaceted challenge for job scheduling algorithms. Beyond the efficient allocation of system resources, schedulers must account for and optimize multiple performance metrics, including job wait time and system throughput. Traditional heuristic-based scheduling algorithms increasingly struggle and lack the efficiency needed to meet the demands and address the complexity and scale of modern HPC systems. Consequently, recent research efforts have focused on leveraging advancements in Artificial Intelligence (AI) and Deep Learning (DL), particularly Reinforcement Learning (RL), to develop more adaptable and intelligent scheduling strategies. Previous RL-based scheduling approaches have explored a range of algorithms, from Deep Q-Networks (DQN) to Proximal Policy Optimization (PPO), and more recently, hybrid methods that integrate Graph Neural Networks (GNNs) with RL techniques. However, a common limitation across these methods is their reliance on relatively small datasets, with few methods being evaluated using large-scale, multi-million-job trace datasets representative of real-world HPC workloads. Moreover, existing RL schedulers face scalability issues due to centralized policy updates, which hinder training efficiency and performance when applied to large datasets. This study introduces a novel RL-based scheduler utilizing Decentralized Distributed Proximal Policy Optimization (DD-PPO) algorithm, which supports large-scale distributed training across multiple workers without requiring parameter synchronization at every step. By eliminating reliance on centralized updates to a shared policy, the DD-PPO scheduler enhances scalability, training efficiency, and sample utilization. Experimental validation using a large real-world dataset containing over 11.5 million job traces collected from petascale HPC systems over six years assesses the influence of dataset scale on training effectiveness and compares DD-PPO performance to traditional and advanced scheduling approaches. The experimental results demonstrate improved scheduling performance in comparison to both heuristic-based schedulers and existing RL-based scheduling algorithms.

AI↗

Deep RL for Fast Long-Horizon Operations Scheduling on NASA's Carruthers Geocorona Observatory Mission

Spacecraft operations scheduling is a highly constrained, long-horizon combinatorial optimization problem that traditionally relies on heuristics, constraint programming, or manual planning. We present a scalable deep reinforcement learning framework developed and deployed for NASA’s Carruthers Geocorona Observatory mission. Our framework introduces a macro-action abstraction known as activity blocks coupled with dynamic action-masking to navigate the intractably large search space and strictly enforce complex power, thermal, and instrument constraints. The resulting architecture generates globally feasible schedules with overwhelming probability, establishes operational trust, and executes a full training cycle in under six hours, circumventing the need for policy robustness by enabling rapid, on-demand retraining. Further, resulting schedules outperform baseline heuristics in scheduled science quality. The deep reinforcement learning framework was deployed as the default operational scheduler for the Carruthers Geocorona Observatory mission from the outset of the mission, demonstrating that deep reinforcement learning can be trusted for real spacecraft operations under complex, evolving constraints.

Geocorona↗

Machine Learning For Planetary Mining Applications

Robotic mining could prove to be an efficient method of mining resources for extended missions on the Moon or Mars. One component of robotic mining is scouting an area for resources to be mined by other robotic systems. Writing controllers for scouting can be difficult due to the need for fault tolerance, inter-agent cooperation, and agent problem solving. Reinforcement learning could solve these problems by enabling the scouts to learn to improve their performance over time. This work is divided into two sections, with each section addressing the use of machine learning in this domain. The first contribution of this work focuses on the application of reinforcement learning to mining mission analysis. Various mission parameters were modified and control policies were learned. Then agent performance was used to assess the effect of the mission parameters on the performance of the mission. The second contribution of this work explores the potential use of reinforcement learning to learn a controller for the scouts. Through learning, these scouts would improve their ability to map their surroundings over time.

Cook, Joshua↗

Harnessing the power of gradient-based simulations for multi-objective optimization in particle accelerators

Abstract Particle accelerator operation requires simultaneous optimization of multiple objectives. Multi-objective optimization (MOO) is particularly challenging due to trade-offs between the objectives. Evolutionary algorithms, such as genetic algorithms (GAs), have been leveraged for many optimization problems, however, they do not apply to complex control problems by design. This paper demonstrates the power of differentiability for solving MOO problems in particle accelerators using a deep differentiable reinforcement learning (DDRL) algorithm. We compare the DDRL algorithm with model-free reinforcement learning (MFRL), GA, and Bayesian optimization (BO) for simultaneous optimization of heat load and trip rates in the continuous electron beam accelerator facility. The underlying problem enforces strict constraints on both individual states and actions as well as cumulative (global) constraints on energy requirements of the beam. Using historical accelerator data, we develop a physics-based surrogate model which is differentiable and allows for back-propagation of gradients. The results are evaluated in the form of a Pareto-front with two objectives. We show that the DDRL outperforms MFRL, BO, and GA on high dimensional problems.

43 PARTICLE ACCELERATORS↗

Deep Reinforcement Learning-Based Control of Energy Storage for Interarea Oscillation Damping

With the increasing electricity consumption and lack of transmission investment, today's power systems are operated much closer to their limits, raising concerns of inter-area oscillations that deteriorate the system stability. Here, this article presents a novel energy storage placement and control approach for enhanced damping of interarea oscillations. Combining the residual analysis and dominant mode analysis, we are able to identify the advantageous locations for placing energy storage that achieve improved damping performance. To overcome the challenges, such as fixed control parameters and insufficient damping, we propose to use a deep reinforcement learning-based approach for energy storage control. A state-of-the-art guided surrogate-gradient-based evolutionary strategy is used to train a learning agent in a robust, efficient, and reproducible manner. Parallel computing is also adopted to speed up the training process. The proposed strategy has been tested on both medium and large-scale systems. The proposed methods have demonstrated their effectiveness in mitigating various interarea oscillations within a timeframe of 20 s, thereby averting system collapse and enhancing power grid stability effectively.

25 ENERGY STORAGE↗

Signal Whisperers: Enhancing Wireless Reception Using DRL-Guided Reflector Arrays

This paper presents a multi-agent reinforcement learning (MARL) approach for controlling adjustable metallic reflector arrays to enhance wireless signal reception in non-line-of-sight (NLOS) scenarios. Unlike conventional reconfigurable intelligent surfaces (RIS) that require complex channel estimation, our system employs a centralized training with decentralized execution (CTDE) paradigm where individual agents corresponding to reflector segments autonomously optimize reflector element orientation in three-dimensional space using spatial intelligence based on user location information. Through extensive ray-tracing simulations with dynamic user mobility, the proposed multi-agent beam-focusing framework demonstrates substantial performance improvements over single-agent reinforcement learning baselines, while maintaining rapid adaptation to user movement within one simulation step. Comprehensive evaluation across varying user densities and reflector configurations validates system scalability and robustness. The results demonstrate the potential of learning-based approaches for adaptive wireless propagation control.

deep reinforcement learning↗

Transient Optimization of a Gas Turbine Engine

Gas turbine engines are the primary power plants for modern commercial aircraft. Transients prompted by significant changes in thrust or power demand are common and unavoidable. Extreme transient scenarios such as those associated with a go-around during a landing attempt are possible and must be accounted for in the design of the engine and its controller. Engine transients tend to cause a reduction in compressor operability margin, which must be addressed by the engine control system and accounted for in the engine design to prevent events such as compressor stall/surge and combustor blow out. Transient operability concerns typically lead to compromises in the engine design that sacrifice efficiency and/or limit responsiveness. Transient operability is typically managed by logic that limits the fuel flow command. If this logic is not optimized, then the potential for valuable performance could be lost. This study presents a strategy for optimizing the transient limit logic and proposes a strategy for updating the control logic over the lifespan of the engine. The results demonstrate significant improvements in transient operability. For example, of the results at sea level static conditions demonstrated a 31% reduction in the usage of the high pressure compressor operability stack during a snap acceleration transient. Furthermore, a reinforcement learning algorithm is demonstrated to modify the transient logic as the engine degrades to minimize response time while respecting a prescribed compressor operability margin limit. A simple demonstration of the reinforcement learning algorithm resulted in a thrust response time reduction of ~11.8%.

transient↗

Transient Optimization of a Gas Turbine Engine

Gas turbine engines are the primary power plants for modern commercial aircraft. Transients prompted by significant changes in thrust or power demand are common and unavoidable. Extreme transient scenarios such as those associated with a go-around during a landing attempt are possible and must be accounted for in the design of the engine and its controller. Engine transients tend to cause a reduction in compressor operability margin, which must be addressed by the engine control system and accounted for in the engine design to prevent events such as compressor stall/surge and combustor blow out. Transient operability concerns typically lead to compromises in the engine design that sacrifice efficiency and/or limit responsiveness. Transient operability is typically managed by logic that limits the fuel flow command. If this logic is not optimized, then the potential for valuable performance could be lost. This study presents a strategy for optimizing the transient limit logic and proposes a strategy for updating the control logic over the lifespan of the engine. The results demonstrate significant improvements in transient operability. For example, of the results at sea level static conditions demonstrated a 31% reduction in the usage of the high pressure compressor operability stack during a snap acceleration transient. Furthermore, a reinforcement learning algorithm is demonstrated to modify the transient logic as the engine degrades to minimize response time while respecting a prescribed compressor operability margin limit. A simple demonstration of the reinforcement learning algorithm resulted in a thrust response time reduction of ~11.8%.

transient↗

Transient Optimization of a Gas Turbine Engine

Gas turbine engines are the primary power plants for modern commercial aircraft. Transients prompted by significant changes in thrust or power demand are common and unavoidable. Extreme transient scenarios such as those associated with a go-around during a landing attempt are possible and must be accounted for in the design of the engine and its controller. Engine transients tend to cause a reduction in compressor operability margin, which must be addressed by the engine control system and accounted for in the engine design to prevent events such as compressor stall/surge and combustor blow out. Transient operability concerns typically lead to compromises in the engine design that sacrifice efficiency and/or limit responsiveness. Transient operability is typically managed by logic that limits the fuel flow command. If this logic is not optimized, then the potential for valuable performance could be lost. This study presents a strategy for optimizing the transient limit logic and proposes a strategy for updating the control logic over the lifespan of the engine. The results demonstrate significant improvements in transient operability. For example, of the results at sea level static conditions demonstrated a 31% reduction in the usage of the high pressure compressor operability stack during a snap acceleration transient. Furthermore, a reinforcement learning algorithm is demonstrated to modify the transient logic as the engine degrades to minimize response time while respecting a prescribed compressor operability margin limit. A simple demonstration of the reinforcement learning algorithm resulted in a thrust response time reduction of ~11.8%.

transient↗

Refining fuzzy logic controllers with machine learning

In this paper, we describe the GARIC (Generalized Approximate Reasoning-Based Intelligent Control) architecture, which learns from its past performance and modifies the labels in the fuzzy rules to improve performance. It uses fuzzy reinforcement learning which is a hybrid method of fuzzy logic and reinforcement learning. This technology can simplify and automate the application of fuzzy logic control to a variety of systems. GARIC has been applied in simulation studies of the Space Shuttle rendezvous and docking experiments. It has the potential of being applied in other aerospace systems as well as in consumer products such as appliances, cameras, and cars.

Berenji, Hamid R.↗

Time-Extended Payoffs for Collectives of Autonomous Agents

A collective is a set of self-interested agents which try to maximize their own utilities, along with a a well-defined, time-extended world utility function which rates the performance of the entire system. In this paper, we use theory of collectives to design time-extended payoff utilities for agents that are both aligned with the world utility, and are "learnable", i.e., the agents can readily see how their behavior affects their utility. We show that in systems where each agent aims to optimize such payoff functions, coordination arises as a byproduct of the agents selfishly pursuing their own goals. A game theoretic analysis shows that such payoff functions have the net effect of aligning the Nash equilibrium, Pareto optimal solution and world utility optimum, thus eliminating undesirable behavior such as agents working at cross-purposes. We then apply collective-based payoff functions to the token collection in a gridworld problem where agents need to optimize the aggregate value of tokens collected across an episode of finite duration (i.e., an abstracted version of rovers on Mars collecting scientifically interesting rock samples, subject to power limitations). We show that, regardless of the initial token distribution, reinforcement learning agents using collective-based payoff functions significantly outperform both natural extensions of single agent algorithms and global reinforcement learning solutions based on "team games".

Tumer, Kagan↗

A Dynamic Pricing Method to Manage the Impact of EV Charging on the Grid Using RL

This work addresses the challenge of managing electrical vehicle (EV) charging loads on distribution feeders with the increase in deployment of fast charging stations. To mitigate the adverse impacts on feeder health, a novel dynamic grid-informed pricing approach is proposed. This approach leverages reinforcement learning (RL) to determine hourly charging prices based on real-time grid conditions. A synthetic environment was developed to train the reinforcement learning agent. A model of an IEEE 34-bus distribution feeder with EV charging stations has been developed in OpenDSS utilizing Caldera for realistic EV charging profiles. Test cases demonstrate that the dynamic pricing strategy achieves higher energy delivery to the EV end user at a lower cost compared to constant pricing methods, while lowering voltage deviations and congestion. This approach offers more granular price adjustments, responding dynamically to feeder conditions and potentially improving grid stability and efficiency. The communication architecture to implement this dynamic pricing method is described. This research contributes to the development of smart grid-informed charging solutions that can reduce the cost of charging to the end user while also helping the grid.

EV charging, dynamic pricing, grid-informed chargi↗

A dynamic pricing method to manage the impact of EV charging on the grid using RL

This work addresses the challenge of managing electrical vehicle (EV) charging loads on distribution feeders with the increase in deployment of fast charging stations. To mitigate the adverse impacts on feeder health, a novel dynamic grid-informed pricing approach is proposed. This approach leverages reinforcement learning (RL) to determine hourly charging prices based on real-time grid conditions. A synthetic environment was developed to train the reinforcement learning agent. A model of an IEEE 34-bus distribution feeder with EV charging stations has been developed in OpenDSS utilizing Caldera for realistic EV charging profiles. Test cases demonstrate that the dynamic pricing strategy achieves higher energy delivery to the EV end user at a lower cost compared to constant pricing methods, while lowering voltage deviations and congestion. This approach offers more granular price adjustments, responding dynamically to feeder conditions and potentially improving grid stability and efficiency. The communication architecture to implement this dynamic pricing method is described. This research contributes to the development of smart grid-informed charging solutions that can reduce the cost of charging to the end user while also helping the grid.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

blastforge

BlastForge is a python module meant to aid reinforcement learning research for geometry optimization research projects. The code is built to use PyTorch as a backend and will contain several model architectures and reinforcement learning training loops as well as helper python functions to evaluate a model’s performance during and after training. BlastForge is meant to be a small, focused python project to study moderate-complexity geometry optimization problems

Hickmann, Kyle [Los Alamos National Laboratory]↗

Advanced Computational Techniques for Improving Resilience of Critical Energy Infrastructure under Cyber-Physical Attacks

In this chapter, we present recent advances in improving the resilience of cyber-physical systems, especially with regards to energy systems. We provide discussions around various types of cyber-physical events that can cause disruptions and new advances in optimization, control, and reinforcement learning (RL) to deal with the challenges posed by such cyber-physical events. The presented methods range from distributed robust optimization, autonomous and coordinated control, reinforcement learning based resilient control and topology reconfiguration in Inter-System resilient control.

Nazir, Mohammad Nawaf [BATTELLE (PACIFIC NW LAB)]↗

ICE-RASSOR: Intelligent Capabilities Enhanced Regolith Advanced Surface Systems Operations Robot

NASA’s Regolith Advanced Surface Systems Operations Robot (RASSOR) is principally designed to mine and deliver regolith for In-Situ Resource Utilization (ISRU)processing. RASSOR’s design enables it to efficiently collect and deposit regolith, return collected material for processing, and myriad related ISRU activities. To reliably perform these operations on the lunar surface, RASSOR software and sensory systems need to be robust and maximize the information extracted from a reduced sensor payload. Herein, we present preliminary findings from the Intelligent Capabilities Enhanced RASSOR project. We apply supervised learning using real data to estimate the soil mass collected without the need for mass flow rate monitors or other explicate sensing techniques. We also create a reduced-order simulation environment to develop autonomous trenching controllers via reinforcement learning and prototype state estimation architectures. Our initial results suggest that excavated regolith mass can be inferred within 2.9% RMS error of full scale, and reinforcement learning for autonomous operations has learned viable trenching strategies and helped identify desirable sensing capabilities, arrangements, and considerations. Future work includes regolith mass estimation during dynamic operation, expanding our simulation to more complex environments, and transfer learning from simulation to hardware.

machine learning↗

ICE-RASSOR: Intelligent Capabilities Enhanced

NASA’s Regolith Advanced Surface Systems Operations Robot (RASSOR) is principally designed to mine and deliver regolith for In-Situ Resource Utilization (ISRU) processing. RAS-SOR’s design enables it to efficiently collect and deposit regolith, return collected material for processing, and myriad related ISRU activities. To reliably perform these operations on the lunar sur-face, RASSOR software and sensory systems need to be robust and maximize the information extracted from on-board sensing. Herein, we present preliminary findings from the Intelligent Capabilities Enhanced RASSOR project. We apply supervised learning using real data to estimate the soil mass collected without the need for mass flow rate monitors or other explicate sensing techniques. We also create a reduced-order simulation environment to develop autonomous trenching controllers via reinforcement learning and proto-type state estimation architectures. Our initial results suggest that excavated regolith mass can be inferred within 2.9% RMS error of full scale, and reinforcement learning for autonomous operations has learned viable trenching strategies and helped identify desirable sensing capabilities, arrangements, and considerations. Future work includes regolith mass estimation during dynamic operation, expanding our simulation to more complex environments, and transfer learning from simulation to hardware.

machine learning↗