Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Deep Reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Autonomous Spacecraft Attitude Control Using Deep Reinforcement Learning

While machine learning and spacecraft autonomy continue to gain research interest, significant work remains to be done in efficiently applying modern machine learning techniques to problems in space ight. This study presents a framework for deriving a discrete neural spacecraft attitude controller using reinforcement learning, a paradigm of machine learning, without the need for high-performance computing. The developed attitude controller is an approximately time-optimal solution to a highly constrained control problem, able to achieve well above industry-standard pointing accuracies. Control examples are also presented of the agent performing large-angle spacecraft slews in the developed simulation environment and future extensions of this work are discussed.

ATAP↗

Towards Autonomous Lunar Resource Excavation via Deep Reinforcement Learning

To support sustainable infrastructure on the Moon, NASA needs to leverage lunar resources for in-situ processing and construction. NASA’s Regolith Advanced Surface Systems Operations Robot (RASSOR) is principally designed to mine and deliver regolith for these tasks. To reliably perform these operations on the lunar surface, RASSOR's sensors and control systems need to be robust and maximize information extracted from a reduced sensor payload. Herein, we present our findings from the Intelligent Capabilities Enhanced RASSOR project. We created reduced-order simulation environments in which we applied reinforcement learning algorithms to learn autonomous trenching controllers and produced state estimation architectures. We developed two simulations: a 2D excavation simulation used to facilitate parameter selection, and a 3D simulation developed using a game physics engine to simulate simplified soil interactions and incorporate robotic agents parameterized by dynamic models. Within these simulations, we learned autonomous excavation routines that exceed excavation efficiency measures as compared against RASSOR's existing control and teleoperation-based methods.

RASSOR↗

Towards Autonomous Lunar Resource Excavation via Deep Reinforcement Learning

To support sustainable infrastructure on the Moon, NASA needs to leverage lunar resources for in-situ processing and construction. NASA’s Regolith Advanced Surface Systems Operations Robot (RASSOR) is principally designed to mine and deliver regolith for these tasks. To reliably perform these operations on the lunar surface, RASSOR's sensors and control systems need to be robust and maximize information extracted from a reduced sensor payload. Herein, we present our findings from the Intelligent Capabilities Enhanced RASSOR project. We created reduced-order simulation environments in which we applied reinforcement learning algorithms to learn autonomous trenching controllers and produced state estimation architectures. We developed two simulations: a 2D excavation simulation used to facilitate parameter selection, and a 3D simulation developed using a game physics engine to simulate simplified soil interactions and incorporate robotic agents parameterized by dynamic models. Within these simulations, we learned autonomous excavation routines that exceed excavation efficiency measures as compared against RASSOR's existing control and teleoperation-based methods.

RASSOR↗

Physics-informed Deep Reinforcement Learning-based Control in Power systems

Incorporating physics information into the deep reinforcement learning (DRL) process is a promising approach for addressing the challenges faced in learning-based control design problems for physical systems. Power grid dynamics, being a physical system, adheres to specific physical laws, constraints, as well as operational and control rules. Therefore, consideration of such physics-based law improves the learning process drastically. In general, traditional grid control schemes rely on rule-based mechanisms that cannot adapt to changing operating conditions. To improve the adaptability and computation time, recent research has seen a surge of DRL-based applications in power grid control. A generic DRL-based control design imposes the system performance requirements through the design of reward functions. In some cases, some of the important physics information is injected through this reward function. However, due to the complex dynamics and large state-action space, learning an optimal DRL policy often becomes challenging. Inspired by the latest developments in general machine learning (ML) research, power system researchers have been investigating more direct ways of incorporating physics knowledge into DRL training. This chapter specifically focuses on these aspects of physics-informed DRL designs in grid control. It discusses the significance, applications, research gaps, and open problems that need to be addressed in future research.

artificial intelligence, machine learning↗

A Sequential Model Predictive and Deep Reinforcement Learning-Based Controller for Distribution System Outage Mitigation under Hurricane Events

This paper proposes a proactive outage mitigation framework for power distribution networks to withstand hurricane-induced disruptions. It leverages Model Predictive Control (MPC) to identify safe lines for proactive switching during hurricanes, minimizing the risk of cascading failures and voltage violations. The switching strategies optimized by MPC are sequentially integrated with a Deep Reinforcement Learning agent using the Advantage Actor-Critic algorithm, enabling dynamic line switching to maximize connected buses and minimize voltage violations in real time. Using a probabilistic hurricane model, the framework predicts line failures and adapts to varying conditions to enhance grid resilience. Simulations on the IEEE 123-bus system demonstrate its effectiveness in maintaining high connectivity and minimizing disruptions. Real-time testing with an RTDS confirms the practicality and reliability of the proposed approach.

Selim, Alaa [University of Connecticut]↗

Deep RL for Fast Long-Horizon Operations Scheduling on NASA's Carruthers Geocorona Observatory Mission

Spacecraft operations scheduling is a highly constrained, long-horizon combinatorial optimization problem that traditionally relies on heuristics, constraint programming, or manual planning. We present a scalable deep reinforcement learning framework developed and deployed for NASA’s Carruthers Geocorona Observatory mission. Our framework introduces a macro-action abstraction known as activity blocks coupled with dynamic action-masking to navigate the intractably large search space and strictly enforce complex power, thermal, and instrument constraints. The resulting architecture generates globally feasible schedules with overwhelming probability, establishes operational trust, and executes a full training cycle in under six hours, circumventing the need for policy robustness by enabling rapid, on-demand retraining. Further, resulting schedules outperform baseline heuristics in scheduled science quality. The deep reinforcement learning framework was deployed as the default operational scheduler for the Carruthers Geocorona Observatory mission from the outset of the mission, demonstrating that deep reinforcement learning can be trusted for real spacecraft operations under complex, evolving constraints.

Geocorona↗

Deep Reinforcement Learning-Based Control of Energy Storage for Interarea Oscillation Damping

With the increasing electricity consumption and lack of transmission investment, today's power systems are operated much closer to their limits, raising concerns of inter-area oscillations that deteriorate the system stability. Here, this article presents a novel energy storage placement and control approach for enhanced damping of interarea oscillations. Combining the residual analysis and dominant mode analysis, we are able to identify the advantageous locations for placing energy storage that achieve improved damping performance. To overcome the challenges, such as fixed control parameters and insufficient damping, we propose to use a deep reinforcement learning-based approach for energy storage control. A state-of-the-art guided surrogate-gradient-based evolutionary strategy is used to train a learning agent in a robust, efficient, and reproducible manner. Parallel computing is also adopted to speed up the training process. The proposed strategy has been tested on both medium and large-scale systems. The proposed methods have demonstrated their effectiveness in mitigating various interarea oscillations within a timeframe of 20 s, thereby averting system collapse and enhancing power grid stability effectively.

25 ENERGY STORAGE↗

ReLIC: Full-Scale Realization of Reinforcement Learning for Infrastructure Control

Prior efforts have shown that deep reinforcement learning (DRL) may provide a new method for controlling networked power systems. Though successful, prior approaches have not yet demonstrated their behavior on systems of realistic scale. This effort examined multiple theoretical and technical approaches to allow a DRL model to operate over a system of 2,000 buses or more. We find that allowing the DRL models to run training episodes in parallel provides near limitless efficiency gains, allowing us to train successful agents to behave on our Kuramoto transmission model of up to 4,000 buses. We further show that we can expand our PowerWorld DRL implementation to systems of up to 25 buses but struggle to go beyond this limit due to PowerWorld’s inability to run multiple instances at once. Finally, we examine a multi-agent approach and find that it performs as well if not better than our existing centralized approach.

97 MATHEMATICS AND COMPUTING↗

GraMeR: Gra ph Me ta R einforcement learning for multi-objective influence maximization

Influence maximization (IM) is a combinatorial problem of identifying a subset of seed nodes in a network (graph), which when activated, provide a maximal spread of influence in the network for a given diffusion model and a budget for seed set size. IM has numerous applications such as viral marketing, epidemic control, sensor placement and other network-related tasks. However, its practical uses are limited due to the computational complexity of current algorithms. Recently, deep reinforcement learning has been leveraged to solve IM in order to ease the computational burden. However, there are serious limitations in current approaches, including narrow IM formulation that only consider influence via spread and ignore self-activation, low scalability to large graphs, and lack of generalizability across graph families leading to a large running time for every test network. In this work, we address these limitations through a unique approach that involves: (1) Formulating a generic IM problem as a Markov decision process that handles both intrinsic and influence activations; (2)incorporating generalizability via meta-learning across graph families. There are previous works that combine deep reinforcement learning with graph neural network, but this work solves a more realistic IM problem and incorporates generalizability across graphs via meta reinforcement learning. Extensive experiments are carried out in various standard networks to validate performance of the proposed Graph Meta Reinforcement learning (GraMeR) framework. Finally, the results indicate that GraMeR is multiple orders faster and generic than conventional approaches when applied on small to medium scale graphs.

97 MATHEMATICS AND COMPUTING↗

Reinforcement Learning Approach to Flight Control Allocation with Distributed Electric Propulsion

The flight control system of the SUSAN Electrofan concept aircraft achieves attitude control using both conventional flight control surfaces and differential thrust through distributed electric propulsion (DEP) from sixteen wing-mounted electric engines. The introduction of eight pairs of wing fans for attitude control creates a highly actuated system. Such a system requires more sophisticated control to operate, especially in the presence of wingfan failures where the loss of a single wingfan can result in a thrust imbalance. This paper investigates the use of deep reinforcement learning (RL) using proximal policy optimization (PPO) to achieve attitude control through a combination of DEP and control surface deflections. First, the paper examines the aircraft undergoing a coordinated turn. Then, it examines the aircraft experiencing a wingfan failure during cruise conditions. It is shown that deep reinforcement learning can be a potential avenue for nonlinear flight control design.

Distributed Electric Propulsion↗

Risk-Aware Reinforcement Learning Framework for User-Centric O-RAN

The evolution of Open Radio Access Networks (O-RAN) presents an opportunity to enhance network performance by enabling dynamic orchestration of configuration and optimization parameters (COPs) through online learning methods. However, leveraging this potential requires overcoming the limitations of traditional cell-centric RAN architectures, which lack the necessary flexibility. On the other hand, despite their recent popularity, the practical deployment of online learning frameworks, such as Deep Reinforcement Learning (DRL)-based COP optimization solutions, remains limited due to their risk of deteriorating network performance during the exploration phase. In this article, we propose and analyze a novel risk-aware DRL framework for user-centric RAN (UC-RAN), which offers both the architectural flexibility and COP optimization to exploit this flexibility. We investigate and identify UC-RAN COPs that can be optimized via a soft actor-critic algorithm implementable as an O-RAN application (rApp) to jointly maximize latency satisfaction, reliability satisfaction, area spectral efficiency, and energy efficiency. We use the offline learning on UC-RAN to reliably accelerate DRL training, thus minimizing the risk of DRL deteriorating cellular network performance. Results show that our proposed solution approaches near-optimal performance in just a few hundred iterations with a decrease in risk score by a factor of ten.

6G and beyond↗

Exploring Transfers Between Earth-Moon Halo Orbits via Multi-Objective Reinforcement Learning

Multi-Reward Proximal Policy Optimization, a multi-objective deep reinforcement learning algorithm, is used to examine the design space of low-thrust trajectories for a SmallSat transferring between two libration point orbits in the Earth-Moon system. Using Multi-Reward Proximal Policy Optimization, multiple policies are simultaneously and efficiently trained on three distinct trajectory design scenarios. Each policy is trained to create a unique control scheme based on the trajectory design scenario and assigned reward function: a unique combination of weights scaling competing objectives that guide the spacecraft to the target mission orbit, incentivize faster flight times, and penalize propellant mass usage. Then, the policies are evaluated on the same set of perturbed initial conditions in each scenario to generate the propellant mass usages, flight times, and state discontinuities from a reference trajectory for each control scheme. This solution space of low-thrust trajectories for a SmallSat is used to examine the multi-objective trade space for the trajectory design scenario. By autonomously constructing the solution space, insights into the required propellant mass, flight time, and transfer geometry are rapidly achieved.

Christopher J Sullivan↗

Exploring Transfers Between Earth-Moon Halo Orbits via Multi-Objective Reinforcement Learning

Multi-Reward Proximal Policy Optimization, a multi-objective deep reinforcement learning algorithm, is used to examine the design space of low-thrust trajectories for a SmallSat transferring between two libration point orbits in the Earth- Moon system. Using Multi-Reward Proximal Policy Optimiza- tion, multiple policies are simultaneously and efficiently trained on three distinct trajectory design scenarios. Each policy is trained to create a unique control scheme based on the trajectory design scenario and assigned reward function: a unique combination of weights scaling competing objectives that guide the spacecraft to the target mission orbit, incentivize faster flight times, and penalize propellant mass usage. Then, the policies are evaluated on the same set of perturbed initial conditions in each scenario to generate the propellant mass usages, flight times, and state discontinuities from a reference trajectory for each control scheme. This solution space of low-thrust trajectories for a SmallSat is used to examine the multi-objective trade space for the trajectory design scenario. By autonomously constructing the solution space, insights into the required propellant mass, flight time, and transfer geometry are rapidly achieved.

Mashiku, Alinda K.↗

Topology-Aware Reinforcement Learning for Voltage Control: Centralized and Decentralized Strategies

Volt-VAR control (VVC) methods based on deep reinforcement learning (DRL) can effectively control distribution grid voltage and minimize power loss by implementing corrective and preventive control measures on the reactive power output of inverter-based distributed energy resources (DERs). However, model-free DRL-based VVC approaches usually cannot capture the important topological feature of the power system since they use a fully-connected network (FCN) to deliver the action. Therefore, this paper proposes a graph convolutional network (GCN)-based DRL approach that can employ the topological information of the network to take better control action for regulating the voltage. Our implementation allows for both centralized and decentralized configurations, utilizing a single agent and multiple agents respectively. Although the centralized GCN-based DRL approach has its advantages of minimizing voltage fluctuation and power loss, it is not suitable for large scale power systems due to its challenges in terms of scalability, computation speed and potential single points of failure. Therefore, these problems can be resolved using the decentralized GCN-based DRL approach. Moreover, to ensure the safe operation of the model, our proposed approach incorporates an exponential barrier function while formulating the reward function for each agent. To validate performance of the proposed approaches, the proposed model is tested on modified IEEE test systems and the performances are measured in terms on voltage fluctuation reduction, minimization of power loss and computational speed. Finally, the results show that the proposed topology-aware approach outperforms the FCN-based DRL approach in terms of reducing voltage fluctuation and minimizing power loss of the network. Moreover, it is shown that the decentralized GCN-based DRL has faster computational speed than other approaches.

42 ENGINEERING↗

User-Centric Communication With Aerial Network for 6G: A Reinforcement Learning Approach

Meeting the diverse needs of user verticals requires innovative cellular architectures that can offer additional degrees of freedom to provide on-demand services. The terrestrial user-centric radio access network (UC-RAN) stands out as an excellent choice for this purpose. However, a drawback of UC-RAN is its tendency to prioritize high-priority verticals, often resulting in a subpar quality of experience for low-priority verticals. This issue is particularly exacerbated in hotspot areas. Here, to address this problem, we introduce an aerial network integrated with terrestrial UC-RAN to provide coverage to users which are not served by the terrestrial network. Furthermore, we analyze the impact of key configuration and optimization parameters (COPs), such as location, transmit power, altitude, and beamwidth of aerial base stations (ABSs) on system key performance indicators (KPIs), such as coverage, latency satisfaction, average spectral efficiency, and energy efficiency. We formulate a robust multiobjective function to maximize these KPIs without biasing toward any specific KPI(s). Finally, we propose a deep reinforcement learning optimization framework based on the state-of-the-art soft actor-critic algorithm to control ABS COPs and optimize system KPIs. Experimental evaluations demonstrate that the proposed optimization framework can converge to near-optimal solutions derived from the pseudo brute force in a few thousand epochs.

6G↗

Design and optimization of a modular hydrogen-based integrated energy system to maximize revenue via nuclear-renewable sources

Here, this paper demonstrates a novel modular distributed framework that uses optimal energy-dispatching strategies to enable greater flexibility and profitability in nuclear-renewable integrated energy systems (NR-IES). Hydrogen is used as a commodity in this framework since its production can improve grid stability and system operational flexibility, decarbonize heavy industry, and create an additional revenue stream for electricity generators, particularly nuclear power plants with high operational expenses. The proposed solution addresses the challenges associated with merging multiple software and services from various domains by using functional mock-up units (FMU) to co-simulate diverse subsystems designed in various platforms. The tightly coupled integrated energy system (IES) is optimized to maximize revenue by utilizing the deep reinforcement learning (DRL) technique to make smart dispatching decisions based on variable electricity prices and the availability of renewable energy. Proximal policy optimization (PPO) algorithm is used in training and testing the DRL agent. Over a period of 120 days, the proposed hydrogen-based IES framework showed about 10% revenue boost compared to a non-hydrogen generating baseline IES while also providing an easily-adoptable framework which can help to improve the flexibility of future generation nuclear power plants.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗