Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Joint Spectrum Access and Power Control in Air-Air Communications - A Deep Reinforcement Learning Based Approach

This paper considers the dynamic spectrum access and power control problem in a single-hop point-to-point Air-Air Communication Network (AACN). Due to spectrum scarcity, we assume the number of Aircraft-to-Aircraft (A2A) communication links is greater than that of the available channels, such that some communication links need to share the same channel, causing co-channel interference. We formulate the joint channel selection and power control optimization problem to maximize the Weighted Sum Spectral Efficiency (WSSE). A distributed and dynamic deep Q learning-based algorithm is proposed to find the optimal solution. Specifically, we design two different policies that are trained by conducting a trial-and-error scheme. Each communication link can achieve the optimal policy by exploiting the local information from its neighbors, and this distributive approach make it scalable to large networks. Finally, our experimental results demonstrate the effectiveness of the proposed solution in various AACN scenarios.

Zhe Wang

Joint Spectrum Access and Power Control in Air-Air Communications - A Deep Reinforcement Learning Based Approach

This paper considers the dynamic spectrum access and power control problem in a single-hop point-to-point Air-Air Communication Network (AACN). Due to spectrum scarcity, we assume the number of Aircraft-to-Aircraft (A2A) communication links is greater than that of the available channels, such that some communication links need to share the same channel, causing co-channel interference. We formulate the joint channel selection and power control optimization problem to maximize the Weighted Sum Spectral Efficiency (WSSE). A distributed and dynamic deep Q learning-based algorithm is proposed to find the optimal solution. Specifically, we design two different policies that are trained by conducting a trial-and-error scheme. Each communication link can achieve the optimal policy by exploiting the local information from its neighbors, and this distributive approach make it scalable to large networks. Finally, our experimental results demonstrate the effectiveness of the proposed solution in various AACN scenarios.

Zhe Wang

Runway Configuration Management with Offline Reinforcement Learning

Runway configuration management (RCM) is a challenging task, and it affects the efficiency of the National Airspace System (NAS) and airport surface operations significantly. Each airport, depending on the geometry, capacity, local climate patterns, etc. has multiple configurations for the runway usage for arriving and departing flights. Many factors such as the incoming/outgoing traffic load, wind direction and speed, convective weather, cloud ceiling and other environmental factors might affect the choice of a runway configuration at any point in time. However, other factors such as safety measures and regulations, noise abatement, capacity of each configuration, and preference of the air traffic controllers (ATCs) can also play a significant role in selecting the configuration. A sub-optimal selection of the runway configuration, or delay in making configuration changes might result in significant increase in taxi times for aircraft on the surface of the airport, fuel and energy use of the aircraft, and maintenance costs. It can also lead to safety concerns, such as an aircraft performing one or more go-arounds before being able to land. All these factors make RCM an extremely important and challenging decision-making process for the ATCs. The current state of practice sets the runway configuration by the ATCs based on relevant information available at the time including weather, traffic, noise abatement, safety bounds, etc. This makes the decision-making process subjective based on the accuracy of the available information and the bias in human decision making. Unfortunately, this approach yields poor results (e.g., significant delays) if the predicted outcomes are uncertain and their relative impact is not well understood. This is especially evident when the uncertainty increases the size of possible predicted outcomes (combinatorial explosion in possible scenarios) that cannot be handled by human reasoning. On the other hand, an automated approach based on machine intelligence can make use of historical data and search through all (or significant amount of) possible scenarios under uncertainty and make well-informed decisions.

Milad Memarzadeh

Multi-objective Reinforcement Learning for Low-thrust Transfer Design Between Libration Point Orbits

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective rein- forcement learning algorithm used to construct low-thrust transfers between pe- riodic orbits in multi-body systems. Previous implementations of MRPPO have relied on a predefined reference transfer to successfully train each policy. In this paper, an algorithmic modification labeled the ‘moving reference’, is introduced to autonomously construct these reference trajectories during training. With this modification, MRPPO is used to recover various low-thrust transfers between two periodic orbits in the Earth-Moon circular restricted three-body problem to solve a multi-objective optimization problem. These results are then compared with the solutions recovered via a gradient descent optimization scheme to validate the performance of MRPPO with the moving reference modification.

Anderson, Rodney L.

Multi-objective Reinforcement Learning for Low-thrust Transfer Design Between Libration Point Orbits

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective rein- forcement learning algorithm used to construct low-thrust transfers between pe- riodic orbits in multi-body systems. Previous implementations of MRPPO have relied on a predefined reference transfer to successfully train each policy. In this paper, an algorithmic modification labeled the ‘moving reference’, is introduced to autonomously construct these reference trajectories during training. With this modification, MRPPO is used to recover various low-thrust transfers between two periodic orbits in the Earth-Moon circular restricted three-body problem to solve a multi-objective optimization problem. These results are then compared with the solutions recovered via a gradient descent optimization scheme to validate the performance of MRPPO with the moving reference modification.

Anderson, Rodney L

Reinforcement Learning for Spacecraft Navigation & Environment Characterization in the Planar-Restricted Two-Body Problem

During mission planning and execution, spacecraft operators must balance data collection and downlink, systems constraints, human factors, and navigation. As missions become increasingly complex and ambitious, these factors become more intricately entwined and conflicted. For example, a spacecraft’s position must be known accurately in order to point to and image a target. Large position errors may cause missed observations or require additional scanning that increases operations complexity and data volume. Some observations require imaging from specific relative geometries which adds orbit control and timing considerations. Adjusting the orbit may allow for optimal observability of environmental parameters and/or enable more efficient sensor coverage, but maneuver execution error adds uncertainty to the current state which impacts both characterization and coverage objectives.

Navigation