Engineering PapersSearch

SEARCH · Engineering Papers

Results for “policy optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Interplanetary Low-Thrust Design Using Proximal Policy Optimization

This paper aims to demonstrate a reinforcement learning technique for developing complex, decision-making policies capable of planning interplanetary transfers.Using Proximal Policy Optimization (PPO), a neural network agent is trained to produce a closed-loop controller capable of transfers between Earth and Mars.The agent is trained in an environment that utilizes a medium fidelity solar electric propulsion model and a real ephemeris model of the Earth and Mars. The results are compared against those generated by the Evolutionary Mission Trajectory Generator (EMTG) tool.

proximal policy optimization

Optimal policies for identification of stochastic linear systems

The problem of designing closed-loop policies for identification of multiinput-multioutput linear discrete-time systems with random time-varying parameters is considered in this paper using a Bayesian approach. A sensitivity index gives a measure of performance for the closed-loop laws. The computation of the optimal laws is shown to be nontrivial, an exercise in stochastic control, but open-loop, affine, and open-loop feedback optimal inputs are shown to yield tractable problems. Numerical examples are given. For time-invariant systems, the criterion considered is shown to be related to the trace of the information matrix associated with the system.

Lopez-Toledo, A. A.

Exploring Transfers Between Earth-Moon Halo Orbits via Multi-Objective Reinforcement Learning

Multi-Reward Proximal Policy Optimization, a multi-objective deep reinforcement learning algorithm, is used to examine the design space of low-thrust trajectories for a SmallSat transferring between two libration point orbits in the Earth-Moon system. Using Multi-Reward Proximal Policy Optimization, multiple policies are simultaneously and efficiently trained on three distinct trajectory design scenarios. Each policy is trained to create a unique control scheme based on the trajectory design scenario and assigned reward function: a unique combination of weights scaling competing objectives that guide the spacecraft to the target mission orbit, incentivize faster flight times, and penalize propellant mass usage. Then, the policies are evaluated on the same set of perturbed initial conditions in each scenario to generate the propellant mass usages, flight times, and state discontinuities from a reference trajectory for each control scheme. This solution space of low-thrust trajectories for a SmallSat is used to examine the multi-objective trade space for the trajectory design scenario. By autonomously constructing the solution space, insights into the required propellant mass, flight time, and transfer geometry are rapidly achieved.

Christopher J Sullivan

Optimum equipment maintenance/replacement policy. Part 2: Markov decision approach

Dynamic programming was utilized as an alternative optimization technique to determine an optimal policy over a given time period. According to a joint effect of the probabilistic transition of states and the sequence of decision making, the optimal policy is sought such that a set of decisions optimizes the long-run expected average cost (or profit) per unit time. Provision of an alternative measure for the expected long-run total discounted costs is also considered. A computer program based on the concept of the Markov Decision Process was developed and tested. The program code listing, the statement of a sample problem, and the computed results are presented.

Charng, T.

Exploring the Low-Thrust Transfer Design Space in an Ephemeris Model via Multi-Objective Reinforcement Learning

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective reinforcement learning algorithm used to train multiple policies to uncover solutions within a multi-objective solution space. MRPPO is used in this paper to train policies to construct low-thrust transfers for a SmallSat from the vicinity of !2 to an !5 short period orbit in the Sun-Earth-Moon system. First, the policies are trained in this scenario in the circular restricted three-body problem. This information is used to initialize the policies before training in a higher-fidelity ephemeris model; a process known as transfer learning. The recovered segments of the solution space will be compared to fundamental dynamical structures to both examine the results of MRPPO in this complex design scenario and explore the effectiveness of transfer learning.

Christopher J Sullivan

Exploring the Low-Thrust Transfer Design Space in an Ephemeris Model via Multi-Objective Reinforcement Learning

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective reinforcement learning algorithm used to train multiple policies to uncover solutions within a multi-objective solution space. MRPPO is used in this paper to train policies to construct low-thrust transfers for a SmallSat from the vicinity of L2 to an L5 short period orbit in the Sun-Earth-Moon system. First, the policies are trained in this scenario in the circular restricted three-body problem. This information is used to initialize the policies before training in a higher-fidelity ephemeris model; a process known as transfer learning. The recovered segments of the solution space will be compared to fundamental dynamical structures to both examine the results of MRPPO in this complex design scenario and explore the effectiveness of transfer learning.

Christopher J. Sullivan

Exploring the Low-Thrust Transfer Design Space in an Ephemeris Model via Multi-Objective Reinforcement Learning

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective reinforcement learning algorithm used to train multiple policies to uncover solutions within a multi-objective solution space. MRPPO is used in this paper to train policies to construct low-thrust transfers for a SmallSat from the vicinity of 𝐿2 to an 𝐿5 short period orbit in the Sun-Earth-Moon system. First, the policies are trained in this scenario in the circular restricted three-body problem. This information is used to initialize the policies before training in a higher-fidelity ephemeris model; a process known as transfer learning. The recovered segments of the solution space will be compared to fundamental dynamical structures to both examine the results of MRPPO in this complex design scenario and explore the effectiveness of transfer learning.

Mashiku, Alinda K.

Optimal analytic rendezvous using Clohessy-Wiltshire equations

The optimal solution time that minimizes the sum of the two applied impulses necessary to rendezvous is obtained analytically for the Clohessy-Wiltshire equations with a linear gravity model assumption. A plume impingement inequality constraint on the solution is examined, and an optimal policy is developed. Numerical tests are conducted to verify the analysis and to illustrate the optimal solution algorithm.

Jezewski, D. J.

An analytic approach to optimal rendezvous using Clohessy-Wiltshire equations

An analytic approach is used to obtain the optimal solution time that minimizes the sum of the two applied impulses necessary to rendezvous for the Clohessy-Wiltshire equations. A plume impingement inequality constraint on the solution is examined, and an optimal policy is developed. Numerical tests are conducted to verify the analysis and to illustrate the optimal solution algorithm.

Jezewski, D. J.

An optimal repartitioning decision policy

A central problem to parallel processing is the determination of an effective partitioning of workload to processors. The effectiveness of any given partition is dependent on the stochastic nature of the workload. The problem of determining when and if the stochastic behavior of the workload has changed enough to warrant the calculation of a new partition is treated. The problem is modeled as a Markov decision process, and an optimal decision policy is derived. Quantification of this policy is usually intractable. A heuristic policy which performs nearly optimally is investigated empirically. The results suggest that the detection of change is the predominant issue in this problem.

Nicol, D. M.

MULTI-OBJECTIVE REINFORCEMENT LEARNING FOR LOW-THRUST TRANSFER DESIGN BETWEEN LIBRATION POINT ORBITS

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective reinforcement learning algorithm used to construct low-thrust transfers between periodic orbits in multi-body systems. Previous implementations of MRPPO have relied on a predefined reference transfer to successfully train each policy. In this paper, an algorithmic modification labeled the ‘moving reference’, is introduced to autonomously construct these reference trajectories during training. With this modification, MRPPO is used to recover various low-thrust transfers between two periodic orbits in the Earth-Moon circular restricted three-body problem to solve a multi-objective optimization problem. These results are then compared with the solutions recovered via a traditional optimization formulation.

algorithm

MULTI-OBJECTIVE REINFORCEMENT LEARNING FOR LOW-THRUST TRANSFER DESIGN BETWEEN LIBRATION POINT ORBITS

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective reinforcement learning algorithm used to construct low-thrust transfers between periodic orbits in multi-body systems. Previous implementations of MRPPO have relied on a predefined reference transfer to successfully train each policy. In this paper, an algorithmic modification labeled the ‘moving reference’, is introduced to autonomously construct these reference trajectories during training. With this modification, MRPPO is used to recover various low-thrust transfers between two periodic orbits in the Earth-Moon circular restricted three-body problem to solve a multi-objective optimization problem. These results are then compared with the solutions recovered via a gradient descent optimization scheme to validate the performance of MRPPO with the moving reference modification.

Christopher J. Sullivan

MULTI-OBJECTIVE REINFORCEMENT LEARNING FOR LOW-THRUST TRANSFER DESIGN BETWEEN LIBRATION POINT ORBITS

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective reinforcement learning algorithm used to construct low-thrust transfers between periodic orbits in multi-body systems. Previous implementations of MRPPO have relied on a predefined reference transfer to successfully train each policy. In this paper, an algorithmic modification labeled the ‘moving reference’, is introduced to autonomously construct these reference trajectories during training. With this modification, MRPPO is used to recover various low-thrust transfers between two periodic orbits in the Earth-Moon circular restricted three-body problem to solve a multi-objective optimization problem. These results are then compared with the solutions recovered via a gradient descent optimization scheme to validate the performance of MRPPO with the moving reference modification

Christopher John Sullivan

Multi-objective Reinforcement Learning for Low-thrust Transfer Design Between Libration Point Orbits

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective rein- forcement learning algorithm used to construct low-thrust transfers between pe- riodic orbits in multi-body systems. Previous implementations of MRPPO have relied on a predefined reference transfer to successfully train each policy. In this paper, an algorithmic modification labeled the ‘moving reference’, is introduced to autonomously construct these reference trajectories during training. With this modification, MRPPO is used to recover various low-thrust transfers between two periodic orbits in the Earth-Moon circular restricted three-body problem to solve a multi-objective optimization problem. These results are then compared with the solutions recovered via a gradient descent optimization scheme to validate the performance of MRPPO with the moving reference modification.

Anderson, Rodney L.

Multi-objective Reinforcement Learning for Low-thrust Transfer Design Between Libration Point Orbits

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective rein- forcement learning algorithm used to construct low-thrust transfers between pe- riodic orbits in multi-body systems. Previous implementations of MRPPO have relied on a predefined reference transfer to successfully train each policy. In this paper, an algorithmic modification labeled the ‘moving reference’, is introduced to autonomously construct these reference trajectories during training. With this modification, MRPPO is used to recover various low-thrust transfers between two periodic orbits in the Earth-Moon circular restricted three-body problem to solve a multi-objective optimization problem. These results are then compared with the solutions recovered via a gradient descent optimization scheme to validate the performance of MRPPO with the moving reference modification.

Anderson, Rodney L

Exploring Transfers Between Earth-Moon Halo Orbits via Multi-Objective Reinforcement Learning

Multi-Reward Proximal Policy Optimization, a multi-objective deep reinforcement learning algorithm, is used to examine the design space of low-thrust trajectories for a SmallSat transferring between two libration point orbits in the Earth- Moon system. Using Multi-Reward Proximal Policy Optimiza- tion, multiple policies are simultaneously and efficiently trained on three distinct trajectory design scenarios. Each policy is trained to create a unique control scheme based on the trajectory design scenario and assigned reward function: a unique combination of weights scaling competing objectives that guide the spacecraft to the target mission orbit, incentivize faster flight times, and penalize propellant mass usage. Then, the policies are evaluated on the same set of perturbed initial conditions in each scenario to generate the propellant mass usages, flight times, and state discontinuities from a reference trajectory for each control scheme. This solution space of low-thrust trajectories for a SmallSat is used to examine the multi-objective trade space for the trajectory design scenario. By autonomously constructing the solution space, insights into the required propellant mass, flight time, and transfer geometry are rapidly achieved.

Mashiku, Alinda K.

Markov Tracking for Agent Coordination

Partially observable Markov decision processes (POMDPs) axe an attractive representation for representing agent behavior, since they capture uncertainty in both the agent's state and its actions. However, finding an optimal policy for POMDPs in general is computationally difficult. In this paper we present Markov Tracking, a restricted problem of coordinating actions with an agent or process represented as a POMDP Because the actions coordinate with the agent rather than influence its behavior, the optimal solution to this problem can be computed locally and quickly. We also demonstrate the use of the technique on sequential POMDPs, which can be used to model a behavior that follows a linear, acyclic trajectory through a series of states. By imposing a "windowing" restriction that restricts the number of possible alternatives considered at any moment to a fixed size, a coordinating action can be calculated in constant time, making this amenable to coordination with complex agents.

Washington, Richard

Life cycle costing with a discount rate

This article studies life cycle costing for a capability needed for the indefinite future, and specifically investigates the dependence of optimal policies on the discount rate chosen. The two costs considered are reprocurement cost and maintenance and operations (M and O) cost. The procurement price is assumed known, and the M and O costs are assumed to be a known function, in fact, a non-decreasing function, of the time since last reprocurement. The problem is to choose the optimum reprocurement time so as to minimize the quotient of the total cost over a reprocurement period divided by the period. Or one could assume a discount rate and try to minimize the total discounted costs into the indefinite future. It is shown that the optimum policy in the presence of a small discount rate hardly depends on the discount rate at all, and leads to essentially the same policy as in the case in which discounting is not considered.

Posner, E. C.