New second-order and first-order algorithms for determining optimal control - A differential dynamic programming approach
Differential dynamic programming to develop second order and first order approximation methods for determining optimal control
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Differential dynamic programming to develop second order and first order approximation methods for determining optimal control
Differential dynamic programming algorithms of second and first order for optimal control, considering nonlinear unconstrained and constrained problems
Recently, Differential Dynamic Programming (DDP) and other similar algorithms have become the solvers of choice when performing non-linear Model Predictive Control (nMPC) with modern robotic devices. The reason is that they have a lower computational cost per iteration when compared with off-the-shelf Non-Linear Programming (NLP) solvers, which enables its online operation. However, they cannot handle constraints, and are known to have poor convergence capabilities. In this paper, we propose a method to solve the optimal control problem with control bounds through a squashing function (i.e., a sigmoid, which is bounded by construction). It has been shown that a naive use of squashing functions damage the convergence rate. To tackle this, we first propose to add a quadratic barrier that avoids the difficulty of the plateau produced by the sigmoid. Second, we add an outer loop that adapts both the sigmoid and the barrier; it makes the optimal control problem with the squashing function converge to the original control-bounded problem. To validate our method, we present simulation results for different types of platforms including a multi-rotor, a biped, a quadruped and a humanoid robot.
Differential dynamic programming (DDP) has been demonstrated as a viable approach to low-thrust trajectory optimization, namely with the recent success of NASA's Dawn mission. The Dawn trajectory was designed with the DDP-based Static/Dynamic Optimal Control algorithm used in the Mystic software.1 Another recently developed method, Hybrid Differential Dynamic Programming (HDDP),2, 3 is a variant of the standard DDP formulation that leverages both first-order and second-order state transition matrices in addition to nonlinear programming (NLP) techniques. Areas of improvement over standard DDP include constraint handling, convergence properties, continuous dynamics, and multi-phase capability. DDP is a gradient based method and will converge to a solution nearby an initial guess. In this study, monotonic basin hopping (MBH) is employed as a stochastic search method to overcome this limitation, by augmenting the HDDP algorithm for a wider search of the solution space.
Differential dynamic programming (DDP) has been demonstrated as a viable approach to low-thrust trajectory optimization, namely with the recent success of NASAs Dawn mission. The Dawn trajectory was designed with the DDP-based Static Dynamic Optimal Control algorithm used in the Mystic software. Another recently developed method, Hybrid Differential Dynamic Programming (HDDP) is a variant of the standard DDP formulation that leverages both first-order and second-order state transition matrices in addition to nonlinear programming (NLP) techniques. Areas of improvement over standard DDP include constraint handling, convergence properties, continuous dynamics, and multi-phase capability. DDP is a gradient based method and will converge to a solution nearby an initial guess. In this study, monotonic basin hopping (MBH) is employed as a stochastic search method to overcome this limitation, by augmenting the HDDP algorithm for a wider search of the solution space.
Differential dynamic programming algorithm for discrete time orbit transfer optimization
Discrete time differential dynamic programming algorithm with application to optimal orbit transfer
The emerging urban air mobility (UAM) sector in aerospace is driving development of unconventional multi-modal vehicle configurations and autonomous flight. The combination of multi-modal vehicle dynamics, complex environment, requirements to deal with flight contingencies in an efficient and safe manner, as well as necessity for precise trajectory following and performance, are the driving influence behind adaptive optimization for system performance. We are interested in trajectory optimization algorithm that would system parameter estimation and identifying the optimal switching time between modes of hybrid dynamical systems. This presentation discusses a parameterized optimal control trajectory optimization algorithm that is an extended and generalized version of Differential Dynamic Programming (DDP), titled Parameterized Differential Dynamic Programming (PDDP). DDP is an efficient trajectory optimization algorithm relying on second order approximations of a system’s dynamics and cost function and has recently been applied to optimize systems with time invariant parameters. Experiments are presented applying PDDP to solve model predictive control (MPC) and moving horizon estimation (MHE) tasks simultaneously. In particular, PDDP is used to determine the optimal transition point between flight regimes of a complex urban air mobility (UAM) class vehicle exhibiting multiple phases of flight and to identify and compensate for actuation faults.
Differential dynamic programming methods for solving bang bang control problems
This paper presents an integration of Differential Dynamic Programming (DDP) with the Optimal Reciprocal Collision Avoidance (ORCA) algorithm as the basis for a new algorithm, titled Combined Bernstein Polynomial Optimal Reciprocal Collision Avoidance DDP (COBRA-DDP), for trajectory replanning and collision avoidance for Urban Air Mobility (UAM) vehicles. State-constrained variants of DDP provide the ability to plan trajectories while avoiding obstacles, but these methods require a large increase in computational time per iteration which hinders the overall speed of the algorithm. ORCA utilizes simplified dynamics to recognize potential collisions along a trajectory and provides an optimal velocity for the avoidance of multiple vehicles. These velocity commands, however, may not result in a dynamically feasible trajectory for DDP to plan around. As such, a Bernstein polynomial curve that considers the general dynamic constraints of the vehicle is generated to approximate a trajectory based on the velocity commands. COBRA-DDP optimizes this suggested trajectory via unconstrained DDP to provide a dynamically feasible trajectory that provides collision avoidance. This new trajectory can be applied to the vehicle or used to warm start the state constrained DDP algorithms to decrease computation time. Its benefits and effectiveness of the algorithm are demonstrated on a UAM Vertical Takeoff and Landing (VTOL) vehicle simulation with highly nonlinear dynamics.
Presentation as part of a AIAA SciTech conference GNC Workshop. The presentation discusses parameterized differential dynamic programming algorithm and its application in the context of urban air mobility and autonomous flight.
Explore the source record for details and available documents.
Nonlinear bang-bang optimal control problems solution based on differential dynamic programming, noting use of Pontryagin adjoint variables
This paper explores two optimal control approaches, widely used in robotics, to establish their viability as real-time trajectory planners for vehicle configurations envisioned for the emerging aviation sector of Urban Air Mobility (UAM). Differential Dynamic Programming (DDP) enables planning over highly nonlinear dynamics using second-order approximations along a nominal trajectory, and displays quadratic convergence to a local solution. Model Predictive Path Integral (MPPI) is a stochastic sampling-based algorithm that can optimize for general cost criteria, including potentially highly nonlinear formulations, and supports parallel computation through the use of modern GPU hardware. In this work, DDP and MPPI were implemented using model predictive control (MPC), and the results indicate they are able to successfully transition the aircraft over different flight envelopes and generate trajectories unique to UAM vehicles.
Low-thrust trajectories about planetary bodies characteristically span a high count of orbital revolutions. Directing the thrust vector over many revolutions presents a challenging optimization problem for any conventional strategy. This paper demonstrates the tractability of low-thrust trajectory optimization about planetary bodies by applying a Sundman transformation to change the independent variable of the spacecraft equations of motion to the eccentric anomaly and performing the optimization with differential dynamic programming. Fuel-optimal geocentric transfers are shown in excess of 1000 revolutions while subject to Earths J2 perturbation and lunar gravity.
Low-thrust trajectories about planetary bodies characteristically span a high count of orbital revolutions. Directing the thrust vector over many revolutions presents a challenging optimization problem for any conventional strategy. This paper demonstrates the tractability of low-thrust trajectory optimization about planetary bodies by applying a Sundman transformation to change the independent variable of the spacecraft equations of motion to the eccentric anomaly and performing the optimization with differential dynamic programming. Fuel-optimal geocentric transfers are shown in excess of 1000 revolutions while subject to Earths J2 perturbation and lunar gravity.
Explore the source record for details and available documents.
Explore the source record for details and available documents.