Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Approximate dynamic programming”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A General Framework for Bounding Approximate Dynamic Programming Schemes

For years, there has been interest in approximation methods for solving dynamic programming problems, because of the inherent complexity in computing optimal solutions characterized by Bellman’s principle of optimality. A wide range of approximate dynamic programming (ADP) methods now exists. It is of great interest to guarantee that the performance of an ADP scheme be at least some known fraction, say ß , of optimal. This letter introduces a general approach to bounding the performance of ADP methods, in this sense, in the stochastic setting. The approach is based on new results for bounding greedy solutions in string optimization problems, where one has to choose a string (ordered set) of actions to maximize an objective function. This bounding technique is inspired by submodularity theory, but submodularity is not required for establishing bounds. Instead, the bounding is based on quantifying certain notions of curvature of string functions; the smaller the curvatures the better the bound. The key insight is that any ADP scheme is a greedy scheme for some surrogate string objective function that coincides in its optimal solution and value with those of the original optimal control problem. The ADP scheme then yields to the bounding technique mentioned above, and the curvatures of the surrogate objective determine the value ß of the bound. The surrogate objective and its curvatures depend on the specific ADP.

discrete event systems↗

Approximate Dynamic Programming With Enhanced Off-Policy Learning for Coordinating Distributed Energy Resources

Herein this paper proposes an innovative approximate dynamic programming (ADP) method for distributed energy resource coordination with the loss of life of battery energy storage system (BESS) explicitly modeled. The dispatch policy is designed to account for both calendrical and cyclical aging effects on BESS, explicitly modeling the impacts of ambient temperature on BESS lifespan. The proposed ADP employs an adaptive critic method and enhanced off-policy deterministic policy gradient (DPG) strategy, addressing the limitations of the on-policy gradient-based ADP approaches, including inadequate exploration, low data usage, and computational complexity. In particular, a customized policy is proposed to guide the algorithm to explore some promising decisions and thereby improve exploration capability and learning efficiency compared to conventional DPG-based learning approaches, which may struggle to find a global optimum due to random noisy action-based exploration or require expert demonstration with extra effort. The proposed method is illustrated using the IEEE 123-node system and compared with the existing ADP methods to prove solution accuracy and demonstrate the effects of incorporating degradation models into control design. Case studies showed that the proposed ADP effectively coordinates DERs with a 10 times smaller optimization gap compared to existing methods, and the incorporation of the BESS life loss model ensures the expected lifespan and results in significant cost savings.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Experimental Validation of Approximate Dynamic Programming Based Optimization and Convergence on Microgrid Applications

Stochastic optimization can better address uncertainties in power system problems. However, when state space and action space become large, many existing approaches become computationally expensive and even infeasible. Approximate dynamic programming (ADP) attracts researchers’ attention as a powerful tool for solving power system optimization problems with reduced computational cost. In this paper, in light of the existing literature, we investigate how the ADP approach with post-decision value function approximation converges to the nearly optimal solution with improved computational speed and experimentally validate the performance of the approach for a microgrid energy optimization problem. The approximation error versus the number of iteration is studied for convergence analysis of the post-decision ADP. A flowchart is provided to illustrate the proposed ADP algorithm for a microgrid energy optimization problem. The performance of ADP and dynamic programming (DP) is compared in terms of optimization error and computational time. It has found that the post-decision ADP approach can achieve competitive optimality with improved computational speed compared to the traditional DP.

Das, Avijit↗

Real-Time Ecodriving Control in Electrified Connected and Autonomous Vehicles Using Approximate Dynamic Programing

Connected and automated vehicles (CAVs), particularly those with a hybrid electric powertrain, have the potential to significantly improve vehicle energy savings in real-world driving conditions. In particular, the ecodriving problem seeks to design optimal speed and power usage profiles based on available information from connectivity and advanced mapping features to minimize the fuel consumption over an itinerary. This paper presents a hierarchical multilayer model predictive control (MPC) approach for improving the fuel economy of a 48 V mild-hybrid powertrain in a connected vehicle environment. Approximate dynamic programing (DP) is used to solve the receding horizon optimal control problem, whose terminal cost is approximated with the base policy obtained from the long-term optimization. The controller was tested virtually (with deterministic and Monte Carlo simulation) across multiple real-world routes, demonstrating energy savings of more than 20%. The controller was then deployed on a test vehicle equipped with a rapid prototyping embedded controller. In-vehicle testing confirm the energy savings obtained in simulation and demonstrate the real-time ability of the controller.

Automation & Control Systems↗

The Heuristic Dynamic Programming Approach in Boost Converters

In this study, a heuristic dynamic programming controller is proposed to control a boost converter. Conventional controllers such as proportional-integral-derivative (PID) or proportional-integral (PI) are designed based on the linearized small-signal model near the operating point. Therefore, the performance of the controller during the start-up, the load change, or the input voltage variation is not optimal since the system model changes by varying the operating point. The heuristic dynamic programming controller optimally controls the boost converter by following the approximate dynamic programming. The advantage of the HDP is that the neural network-based characteristic of the proposed controller enables boost converters to easily cope with large disturbances. An HDP with a well-trained critic and action networks can perform as an optimal controller for the boost converter. To compare the effectiveness of the traditional PI-based and the HDP boost converter, the simulation results are provided.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Data-based and secure switched cyber–physical systems

In this work, we develop a completely model-free moving target defense framework for the detection and mitigation of sensor and/or actuator attacks in cyber–physical systems with dynamics that evolve in discrete-time. We incorporate an intrusion detection mechanism based on an approximate dynamic programming technique that learns the policies for optimal regulation and optimal tracking while simultaneously defending against actuator and sensor attacks in a model-free fashion. Switching rules are leveraged to force proactive and reactive defense mechanisms as well as, guarantee the stability of the equilibrium point. Finally, as a case study, we apply the proposed moving target defense framework to a DC–DC converter that is used in electric vehicles.

42 ENGINEERING↗

Integrated Optimization of Powertrain Energy Management and Vehicle Motion Control for Autonomous Hybrid Electric Vehicles

Hybrid Electric Vehicles (HEVs) and autonomous vehicles have been widely studied recently for on-road transportation. In the study of autonomous HEVs, the control of the vehicle's external dynamics and powertrain dynamics are often treated separately. Optimizing these two problems together can significantly improve fuel economy. In this paper, an autonomous HEV following a leader is considered. First, the augmented model to integrate the abovementioned dynamics is presented. Second, the optimization problem is defined to find the optimum fuel consumption of the follower in pursuit of a leader in a drive cycle. A customized control strategy based on Approximate Dynamic Programming (ADP) is then explored in which the optimal cost-to-go at each time step is approximated using neural networks. Also, the accuracy of the optimization solution is enhanced by applying the concept of the reachable sets. At last, three case studies show that the examined integrated control strategy outperforms the one with the separated optimization method by an additional 7.4%, 4.6%, and 11.8% improvement in fuel consumption, respectively.

33 ADVANCED PROPULSION SYSTEMS↗

Optimal, centralized dynamic curbside parking space zoning

In this paper we formulate a dynamic mixed integer program for optimally zoning curbside parking spaces subject to transportation policy-inspired constraints and regularization terms. First, we illustrate how given some objective of curb zoning valuation as a function of zone type (paid parking, bus stop, etc.), dynamically rezoning involves unrolling this optimization program over a fixed time horizon. Second, we implement two different solution methods given an example curb zoning valuation. In the first method, we solve long horizon dynamic zoning problems via approximate dynamic programming. In the second method, we employ Dantzig-Wolfe decomposition to break-up the mixed-integer program into a master problem and several sub-problems that can be solved in parallel. This speeds up the computational solve-time of the MIP considerably. We present simulation results and comparisons of the different employed techniques on vehicle arrival-rate data obtained for a neighborhood in downtown Seattle, Washington, USA.

Nazir, Mohammad Nawaf↗

Optimal Coordination of Distributed Energy Resources Using Deep Deterministic Policy Gradient

Recent studies showed that reinforcement learning (RL) is a promising approach for coordination and control of distributed energy resources (DER) under uncertainties. Many existing RL approaches, including Q-learning and approximate dynamic programming, are based on lookup table methods, which become inefficient when the problem size is large and infeasible when continuous states and actions are involved. In addition, when modeling battery energy storage system (BESS), the loss of life is not reasonably considered into the decision-making process. This paper proposes an innovative deep RL method for DER coordination considering BESS degradation. The proposed deep RL is designed based on an adaptive actor-critic architecture and employs an off-policy deterministic policy gradient method for determining the dispatch operation that minimizes the operation cost and BESS life loss. Case studies were performed to validate the proposed method and demonstrate the effects of incorporating degradation models into control design.

Das, Avijit↗

Adaptive critic design-based reinforcement learning approach in controlling virtual inertia-based grid-connected inverters

In this report, an adaptive critic design (ACD) approach is proposed to control the phase and voltage of a grid-connected virtual synchronous generator (VSG). The penetration of fast responding inertia-less power converters significantly affect the stability of the power system, especially weak systems such as micro grids. The concept of virtual inertia addresses this concern by virtually emulating the behavior of a synchronous generator. However, the conventional VSG is designed based on two conditions: (i) fixed operating point and (ii) inductive grid connections. The performance of VSGs in low-voltage semi-resistive microgrids is far from optimal. To overcome the aforementioned concerns, a heuristic dynamic programing (HDP) approach is proposed to optimally control grid-connected VSGs. The neural-network-based inherence of the HDP enables the proposed technique to adapt to any impedance angle. The HDP controller includes two subnetworks: (i) the action network that controls the system optimally and (ii) the critic network, which evaluates the effectiveness of the action network. The simulation and experimental results are provided to evaluate the effectiveness of the proposed technique. As shown, the HDP-based approach illustrates a better performance in comparison with the conventional PI-based VSG in various operating conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

How good are learning-based control v.s. model-based control for load shifting? Investigations on a single zone building energy system

Both model predictive control (MPC) and deep reinforcement learning control (DRL) have been presented as a way to approximate the true optimality of a dynamic programming problem, and these two have shown significant operational cost saving potentials for building energy systems. Furthermore, there is still a lack of in-depth quantitative studies on their approximation levels to the true optimality, especially in the building energy domain. To fill in the gap, this paper provides a numerical framework that enables the evaluation of the optimality levels of different controllers for building energy systems. This framework is then used to comprehensively compare the optimal control performance of both MPC and DRL controllers with given computation budgets for a single zone fan coil unit system. Note the optimality is estimated based on a user-specific selection of trade-off weights among energy costs, thermal comfort and control slew rates. Compared with the best optimality we can find through expensive optimization simulations, the best DRL agent can maximally approximate the optimality by 96.54%, which outperforms the best MPC whose optimality level is 90.11%. However, due to the stochasticity, the DRL agent is only expected to approximate the optimality by 90.42%, which is almost equivalent to the best MPC. Except for Proximal Policy Optimization (PPO), all DRL agents can have a better approximation to the optimality than the best MPC, and are expected to have better approximation than the MPC with a prediction horizon of 32 steps (15 min per step). In terms of reducing energy cost and thermal discomfort, MPC can outperform the rule-based control (RBC) by 18.47%–25.44%. DRL can be expected to outperform RBC by 18.95%–25.65% ,and the best DRL control policy can outperform RBC by 20.29%–29.72%. Although the comparison of the optimality level is performed in a perfect setting, e.g., MPC assumes perfect models, and DRL assumes a perfect offline training process and online deployment process, this can shed insight on their capabilities of approximating to the original dynamic programming problem.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Hybrid data-driven and model-informed online tool wear detection in milling machines

Precision machining tool wear is responsible for low product throughput and quality. Monitoring the tool wear online is vital to prevent degradation in machining quality. However, direct real-time tool wear measurement is not practical. This paper presents residual-based anomaly detection models, combining a hybrid model comprised of a physics-based model and a data-driven model (a decision tree or a neural network) to predict signals of interest (e.g., power or forces) under nominal conditions, followed by Page’s cumulative sum test for detecting tool wear on-line using the computer numerical control machine measurements. The most informative features are ranked using dynamic programming and its approximation variants from real-time measurements and machine settings, such as the width of cut, depth of cut, feed rate and spindle speed, that serve as inputs to the predictive models. The baseline nominal model is incrementally updated with experimental data via a gradient boosted adaptation model to generate the residuals that account for discrepancies between the actual machine data under normal conditions and the baseline nominal model predictions. The hybrid model is validated against 20 Mazak milling machine experimental tests and one Haas run-to-failure experiment. The proposed anomaly detector is applied to synthetic data from simulations of the physics-based model at different operating conditions, measurement noise levels, and tool wear levels, and the methods were able to achieve an overall 92% accuracy in data with 1% noise. The anomaly detection methods based on hybrid model reduced the false alarms of either the data-driven or physical-based models alone, and are found to be capable of good online detection of tool wear.

42 ENGINEERING↗

Control of Fractional Diffusion Problems via Dynamic Programming Equations

In this study, we explore the approximation of feedback control of integro-differential equations containing a fractional Laplacian term. To obtain feedback control for the state variable of this nonlocal equation, we use the Hamilton–Jacobi–Bellman equation. It is well known that this approach suffers from the curse of dimensionality, and to mitigate this problem we couple semi-Lagrangian schemes for the discretization of the dynamic programming principle with the use of Shepard approximation. This coupling enables approximation of high-dimensional problems. Numerical convergence toward the solution of the continuous problem is provided together with linear and nonlinear examples. The robustness of the method with respect to disturbances of the system is illustrated by comparisons with an open-loop control approach.

97 MATHEMATICS AND COMPUTING↗

Management of Risk and Uncertainty Through Optimized Co-Operation of Transmission Systems and Microgrids with Responsive Loads (Final Report)

The evolution of the power system to the reliable, efficient and sustainable system of the future will involve development of both demand- and supply-side technology and operations. Ambitious national and state-level goals around the decarbonization of electricity relies on the integration of very high levels of renewable resources, most of which are variable and intermittent. The use of demand response is an ideal approach to counterbalance the intermittency of renewable generation and brings the consumer into the spotlight. Until recently, very little research had been conducted on the co-optimization of these two systems due to computational limitations. However, advances in computational capabilities, and the judicious use of decomposition methods and innovative approximation methods for high-dimension dynamic programming made this goal a viable objective for this project, leading to a fundamental shift in the ability to integrate and fully utilize demand-side resources. To this end, the modeling framework developed introduces a novel co-optimization framework, to include the operations of both the transmission and distribution systems (or microgrids) in operational decision making. This framework was used to analyze renewable and distributed generation along with responsive demand and to compare the capability of co-optimized systems to perform with higher levels of variable renewables. Results show that the use of a bi-level optimization approach is an appropriate structure, capable of co-optimizing a transmission system with multiple distribution systems and microgrids. While increasing the number of connected systems provides increasing flexibility for renewables integration this can also the economic benefits to the low-voltage subsystems with each additional system connected. Comparison of a traditional single-level decision structure with the co-optimization approach illustrates a reduction in overall system cost under co-optimization, while specific cost allocations to transmission and distribution systems are changed.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Management of Risk and Uncertainty Through Optimized Co-Operation of Transmission Systems and Microgrids With Responsive Loads (Final Report)

The evolution of the power system to the reliable, efficient and sustainable system of the future will involve development of both demand- and supply-side technology and operations. Ambitious national and state-level goals around the decarbonization of electricity relies on the integration of very high levels of renewable resources, most of which are variable and intermittent. The use of demand response is an ideal approach to counterbalance the intermittency of renewable generation and brings the consumer into the spotlight. Though individual consumers are interconnected at the low-voltage distribution system, these resources are typically modeled as variables at the transmission network level. Demand-side participation cannot be leveraged effectively without explicitly including the distribution system dynamics in the optimization-based wholesale market operations. This project grew from a vision for co-optimized interaction of distribution systems, or microgrids, with the high-voltage transmission system. In this framework, microgrids encompass consumers, distributed renewables and storage. The energy management system of the lower voltage system (distribution or microgrid) can also sell (buy) excess (necessary) energy from the transmission system. Until recently, very little research had been conducted on the co-optimization of these two systems due to computational limitations. However, advances in computational capabilities, and the judicious use of decomposition methods and innovative approximation methods for high-dimension dynamic programming made this goal a viable objective for this project, leading to a fundamental shift in the ability to integrate and fully utilize demand-side resources. To this end, the modeling framework developed introduces a novel co-optimization framework, to include the operations of both the transmission and distribution systems (or microgrids) in operational decision making. This framework was used to analyze renewable and distributed generation along with responsive demand and to compare the capability of co-optimized systems to perform with higher levels of variable renewables. An ideal microgrid is defined as an electric entity capable of operating in both interconnected (with the high-voltage grid) and islanded mode. As such, the microgrid should incorporate generating units (traditional units and intermittent) and if needed, exchange power with the high-voltage grid. The interplay between the microgrid and high-voltage grid motivated the development of the co-optimization approach to ensure efficient performance of the interconnected network. Results show that the use of a bi-level optimization approach is an appropriate structure, capable of co-optimizing a transmission system with multiple distribution systems and microgrids. While increasing the number of connected systems provides increasing flexibility for renewables integration this can also the economic benefits to the low-voltage subsystems with each additional system connected. Comparison of a traditional single-level decision structure with the co-optimization approach illustrates a reduction in overall system cost under co-optimization, while specific cost allocations to transmission and distribution systems are changed.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Physics-Informed Graph Neural Networks for Collaborative Dynamic Reconfiguration and Voltage Regulation in Unbalanced Distribution Systems

Network reconfiguration has long been employed as a strategic approach to minimize power distribution system losses and effectively regulate voltage levels. Tap-changing voltage regulators are also critical for controlling bus voltages, especially in accommodating the increasing integration of distributed energy resources (DERs) with intermittent outputs. This paper introduces novel methodologies to address the challenges of dynamic reconfiguration and optimal tap setting in unbalanced three-phase distribution systems. We propose an approximated mixed-integer quadratically constrained program (MIQCP) to model dynamic reconfiguration, along with a pioneering formulation for voltage regulator (VR) tap-setting based on Special Ordered Set type 1 (SOS1). To mitigate computational complexity, we propose a physics-informed spatial-temporal graph convolutional network (STGCN) with an integrated link classifier. The proposed approach enables efficient solution generation by fixing specific variables in the MIQCP instance and solving the simplified sub-MIP using an MIP solver. Numerical studies demonstrate the superior prediction accuracy of our STGCN model compared to baseline neural network models, resulting in reduced DER curtailment and voltage deviation with shorter computation time.

dynamic reconfiguration↗