Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model-free control”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

103 records · Page 6

Role of reinforcement learning for risk-based robust control of cyber-physical energy systems

Critical infrastructures such as cyber-physical energy systems (CPS-E) integrate information flow and physical operations that are vulnerable to natural and targeted failures. Safe, secure, and reliable operation and control of CPS-E is critical to ensure societal well-being and economic prosperity. Automated control is key for real-time operations and may be mathematically cast as a sequential decision-making problem under uncertainty. Emergence of data-driven techniques for decision making under uncertainty, such as reinforcement learning (RL), have led to promising advances for addressing sequential decision-making problems for risk-based robust CPS-E control. However, existing research challenges include understanding the applicability of RL methods across diverse CPS-E applications, addressing the effect of risk preferences across multiple RL methods, and development of open-source domain-aware simulation environments for RL experimentation within a CPS-E context. This article systematically analyzes the applicability of four types of RL methods (model-free, model-based, hybrid model-free and model-based, and hierarchical) for risk-based robust CPS-E control. Problem features and solution stability for the RL methods are also discussed. We demonstrate and compare the performance of multiple RL methods under different risk specifications (risk-averse, risk-neutral, and risk-seeking) through the development and application of an open-source simulation environment. Motivating numerical simulation examples include representative single-zone and multizone building control use cases. Finally, six key insights for future research and broader adoption of RL methods are identified, with specific emphasis on problem features, algorithmic explainability, and solution stability.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Distributed Reinforcement Learning Yaw Control Approach for Wind Farm Energy Capture Maximization: Preprint

In this paper, we present a reinforcement-learning based distributed approach to wind farm energy capture maximization using yaw control, also known as wake steering. In order to maximize the power output of a wind farm, it is often necessary for individual turbines to decrease their own power output through yaw misalignment so as to deflect their wakes away from downstream turbines. Although using model-based methods to achieve yaw misalignment is one option, a model-free method might be better suited to incorporate factors that are difficult to model or changing conditions. We propose an algorithm that adapts concepts of temporal difference reinforcement learning distributed to a multi-agent environment that allows individual turbines to act so as to optimize overall wind farm output and react to unforeseen disturbances.

controls↗

Data-based and secure switched cyber–physical systems

In this work, we develop a completely model-free moving target defense framework for the detection and mitigation of sensor and/or actuator attacks in cyber–physical systems with dynamics that evolve in discrete-time. We incorporate an intrusion detection mechanism based on an approximate dynamic programming technique that learns the policies for optimal regulation and optimal tracking while simultaneously defending against actuator and sensor attacks in a model-free fashion. Switching rules are leveraged to force proactive and reactive defense mechanisms as well as, guarantee the stability of the equilibrium point. Finally, as a case study, we apply the proposed moving target defense framework to a DC–DC converter that is used in electric vehicles.

42 ENGINEERING↗

Model-free distributed learning

Model-free learning for synchronous and asynchronous quasi-static networks is presented. The network weights are continuously perturbed, while the time-varying performance index is measured and correlated with the perturbation signals; the correlation output determines the changes in the weights. The perturbation may be either via noise sources or orthogonal signals. The invariance to detailed network structure mitigates large variability between supposedly identical networks as well as implementation defects. This local, regular, and completely distributed mechanism requires no central control and involves only a few global signals. Thus it allows for integrated on-chip learning in large analog and optical networks.

Dembo, Amir↗

Model-Free Primal-Dual Methods for Network Optimization with Application to Real-Time Optimal Power Flow: Preprint

This paper examines the problem of real-time optimization of networked systems and develops online algorithms that steer the system towards the optimal trajectory without explicit knowledge of the system model. The problem is modeled as a dynamic optimization problem with time-varying performance objectives and engineering constraints. The design of the algorithms leverages the online zero-order primal-dual projected-gradient method. In particular, the primal step that involves the gradient of the objective function (and hence requires networked systems model) is replaced by its zero-order approximation with two function evaluations using a deterministic perturbation signal. The evaluations are performed using the measurements of the system output, hence giving rise to a feedback interconnection, with the optimization algorithm serving as a feedback controller. The paper provides some insights on the stability and tracking properties of this interconnection. Finally, the paper applies this methodology to a real-time optimal power flow problem in power systems, and shows its efficacy on the IEEE 37-node distribution test feeder for reference power tracking and voltage regulation.

61 RADIATION PROTECTION AND DOSIMETRY↗

Automated Escape Guidance Algorithms for An Escape Vehicle

An escape vehicle was designed to provide an emergency evacuation for crew members living on a space station. For maximum escape capability, the escape vehicle needs to have the ability to safely evacuate a station in a contingency scenario such as an uncontrolled (e.g., tumbling) station. This emergency escape sequence will typically be divided into three events: The fust separation event (SEP1), the navigation reconstruction event, and the second separation event (SEP2). SEP1 is responsible for taking the spacecraft from its docking port to a distance greater than the maximum radius of the rotating station. The navigation reconstruction event takes place prior to the SEP2 event and establishes the orbital state to within the tolerance limits necessary for SEP2. The SEP2 event calculates and performs an avoidance burn to prevent station recontact during the next several orbits. This paper presents the tools and results for the whole separation sequence with an emphasis on the two separation events. The fust challenge includes collision avoidance during the escape sequence while the station is in an uncontrolled rotational state, with rotation rates of up to 2 degrees per second. The task of avoiding a collision may require the use of the Vehicle's de-orbit propulsion system for maximum thrust and minimum dwell time within the vicinity of the station vicinity. The thrust of the propulsion system is in a single direction, and can be controlled only by the attitude of the spacecraft. Escape algorithms based on a look-up table or analytical guidance can be implemented since the rotation rate and the angular momentum vector can be sensed onboard and a-priori knowledge of the position and relative orientation are available. In addition, crew intervention has been provided for in the event of unforeseen obstacles in the escape path. The purpose of the SEP2 burn is to avoid re-contact with the station over an extended period of time. Performing this maneuver properly requires knowledge of the orbital state, which is obtained during the navigation state reconstruction event. Since the direction of the delta-v of the SEPI maneuver is a random variable with respect to the Local Vertical Local Horizontal (LVLH) coordinate system, calculating the required SEP2 burn is a challenge. This problem was solved using a neural network as a model-free function approximation technique.

Flanary, Ronald↗

Implications of stop-and-go traffic on training learning-based car-following control

Learning-based car-following control (LCC) of connected and autonomous vehicles (CAVs) is gaining significant attention with the advancement of computing power and data accessibility. While the flexibility and large model capacity of model-free architecture enable LCC to potentially outperform the model-based car-following (CF) model in improving traffic efficiency and mitigating congestion, the generalizability of LCC for traffic conditions different from the training environment/dataset is not well-understood. Herein, this study seeks to explore the impact of stop-and-go traffic in the training dataset on the generalizability of LCC. It uses the characteristics of lead vehicle trajectories to describe stop-and-go traffic, and links the theory of identifiability (i.e., obtaining a unique parameter estimation result using sensor measurements) to the generalizability of behavior cloning (BC) and policy-based deep reinforcement learning (DRL). Correspondingly, the study shows theoretically that: (i) stop-and-go traffic can enable the property of identifiability and enhance the control performance of BC-based LCC in different traffic conditions; (ii) stop-and-go traffic is not necessary for DRL-based LCC to generalize to different traffic conditions; (iii) DRL-based LCC trained with only constant-speed lead vehicle trajectories (not sufficient to ensure identifiability) can be generalized to different traffic conditions; and (iv) stop-and-go traffic increases variance in the training dataset, which improves the convergence of parameter estimation while negatively impacting the convergence of DRL to the optimal control policy. Numerical experiments validate the above findings, illustrating that BC-based LCC entails comprehensive training datasets for generalizing to different traffic conditions, while DRL-based LCC can achieve generalization with simple free-flow traffic training environments. This further suggests DRL as a more promising and cost-effective LCC approach to reduce operational costs, mitigate traffic congestion, and enhance safety and mobility, which can accelerate the deployment and acceptance of CAVs.

33 ADVANCED PROPULSION SYSTEMS↗

Towards robust autonomous impedance spectroscopy analysis: A calibrated hierarchical Bayesian approach for electrochemical impedance spectroscopy (EIS) inversion

Distribution-based analyses, such as the distribution of relaxation times (DRT) and the distribution of diffusion times (DDT), present model-free alternatives to equivalent circuit modeling for analysis of electrochemical impedance spectroscopy (EIS) data. However, reconstructing such distributions from noisy impedance data is an ill-posed problem that must be solved with specialized inversion algorithms, requiring careful control and tuning. Furthermore, most inversion algorithms developed to date can only solve problems of limited complexity. Herein, we present a new hierarchical Bayesian method for EIS inversion, leveraging efficient algorithms for optimization and Hamiltonian Monte Carlo (HMC) sampling to solve models of arbitrary complexity. We overcome the challenge of ad-hoc parameter tuning by encoding intrinsic characteristics of the DRT and DDT into flexible prior distributions and “pre-calibrating” the model to simulated data. This approach is versatile, highly robust to noise, and provides quantitative estimates of both the error structure of the data and the uncertainty in the recovered distributions. The model is validated with simulated data to demonstrate accurate recovery of the DRT and the DDT. The method also shows promise for simultaneous recovery of multiple distributions, raising the intriguing possibility of semi-autonomous EIS analysis and ad-hoc model construction. Finally, the practical utility of the method is illustrated with experimental data. Throughout, we draw comparisons to several recently published EIS inversion methodologies.

36 MATERIALS SCIENCE↗

Model-based and Model-free Designs for an Extended Continuous-time LQR with Exogenous Inputs

We present an extended linear quadratic regulator (LQR) design for continuous-time linear time-invariant (LTI) systems in the presence of exogenous inputs. We first propose a model-based solution with cost minimization guarantees for states and inputs using dynamic programming (DP). The control law consists of a combination of the optimal state feedback and an additional optimal term dependent on the exogenous inputs. The control gains for the two components are obtained by solving a set of matrix differential equations. We provide these solutions for both finite horizons and steady-state cases. In the second part of the paper, we formulate a reinforcement learning (RL) based algorithm which does not need any model information except the input matrix, and can compute an approximate steady-state LQR gain using measurements of the states, the control inputs, and the exogenous inputs. Both model-based and data-driven optimal control algorithms are tested with a numerical example under different exogenous inputs showcasing the effectiveness of the designs.

Mukherjee, Sayak↗

A Modular and Transferable Reinforcement Learning Framework for the Fleet Rebalancing Problem

Mobility on demand (MoD) systems show great promise in realizing flexible and efficient urban transportation. However, significant technical challenges arise from operational decision making associated with MoD vehicle dispatch and fleet rebalancing. For this reason, operators tend to employ simplified algorithms that have been demonstrated to work well in a particular setting. To help bridge the gap between novel and existing methods, we propose a modular framework for fleet rebalancing based on model-free reinforcement learning (RL) that can leverage an existing dispatch method to minimize system cost. In particular, by treating dispatch as part of the environment dynamics, a centralized agent can learn to intermittently direct the dispatcher to reposition free vehicles and mitigate against fleet imbalance. We formulate RL state and action spaces as distributions over a grid partitioning of the operating area, making the framework scalable and avoiding the complexities associated with multiagent RL. Numerical experiments, using real-world trip and network data, demonstrate that RL reduces waiting time by 28% to 38% for the same-day evaluation, 17% to 44% for cross-day evaluation, and 22% to 25% for cross-season evaluation compared with no rebalancing scenarios. This approach has several distinct advantages over baseline methods including: improved system cost; high degree of adaptability to the selected dispatch method; and the ability to perform scale-invariant transfer learning between problem instances with similar vehicle and request distributions.

33 ADVANCED PROPULSION SYSTEMS↗

Physics-Informed Evolutionary Strategy Based Control for Mitigating Delayed Voltage Recovery

Here, in this work we propose a novel data-driven, real-time power system voltage stability control method based on the physics-informed guided meta evolutionary strategy (ES). The main objective is to quickly provide an adaptive control strategy to secure system voltage stability. The problem is challenging due to the high-dimensional feature of the power system model and the fast-changing and uncertain nature of power system operation scenarios. To this end, a model-free and derivative-free guided ES method is applied. The method is further combined with a meta-learning strategy to make the learnt control policy automatically adapted to unseen operation conditions and fault scenarios, which is highly desired for real-time emergency control. Last but not least, physical knowledge is embedded in the above method through a trainable action mask technique to rule out unnecessary load shedding actions for better learning and control performance. Case studies on the IEEE 300-bus system and comparisons with other state-of-the-art benchmark methods verify the superiority of the proposed physics-informed guided meta ES method in realizing fast and adaptive power system voltage stability control.

42 ENGINEERING↗

Achieving Cyber-Resilience for Power Systems using a Learning, Model-Assisted Blockchain Framework

The secure integration and management of distributed energy resources (DER) and power aggregators in the electric grid requires secure communications and a physics-aware Command and Control (C2) strategy. A Blockchain (BC)-based overlay network was developed to provide a security layer for the existing power grid network that mitigates risks in current and legacy network and C2 protocols. By integrating a Model-Assisted Machine Learning (MAML) framework with a Secure Blockchain Overlay Network (SBON) a defense-in-depth strategy was achieved. In our approach, the MAML framework leveraged a smart contract framework to gather network data and learn the dynamics of DER to develop detection strategies for attacks targeting sensors and actuators used by DER. The MAML framework learned dynamical systems models for individual DERs to detect sensor attacks. For DER we utilized a Digital Twin (DT) to accelerate the learning process for a model resistant to stealthy attacks. The project created DT for PV inverters and BESS. The DTs were coupled with a model-assisted, data-driven learning of DER behavior. Specifically, we evaluated architectures for model-based learning with model-free fine-tuning. Additionally, differential privacy techniques were used to obfuscate data, while still allowing the computation of attack detection results based on obfuscated data. The SBON developed leverages a private permissioned blockchain network orchestrated with the Hyperledger Fabric framework. To connect the cyber world, which orchestrates the blockchain fabric, and the physical world where the power network resides, we developed a system implementation to enable the secure interaction of the physical world and the abstracted blockchain.

97 MATHEMATICS AND COMPUTING↗

Model-Free Primal-Dual Methods for Network Optimization with Application to Real-Time Optimal Power Flow

This paper examines the problem of real-time optimization of networked systems and develops online algorithms that steer the system towards the optimal trajectory without explicit knowledge of the system model. The problem is modeled as a dynamic optimization problem with time-varying performance objectives and engineering constraints. The design of the algorithms leverages the online zero-order primal-dual projected-gradient method. In particular, the primal step that involves the gradient of the objective function (and hence requires a networked systems model) is replaced by its zero-order approximation with two function evaluations using a deterministic perturbation signal. The evaluations are performed using the measurements of the system output, hence giving rise to a feedback interconnection, with the optimization algorithm serving as a feedback controller. The paper provides some insights on the stability and tracking properties of this interconnection. Finally, the paper applies this methodology to a real-time optimal power flow problem in power systems, and shows its efficacy on the IEEE 37-node distribution test feeder for reference power tracking and voltage regulation.

61 RADIATION PROTECTION AND DOSIMETRY↗