Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model-free control”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Reinforcement-Learning-Based Smart Water Heater Control: An Actual Deployment

Utilizing smart control algorithms for electric water heaters (EWHs) is essential for fully harnessing the demand response (DR) potential of EWHs. For this reason, the use of reinforcement learning (RL) algorithms for EWHs has received increasing attention in recent years. However, existing RL approaches are either simulation-based or use pretrained RL agents. To this end, this paper presents the real-world deployment of a set of model-free RL approaches that aim to minimize the electricity cost of a EWH under a time-of-use electricity pricing policy using standard DR commands (e.g., shed, load up). The experiment results showed that the RL agents can help save electricity cost in the range of 11% to 14% compared to the baseline operation. This study demonstrated that RL-based EWH controllers can be deployed in real world without any prior training and can still save electricity cost.

deep learning↗

Topology-Aware Reinforcement Learning for Voltage Control: Centralized and Decentralized Strategies

Volt-VAR control (VVC) methods based on deep reinforcement learning (DRL) can effectively control distribution grid voltage and minimize power loss by implementing corrective and preventive control measures on the reactive power output of inverter-based distributed energy resources (DERs). However, model-free DRL-based VVC approaches usually cannot capture the important topological feature of the power system since they use a fully-connected network (FCN) to deliver the action. Therefore, this paper proposes a graph convolutional network (GCN)-based DRL approach that can employ the topological information of the network to take better control action for regulating the voltage. Our implementation allows for both centralized and decentralized configurations, utilizing a single agent and multiple agents respectively. Although the centralized GCN-based DRL approach has its advantages of minimizing voltage fluctuation and power loss, it is not suitable for large scale power systems due to its challenges in terms of scalability, computation speed and potential single points of failure. Therefore, these problems can be resolved using the decentralized GCN-based DRL approach. Moreover, to ensure the safe operation of the model, our proposed approach incorporates an exponential barrier function while formulating the reward function for each agent. To validate performance of the proposed approaches, the proposed model is tested on modified IEEE test systems and the performances are measured in terms on voltage fluctuation reduction, minimization of power loss and computational speed. Finally, the results show that the proposed topology-aware approach outperforms the FCN-based DRL approach in terms of reducing voltage fluctuation and minimizing power loss of the network. Moreover, it is shown that the decentralized GCN-based DRL has faster computational speed than other approaches.

42 ENGINEERING↗

Convex Q-Learning in Continuous Time with Application to Dispatch of Distributed Energy Resources

Convex Q-learning is a recent approach to reinforcement learning, motivated by the possibility of a firmer theory for convergence, and the possibility of making use of greater a priori knowledge regarding policy or value function structure. This paper explores algorithm design in the continuous time domain, with a finite-horizon optimal control objective. The main contributions are (i) The new Q-ODE: a model-free characterization of the Hamilton-Jacobi-Bellman equation. (ii) A formulation of Convex Q-learning that avoids approximations appearing in prior work. The Bellman error used in the algorithm is defined by filtered measurements, which is necessary in the presence of measurement noise. (iii) Convex Q-learning with linear function approximation is a convex program. It is shown that the constraint region is bounded, subject to an exploration condition on the training input. (iv) The theory is illustrated in application to resource allocation for distributed energy resources, for which the theory is ideally suited.

Lu, Fan↗

Cooperative Merging via Online Speed Replanning: A Model-Free Approach With Vehicle-to-Vehicle Communication Packet Drop Compensation

On-ramp merging is a critical bottleneck in freeway traffic flow, contributing to congestion, accidents, and excessive fuel consumption. Although traditional ramp metering provides macroscopic control, it lacks the granularity for optimizing an individual vehicle’s trajectory. Cooperative merging, enabled by connected and automated vehicles, can potentially enhance traffic efficiency, safety, and fuel economy. However, existing research often neglects the influence of heterogeneous vehicle dynamics, unreliable vehicle-to-vehicle (V2V) communication, and real-time implementation challenges. Here, this paper introduces novel model-free online speed planners for cooperative on-ramp merging. The planners address these limitations by being agnostic to vehicle dynamics, effectively compensating for V2V communication packet drops and incurring only a light computational burden. Comprehensive evaluation, conducted on a real-time traffic-vehicle-communication co-simulation platform integrating high-fidelity vehicle dynamics, a traffic simulator, and recorded V2V communication footprints, demonstrates the effectiveness of the proposed speed planners. Simulation results reveal that the proposed method yields accurate tracking of desired speed and inter-vehicle distance, maintaining low fuel consumption even under high packet drop ratios, and demonstrating real-time implementation efficiency.

Wang, Zejiang [Univ. of Texas at Dallas, Richardso↗

Learning Constrained Parametric Differentiable Predictive Control Policies With Guarantees

We present differentiable predictive control (DPC), a method for offline learning of constrained neural control policies for nonlinear dynamical systems with performance guarantees. We show that the sensitivities of the parametric optimal control problem can be used to obtain direct policy gradients. Specifically, we employ automatic differentiation (AD) to efficiently compute the sensitivities of the model predictive control (MPC) objective function and constraints penalties. To guarantee safety upon deployment, we derive probabilistic guarantees on closed-loop stability and constraint satisfaction based on indicator functions and Hoeffding’s inequality. We empirically demonstrate that the proposed method can learn neural control policies for various parametric optimal control tasks. In particular, we show that the proposed DPC method can stabilize systems with unstable dynamics, track time-varying references, and satisfy nonlinear state and input constraints. Our DPC method has practical time savings compared to alternative approaches for fast and memory-efficient controller design. Specifically, DPC does not depend on a supervisory controller as opposed to approximate MPC based on imitation learning. We demonstrate that, without losing performance, DPC is scalable with greatly reduced demands on memory and computation compared to implicit and explicit MPC while being more sample efficient than model-free reinforcement learning (RL) algorithms.

97 MATHEMATICS AND COMPUTING↗

Model-free stabilization via Extremum Seeking using a cost neural estimator

In this paper, a fully model-free architecture for vertical stabilization of thermonuclear plasmas in tokamak experimental reactors is presented. For the first time, an Extremum Seeking control algorithm is combined with neural networks to estimate the Lyapunov function to be minimized, resulting in a fully data-driven control architecture. The performance of different neural networks are compared. Specifically, Multilayer Perceptrons and Extreme Learning Machines are considered. The proposed architecture is tested in simulation to show that it can counteract relevant plasma disturbances, resulting in a significant improvement in terms of the achievable operative space compared to the Extremum Seeking algorithm, which still relies on model-based cost estimator.

42 ENGINEERING↗

An innovative heterogeneous transfer learning framework to enhance the scalability of deep reinforcement learning controllers in buildings with integrated energy systems

Deep Reinforcement Learning (DRL)-based control shows enhanced performance in the management of integrated energy systems when compared with Rule-Based Controllers (RBCs), but it still lacks scalability and generalisation due to the necessity of using tailored models for the training process. Transfer Learning (TL) is a potential solution to address this limitation. However, existing TL applications in building control have been mostly tested among buildings with similar features, not addressing the need to scale up advanced control in real-world scenarios with diverse energy systems. This paper assesses the performance of an online heterogeneous TL strategy, comparing it with RBC and offline and online DRL controllers in a simulation setup using EnergyPlus and Python. The study tests the transfer in both transductive and inductive settings of a DRL policy designed to manage a chiller coupled with a Thermal Energy Storage (TES). The control policy is pre-trained on a source building and transferred to various target buildings characterised by an integrated energy system including photovoltaic and battery energy storage systems, different building envelope features, occupancy schedule and boundary conditions (e.g., weather and price signal). The TL approach incorporates model slicing, imitation learning and fine-tuning to handle diverse state spaces and reward functions between source and target buildings. Results show that the proposed methodology leads to a reduction of 10% in electricity cost and between 10% and 40% in the mean value of the daily average temperature violation rate compared to RBC and online DRL controllers. Moreover, online TL maximises self-sufficiency and self-consumption by 9% and 11% with respect to RBC. Conversely, online TL achieves worse performance compared to offline DRL in either transductive or inductive settings. However, offline Deep Reinforcement Learning (DRL) agents should be trained at least for 15 episodes to reach the same level of performance as the online TL. Therefore, the proposed online TL methodology is effective, completely model-free and it can be directly implemented in real buildings with satisfying performance.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A Generation-Storage Coordination Dispatch Strategy for Power System Based on Causal Reinforcement Learning

In the backdrop of global energy transformation, power systems integrating high proportions of renewable energy sources are facing unprecedented challenges in operational stability and dispatch efficiency. To address these challenges, this study introduces a generation-storage coordination real-time dispatch strategy based on Causal Power System Dynamic Reinforcement Learning (CPSDRL). Diverging from traditional reinforcement learning approaches, CPSDRL innovatively incorporates causal inference within the state prediction model - the crux of model-based reinforcement learning - thereby establishing the Power Causal Dynamic Model (PCDM). Assisted by the prior knowledge of power systems, the model significantly enhances prediction accuracy and reliability through a two-stage training process. Utilizing PCDM, this study further applies a direct policy search algorithm to optimize the real-time dispatch strategy. Experimental results indicate that the proposed method improves the stability of generation-storage coordination real-time dispatch and exhibits competitive advantages in sample efficiency and computational speed, compared to traditional model-based and model-free reinforcement learning algorithms. This method is expected to enhance the practicality and adaptability of causal reinforcement learning techniques in power system scheduling and control.

causal reinforcement learning↗

Multi-agent voltage control in distribution systems using GAN-DRL-based approach

Active distribution grids can experience voltage fluctuations and violations due to the high penetration of variable distributed energy resources (DERs). These problems might occur because of the uncertain and variable generation natures of these resources, especially solar photovoltaic resources, during panel shadowing scenarios. Volt-VAR control (VVC) is an efficient method that controls the reactive power set-points of the inverters to regulate the voltage of distribution grids. Although several VVC approaches have been proposed recently, the performance of these approaches degrades significantly if behind-the-meter solar generation data are unobservable/missing. Therefore, it is necessary to impute missing/unobservable PV data accurately to be utilized in VVC approaches. Further, this paper proposes a model-free, data-driven, centrally trained, and decentrally executed multi-agent deep reinforcement learning-based VVC architecture to regulate the voltage of distribution networks. A generative adversarial network (GAN) is incorporated to impute the unobservable PV data accurately, which improves the performance of the proposed control architecture. The proposed multi-agent-soft-actor–critic algorithm (MASAC)-based VVC technique utilizes the actual PV dataset as well as the imputed dataset from the GAN framework to learn the optimal coordinated control policy for controlling the optimal reactive power set-points of PV inverters. The effectiveness of the proposed approach is analyzed on a modified IEEE 34-bus test case with added PV inverters. The results are compared and analyzed with a base case model with no VVC and VVC with a local droop control approach, genetic algorithm optimization, and a centralized soft actor–critic-based approach. Moreover, the performance of the proposed approach is compared with that of a multi-agent VVC framework without using the PV generation data and load information as the system state. The results illustrate that the proposed method with more state input improves the voltage profile and reduces the power loss of the network across various loading and PV generation scenarios.

14 SOLAR ENERGY↗

Online Model-Free DER Dispatch Via Adaptive Voltage Sensitivity Estimation and Chance Constrained Programming

This paper proposes an online data-driven distributed energy resource management system (DERMS) for distribution system optimal DER dispatch as well as voltage regulation. Here, the key innovation is to leverage the Local Sensitivity Factor (LSF) for transforming the DER control into a computationally efficient linear programming (LP) problem. By taking real-time measurements, the estimation of LSF eliminates the need for an accurate distribution system model as well as full nodal load information, which is difficult to achieve in practice. A robust recursive least squares method is also developed to ensure the robust estimation of LSF, which is initialized using reasonable values from model-derived LSFs. This allows the system to adapt to changing operational conditions effectively. A scenario-based, chance-constrained framework is further employed to ensure voltage remains within acceptable limits in the presence of measurement and estimation uncertainties. Test results on a real-world, 759-node distribution network located in western Colorado, U.S., validate the effectiveness and robustness of the proposed control approach and demonstrate its superior performance as compared to alternative methods.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Harnessing the power of gradient-based simulations for multi-objective optimization in particle accelerators

Abstract Particle accelerator operation requires simultaneous optimization of multiple objectives. Multi-objective optimization (MOO) is particularly challenging due to trade-offs between the objectives. Evolutionary algorithms, such as genetic algorithms (GAs), have been leveraged for many optimization problems, however, they do not apply to complex control problems by design. This paper demonstrates the power of differentiability for solving MOO problems in particle accelerators using a deep differentiable reinforcement learning (DDRL) algorithm. We compare the DDRL algorithm with model-free reinforcement learning (MFRL), GA, and Bayesian optimization (BO) for simultaneous optimization of heat load and trip rates in the continuous electron beam accelerator facility. The underlying problem enforces strict constraints on both individual states and actions as well as cumulative (global) constraints on energy requirements of the beam. Using historical accelerator data, we develop a physics-based surrogate model which is differentiable and allows for back-propagation of gradients. The results are evaluated in the form of a Pareto-front with two objectives. We show that the DDRL outperforms MFRL, BO, and GA on high dimensional problems.

43 PARTICLE ACCELERATORS↗

SDSS-IV MaNGA: How the Stellar Populations of Passive Central Galaxies Depend on Stellar and Halo Mass

We analyze spatially resolved and co-added SDSS-IV MaNGA spectra with signal-to-noise ratio ~100 from 2200 passive central galaxies (z ~ 0.05) to understand how central galaxy assembly depends on stellar mass (M*) and halo mass (M h ). We control for systematic errors in M h by employing a new group catalog from Tinker and the widely used Yang et al. catalog. At fixed M*, the strengths of several stellar absorption features vary systematically with M h . Completely model-free, this is one of the first indications that the stellar populations of centrals with identical M* are affected by the properties of their host halos. To interpret these variations, we applied full spectral fitting with the code alf. At fixed M*, centrals in more massive halos are older, show lower [Fe/H], and have higher [Mg/Fe] with 3.5σ confidence. We conclude that halos not only dictate how much M* galaxies assemble but also modulate their chemical enrichment histories. Turning to our analysis at fixed M h , high-M* centrals are older, show lower [Fe/H], and have higher [Mg/Fe] for M h > 10 12 h –1 M⊙ with confidence >4σ. While massive passive galaxies are thought to form early and rapidly, our results are among the first to distinguish these trends at fixed M h . They suggest that high-M* centrals experienced unique early formation histories, either through enhanced collapse and gas fueling or because their halos were early forming and highly concentrated, a possible signal of galaxy assembly bias.

79 ASTRONOMY AND ASTROPHYSICS↗

Deep Reinforcement Learning Based Smart Water Heater Control for Reducing Electricity Consumption and Carbon Emission

Water heating is the third largest electricity consumer in U.S. households, after space heating and cooling. Thus, water heaters represent a significant potential for reducing electricity consumption and associated CO2 emissions of residential buildings. To this end, this study proposes a model-free deep reinforcement learning (RL) approach that aims to minimize the electricity consumption and the CO2 emissions of a heat pump water heater without affecting user comfort. In this approach, a set of RL agents focusing on either electricity saving or emission reduction, with different look ahead periods, were trained using the deep Q-networks (DQN) algorithm and their performance was tested on different hot water usage and Marginal Operating Emissions Rate (MOER) profiles. The testing results showed that the RL agents that focus on electricity saving can save electricity in the range of 12–22% by operating the water heater with maximum heat pump efficiency and minimum electric element utilization. On the other hand, the RL agents that focus on emission reduction reduced emissions in the range of 18–37% by making use of the variable MOER values. These RL agents used the heat pump and/or an element when the MOER values are low due to the availability of renewable energy sources (e.g., solar and wind) and mostly avoided the periods of carbon-intensive periods. Overall, these results showed that the proposed RL approach can help minimize the electricity consumption and the CO2 emissions of a heat pump water heater without having any prior knowledge about the device.

Amasyali, Kadir↗

Adaptive Control of Distributed Energy Resources for Distribution Grid Voltage Stability

Volt-VAR and Volt-Watt functionality in photovoltaic (PV) smart inverters provide mechanisms to ensure system voltage magnitudes and power factors remain within acceptable limits. However, these control functions can become unstable, introducing oscillations in system voltages when not appropriately configured or maliciously altered during a cyberattack. In the event that Volt-VAR and Volt-Watt control functions in a portion of PV smart inverters in a distribution grid are unstable, the proposed adaptation scheme utilizes the remaining and stably-behaving PV smart inverters and other Distributed Energy Resources to mitigate the effect of the instability. The adaptation mechanism is entirely decentralized, model-free, communication-free, and requires virtually no external configuration. Here we provide a derivation of the adaptive control approach and validate the algorithm in experiments on the IEEE 37 and 8500 node test feeders.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Role of reinforcement learning for risk-based robust control of cyber-physical energy systems

Critical infrastructures such as cyber-physical energy systems (CPS-E) integrate information flow and physical operations that are vulnerable to natural and targeted failures. Safe, secure, and reliable operation and control of CPS-E is critical to ensure societal well-being and economic prosperity. Automated control is key for real-time operations and may be mathematically cast as a sequential decision-making problem under uncertainty. Emergence of data-driven techniques for decision making under uncertainty, such as reinforcement learning (RL), have led to promising advances for addressing sequential decision-making problems for risk-based robust CPS-E control. However, existing research challenges include understanding the applicability of RL methods across diverse CPS-E applications, addressing the effect of risk preferences across multiple RL methods, and development of open-source domain-aware simulation environments for RL experimentation within a CPS-E context. This article systematically analyzes the applicability of four types of RL methods (model-free, model-based, hybrid model-free and model-based, and hierarchical) for risk-based robust CPS-E control. Problem features and solution stability for the RL methods are also discussed. We demonstrate and compare the performance of multiple RL methods under different risk specifications (risk-averse, risk-neutral, and risk-seeking) through the development and application of an open-source simulation environment. Motivating numerical simulation examples include representative single-zone and multizone building control use cases. Finally, six key insights for future research and broader adoption of RL methods are identified, with specific emphasis on problem features, algorithmic explainability, and solution stability.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Model-free distributed learning

Model-free learning for synchronous and asynchronous quasi-static networks is presented. The network weights are continuously perturbed, while the time-varying performance index is measured and correlated with the perturbation signals; the correlation output determines the changes in the weights. The perturbation may be either via noise sources or orthogonal signals. The invariance to detailed network structure mitigates large variability between supposedly identical networks as well as implementation defects. This local, regular, and completely distributed mechanism requires no central control and involves only a few global signals. Thus it allows for integrated on-chip learning in large analog and optical networks.

Dembo, Amir↗

Automated Escape Guidance Algorithms for An Escape Vehicle

An escape vehicle was designed to provide an emergency evacuation for crew members living on a space station. For maximum escape capability, the escape vehicle needs to have the ability to safely evacuate a station in a contingency scenario such as an uncontrolled (e.g., tumbling) station. This emergency escape sequence will typically be divided into three events: The fust separation event (SEP1), the navigation reconstruction event, and the second separation event (SEP2). SEP1 is responsible for taking the spacecraft from its docking port to a distance greater than the maximum radius of the rotating station. The navigation reconstruction event takes place prior to the SEP2 event and establishes the orbital state to within the tolerance limits necessary for SEP2. The SEP2 event calculates and performs an avoidance burn to prevent station recontact during the next several orbits. This paper presents the tools and results for the whole separation sequence with an emphasis on the two separation events. The fust challenge includes collision avoidance during the escape sequence while the station is in an uncontrolled rotational state, with rotation rates of up to 2 degrees per second. The task of avoiding a collision may require the use of the Vehicle's de-orbit propulsion system for maximum thrust and minimum dwell time within the vicinity of the station vicinity. The thrust of the propulsion system is in a single direction, and can be controlled only by the attitude of the spacecraft. Escape algorithms based on a look-up table or analytical guidance can be implemented since the rotation rate and the angular momentum vector can be sensed onboard and a-priori knowledge of the position and relative orientation are available. In addition, crew intervention has been provided for in the event of unforeseen obstacles in the escape path. The purpose of the SEP2 burn is to avoid re-contact with the station over an extended period of time. Performing this maneuver properly requires knowledge of the orbital state, which is obtained during the navigation state reconstruction event. Since the direction of the delta-v of the SEPI maneuver is a random variable with respect to the Local Vertical Local Horizontal (LVLH) coordinate system, calculating the required SEP2 burn is a challenge. This problem was solved using a neural network as a model-free function approximation technique.

Flanary, Ronald↗

Implications of stop-and-go traffic on training learning-based car-following control

Learning-based car-following control (LCC) of connected and autonomous vehicles (CAVs) is gaining significant attention with the advancement of computing power and data accessibility. While the flexibility and large model capacity of model-free architecture enable LCC to potentially outperform the model-based car-following (CF) model in improving traffic efficiency and mitigating congestion, the generalizability of LCC for traffic conditions different from the training environment/dataset is not well-understood. Herein, this study seeks to explore the impact of stop-and-go traffic in the training dataset on the generalizability of LCC. It uses the characteristics of lead vehicle trajectories to describe stop-and-go traffic, and links the theory of identifiability (i.e., obtaining a unique parameter estimation result using sensor measurements) to the generalizability of behavior cloning (BC) and policy-based deep reinforcement learning (DRL). Correspondingly, the study shows theoretically that: (i) stop-and-go traffic can enable the property of identifiability and enhance the control performance of BC-based LCC in different traffic conditions; (ii) stop-and-go traffic is not necessary for DRL-based LCC to generalize to different traffic conditions; (iii) DRL-based LCC trained with only constant-speed lead vehicle trajectories (not sufficient to ensure identifiability) can be generalized to different traffic conditions; and (iv) stop-and-go traffic increases variance in the training dataset, which improves the convergence of parameter estimation while negatively impacting the convergence of DRL to the optimal control policy. Numerical experiments validate the above findings, illustrating that BC-based LCC entails comprehensive training datasets for generalizing to different traffic conditions, while DRL-based LCC can achieve generalization with simple free-flow traffic training environments. This further suggests DRL as a more promising and cost-effective LCC approach to reduce operational costs, mitigate traffic congestion, and enhance safety and mobility, which can accelerate the deployment and acceptance of CAVs.

33 ADVANCED PROPULSION SYSTEMS↗