Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model-free control”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

An innovative heterogeneous transfer learning framework to enhance the scalability of deep reinforcement learning controllers in buildings with integrated energy systems

Deep Reinforcement Learning (DRL)-based control shows enhanced performance in the management of integrated energy systems when compared with Rule-Based Controllers (RBCs), but it still lacks scalability and generalisation due to the necessity of using tailored models for the training process. Transfer Learning (TL) is a potential solution to address this limitation. However, existing TL applications in building control have been mostly tested among buildings with similar features, not addressing the need to scale up advanced control in real-world scenarios with diverse energy systems. This paper assesses the performance of an online heterogeneous TL strategy, comparing it with RBC and offline and online DRL controllers in a simulation setup using EnergyPlus and Python. The study tests the transfer in both transductive and inductive settings of a DRL policy designed to manage a chiller coupled with a Thermal Energy Storage (TES). The control policy is pre-trained on a source building and transferred to various target buildings characterised by an integrated energy system including photovoltaic and battery energy storage systems, different building envelope features, occupancy schedule and boundary conditions (e.g., weather and price signal). The TL approach incorporates model slicing, imitation learning and fine-tuning to handle diverse state spaces and reward functions between source and target buildings. Results show that the proposed methodology leads to a reduction of 10% in electricity cost and between 10% and 40% in the mean value of the daily average temperature violation rate compared to RBC and online DRL controllers. Moreover, online TL maximises self-sufficiency and self-consumption by 9% and 11% with respect to RBC. Conversely, online TL achieves worse performance compared to offline DRL in either transductive or inductive settings. However, offline Deep Reinforcement Learning (DRL) agents should be trained at least for 15 episodes to reach the same level of performance as the online TL. Therefore, the proposed online TL methodology is effective, completely model-free and it can be directly implemented in real buildings with satisfying performance.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A Generation-Storage Coordination Dispatch Strategy for Power System Based on Causal Reinforcement Learning

In the backdrop of global energy transformation, power systems integrating high proportions of renewable energy sources are facing unprecedented challenges in operational stability and dispatch efficiency. To address these challenges, this study introduces a generation-storage coordination real-time dispatch strategy based on Causal Power System Dynamic Reinforcement Learning (CPSDRL). Diverging from traditional reinforcement learning approaches, CPSDRL innovatively incorporates causal inference within the state prediction model - the crux of model-based reinforcement learning - thereby establishing the Power Causal Dynamic Model (PCDM). Assisted by the prior knowledge of power systems, the model significantly enhances prediction accuracy and reliability through a two-stage training process. Utilizing PCDM, this study further applies a direct policy search algorithm to optimize the real-time dispatch strategy. Experimental results indicate that the proposed method improves the stability of generation-storage coordination real-time dispatch and exhibits competitive advantages in sample efficiency and computational speed, compared to traditional model-based and model-free reinforcement learning algorithms. This method is expected to enhance the practicality and adaptability of causal reinforcement learning techniques in power system scheduling and control.

causal reinforcement learning↗

Multi-agent voltage control in distribution systems using GAN-DRL-based approach

Active distribution grids can experience voltage fluctuations and violations due to the high penetration of variable distributed energy resources (DERs). These problems might occur because of the uncertain and variable generation natures of these resources, especially solar photovoltaic resources, during panel shadowing scenarios. Volt-VAR control (VVC) is an efficient method that controls the reactive power set-points of the inverters to regulate the voltage of distribution grids. Although several VVC approaches have been proposed recently, the performance of these approaches degrades significantly if behind-the-meter solar generation data are unobservable/missing. Therefore, it is necessary to impute missing/unobservable PV data accurately to be utilized in VVC approaches. Further, this paper proposes a model-free, data-driven, centrally trained, and decentrally executed multi-agent deep reinforcement learning-based VVC architecture to regulate the voltage of distribution networks. A generative adversarial network (GAN) is incorporated to impute the unobservable PV data accurately, which improves the performance of the proposed control architecture. The proposed multi-agent-soft-actor–critic algorithm (MASAC)-based VVC technique utilizes the actual PV dataset as well as the imputed dataset from the GAN framework to learn the optimal coordinated control policy for controlling the optimal reactive power set-points of PV inverters. The effectiveness of the proposed approach is analyzed on a modified IEEE 34-bus test case with added PV inverters. The results are compared and analyzed with a base case model with no VVC and VVC with a local droop control approach, genetic algorithm optimization, and a centralized soft actor–critic-based approach. Moreover, the performance of the proposed approach is compared with that of a multi-agent VVC framework without using the PV generation data and load information as the system state. The results illustrate that the proposed method with more state input improves the voltage profile and reduces the power loss of the network across various loading and PV generation scenarios.

14 SOLAR ENERGY↗

Online Model-Free DER Dispatch Via Adaptive Voltage Sensitivity Estimation and Chance Constrained Programming

This paper proposes an online data-driven distributed energy resource management system (DERMS) for distribution system optimal DER dispatch as well as voltage regulation. Here, the key innovation is to leverage the Local Sensitivity Factor (LSF) for transforming the DER control into a computationally efficient linear programming (LP) problem. By taking real-time measurements, the estimation of LSF eliminates the need for an accurate distribution system model as well as full nodal load information, which is difficult to achieve in practice. A robust recursive least squares method is also developed to ensure the robust estimation of LSF, which is initialized using reasonable values from model-derived LSFs. This allows the system to adapt to changing operational conditions effectively. A scenario-based, chance-constrained framework is further employed to ensure voltage remains within acceptable limits in the presence of measurement and estimation uncertainties. Test results on a real-world, 759-node distribution network located in western Colorado, U.S., validate the effectiveness and robustness of the proposed control approach and demonstrate its superior performance as compared to alternative methods.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Harnessing the power of gradient-based simulations for multi-objective optimization in particle accelerators

Abstract Particle accelerator operation requires simultaneous optimization of multiple objectives. Multi-objective optimization (MOO) is particularly challenging due to trade-offs between the objectives. Evolutionary algorithms, such as genetic algorithms (GAs), have been leveraged for many optimization problems, however, they do not apply to complex control problems by design. This paper demonstrates the power of differentiability for solving MOO problems in particle accelerators using a deep differentiable reinforcement learning (DDRL) algorithm. We compare the DDRL algorithm with model-free reinforcement learning (MFRL), GA, and Bayesian optimization (BO) for simultaneous optimization of heat load and trip rates in the continuous electron beam accelerator facility. The underlying problem enforces strict constraints on both individual states and actions as well as cumulative (global) constraints on energy requirements of the beam. Using historical accelerator data, we develop a physics-based surrogate model which is differentiable and allows for back-propagation of gradients. The results are evaluated in the form of a Pareto-front with two objectives. We show that the DDRL outperforms MFRL, BO, and GA on high dimensional problems.

43 PARTICLE ACCELERATORS↗

SDSS-IV MaNGA: How the Stellar Populations of Passive Central Galaxies Depend on Stellar and Halo Mass

We analyze spatially resolved and co-added SDSS-IV MaNGA spectra with signal-to-noise ratio ~100 from 2200 passive central galaxies (z ~ 0.05) to understand how central galaxy assembly depends on stellar mass (M*) and halo mass (M h ). We control for systematic errors in M h by employing a new group catalog from Tinker and the widely used Yang et al. catalog. At fixed M*, the strengths of several stellar absorption features vary systematically with M h . Completely model-free, this is one of the first indications that the stellar populations of centrals with identical M* are affected by the properties of their host halos. To interpret these variations, we applied full spectral fitting with the code alf. At fixed M*, centrals in more massive halos are older, show lower [Fe/H], and have higher [Mg/Fe] with 3.5σ confidence. We conclude that halos not only dictate how much M* galaxies assemble but also modulate their chemical enrichment histories. Turning to our analysis at fixed M h , high-M* centrals are older, show lower [Fe/H], and have higher [Mg/Fe] for M h > 10 12 h –1 M⊙ with confidence >4σ. While massive passive galaxies are thought to form early and rapidly, our results are among the first to distinguish these trends at fixed M h . They suggest that high-M* centrals experienced unique early formation histories, either through enhanced collapse and gas fueling or because their halos were early forming and highly concentrated, a possible signal of galaxy assembly bias.

79 ASTRONOMY AND ASTROPHYSICS↗

Deep Reinforcement Learning Based Smart Water Heater Control for Reducing Electricity Consumption and Carbon Emission

Water heating is the third largest electricity consumer in U.S. households, after space heating and cooling. Thus, water heaters represent a significant potential for reducing electricity consumption and associated CO2 emissions of residential buildings. To this end, this study proposes a model-free deep reinforcement learning (RL) approach that aims to minimize the electricity consumption and the CO2 emissions of a heat pump water heater without affecting user comfort. In this approach, a set of RL agents focusing on either electricity saving or emission reduction, with different look ahead periods, were trained using the deep Q-networks (DQN) algorithm and their performance was tested on different hot water usage and Marginal Operating Emissions Rate (MOER) profiles. The testing results showed that the RL agents that focus on electricity saving can save electricity in the range of 12–22% by operating the water heater with maximum heat pump efficiency and minimum electric element utilization. On the other hand, the RL agents that focus on emission reduction reduced emissions in the range of 18–37% by making use of the variable MOER values. These RL agents used the heat pump and/or an element when the MOER values are low due to the availability of renewable energy sources (e.g., solar and wind) and mostly avoided the periods of carbon-intensive periods. Overall, these results showed that the proposed RL approach can help minimize the electricity consumption and the CO2 emissions of a heat pump water heater without having any prior knowledge about the device.

Amasyali, Kadir↗

Multiarea Distribution System State Estimation via Distributed Tensor Completion

Here, this paper proposes a model-free distribution system state estimation method based on tensor completion using canonical polyadic decomposition. In particular, we consider a setting where the network is divided into multiple areas. The measured physical quantities at buses located in the same area are processed by an area controller. A three-way tensor is constructed to collect these measured quantities. The measurements are analyzed locally to recover the full state information of the network. A distributed closed-form iterative algorithm based on the alternating direction method of multipliers is developed to obtain the low-rank factors of the whole network state tensor where information exchange happens only between neighboring areas. The convergence properties of the distributed algorithm and the sufficient conditions on the number of samples for each smaller network that guarantee the identifiability of the factors of the state tensor are presented. To demonstrate the efficacy of the proposed algorithm and to check the identifiability conditions, numerical simulations are carried out using the IEEE 123-bus system and a large-scale real utility feeder.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Adaptive Control of Distributed Energy Resources for Distribution Grid Voltage Stability

Volt-VAR and Volt-Watt functionality in photovoltaic (PV) smart inverters provide mechanisms to ensure system voltage magnitudes and power factors remain within acceptable limits. However, these control functions can become unstable, introducing oscillations in system voltages when not appropriately configured or maliciously altered during a cyberattack. In the event that Volt-VAR and Volt-Watt control functions in a portion of PV smart inverters in a distribution grid are unstable, the proposed adaptation scheme utilizes the remaining and stably-behaving PV smart inverters and other Distributed Energy Resources to mitigate the effect of the instability. The adaptation mechanism is entirely decentralized, model-free, communication-free, and requires virtually no external configuration. Here we provide a derivation of the adaptive control approach and validate the algorithm in experiments on the IEEE 37 and 8500 node test feeders.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Adaptive Control Algorithm to Adjust Settings in Photovoltaic Inverters for Electric Grid Cybersecurity (DERAC) v0.1

Volt-VAR and Volt-Watt functionality in photovoltaic (PV) smart inverters provide mechanisms to ensure system voltage magnitudes and power factors remain within acceptable limits. However, these control functions can become unstable, introducing oscillations in system voltages when not appropriately configured or maliciously altered during a cyberattack. In the event that Volt-VAR and Volt-Watt control functions in a portion of PV smart inverters in a distribution grid are unstable, the proposed adaptation scheme utilizes the remaining and stably-behaving PV smart inverters and other Distributed Energy Resources to mitigate the effect of the instability in real-time. The adaptation mechanism is entirely decentralized, model-free, communication-free, and requires virtually no external configuration. This repository provides code to simulate the algorithm in experiments on the IEEE 37 node test feeder.

Arnold, Daniel↗

Deep Reinforcement Learning for Autonomous Water Heater Control

Electric water heaters represent 14% of the electricity consumption in residential buildings. An average household in the United States (U.S.) spends about USD 400–600 (0.45 ¢/L–0.68 ¢/L) on water heating every year. In this context, water heaters are often considered as a valuable asset for Demand Response (DR) and building energy management system (BEMS) applications. To this end, this study proposes a model-free deep reinforcement learning (RL) approach that aims to minimize the electricity cost of a water heater under a time-of-use (TOU) electricity pricing policy by only using standard DR commands. In this approach, a set of RL agents, with different look ahead periods, were trained using the deep Q-networks (DQN) algorithm and their performance was tested on an unseen pair of price and hot water usage profiles. The testing results showed that the RL agents can help save electricity cost in the range of 19% to 35% compared to the baseline operation without causing any discomfort to end users. Additionally, the RL agents outperformed rule-based and model predictive control (MPC)-based controllers and achieved comparable performance to optimization-based control.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Role of reinforcement learning for risk-based robust control of cyber-physical energy systems

Critical infrastructures such as cyber-physical energy systems (CPS-E) integrate information flow and physical operations that are vulnerable to natural and targeted failures. Safe, secure, and reliable operation and control of CPS-E is critical to ensure societal well-being and economic prosperity. Automated control is key for real-time operations and may be mathematically cast as a sequential decision-making problem under uncertainty. Emergence of data-driven techniques for decision making under uncertainty, such as reinforcement learning (RL), have led to promising advances for addressing sequential decision-making problems for risk-based robust CPS-E control. However, existing research challenges include understanding the applicability of RL methods across diverse CPS-E applications, addressing the effect of risk preferences across multiple RL methods, and development of open-source domain-aware simulation environments for RL experimentation within a CPS-E context. This article systematically analyzes the applicability of four types of RL methods (model-free, model-based, hybrid model-free and model-based, and hierarchical) for risk-based robust CPS-E control. Problem features and solution stability for the RL methods are also discussed. We demonstrate and compare the performance of multiple RL methods under different risk specifications (risk-averse, risk-neutral, and risk-seeking) through the development and application of an open-source simulation environment. Motivating numerical simulation examples include representative single-zone and multizone building control use cases. Finally, six key insights for future research and broader adoption of RL methods are identified, with specific emphasis on problem features, algorithmic explainability, and solution stability.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Model-free distributed learning

Model-free learning for synchronous and asynchronous quasi-static networks is presented. The network weights are continuously perturbed, while the time-varying performance index is measured and correlated with the perturbation signals; the correlation output determines the changes in the weights. The perturbation may be either via noise sources or orthogonal signals. The invariance to detailed network structure mitigates large variability between supposedly identical networks as well as implementation defects. This local, regular, and completely distributed mechanism requires no central control and involves only a few global signals. Thus it allows for integrated on-chip learning in large analog and optical networks.

Dembo, Amir↗

Automated Escape Guidance Algorithms for An Escape Vehicle

An escape vehicle was designed to provide an emergency evacuation for crew members living on a space station. For maximum escape capability, the escape vehicle needs to have the ability to safely evacuate a station in a contingency scenario such as an uncontrolled (e.g., tumbling) station. This emergency escape sequence will typically be divided into three events: The fust separation event (SEP1), the navigation reconstruction event, and the second separation event (SEP2). SEP1 is responsible for taking the spacecraft from its docking port to a distance greater than the maximum radius of the rotating station. The navigation reconstruction event takes place prior to the SEP2 event and establishes the orbital state to within the tolerance limits necessary for SEP2. The SEP2 event calculates and performs an avoidance burn to prevent station recontact during the next several orbits. This paper presents the tools and results for the whole separation sequence with an emphasis on the two separation events. The fust challenge includes collision avoidance during the escape sequence while the station is in an uncontrolled rotational state, with rotation rates of up to 2 degrees per second. The task of avoiding a collision may require the use of the Vehicle's de-orbit propulsion system for maximum thrust and minimum dwell time within the vicinity of the station vicinity. The thrust of the propulsion system is in a single direction, and can be controlled only by the attitude of the spacecraft. Escape algorithms based on a look-up table or analytical guidance can be implemented since the rotation rate and the angular momentum vector can be sensed onboard and a-priori knowledge of the position and relative orientation are available. In addition, crew intervention has been provided for in the event of unforeseen obstacles in the escape path. The purpose of the SEP2 burn is to avoid re-contact with the station over an extended period of time. Performing this maneuver properly requires knowledge of the orbital state, which is obtained during the navigation state reconstruction event. Since the direction of the delta-v of the SEPI maneuver is a random variable with respect to the Local Vertical Local Horizontal (LVLH) coordinate system, calculating the required SEP2 burn is a challenge. This problem was solved using a neural network as a model-free function approximation technique.

Flanary, Ronald↗

Implications of stop-and-go traffic on training learning-based car-following control

Learning-based car-following control (LCC) of connected and autonomous vehicles (CAVs) is gaining significant attention with the advancement of computing power and data accessibility. While the flexibility and large model capacity of model-free architecture enable LCC to potentially outperform the model-based car-following (CF) model in improving traffic efficiency and mitigating congestion, the generalizability of LCC for traffic conditions different from the training environment/dataset is not well-understood. Herein, this study seeks to explore the impact of stop-and-go traffic in the training dataset on the generalizability of LCC. It uses the characteristics of lead vehicle trajectories to describe stop-and-go traffic, and links the theory of identifiability (i.e., obtaining a unique parameter estimation result using sensor measurements) to the generalizability of behavior cloning (BC) and policy-based deep reinforcement learning (DRL). Correspondingly, the study shows theoretically that: (i) stop-and-go traffic can enable the property of identifiability and enhance the control performance of BC-based LCC in different traffic conditions; (ii) stop-and-go traffic is not necessary for DRL-based LCC to generalize to different traffic conditions; (iii) DRL-based LCC trained with only constant-speed lead vehicle trajectories (not sufficient to ensure identifiability) can be generalized to different traffic conditions; and (iv) stop-and-go traffic increases variance in the training dataset, which improves the convergence of parameter estimation while negatively impacting the convergence of DRL to the optimal control policy. Numerical experiments validate the above findings, illustrating that BC-based LCC entails comprehensive training datasets for generalizing to different traffic conditions, while DRL-based LCC can achieve generalization with simple free-flow traffic training environments. This further suggests DRL as a more promising and cost-effective LCC approach to reduce operational costs, mitigate traffic congestion, and enhance safety and mobility, which can accelerate the deployment and acceptance of CAVs.

33 ADVANCED PROPULSION SYSTEMS↗

Model-based and Model-free Designs for an Extended Continuous-time LQR with Exogenous Inputs

We present an extended linear quadratic regulator (LQR) design for continuous-time linear time-invariant (LTI) systems in the presence of exogenous inputs. We first propose a model-based solution with cost minimization guarantees for states and inputs using dynamic programming (DP). The control law consists of a combination of the optimal state feedback and an additional optimal term dependent on the exogenous inputs. The control gains for the two components are obtained by solving a set of matrix differential equations. We provide these solutions for both finite horizons and steady-state cases. In the second part of the paper, we formulate a reinforcement learning (RL) based algorithm which does not need any model information except the input matrix, and can compute an approximate steady-state LQR gain using measurements of the states, the control inputs, and the exogenous inputs. Both model-based and data-driven optimal control algorithms are tested with a numerical example under different exogenous inputs showcasing the effectiveness of the designs.

Mukherjee, Sayak↗

A Modular and Transferable Reinforcement Learning Framework for the Fleet Rebalancing Problem

Mobility on demand (MoD) systems show great promise in realizing flexible and efficient urban transportation. However, significant technical challenges arise from operational decision making associated with MoD vehicle dispatch and fleet rebalancing. For this reason, operators tend to employ simplified algorithms that have been demonstrated to work well in a particular setting. To help bridge the gap between novel and existing methods, we propose a modular framework for fleet rebalancing based on model-free reinforcement learning (RL) that can leverage an existing dispatch method to minimize system cost. In particular, by treating dispatch as part of the environment dynamics, a centralized agent can learn to intermittently direct the dispatcher to reposition free vehicles and mitigate against fleet imbalance. We formulate RL state and action spaces as distributions over a grid partitioning of the operating area, making the framework scalable and avoiding the complexities associated with multiagent RL. Numerical experiments, using real-world trip and network data, demonstrate that RL reduces waiting time by 28% to 38% for the same-day evaluation, 17% to 44% for cross-day evaluation, and 22% to 25% for cross-season evaluation compared with no rebalancing scenarios. This approach has several distinct advantages over baseline methods including: improved system cost; high degree of adaptability to the selected dispatch method; and the ability to perform scale-invariant transfer learning between problem instances with similar vehicle and request distributions.

33 ADVANCED PROPULSION SYSTEMS↗

Physics-Informed Evolutionary Strategy Based Control for Mitigating Delayed Voltage Recovery

Here, in this work we propose a novel data-driven, real-time power system voltage stability control method based on the physics-informed guided meta evolutionary strategy (ES). The main objective is to quickly provide an adaptive control strategy to secure system voltage stability. The problem is challenging due to the high-dimensional feature of the power system model and the fast-changing and uncertain nature of power system operation scenarios. To this end, a model-free and derivative-free guided ES method is applied. The method is further combined with a meta-learning strategy to make the learnt control policy automatically adapted to unseen operation conditions and fault scenarios, which is highly desired for real-time emergency control. Last but not least, physical knowledge is embedded in the above method through a trainable action mask technique to rule out unnecessary load shedding actions for better learning and control performance. Case studies on the IEEE 300-bus system and comparisons with other state-of-the-art benchmark methods verify the superiority of the proposed physics-informed guided meta ES method in realizing fast and adaptive power system voltage stability control.

42 ENGINEERING↗