Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Markov decision processes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Federated Deep Reinforcement Learning for Decentralized VVO of BTM DERs

The future of grid control requires a hybrid approach combining centralized and decentralized methods to fully utilize the potential of smart edge devices with artificial intelligence (AI) capabilities. This paper aims to develop and evaluate a federated deep reinforcement learning (FDRL) framework for decentralized adaptive volt-var optimization (VVO) of behind-the-meter (BTM) distributed energy resources (DERs). First, this paper models a single deep reinforcement learning (DRL) agent using the Markov Decision Process (MDP) framework for decentralized adaptive VVO of BTM DERs. Two DRL algorithms, soft actor-critic (SAC) and twin-delayed deep deterministic policy gradient (TD3), are compared for their effectiveness in optimizing VVO. Results show that TD3 outperforms SAC, achieving a 71.3% improvement in mean reward. Finally, the DRL agent is deployed within the FDRL framework, using the Flower platform, to enhance learning, provide adaptive control, and ensure data privacy for BTM DERs.

Ravi, Abhijith↗

Safe Deep Reinforcement Learning for Robust Frequency and Voltage-Constrained Networked Microgrid Restoration

Here, this paper proposes a safe soft actor-critic reinforcement learning (RL) algorithm–based controller for networked microgrid restoration. It formulates the post black-start start as a finite-horizon constrained Markov decision process. The RL agent co-optimizes real and reactive power set-points for both grid-forming and grid-following inverters under explicit voltage and frequency constraints, while enforcing proper power sharing via the Mean Active Power Sharing Index (MPSI) and Mean Reactive Power Sharing Index (MQSI). Numerical results obtained on the IEEE 123-bus distribution system show that the proposed method achieves a mean voltage build-up time of 0.01 s without breaching the 5% sharing-violation budget under various load scenarios, considering MPSI and MQSI indices. These findings demonstrate that the proposed method yields fast and safe black-start schedules without resorting to heuristic penalties.

Selim, Alaa [Dartmouth College, Hanover, NH (Unite↗

Distributional Deep Reinforcement Learning-Based Emergency Frequency Control

Emergency frequency control is one of the most critical approaches to maintain power system stability after major disturbances. With the increasing number of grid-connected renewable energy sources, existing model-based methods of frequency control are facing up with challenges of computational speed and scalability for large-scale systems. In this paper, the emergency frequency control problem is formulated as a Markov Decision Process (MDP) and solved through a novel Distributional Deep Reinforcement Learning (DDRL) method, namely the distributional soft actor critic (DSAC) method. Compared with other RL methods that only estimate the mean value, the proposed DSAC model estimates the distribution of value function over returns. This advancement can lead to more insights and knowledge for the agent, with the benefit of a much faster and more stable learning process, and the improved frequency control performance. Here, the simulation results on IEEE 39-bus and IEEE 118-bus systems demonstrate the effectiveness and robustness of proposed models, as well as the advantage compared to other state-of-the-art DRL algorithms.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Deep Reinforcement Learning based Model-free On-line Dynamic Multi-Microgrid Formation to Enhance Resilience

Multi-microgrid formation (MMGF) is a promising solution for enhancing power system resilience. This paper proposes a new deep reinforcement learning (RL) based model-free on-line dynamic MMGF scheme. Additionally, the dynamic MMGF problem is formulated as a Markov decision process, and a complete deep RL framework is specially designed for the topologytransformable micro-grids. In order to reduce the large action space caused by flexible switch operations, a topology transformation method is proposed and an action-decoupling Q-value is applied. Then, a convolutional neural network (CNN) based multi-buffer double deep Q-network (CM-DDQN) is developed to further improve the learning ability of the original DQN method. The proposed deep RL method provides real-time computing to support the on-line dynamic MMGF scheme, and the scheme handles a long-term resilience enhancement problem using an adaptive on-line MMGF to defend changeable conditions. The effectiveness of the proposed method is validated using a 7-bus system and the IEEE 123-bus system. The results show strong learning ability, timely response for varying system conditions and convincing resilience enhancement.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Learning Robust Marking Policies for Adaptive Mesh Refinement

Here in this work, we revisit the marking decisions made in the standard adaptive finite element method (AFEM). Experience shows that a naïve marking policy leads to inefficient use of computational resources for adaptive mesh refinement (AMR). Consequently, using AMR in practice often involves ad-hoc or time-consuming offline parameter tuning to set appropriate parameters for the marking subroutine. To address these practical concerns, we recast AMR as a Markov decision process in which refinement parameters can be selected on-the-fly at run time, without the need for pre-tuning by expert users. In this new paradigm, the refinement parameters are also chosen adaptively via a marking policy that can be optimized using methods from reinforcement learning. We use the Poisson equation to demonstrate our techniques on h- and hp-refinement benchmark problems, and our experiments suggest that superior marking policies remain undiscovered for many classical AFEM applications. Furthermore, an unexpected observation from this work is that marking policies trained on one family of PDEs are sometimes robust enough to perform well on problems far outside the training family. For illustration, we show that a simple hp-refinement policy trained on 2D domains with only a single re-entrant corner can be deployed on far more complicated 2D domains, and even 3D domains, without significant performance loss. For reproduction and broader adoption, we accompany this work with an open-source implementation of our methods.

97 MATHEMATICS AND COMPUTING↗

Automated Adversary-in-the-Loop Cyber-Physical Defense Planning

Security of cyber-physical systems (CPS) continues to pose new challenges due to the tight integration and operational complexity of the cyber and physical components. To address these challenges, this article presents a domain-aware, optimization-based approach to determine an effective defense strategy for CPS in an automated fashion—by emulating a strategic adversary in the loop that exploits system vulnerabilities, interconnection of the CPS, and the dynamics of the physical components. Our approach builds on an adversarial decision-making model based on a Markov Decision Process (MDP) that determines the optimal cyber (discrete) and physical (continuous) attack actions over a CPS attack graph. The defense planning problem is modeled as a non-zero-sum game between the adversary and defender. We use a model-free reinforcement learning method to solve the adversary’s problem as a function of the defense strategy. We then employ Bayesian optimization (BO) to find an approximate best-response for the defender to harden the network against the resulting adversary policy. This process is iterated multiple times to improve the strategy for both players. We demonstrate the effectiveness of our approach on a ransomware-inspired graph with a smart building system as the physical process. Numerical studies show that our method converges to a Nash equilibrium for various defender-specific costs of network hardening.

97 MATHEMATICS AND COMPUTING↗

Hybrid-RL-MPC4CLR (Hybird-Reinforcement-Learning-Model-Predictive-Control-for-Reserve-Policy-Assisted-Critical-Load-Restoration-in-Distribution-Grids)

Hybrid-RL-MPC4CLR was developed as a hybrid controller for active distribution grid critical load restoration, combining deep reinforcement learning (RL) and model predictive control (MPC) aiming at maximizing total restored load following an extreme event. The RL determines a policy for quantifying operating reserve requirements, thereby hedging against uncertainty, while the MPC models grid operations incorporating the RL policy actions (i.e., reserve requirements), renewable (wind and solar) power predictions, and load demand forecasts. The developers formulated the reserve requirement determination problem as a sequential decision-making problem based on the Markov Decision Process (MDP) and design an RL learning environment based on the OpenAI Gym framework and MPC simulation. The RL agent reward and MPC objective function aim to maximize and monotonically increase total restored load and minimize load shedding and renewable power curtailment. The software is developed using various software packages in Python. The MPC's optimal power flow (OPF) model is implemented using the Pyomo package, the RL simulation environment is implemented using the MPC simulation with various scenarios of renewable energy and load demand profiles and power outage beginning times, based on the OpenAI Gym framework. The RL agent training is performed using the RLlib Ray package. The RL algorithm is trained offline using historical forecasts of renewable generation and load demand profiles. Simulation analysis and performance tests are conducted using a modified IEEE 13-bus distribution test feeder containing wind turbine, photovoltaic, microturbine, and battery.

Eseye, Abinet Tesfaye↗

Robotic Planning under Uncertainty in Spatiotemporal Environments in Expeditionary Science

In the expeditionary sciences, spatiotemporally varying environments -- hydrothermal plumes, algal blooms, lava flows, or animal migrations -- are ubiquitous. Mobile robots are uniquely well-suited to study these dynamic, mesoscale natural environments. We formalize expeditionary science as a sequential decision-making problem, modeled using the language of partially-observable Markov decision processes (POMDPs). Solving the expeditionary science POMDP under real-world constraints requires efficient probabilistic modeling and decision-making in problems with complex dynamics and observational models. Previous work in informative path planning, adaptive sampling, and experimental design have shown compelling results, largely in static environments, using data-driven models and information-based rewards. However, these methodologies do not trivially extend to expeditionary science in spatiotemporal environments: they generally do not make use of scientific knowledge such as equations of state dynamics, they focus on information gathering as opposed to scientific task execution, and they make use of decision-making approaches that scale poorly to large, continuous problems with long planning horizons and real-time operational constraints. In this work, we discuss these and other challenges related to probabilistic modeling and decision-making in expeditionary science, and present some of our preliminary work that addresses these gaps. We ground our results in a real expeditionary science deployment of an autonomous underwater vehicle (AUV) in the deep ocean for hydrothermal vent discovery and characterization. Our concluding thoughts highlight remaining work to be done, and the challenges that merit consideration by the reinforcement learning and decision-making community.

Preston, Victoria↗

A Hybrid Reinforcement Learning-MPC Approach for Distribution System Critical Load Restoration: Preprint

This paper proposes a hybrid control approach for distribution system critical load restoration, combining deep reinforcement learning (RL) and model predictive control (MPC) aiming at maximizing total restored load following an extreme event. RL determines a policy for quantifying operating reserve requirements, thereby hedging against uncertainty, while MPC models grid operations incorporating RL policy actions, i.e., the reserve requirement, renewable (wind and solar) power predictions, and load demand forecasts. We formulate the reserve requirement determination problem as a sequential decision making problem based on the Markov Decision Process (MDP) and design an RL learning environment based on the OpenAI Gym framework and MPC. The RL agent reward and MPC objective function aim to maximize and monotonically increase total restored load and minimize load shedding and renewable power curtailment. The RL algorithm is trained off-line using historical forecast of renewable generation and load demand. The method is tested using a modified IEEE 13-bus distribution test feeder containing wind turbine, photovoltaic, microturbine and battery. Case studies demonstrated that the proposed method outperforms other operating reserve determination methods.

distribution system↗

A Hybrid Reinforcement Learning-MPC Approach for Distribution System Critical Load Restoration

This paper proposes a hybrid control approach for distribution system critical load restoration, combining deep reinforcement learning (RL) and model predictive control (MPC) aiming at maximizing total restored load following an extreme event. RL determines a policy for quantifying operating reserve requirements, thereby hedging against uncertainty, while MPC models grid operations incorporating RL policy actions (i.e., reserve requirements), renewable (wind and solar) power predictions, and load demand forecasts. We formulate the reserve requirement determination problem as a sequential decision-making problem based on the Markov Decision Process (MDP) and design an RL learning environment based on the OpenAI Gym framework and MPC simulation. The RL agent reward and MPC objective function aim to maximize and monotonically increase total restored load and minimize load shedding and renewable power curtailment. The RL algorithm is trained offline using a historical forecast of renewable generation and load demand. The method is tested using a modified IEEE 13-bus distribution test feeder containing wind turbine, photovoltaic, microturbine, and battery. Case studies demonstrated that the proposed method outperforms other policies with static operating reserves.

distribution system↗

A Novel Multi-Agent Deep Reinforcement Learning-enabled Distributed Power Allocation Scheme for mmWave Cellular Networks

We consider the power allocation problem over shared spectrum for millimeter-Wave (mmWave) cellular downlink. Existing approaches usually find sub-optimal solutions by solving a non-convex optimization which leads to scalability issues due to centralized control. Therefore, distributed and adaptive approaches are desirable. Recently, model-free Deep Reinforcement Learning (DRL) has achieved success in such wireless resource management tasks. By modeling the radio environment as a Markov Decision Process (MDP) with the base stations (BSs) being the agents, power allocation can be automated at the agent level with comparable throughput performance to conventional centralized schemes. The multi-agent setting presents new challenges as the radio environment is impacted by the joint actions of the agents and is no longer stationary from any individual agent’s perspective. Existing literature bypasses this non-stationarity violation by ignoring it which may cause performance degradation. To tackle this issue, we propose a distributed continuous power allocation scheme based on a modified version of multi-agent Deep Deterministic Policy Gradient (MADDPG) that is tailored for the distributed multiple-agent setting. The proposed scheme employs a centralized-training distributed-execution framework where Q-functions are trained over subsets of BSs while each BS determines its transmit power based only on its own local observation. It admits constant per-BS communication and computation complexity and is thus scalable to large networks. Numerical evaluation shows that the proposed scheme adapts well to a wide range of interference conditions and can achieve comparable or better performance than several state-of-the-art non-learning approaches.

99 GENERAL AND MISCELLANEOUS↗

Learning infinite-horizon average-reward restless multi-action bandits via index awareness

We consider the online restless bandits with average-reward and multiple actions, where the state of each arm evolves according to a Markov decision process (MDP), and the reward of pulling an arm depends on both the current state of the corresponding MDP and the action taken. Since finding the optimal control is typically intractable for restless bandits, existing learning algorithms are often computationally expensive or with a regret bound that is exponential in the number of arms and states. In this paper, we advocate \textit{index-aware reinforcement learning} (RL) solutions to design RL algorithms operating on a much smaller dimensional subspace by exploiting the inherent structure in restless bandits. Specifically, we first propose novel index policies to address dimensionality concerns, which are provably optimal. We then leverage the indices to develop two low-complexity index-aware RL algorithms, namely, (i) GM-R2MAB, which has access to a generative model; and (ii) UC-R2MAB, which learns the model using an upper confidence style online exploitation method. We prove that both algorithms achieve a sub-linear regret that is only polynomial in the number of arms and states. A key differentiator between our algorithms and existing ones stems from the fact that our RL algorithms contain a novel exploitation that leverages our proposed provably optimal index policies for decision-makings.

Xiong, Guojun↗

Variational actor-critic algorithms,

We introduce a class of variational actor-critic algorithms based on a variational formulation over both the value function and the policy. The objective function of the variational formulation consists of two parts: one for maximizing the value function and the other for minimizing the Bellman residual. Besides the vanilla gradient descent with both the value function and the policy updates, we propose two variants, the clipping method and the flipping method, in order to speed up the convergence. We also prove that, when the prefactor of the Bellman residual is sufficiently large, the fixed point of the algorithm is close to the optimal policy.

97 MATHEMATICS AND COMPUTING↗

Deep Reinforcement Learning for Distribution System Cyber Attack Defense with DERs

The use of smart inverter capabilities of distributed energy resources (DERs) enhances the grid reliability but in the meanwhile exhibits more vulnerabilities to cyber-attacks. This paper proposes a deep reinforcement learning (DRL)-based defense approach. The defense problem is reformulated as a Markov decision making process to control DERs and minimizing load shedding to address the voltage violations caused by cyber-attacks. The original soft actor-critic (SAC) method for continuous actions has been extended to handle discrete and continuous actions for controlling DERs' setpoints and loadshedding scenarios. Numerical comparison results with other control approaches, such as Volt-VAR and Volt-Watt on the modified IEEE 33-node, show that the proposed method can achieve better voltage regulation and have less power losses in the presence of cyber-attacks.

active distribution systems↗

Adaptive Deep Reinforcement Learning Algorithm for Distribution System Cyber Attack Defense With High Penetration of DERs

With grid modernization, smart inverters are increasingly used to execute advanced controls for distribution network reliability. However, this also increases the cyber-attack space. Here this paper focuses on the defense approaches to restore the system to normal operation circumstances in the presence of cyber-attacks. A unique deep reinforcement learning (DRL) method is developed to minimize voltage violations and reduce power losses for impacted feeders. The defense problem is reformulated as a Markov decision-making process to dynamically control DERs while minimizing load shedding. This is achieved via an improved soft actor-critic (SAC)-based DRL algorithm, which can govern DER set points and load-shedding scenarios in discrete and continuous modes via the auto-tune entropy and Gaussian policy features. Numerical comparison results on the modified IEEE 123-node system with other control approaches, such as Volt-VAR (VV), Volt-Watt (VW), and model predictive control (MPC) show that the proposed method can eliminate voltage violations and provide feasible control actions that perform complete mitigation of cyber-threats.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Reinforcement learning for adaptive maintenance policy optimization under imperfect knowledge of the system degradation model and partial observability of system states

Maintenance policy optimization usually is faced with challenges that arise from an imperfect knowledge of system degradation models and from the partial observability of system degradation states. Here, this paper proposes a reinforcement learning method to address these two challenges for a class of maintenance problems with Markov degradation processes. The reinforcement learning approach consists of a learning component and a planning component. Using sequentially collected observations, at each step of decision-making the learning component improves the knowledge of system degradation in terms of the probability distributions of the transition rates based on sequential Bayesian inference. Using the updated transition rates, at each step of decision-making the maintenance policy optimization problem is then formulated as a partially observable Markov decision problem, and the planning component computes the optimal maintenance policy that maximizes the expected cumulative reward. The proposed method is illustrated using a numerical example with repair and inspection maintenance actions. The result shows that as more observations are collected, the learning component progressively learns the true system degradation process, and the planning component adjusts the optimal maintenance policy accordingly as well, which leads to increased reward.

42 ENGINEERING↗