Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Markov decision process”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

A Novel Multi-Agent Deep Reinforcement Learning-enabled Distributed Power Allocation Scheme for mmWave Cellular Networks

We consider the power allocation problem over shared spectrum for millimeter-Wave (mmWave) cellular downlink. Existing approaches usually find sub-optimal solutions by solving a non-convex optimization which leads to scalability issues due to centralized control. Therefore, distributed and adaptive approaches are desirable. Recently, model-free Deep Reinforcement Learning (DRL) has achieved success in such wireless resource management tasks. By modeling the radio environment as a Markov Decision Process (MDP) with the base stations (BSs) being the agents, power allocation can be automated at the agent level with comparable throughput performance to conventional centralized schemes. The multi-agent setting presents new challenges as the radio environment is impacted by the joint actions of the agents and is no longer stationary from any individual agent’s perspective. Existing literature bypasses this non-stationarity violation by ignoring it which may cause performance degradation. To tackle this issue, we propose a distributed continuous power allocation scheme based on a modified version of multi-agent Deep Deterministic Policy Gradient (MADDPG) that is tailored for the distributed multiple-agent setting. The proposed scheme employs a centralized-training distributed-execution framework where Q-functions are trained over subsets of BSs while each BS determines its transmit power based only on its own local observation. It admits constant per-BS communication and computation complexity and is thus scalable to large networks. Numerical evaluation shows that the proposed scheme adapts well to a wide range of interference conditions and can achieve comparable or better performance than several state-of-the-art non-learning approaches.

99 GENERAL AND MISCELLANEOUS↗

Learning infinite-horizon average-reward restless multi-action bandits via index awareness

We consider the online restless bandits with average-reward and multiple actions, where the state of each arm evolves according to a Markov decision process (MDP), and the reward of pulling an arm depends on both the current state of the corresponding MDP and the action taken. Since finding the optimal control is typically intractable for restless bandits, existing learning algorithms are often computationally expensive or with a regret bound that is exponential in the number of arms and states. In this paper, we advocate \textit{index-aware reinforcement learning} (RL) solutions to design RL algorithms operating on a much smaller dimensional subspace by exploiting the inherent structure in restless bandits. Specifically, we first propose novel index policies to address dimensionality concerns, which are provably optimal. We then leverage the indices to develop two low-complexity index-aware RL algorithms, namely, (i) GM-R2MAB, which has access to a generative model; and (ii) UC-R2MAB, which learns the model using an upper confidence style online exploitation method. We prove that both algorithms achieve a sub-linear regret that is only polynomial in the number of arms and states. A key differentiator between our algorithms and existing ones stems from the fact that our RL algorithms contain a novel exploitation that leverages our proposed provably optimal index policies for decision-makings.

Xiong, Guojun↗

Variational actor-critic algorithms,

We introduce a class of variational actor-critic algorithms based on a variational formulation over both the value function and the policy. The objective function of the variational formulation consists of two parts: one for maximizing the value function and the other for minimizing the Bellman residual. Besides the vanilla gradient descent with both the value function and the policy updates, we propose two variants, the clipping method and the flipping method, in order to speed up the convergence. We also prove that, when the prefactor of the Bellman residual is sufficiently large, the fixed point of the algorithm is close to the optimal policy.

97 MATHEMATICS AND COMPUTING↗

Deep Reinforcement Learning for Distribution System Cyber Attack Defense with DERs

The use of smart inverter capabilities of distributed energy resources (DERs) enhances the grid reliability but in the meanwhile exhibits more vulnerabilities to cyber-attacks. This paper proposes a deep reinforcement learning (DRL)-based defense approach. The defense problem is reformulated as a Markov decision making process to control DERs and minimizing load shedding to address the voltage violations caused by cyber-attacks. The original soft actor-critic (SAC) method for continuous actions has been extended to handle discrete and continuous actions for controlling DERs' setpoints and loadshedding scenarios. Numerical comparison results with other control approaches, such as Volt-VAR and Volt-Watt on the modified IEEE 33-node, show that the proposed method can achieve better voltage regulation and have less power losses in the presence of cyber-attacks.

active distribution systems↗

Adaptive Deep Reinforcement Learning Algorithm for Distribution System Cyber Attack Defense With High Penetration of DERs

With grid modernization, smart inverters are increasingly used to execute advanced controls for distribution network reliability. However, this also increases the cyber-attack space. Here this paper focuses on the defense approaches to restore the system to normal operation circumstances in the presence of cyber-attacks. A unique deep reinforcement learning (DRL) method is developed to minimize voltage violations and reduce power losses for impacted feeders. The defense problem is reformulated as a Markov decision-making process to dynamically control DERs while minimizing load shedding. This is achieved via an improved soft actor-critic (SAC)-based DRL algorithm, which can govern DER set points and load-shedding scenarios in discrete and continuous modes via the auto-tune entropy and Gaussian policy features. Numerical comparison results on the modified IEEE 123-node system with other control approaches, such as Volt-VAR (VV), Volt-Watt (VW), and model predictive control (MPC) show that the proposed method can eliminate voltage violations and provide feasible control actions that perform complete mitigation of cyber-threats.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Decentralized control of Markovian decision processes: Existence Sigma-admissable policies

The problem of formulating and analyzing Markov decision models having decentralized information and decision patterns is examined. Included are basic examples as well as the mathematical preliminaries needed to understand Markov decision models and, further, to superimpose decentralized decision structures on them. The notion of a variance admissible policy for the model is introduced and it is proved that there exist (possibly nondeterministic) optional policies from the class of variance admissible policies. Directions for further research are explored.

Greenland, A.↗

Reinforcement learning for adaptive maintenance policy optimization under imperfect knowledge of the system degradation model and partial observability of system states

Maintenance policy optimization usually is faced with challenges that arise from an imperfect knowledge of system degradation models and from the partial observability of system degradation states. Here, this paper proposes a reinforcement learning method to address these two challenges for a class of maintenance problems with Markov degradation processes. The reinforcement learning approach consists of a learning component and a planning component. Using sequentially collected observations, at each step of decision-making the learning component improves the knowledge of system degradation in terms of the probability distributions of the transition rates based on sequential Bayesian inference. Using the updated transition rates, at each step of decision-making the maintenance policy optimization problem is then formulated as a partially observable Markov decision problem, and the planning component computes the optimal maintenance policy that maximizes the expected cumulative reward. The proposed method is illustrated using a numerical example with repair and inspection maintenance actions. The result shows that as more observations are collected, the learning component progressively learns the true system degradation process, and the planning component adjusts the optimal maintenance policy accordingly as well, which leads to increased reward.

42 ENGINEERING↗

Performance evaluation of fault-tolerant systems with application to the IUS

A new method for quantitatively evaluating the performance of fault-tolerant systems is developed and applied to an example. The method assumes that the random failure and diagnostic decision behavior of the system can be modeled by a finite state Markov process. A performance value must be assigned to each of the states of the model. The method then generates the moments of the probability mass function of the cumulative performance and uses these to generate a maximum entropy approximation to the performance PMF. Some computational considerations are discussed. The method is applied to a typical mission for the Inertial Upper Stage to examine the attitude accuracy performance of the inertial system.

Missana, J.-O.↗

Metrics for Labeled Markov Systems

Partial Labeled Markov Chains are simultaneously generalizations of process algebra and of traditional Markov chains. They provide a foundation for interacting discrete probabilistic systems, the interaction being synchronization on labels as in process algebra. Existing notions of process equivalence are too sensitive to the exact probabilities of various transitions. This paper addresses contextual reasoning principles for reasoning about more robust notions of "approximate" equivalence between concurrent interacting probabilistic systems. The present results indicate that:We develop a family of metrics between partial labeled Markov chains to formalize the notion of distance between processes. We show that processes at distance zero are bisimilar. We describe a decision procedure to compute the distance between two processes. We show that reasoning about approximate equivalence can be done compositionally by showing that process combinators do not increase distance. We introduce an asymptotic metric to capture asymptotic properties of Markov chains; and show that parallel composition does not increase asymptotic distance.

Desharnais, Josee↗

Detection of digital FSK using a phase-locked loop

A theory is presented for the design of a digital FSK receiver which employs a phase-locked loop to set up the desired matched filter as the arriving signal frequency switches. The developed mathematical model makes it possible to establish the error probability performance of systems which employ a class of digital FM modulations. The noise mechanism which accounts for decision errors is modeled on the basis of the Meyr distribution and renewal Markov process theory.

Lindsey, W. C.↗

Perseveration effects in detection tasks with correlated decision intervals

An investigation of the behavior of the human decisionmaker is described for a task related to the problem of a pilot using a traffic situation display to avoid collisions. This sequential signal detection task is characterized by highly correlated signals with time varying strength. Experimental results are presented and the behavior of the observers is analyzed using the theory of Markov processes and classical signal detection theory. Mathematical models are developed which describe the main result of the experiment: that correlation in sequential signals induced perseveration in the observer response and a strong tendency to repeat their previous decision, even when they were wrong.

Gai, E. G.↗

Quantum Decision Maker Theory and Simulation

A quantum device simulating the human decision making process is introduced. It consists of quantum recurrent nets generating stochastic processes which represent the motor dynamics, and of classical neural nets describing the evolution of probabilities of these processes which represent the mental dynamics.

Quantum↗

Machine Learning and Economic Models to Enable Risk-Informed Condition Based Maintenance of a Nuclear Plant Asset

The primary objective of this research is to address challenges in the implementation of risk-informed, condition-based predictive maintenance (PdM), which reduces operating costs while still maintaining the safety and reliability of commercial nuclear power plants (NPPs). To achieve the objective, risk models are being developed by taking advantage of advancements in data analytics, deep learning, machine learning (ML), and artificial intelligence (AI). The notable outcomes presented in the report include ? Development of a ML models using heterogeneous plant process and vibration data collected at different spatial and temporal resolutions from the Salem?s CWS to diagnose a circulating water pump (CWP) failure based on salient fault signatures. The developed diagnostic models are extendable to other faults associated with CWPs and CWP motors given associated fault signatures. ? Development of a natural language processing (NLP) technique to automatically classify the WO data into different categories. The developed NLP technique was validated on independent WO data. This automates the tedious and time-consuming activity of mining and classifying WOs by subject matter experts. ? Estimation of mean time between downtime (i.e., time duration between time instances when 1 or more CWPs are not available) and developed an approach to establish reliability of CWS components using unstructured WO data along with CWS plant process data. ? Formulation of economic model based on Markov chain models. The parameters of associated with the transition rate between different states of Markov chain models were estimated using WO data. The economic model formulation and discussion captures both time-independent and time-dependent parameter variation, leading to risk-informed decision-making.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗