Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “stochastic decision-making”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Bill Savings vs. Backup Power: Evaluating operational tradeoffs for home solar+storage systems [Slides]

Adoption of residential solar photovoltaic+energy storage systems (PVESS) is driven by both bill savings opportunities and customer demand for backup power. Prior work by this team (Gorman et al., 2022; Gorman et al., 2023) explored PVESS backup power capabilities during long-duration power interruptions (e.g., due to severe weather events), when customers are assumed to be able to anticipate the event and charge their batteries in advance. In many cases, however, power interruptions are unpredictable (and often relatively short); for those types of events, a customer will typically set its battery to maintain some minimum capacity in reserve in case of an interruption, which reduces the capacity available for managing utility bills. This study evaluates this operational tradeoff to help customers and installers configure backup reserve settings, and to inform decision-making more generally about the customer value of backup power services compared to utility bill savings. This study utilizes Berkeley Lab’s PRESTO tool to produce stochastic simulations of (predominantly short-duration) power interruption events, and builds on an earlier case-study demonstrating PVESS backup performance during short-duration interruptions (Baik et al., 2023).

14 SOLAR ENERGY↗

An Online Approach to Solve the Dynamic Vehicle Routing Problem with Stochastic Trip Requests for Paratransit Services

Many transit agencies operating paratransit and microtransit services have to respond to trip requests that arrive in real-time, which entails solving hard combinatorial and sequential decision-making problems under uncertainty. To avoid decisions that lead to significant inefficiency in the long term, vehicles should be allocated to requests by optimizing a non-myopic utility function or by batching requests together and optimizing a myopic utility function. While the former approach is typically offline, the latter can be performed online. We point out two major issues with such approaches when applied to paratransit services in practice. First, it is difficult to batch paratransit requests together as they are temporally sparse. Second, the environment in which transit agencies operate changes dynamically (e.g., traffic conditions can change over time), causing the estimates that are learned offline to become stale. To address these challenges, we propose a fully online approach to solve the dynamic vehicle routing problem (DVRP) with time windows and stochastic trip requests that is robust to changing environmental dynamics by construction. We focus on scenarios where requests are relatively sparse—our problem is motivated by applications to paratransit services. We formulate DVRP as a Markov decision process and use Monte Carlo tree search to evaluate actions for any given state. Accounting for stochastic requests while optimizing a non-myopic utility function is computationally challenging; indeed, the action space for such a problem is intractably large in practice. To tackle the large action space, we leverage the structure of the problem to design heuristics that can sample promising actions for the tree search. Our experiments using real-world data from our partner agency show that the proposed approach outperforms existing state-of-the-art approaches both in terms of performance and robustness.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Survey on stochastic distribution systems: A full probability density function control theory with potential applications

Complex systems seen either in general engineering practice or economics are subjected to ever increased uncertainties that are mostly represented as random variables or parameters, and the characteristics of random variables are represented by their probability density functions (PDFs). Controlling their PDFs means to shape their stochastic distributions and in general it would provide a full treatment for system analysis and operational control and optimization. This leads to the development of stochastic distribution control (SDC) systems theory in the past decades, where the original aim of the controller design is to realize a shape control of the distributions of certain random variables in their PDFs sense for some engineering processes. Indeed, once the PDFs of these random variables or parameters are used to describe their distribution characters, the control task is to obtain control signals so that the output PDFs of stochastic systems are made to follow their target PDFs. The subject of SDC was initially originated for non-Gaussian stochastic control systems design but has found a wide spectrum of applications in general systems in terms of data-driven modeling, analysis, signal processing (filtering), data mining via multivariable statistics, decision-making (optimization) for systems subjected to uncertainties and even in economics. In this context, SDC constitutes an effective primer tool for complex system analysis, control and operational optimizations. In this review paper, a detailed survey of the developments on the research of SDC systems will be made together with their wide spectrum applications and future perspectives.

42 ENGINEERING↗

A Multi-Stage Stochastic Risk Assessment With Markovian Representation of Renewable Power

Probabilistic forecasts provide a distribution of possible outputs and so can capture the uncertainty and variability of Variable Renewable Energy (VRE). However, taking advantage of uncertainty information has practical challenges that make it difficult to integrate probabilistic forecasting into control room decision-making. This paper proposes a novel use-case for probabilistic forecasts by incorporating them into the hour-ahead operations for situational awareness via a risk-averse multi-stage stochastic program. We employ a Markovian representation of the probabilistic forecasts that enables the formulation of the multi-stage problem and avoids a scenario generation phase. We test the model on a realistically sized system to assess risk and showcase the capability of using probabilistic renewable forecast as input to produce probabilistic output forecasts of future system states. The results show that the model can capture time consistency in the reserves and Area Control Error (ACE) forecast. The solution times are adequate for risk profiling in hour-ahead timescales.

forecasting↗

Exploration of Advanced Probabilistic and Stochastic Design Methods

The primary objective of the three year research effort was to explore advanced, non-deterministic aerospace system design methods that may have relevance to designers and analysts. The research pursued emerging areas in design methodology and leverage current fundamental research in the area of design decision-making, probabilistic modeling, and optimization. The specific focus of the three year investigation was oriented toward methods to identify and analyze emerging aircraft technologies in a consistent and complete manner, and to explore means to make optimal decisions based on this knowledge in a probabilistic environment. The research efforts were classified into two main areas. First, Task A of the grant has had the objective of conducting research into the relative merits of possible approaches that account for both multiple criteria and uncertainty in design decision-making. In particular, in the final year of research, the focus was on the comparison and contrasting between three methods researched. Specifically, these three are the Joint Probabilistic Decision-Making (JPDM) technique, Physical Programming, and Dempster-Shafer (D-S) theory. The next element of the research, as contained in Task B, was focused upon exploration of the Technology Identification, Evaluation, and Selection (TIES) methodology developed at ASDL, especially with regards to identification of research needs in the baseline method through implementation exercises. The end result of Task B was the documentation of the evolution of the method with time and a technology transfer to the sponsor regarding the method, such that an initial capability for execution could be obtained by the sponsor. Specifically, the results of year 3 efforts were the creation of a detailed tutorial for implementing the TIES method. Within the tutorial package, templates and detailed examples were created for learning and understanding the details of each step. For both research tasks, sample files and tutorials are attached in electronic form with the enclosed CD.

Marvis, Dimitri N.↗

Exploration of Advanced Probabilistic and Stochastic Design Methods

The primary objective of the three year research effort was to explore advanced, non-deterministic aerospace system design methods that may have relevance to designers and analysts. The research pursued emerging areas in design methodology and leverage current fundamental research in the area of design decision-making, probabilistic modeling, and optimization. The specific focus of the three year investigation was oriented toward methods to identify and analyze emerging aircraft technologies in a consistent and complete manner, and to explore means to make optimal decisions based on this knowledge in a probabilistic environment. The research efforts were classified into two main areas. First, Task A of the grant has had the objective of conducting research into the relative merits of possible approaches that account for both multiple criteria and uncertainty in design decision-making. In particular, in the final year of research, the focus was on the comparison and contrasting between three methods researched. Specifically, these three are the Joint Probabilistic Decision-Making (JPDM) technique, Physical Programming, and Dempster-Shafer (D-S) theory. The next element of the research, as contained in Task B, was focused upon exploration of the Technology Identification, Evaluation, and Selection (TIES) methodology developed at ASDL, especially with regards to identification of research needs in the baseline method through implementation exercises. The end result of Task B was the documentation of the evolution of the method with time and a technology transfer to the sponsor regarding the method, such that an initial capability for execution could be obtained by the sponsor. Specifically, the results of year 3 efforts were the creation of a detailed tutorial for implementing the TIES method. Within the tutorial package, templates and detailed examples were created for learning and understanding the details of each step. For both research tasks, sample files and tutorials are attached in electronic form with the enclosed CD.

Mavris, Dimitri N.↗

A Hybrid Energy System Workflow for Energy Portfolio Optimization

This manuscript develops a workflow, driven by data analytics algorithms, to support the optimization of the economic performance of an Integrated Energy System. The goal is to determine the optimum mix of capacities from a set of different energy producers (e.g., nuclear, gas, wind and solar). A stochastic-based optimizer is employed, based on Gaussian Process Modeling, which requires numerous samples for its training. Each sample represents a time series describing the demand, load, or other operational and economic profiles for various types of energy producers. These samples are synthetically generated using a reduced order modeling algorithm that reads a limited set of historical data, such as demand and load data from past years. Numerous data analysis methods are employed to construct the reduced order models, including, for example, the Auto Regressive Moving Average, Fourier series decomposition, and the peak detection algorithm. All these algorithms are designed to detrend the data and extract features that can be employed to generate synthetic time histories that preserve the statistical properties of the original limited historical data. The optimization cost function is based on an economic model that assesses the effective cost of energy based on two figures of merit: the specific cash flow stream for each energy producer and the total Net Present Value. An initial guess for the optimal capacities is obtained using the screening curve method. The results of the Gaussian Process model-based optimization are assessed using an exhaustive Monte Carlo search, with the results indicating reasonable optimization results. The workflow has been implemented inside the Idaho National Laboratory’s Risk Analysis and Virtual Environment (RAVEN) framework. The main contribution of this study addresses several challenges in the current optimization methods of the energy portfolios in IES: First, the feasibility of generating the synthetic time series of the periodic peak data; Second, the computational burden of the conventional stochastic optimization of the energy portfolio, associated with the need for repeated executions of system models; Third, the inadequacies of previous studies in terms of the comparisons of the impact of the economic parameters. The proposed workflow can provide a scientifically defendable strategy to support decision-making in the electricity market and to help energy distributors develop a better understanding of the performance of integrated energy systems.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Operation Optimization using Reinforcement Learning with Integrated Artificial Reasoning Framework

In large and complex systems, operational decision-making requires a systematic analysis with a vast amount of data from both process parameters and component status monitoring. In this paper, we present an integrated artificial reasoning approach for system state transition models that can help operational decision-making with explainable and traceable reasoning. The integrated artificial reasoning framework is a physics-based approach of defining the system structure in a Bayesian network, so we leveraged it in a Markov decision process (MDP) for finding optimal operational solutions. In our proposed framework, the MDP is implemented on a dynamic Bayesian network (DBN), which represents causalities in a system. The multilevel flow modeling was utilized in order to extract these causalities in a more efficient and objective manner. Since multilevel flow modeling is based on the fundamental energy and mass conservation laws, the target system is decomposed into several mass, energy, and information structures, which serve as the basis for a DBN. The MDP consists of the processes of finding a solution for the Bellman equation, which can be derived from the conditional probability equations of the constructed DBN. System operators can capture stochastic system dynamics as multiple subsystem state transitions based on their physical relations and uncertainties coming from the component degradation process or random failures. We analyzed a simplified example system to illustrate finding an optimal operational policy with this approach.

99 GENERAL AND MISCELLANEOUS↗

Advancing process-based flood frequency analysis for assessing flood hazard and population flood exposure

Recent studies have showcased the use of process-based hydrological models with Stochastic Storm Transposition (SST) techniques to conduct Flood Frequency Analysis (FFA). This framework, referred hereby FFA-SST, has proved to be a robust strategy to estimate peak flows of specific annual exceedance probability (e.g., 100-year peak flow) that can reflect natural and anthropogenic disturbances, including changes in land use and meteorological patterns. With the objective of advancing the FFA-SST framework, this study presents for the first time the use of an Integrated Surface-Subsurface Hydrological Model (ISSHM) to conduct FFA-SST by extending the analysis from peak flow responses to flood extent, enabling a unique view and analysis of flood hazard and population flood exposure at the basin scale. As a proof-of-concept, we used the ISSHM, Advanced Terrestrial Simulator (Amanzi-ATS) model, and the SST model, RainyDay, to conduct FFA-SST by simulating the flood response to 5,000 annual synthetic storm events in a 2,227 $km^2$ Southeast Texas watershed. We demonstrate that ATS, without site-specific calibration, provides a robust process-based representation of peak flows, flood extent, streamflow, evapotranspiration, soil moisture content, and water storage changes. Our results and analyses, covering frequency curves up to a 500-year return period for peak flows, basin inundation fractions, and the number of people exposed to flooding, offer a unique perspective to analyze flood impacts across spatial scales. Overall, this study provides critical insights for flood risk management by extending the FFA-SST framework to include both flood hazard and population flood exposure analyses at the basin scale. Such an approach will empower stakeholders and disaster emergency agencies with a more comprehensive understanding of flood impacts across the entire basin domain, facilitating informed decision-making for flood risk assessment and management.

58 GEOSCIENCES↗

Quantifying Uncertainty of Deep Reinforcement Learning Based Decision Making for Operations and Maintenance of Nuclear Power Plant

This paper summarizes research that integrates condition monitoring and prognostics with decision making for nuclear power plant operations and maintenance. As part of this research, we have developed an online asset management tool to help reduce life-cycle maintenance and repair costs. Using the latest advancements in condition monitoring, supply chain analytics, and deep reinforcement learning, we have created a predictive maintenance tool that can optimize the maintenance and spare-part management of a repairable nuclear system. To demonstrate these methods, preliminary studies were conducted on a simple, representative maintenance system undergoing a stochastic degradation process that requires repairs or replacement to continue operation. Through Monte Carlo simulations, we were able to reduce maintenance spending by approximately 50% compared to optimized, time-based maintenance strategies. Not only does the decision maker reduce the average life-cycle costs, it also minimizes the chance of high cost scenarios, lowering the variance of the expected cost distributions, and reducing overall financial risk. Furthermore, this work also studies the ability of the decision maker to handle various levels of noise from observation uncertainty. By introducing uncertainty into the decision-making process, we have quantified the robustness and resiliency of the decision maker, as well as identified necessary levels of observability to demonstrate cost effectiveness.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Data-driven optimization of mixed-integer bi-level multi-follower integrated planning and scheduling problems under demand uncertainty

The coordination of interconnected elements across the different layers of the supply chain is essential for all industrial processes and the key to optimal decision-making. Yet, the modeling and optimization of such interdependent systems are still burdensome. Here we address the simultaneous modeling and optimization of medium-term planning and short-term scheduling problems under demand uncertainty using mixed-integer bi-level multi-follower programming and data-driven optimization. Bi-level multi-follower programs model the natural hierarchy between different layers of supply chain management holistically, while scenario analysis and data-driven optimization allow us to retrieve the guaranteed feasible solutions of the integrated formulation under various demand considerations. We address the data-driven optimization of this challenging class of problems using the DOMINO framework, which was initially developed to solve single-leader single-follower bi-level optimization problems to guaranteed feasibility. This framework is extended to solve single-leader multi-follower stochastic formulations and its performance is characterized by well-known single and multi-product process scheduling case studies. Through our data-driven algorithmic approach, we present guaranteed feasible solutions to linear and nonlinear mixed-integer bi-level formulations of simultaneous planning and scheduling problems and further characterize the effects of the scheduling level complexity on the solution performance, which spans over several hundred continuous and binary variables, and thousands of constraints.

42 ENGINEERING↗

Statistical Uncertainty of Inhalation Dose Coefficients in Consequence Management: Propagated Dose Uncertainty in ICRP 66 Human Respiratory Tract Model

Reference inhalation dose models rely on deterministic biokinetics and reference computational phantoms, limiting their applicability to the variability present in population-specific exposures encountered in emergency response scenarios. Here, this study introduces REDCAL, a Python-based computational framework developed to propagate uncertainty in inhalation dose coefficients using the International Commission on Radiological Protection (ICRP) Publication 66 Human Respiratory Tract Model. REDCAL integrates ICRP deposition and clearance models, systemic biokinetics, and governing physics principles, and leverages Sandia National Laboratories’ Dakota toolkit for uncertainty quantification via Latin Hypercube Sampling. REDCAL was validated against DCAL, with biokinetic retention results differing by less than 1% and effective dose coefficients by less than 2% across all tested radionuclides. Stochastic sampling introduced variability in dose coefficients, with geometric standard deviations (GSD) in committed effective dose coefficients (CEDC) ranging from 1.0 to 1.5, based on lognormal distribution fits. Analysis demonstrated that variations in the activity median aerodynamic diameter (AMAD) notably influenced the computed CEDC values. Smaller particles (<1 µm) increased doses by 20–30% due to deeper lung deposition and prolonged retention for alpha emitting radionuclides, such as 241 Am and 239 Pu. Radionuclides with fast clearance, such as 133 I, demonstrated a dose reduction exceeding 50%, as AMAD increased beyond 5 µm due to upper airway deposition and rapid mucociliary clearance. The greatest GSD among the radionuclides reported in this study was for 241 Am. In most cases, the largest GSDs in the CEDC were associated with larger particle sizes, an expected outcome, as ICRP Publication 66 defines GSD in particle size as a function of AMAD, resulting in an extended tail of the lognormal distribution. The findings support improved inhalation dose assessments and enhance consequence management strategies for the U.S. Federal Radiological Monitoring and Assessment Center by quantifying uncertainty in dose coefficients and strengthening decision-making for emergency response scenarios.

Biokinetic Modeling↗

AEOLUS: Advances in Experimental Design, Optimal Control, and Learning for Uncertain Complex Systems

Sustained advances in the mathematics of modeling and simulation have resulted in the capability today for routine simulation of a number of large scale complex DOE-relevant systems. As remarkable as this capability for solving the so-called forward problem is, it is typically only the first step-an inner loop within an outer loop that explores the simulation model's parameter space and decision space to characterize uncertainty in the model's predictions, learn unknown model parameters from data, design the most informative experiments, determine optimal control strategies, and create optimal designs. Broadly, what unifies all of these outer loop problems is that they are, in one form or another, optimization problems over parameter/control/design space that are constrained by complex uncertain models. To fully realize the power of scientific simulation as a basis for scientific discovery, technological innovation, and rational decision-making, it is imperative to move beyond simulation to tackle the outer loop of optimization for learning from data, experimental design, and control with complex uncertain models. When the models under consideration are large-scale and complex, and when the optimization variable and uncertain parameter spaces are high (or infinite) dimensional, this constitutes a grand challenge of the highest order, and is intractable with conventional methods. To overcome these challenges, the AEOLUS Center was established to develop a unified mathematical, computational, and statistical framework for (1) Learning predictive models from complex data via Bayesian inference and optimization, and (2) Optimizing experiments, processes, and designs using the resulting uncertain models. These problems are intractable with conventional methods, for several reasons: (1) The simulation problems that govern the inner loops of the optimization problems are expensive to execute (due to severe nonlinearity, heterogeneity, multiphysics/multiscale coupling); (2) The optimization variable and uncertain parameter spaces are high dimensional, often stemming from discretizations of infinite dimensional fields such as initial conditions, sources, or material properties. We argue that the key to overcoming these challenges is to develop new mathematical, computational, and statistical methods that exploit the structure of the Bayesian inference and optimization problems mediated by their underlying complex uncertain models. This structure includes the regularity, sparsity, geometry, low intrinsic dimensionality, and multifidelity nature of the maps from uncertain parameter/optimization variable spaces to the specific objectives targeted: Bayesian inference, optimal experimental design, and optimal control design. Black box methods developed as generic tools are incapable of exploiting this structure. To be successful, we must create, integrate, and cross-fertilize ideas across multiple areas of applied math--including approximation theory, Bayesian inference, data science, experimental design, information theory, machine learning, model reduction, optimal control theory, parallel algorithms, PDE-constrained optimization, randomized algorithms, stochastic optimization, and uncertainty quantification--all while exploiting the structure of the problems at hand. With this goal in mind, we have marshaled a team of leading authorities in these areas. While the methods we develop will be broadly applicable across a wide spectrum of DOE problems in which experiments inform models and the systems those models describe must be optimized under uncertainty, we have chosen a specific area, advanced manufacturing and materials, to drive our work. AMM is characterized by complex models across multiple scales, and is a rich source of challenging problems in inference, experimental design, and optimal control, requiring multifaceted and integrated advances in applied mathematics. As such, AMM serves as an excellent vehicle to motivate and demonstrate the advances in applied mathematics developed by our center.

97 MATHEMATICS AND COMPUTING↗

Decision-making based on Markov decision process in integrated artificial reasoning framework—Part I: Theory

This paper presents a decision-making framework based on an integrated artificial reasoning framework and Markov decision process (MDP). The integrated artificial reasoning framework provides a physics-based approach that converts system information into state transition models, and the analysis result will be represented by the transition probabilities that can be used with an MDP to find a traceable and explainable optimal pathway. A dynamic Bayesian network (DBN) is well suited for representing the structure of an MDP. The causality information among process variables (or among subsystems) is mathematically represented in a DBN by the conditional probabilities of the node’s states provided different probabilities of the parent node’s states. To define node states in a physically understandable manner, we used multilevel flow modeling (MFM). An MFM follows the fundamental energy and mass conservation laws and supports the selection of process variables that represent the system of interest so that causal relations among process variables are properly captured. An MFM-based DBN supports developing state transition models in an MDP to capture the effect of process variables of system having physical relations. The operators of the target system can capture stochastic system dynamics as multiple subsystem state transitions based on their physical relations and uncertainties coming from component degradation or random failures. We analyzed a simplified exemplary system to illustrate an optimal operational policy using the suggested approach.

Markov decision process↗

Resilience-based Recovery Scheduling of Transportation Network in Mixed Traffic Environment: A Deep-Ensemble-Assisted Active Learning Approach

Devising effective post-hazard recovery strategies is critical in enhancing the resilience of transportation networks (TNs). However, existing work does not consider the multiclass users’ travel behavior in network functionality quantification and the metaheuristic solution procedures often suffer from extensive computational burden due to the exploration need in large solution space and the expensive functionality quantification. Here, this study develops a bilevel decision-making framework for the resilience-based recovery scheduling of the TN in a mixed traffic environment with connected and autonomous vehicles (CAVs) and human-driven vehicles (HDVs). The lower level quantifies the TN's functionality over time considering different travel behavior of CAV and HDV users arisen from their different levels of information perception. The upper level presents a novel deep-ensemble-assisted active learning approach to balance optimization performance and computational cost. This framework can help decision makers better quantify the TN's functionality to support effective recovery scheduling of TN with different mixed traffic scenarios ranging from HDV-only to future CAV-dominant traffic. The optimization approach bears the potential to be extended to solving general large-scale network recovery scheduling problems effectively and efficiently. The proposed methodology is demonstrated using a real-world traffic network in Southern California under earthquake considering deterministic and stochastic repair durations.

42 ENGINEERING↗

Optimization problems governed by systems of PDEs with uncertainties

This paper reviews current theoretical and numerical approaches to optimization problems governed by partial differential equations (PDEs) that depend on random variables or random fields. Such problems arise in many engineering, science, economics and societal decision-making tasks. This paper focuses on problems in which the governing PDEs are parametrized by the random variables/fields, and the decisions are made at the beginning and are not revised once uncertainty is revealed. Examples of such problems are presented to motivate the topic of this paper, and to illustrate the impact of different ways to model uncertainty in the formulations of the optimization problem and their impact on the solution. A linear–quadratic elliptic optimal control problem is used to provide a detailed discussion of the set-up for the risk-neutral optimization problem formulation, study the existence and characterization of its solution, and survey numerical methods for computing it. Different ways to model uncertainty in the PDE-constrained optimization problem are surveyed in an abstract setting, including risk measures, distributionally robust optimization formulations, probabilistic functions and chance constraints, and stochastic orders. Furthermore, approximation-based optimization approaches and stochastic methods for the solution of the large-scale PDE-constrained optimization problems under uncertainty are described. Some possible future research directions are outlined.

Heinkenschloss, Matthias [Rice Univ., Houston, TX ↗

Markov Decision Process based Trajectory Planning for UAVs under Uncertain Wind Conditions

In this paper we propose a Markov Decision Process (MDP) algorithm for path-planning of Unmanned Aviation Vehicles (UAVs) under varying wind conditions. Solutions to path-planning for UAVs are becoming increasingly necessary as autonomous UAVs continue to enter commercial and government spaces. Path-planning is inherently challenging, as UAVs needs to account for dynamically changing flying conditions such as weather, obstacle or no-fly zones, degraded vehicle health and off-nominal battery power consumption. Machine learning methods such as Markov Decision Process (MDPs) have the potential to revolutionize how vehicles navigate in such uncertain environments. Previous papers have demonstrated the use of MDPs to optimize UAV path-planning for energy consumption under time-varying wind distribution. In this study, UAV trajectories from a pre-determined waypoint to target cell, will be computed on a 7X7 grid environment by optimizing parameters for mission assurance and safety limits in addition to the energy consumption, and operation time. The UAV navigates the grid by taking actions to move in either of the eight cardinal and intercardinal directions, under constant thrust profile. The next state of the UAV is calculated by considering its action, transition probability, obstacle cells and the wind speed magnitude and direction. Both constant and stochastic wind will be considered in this paper, the parameters being extracted from real wind measurements in proximity to an experimental UAV flight. One of the studies to be demonstrated in this paper is that as the unmanned airspace gets more complex with multiple vehicles and environmental uncertainties, trade-offs between energy consumption, operation time, risk tolerance, and mission assurance needs to be made. Further, MDPs are capable of fast computation of UAV trajectories under varying wind, hence making them suitable for in-flight path planners.

decision-making↗

Operation Optimization Using Reinforcement Learning with Integrated Artificial Reasoning Framework

In large and complex systems, operational decision-making requires a systematic analysis with a vast amount of data from both process parameters and component status monitoring. In this paper, we present an integrated artificial reasoning approach for system state transition models that can help operational decision-making with explainable and traceable reasoning. The integrated artificial reasoning framework is a physics-based approach of defining the system structure in a Bayesian network, so we leveraged it in a Markov decision process (MDP) for finding optimal operational solutions. In our proposed framework, the MDP is implemented on a dynamic Bayesian network (DBN), which represents causalities in a system. The multilevel flow modeling was utilized in order to extract these causalities in a more efficient and objective manner. Since multilevel flow modeling is based on the fundamental energy and mass conservation laws, the target system is decomposed into several mass, energy, and information structures, which serve as the basis for a DBN. The MDP consists of the processes of finding a solution for the Bellman equation, which can be derived from the conditional probability equations of the constructed DBN. System operators can capture stochastic system dynamics as multiple subsystem state transitions based on their physical relations and uncertainties coming from the component degradation process or random failures. We analyzed a simplified example system to illustrate finding an optimal operational policy with this approach.

Kim, Junyung↗