Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “conditional probability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Quasar Identification Using Multivariate Probability Density Estimated from Nonparametric Conditional Probabilities

Nonparametric estimation for a probability density function that describes multivariate data has typically been addressed by kernel density estimation (KDE). A novel density estimator recently developed by Farmer and Jacobs offers an alternative high-throughput automated approach to univariate nonparametric density estimation based on maximum entropy and order statistics, improving accuracy over univariate KDE. This article presents an extension of the single variable case to multiple variables. The univariate estimator is used to recursively calculate a product array of one-dimensional conditional probabilities. In combination with interpolation methods, a complete joint probability density estimate is generated for multiple variables. Good accuracy and speed performance in synthetic data are demonstrated by a numerical study using known distributions over a range of sample sizes from 100 to 10 6 for two to six variables. Performance in terms of speed and accuracy is compared to KDE. The multivariate density estimate developed here tends to perform better as the number of samples and/or variables increases. As an example application, measurements are analyzed over five filters of photometric data from the Sloan Digital Sky Survey Data Release 17. The multivariate estimation is used to form the basis for a binary classifier that distinguishes quasars from galaxies and stars with up to 94% accuracy.

79 ASTRONOMY AND ASTROPHYSICS↗

Correlation energy of the uniform electron gas determined by ground-state conditional probability density functional theory

Conditional-probability density functional theory (CP-DFT) is a formally exact method for finding correlation energies from Kohn-Sham DFT without evaluating an explicit energy functional. We present details on how to generate accurate exchange-correlation energies for the ground-state uniform gas. We also use the exchange hole in a CP antiparallel spin calculation to extract the high-density limit. We give a highly accurate analytic solution to the Thomas-Fermi model for this problem, showing its performance relative to Kohn-Sham and may be useful at high temperatures. We explore several approximations to the CP potential. Furthermore, results are compared to accurate parameterizations for both exchange-correlation energies and holes.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Conditional probability density functional theory

Here we present conditional probability (CP) density functional theory (DFT) as a formally exact theory. In essence, CP-DFT determines the ground-state energy of a system by finding the CP density from a series of independent Kohn-Sham (KS) DFT calculations. By directly calculating CP densities, we bypass the need for an approximate XC energy functional. In this paper we discuss and derive several key properties of the CP density and corresponding CP-KS potential. Illustrative examples are used throughout to help guide the reader through the various concepts and theory presented. We explore a suitable CP-DFT approximation and discuss exact conditions, limitations, and results for selected examples.

36 MATERIALS SCIENCE↗

Quantifying conditional probabilities of fish-turbine encounters and impacts

Tidal turbines are one source of marine renewable energy but development of tidal power is hampered by uncertainties in fish-turbine interaction impacts. Current knowledge gaps exist in efforts to quantify risks, as empirical data and modeling studies have characterized components of fish approach and interaction with turbines, but a comprehensive model that quantifies conditional occurrence probabilities of fish approaching and then interacting with a turbine in sequential steps is lacking. We combined empirical acoustic density measurements of Pacific herring ( Clupea pallasii ) and when data limited, published probabilities in an impact probability model that includes approach, entrainment, interactions, and avoidance of fish with axial or cross-flow tidal turbines. Interaction impacts include fish collisions with stationary turbine components, blade strikes by rotating blades, and/or a collision followed by a blade strike. Impact probabilities for collision followed by a blade strike were lowest with estimates ranging from 0.0000242 to 0.0678, and highest for blade strike ranging from 0.000261 to 0.40. Maximum probabilities occurred for a cross-flow turbine at night with no active or passive avoidance. Estimates were lowest when probabilities were conditional on sequential events, and when active and passive avoidance was included for an axial-flow turbine during the day. As expected, conditional probabilities were typically lower than analogous independent events and literature values. Estimating impact probabilities for Pacific herring in Admiralty Inlet, Washington, United States for two device types illustrates utilization of existing data and simultaneously identifies data gaps needed to fully calculate empirical-based probabilities for any site-species combination.

collision risk↗

Conditional Pseudo-Reversible Normalizing Flow for Surrogate Modeling in Quantifying Uncertainty Propagation

We introduce a conditional pseudo-reversible normalizing flow (PR-NF) that directly learns conditional probability distributions from noisy physical models to efficiently quantify both forward and inverse uncertainty propagation. Traditional surrogate modeling approaches approximate only the deterministic component of physical models, requiring separate noise characterization and computationally expensive sampling methods for inverse problems. Here, in this work, we develop the conditional PR-NF model to directly learn and efficiently generate samples from the conditional probability density functions (PDFs). The training process utilizes dataset consisting of input-output pairs without requiring prior knowledge about the noise and the function. Once trained, our model efficiently generates samples from conditional PDFs for any input within the training domain. Moreover, the pseudo-reversibility feature allows for the use of fully connected neural network architectures, which simplifies the implementation and enables theoretical analysis. We provide a rigorous convergence analysis of the conditional PR-NF model, showing its ability to converge to the target conditional PDF using the Kullback−Leibler divergence. To demonstrate the effectiveness of our method, we apply it to several benchmark tests and a real-world geologic carbon storage problem.

97 MATHEMATICS AND COMPUTING↗

Identifying Adversarial Cyber-Activity in Operational Technology Environments Using Bayesian Networks

Critical infrastructure and other operational technology (OT) environments face increasing cybersecurity risks from adversarial behavior. This paper describes the development of a risk model using a Bayesian network to enhance the comprehension of observable cyber events caused by malicious activity in OT environments. The core of the Bayesian network is a process model that describes the stages of adversary behavior. The remainder of the model is based on the MITRE ATT&CK® for Industrial Control Systems (ICS) taxonomy, which includes tactics and techniques that may be used by the adversary. The observables provide evidence for adversary behavior through the intermediary technique and tactic nodes. One challenge in constructing this model is a lack of open-source data from cyber-attacks on OT systems. This paper discusses learning from limited data, the elicitation of expert opinion to construct the conditional probability tables when data is scarce, and the refinement of the most difficult conditional probabilities tables using several forms of sensitivity analyses. Finally, the Bayesian network is demonstrated using two historical case studies: the DarkSide ransomware attack on the Colonial Pipeline and the destructive cyberattack targeting the ThyssenKrupp blast furnace. Index Terms—Cybersecurity, industrial control systems, operational technology

97 - MATHEMATICS AND COMPUTING↗

A New Framework for Interstellar Medium Emission Line Models: Connecting Multiscale Simulations across Cosmological Volumes

The James Webb Space Telescope (JWST) and Atacama Large Millimeter/submillimeter Array have detected emission lines from the ionized interstellar medium (ISM) in some of the first galaxies at z ≳ 6. These measurements present an opportunity to better understand galaxy assembly histories and may allow important tests of state-of-the-art galaxy formation simulations. It is challenging, however, to model these lines in their proper cosmological context. In order to meet this challenge, we introduce a novel subgrid line emission modeling framework. The framework uses the high-z zoom-in simulation suite from the Feedback in Realistic Environments (FIRE) collaboration. The line emission signals from H II regions within each simulated FIRE galaxy are modeled using the semianalytic HIIL INES code. A machine learning approach is then used to determine the conditional probability distribution for the line luminosity to stellar-mass ratio from the H II regions around each simulated stellar particle. This conditional probability distribution can then be applied to predict the line luminosities around stellar particles in lower-resolution, yet larger volume cosmological simulations. As an example, we apply this approach to the IllustrisTNG simulations at z = 6. The resulting predictions for the [O II ], [O III ], and Balmer line luminosities as a function of star formation rate agree well with current observations. Our predictions differ, however, from related works in the literature, which lack detailed subgrid ISM models. This highlights the importance of our multiscale simulation modeling framework. Finally, we provide forecasts for future line luminosity function measurements from the JWST and quantify the cosmic variance in such surveys.

(ISM:) H II regions↗

Operation Optimization using Reinforcement Learning with Integrated Artificial Reasoning Framework

In large and complex systems, operational decision-making requires a systematic analysis with a vast amount of data from both process parameters and component status monitoring. In this paper, we present an integrated artificial reasoning approach for system state transition models that can help operational decision-making with explainable and traceable reasoning. The integrated artificial reasoning framework is a physics-based approach of defining the system structure in a Bayesian network, so we leveraged it in a Markov decision process (MDP) for finding optimal operational solutions. In our proposed framework, the MDP is implemented on a dynamic Bayesian network (DBN), which represents causalities in a system. The multilevel flow modeling was utilized in order to extract these causalities in a more efficient and objective manner. Since multilevel flow modeling is based on the fundamental energy and mass conservation laws, the target system is decomposed into several mass, energy, and information structures, which serve as the basis for a DBN. The MDP consists of the processes of finding a solution for the Bellman equation, which can be derived from the conditional probability equations of the constructed DBN. System operators can capture stochastic system dynamics as multiple subsystem state transitions based on their physical relations and uncertainties coming from the component degradation process or random failures. We analyzed a simplified example system to illustrate finding an optimal operational policy with this approach.

99 GENERAL AND MISCELLANEOUS↗

Path integrals, complex probabilities and the discrete Weyl representation

Abstract A discrete formulation of the real-time path integral as the expectation value of a functional of paths with respect to a complex probability on a sample space of discrete valued paths is explored. The formulation in terms of complex probabilities is motivated by a recent reinterpretation of the real-time path integral as the expectation value of a potential functional with respect to a complex probability distribution on cylinder sets of paths. The discrete formulation in this work is based on a discrete version of the Weyl algebra that can be applied to any observable with a finite number of outcomes. The origin of the complex probability in this work is the completeness relation. In the discrete formulation the complex probability exactly factors into products of conditional probabilities and exact unitarity is maintained at each level of approximation. The approximation of infinite dimensional quantum systems by discrete systems is discussed. The method is illustrated by applying it to scattering theory and quantum field theory. The implications of these applications for quantum computing is discussed.

Physics↗

Graph-Augmented Normalizing Flows for Anomaly Detection of Multiple Time Series

Anomaly detection is a widely studied task for a broad variety of data types; among them, multiple time series appear frequently in applications, including for example, power grids and traffic networks. Detecting anomalies for multiple time series, however, is a challenging subject, owing to the intricate interdependencies among the constituent series. We hypothesize that anomalies occur in low density regions of a distribution and explore the use of normalizing flows for unsupervised anomaly detection, because of their superior quality in density estimation. Moreover, we propose a novel flow model by imposing a Bayesian network among constituent series. A Bayesian network is a directed acyclic graph (DAG) that models causal relationships; it factorizes the joint probability of the series into the product of easy-to-evaluate conditional probabilities. We call such a graph-augmented normalizing flow approach GANF and propose joint estimation of the DAG with flow parameters. We conduct extensive experiments on real-world datasets and demonstrate the effectiveness of GANF for density estimation, anomaly detection, and identification of time series distribution drift.

Dai, Enyan↗

Probabilistic locked mode predictor in the presence of a resistive wall and finite island saturation in tokamaks

We present a framework for estimating the probability of locking to an error field in a rotating tokamak plasma. This leverages machine learning methods trained on data from a mode-locking model, including an error field, resistive magnetohydrodynamics modeling of the plasma, a resistive wall, and an external vacuum region, leading to a fifth-order ordinary differential equation (ODE) system. It is an extension of the model without a resistive wall introduced by Akçay et al. [Phys. Plasmas 28, 082106 (2021)]. Tearing mode saturation by a finite island width is also modeled. We vary three pairs of control parameters in our studies: the momentum source plus either the error field, the tearing stability index, or the island saturation term. The order parameters are the time-asymptotic values of the five ODE variables. Normalization of them reduces the system to 2D and facilitates the classification into locked (L) or unlocked (U) states, as illustrated by Akçay et al., [Phys. Plasmas 28, 082106 (2021)]. This classification splits the control space into three regions: L̂, with only L states; Û, with only U states; and a hysteresis (hysteretic) region Ĥ, with both L and U states. In regions L̂ and Û, the cubic equation of torque balance yields one real root. Region Ĥ has three roots, allowing bifurcations between the L and U states. The classification of the ODE solutions into L/U is used to estimate the locking probability, conditional on the pair of the control parameters, using a neural network. We also explore estimating the locking probability for a sparse dataset, using a transfer learning method based on a dense model dataset.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Generative unfolding with distribution mapping

Machine learning enables unbinned, highly-differential cross section measurements. A recent idea uses generative models to morph a starting simulation into the unfolded data. We show how to extend two morphing techniques, Schrödinger Bridges and Direct Diffusion, in order to ensure that the models learn the correct conditional probabilities. This brings distribution mapping (DM) to a similar level of accuracy as the state-of-the-art conditional generative unfolding methods. Numerical results are presented with a standard benchmark dataset of single jet substructure as well as for a new dataset describing a 22-dimensional phase space of Z+2 -jets.

Butter, Anja↗

Accelerating template generation in resonant anomaly detection searches with optimal transport

We introduce Resonant Anomaly Detection with Optimal Transport (RAD-OT), a method for generating signal templates in resonant anomaly detection searches. RAD-OT leverages the fact that the samples from the conditional probability density of the target features vary approximately linearly along the optimal transport path connecting the resonant feature. This does not assume that the conditional density itself is linear with the resonant feature, allowing RAD-OT to efficiently capture multimodal relationships, changes in resolution, etc. By solving the optimal transport problem, RAD-OT can quickly build a template by interpolating between the background distributions in two sideband regions. We demonstrate the performance of RAD-OT using the LHC Olympics R&D dataset, where we find comparable sensitivity and improved stability with respect to deep learning-based approaches.

Automation↗

Geostatistical Mapping of Salinity Conditioned on Borehole Logs, Montebello Oil Field, California

We present a geostatistics-based stochastic salinity estimation framework for the Montebello Oil Field that capitalizes on available total dissolved solids (TDS) data from groundwater samples as well as electrical resistivity (ER) data from borehole logging. Data from TDS samples (n = 4924) was coded into an indicator framework based on falling below four selected thresholds (500, 1000, 3000, and 10,000 mg/L). Collocated TDS-ER data from the surrounding groundwater basin were then employed to produce a kernel density estimator to establish conditional probabilities for ER data (n = 8 boreholes) falling below the selected TDS thresholds within the Montebello Oil Field area. Directional variograms were estimated from these indicator coded data, and 500 TDS realizations from conditional indicator simulation were generated for the subsurface region above the Montebello Oil Field reservoir. Simulations were summarized as 3D maps of median TDS, most likely salinity class, and probability for exceeding each of the specified TDS thresholds. Results suggested TDS was below 500 mg/L in most of the study area, with a trend toward higher values (500 to 1000 mg/L) to the southwest; consistent with the average regional groundwater flow direction. Discrete localized zones of TDS greater than 1000 mg/L were observed, with one of these zones in the greater than 10,000 mg/L range; however, these areas were not prevalent. The probabilistic approach used here is adaptable and is readily modified to include additional data and types and can be employed in time-lapse salinity modeling through Bayesian updating.

54 ENVIRONMENTAL SCIENCES↗

Multisource Data Fusion Outage Location in Distribution Systems via Probabilistic Graphical Models

Efficient outage location is critical to enhancing the resilience of power distribution systems. However, accurate outage location requires combining massive evidence received from diverse data sources, including smart meter (SM) last gasp signals, customer trouble calls, social media messages, weather data, vegetation information, and physical parameters of the network. This is a computationally complex task due to the high dimensionality of data in distribution grids. In this paper, we propose a multi-source data fusion approach to locate outage events in partially observable distribution systems using Bayesian networks (BNs). A novel aspect of the proposed approach is that it takes multi-source evidence and the complex structure of distribution systems into account using a probabilistic graphical method. Our method can radically reduce the computational complexity of outage location inference in high-dimensional spaces. The graphical structure of the proposed BN is established based on the network’s topology and the causal relationship between random variables, such as the states of branches/customers and evidence. Utilizing this graphical model, accurate outage locations are obtained by leveraging a Gibbs sampling (GS) method, to infer the probabilities of de-energization for all branches. Compared with commonly-used exact inference methods that have exponential complexity in the size of the BN, GS quantifies the target conditional probability distributions in a timely manner. As a result, a case study of several real-world distribution systems is presented to validate the proposed method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Artificial Neural Networks for In-Cycle Prediction of Knock Events

Downsized turbocharged engines have been increasingly popular in modern light-duty vehicles due to their fuel efficiency benefits. However, high power density in such engines is achieved thanks to high in-cylinder pressure and temperature conditions that increase knock propensity. Next-cycle control has been studied as a method to reduce the damaging effects of knock by operating the engine in a low knock probability condition. This exploratory study looks at the feasibility of in-cycle knock prediction as a tool for advanced knock control algorithms. A methodology is proposed to 1) choose in-cycle features of the pressure trace that highly correlate with knock events and 2) train artificial neural networks to predict in-cycle knock events before knock onset. The methodology was validated at different operating conditions and different levels of generalization. Precision and recall were used as metrics to evaluate the binary classifier. However, the Fowlkes-Mallows (FM) index was used to compare the result of the clustering algorithm at different operating conditions. The results showed a maximum FM index of 0.7 when the prediction was done at knock onset and a minimum FM index of 0.45 when the prediction was done at spark timing.

42 ENGINEERING↗

Prioritizing Uncertainties in Hydrogen Contribution to Risk in Post-Crash Outcomes for Rail

This report presents analysis from Sandia National Laboratories predicting contributions to risk associated with the use of hydrogen technology for rail. Event sequence diagrams are used to describe possible accident scenarios and progressions. Initiating event frequencies and branch event probabilities for each scenario are quantified with uncertainty using distributions fit to Federal Railroad Administration and U.S. Department of Transportation Pipeline and Hazardous Materials Safety Administration data on applicable accidents from 2000 to 2020. Uncertainty is propagated through the event sequence diagram to estimate the frequency and conditional probability of accident end states. The analysis identifies four scenarios with significant contributions to risk from hydrogen that are predicted to occur relatively frequently, which may inform priorities for reducing uncertainty. These scenarios are 1) overpressure events resulting from collisions with hydrogen release due to mechanical damage and delayed ignition, 2) jet fire events resulting from collisions with hydrogen release due to mechanical damage and immediate ignition, 3) jet fires resulting from fire or explosion initiating events involving the hydrogen tank and correct operation of the thermally-activated pressure relief device (TPRD) subsequent to the thermal insult, and 4) pressure burst resulting from fire or explosion initiating events involving the hydrogen tank and failure of the TPRD. Delayed and immediate hydrogen ignition probabilities are identified as being highly uncertain and potential candidates for reducing conservatism in the predicted frequencies for these two scenarios.

08 HYDROGEN↗

Decision-making based on Markov decision process in integrated artificial reasoning framework—Part I: Theory

This paper presents a decision-making framework based on an integrated artificial reasoning framework and Markov decision process (MDP). The integrated artificial reasoning framework provides a physics-based approach that converts system information into state transition models, and the analysis result will be represented by the transition probabilities that can be used with an MDP to find a traceable and explainable optimal pathway. A dynamic Bayesian network (DBN) is well suited for representing the structure of an MDP. The causality information among process variables (or among subsystems) is mathematically represented in a DBN by the conditional probabilities of the node’s states provided different probabilities of the parent node’s states. To define node states in a physically understandable manner, we used multilevel flow modeling (MFM). An MFM follows the fundamental energy and mass conservation laws and supports the selection of process variables that represent the system of interest so that causal relations among process variables are properly captured. An MFM-based DBN supports developing state transition models in an MDP to capture the effect of process variables of system having physical relations. The operators of the target system can capture stochastic system dynamics as multiple subsystem state transitions based on their physical relations and uncertainties coming from component degradation or random failures. We analyzed a simplified exemplary system to illustrate an optimal operational policy using the suggested approach.

Markov decision process↗