Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Probabilistic Graphical Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A probabilistic graphical model foundation for enabling predictive digital twins at scale

A unifying mathematical formulation is needed to move from one-off digital twins built through custom implementations to robust digital twin implementations at scale. This work proposes a probabilistic graphical model as a formal mathematical representation of a digital twin and its associated physical asset. We create an abstraction of the asset–twin system as a set of coupled dynamical systems, evolving over time through their respective state spaces and interacting via observed data and control inputs. The formal definition of this coupled system as a probabilistic graphical model enables us to draw upon well-established theory and methods from Bayesian statistics, dynamical systems and control theory. The declarative and general nature of the proposed digital twin model make it rigorous yet flexible, enabling its application at scale in a diverse range of application areas. Here, we demonstrate how the model is instantiated to enable a structural digital twin of an unmanned aerial vehicle (UAV). The digital twin is calibrated using experimental data from a physical UAV asset. Its use in dynamic decision-making is then illustrated in a synthetic example where the UAV undergoes an in-flight damage event and the digital twin is dynamically updated using sensor data. The graphical model foundation ensures that the digital twin calibration and updating process is principled, unified and able to scale to an entire fleet of digital twins.

42 ENGINEERING↗

Active learning of chemical reaction networks via probabilistic graphical models and Boolean reaction circuits

Discerning networks of many reactions among multiple interconverting species is challenging. Here, we present a reaction network identification methodology. Our methodology enumerates all stoichiometrically and chemically feasible reactions and requires statistical evidence from effluent concentrations for the inclusion or exclusion of each from the reaction network, contrasting with the commonly seen incremental approach and other work of relying heavily upon chemical intuition and assuming the reactions occurring. Using graph theory alongside an active learning design of experiments that propose maximally informative feeds, we identify the underlying reaction network with minimal laboratory runs. Here, we introduce chemistry-probabilistic graphical modeling and Boolean reaction circuits to statistically quantify which reactions occur from effluent concentrations. Our methodology accurately discerns active reactions, as showcased upon a laboratory network of cross-ketonization of furoic and lauric acid and validated upon simulated networks of thermal and CO 2 -assisted ethane dehydrogenation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multisource Data Fusion Outage Location in Distribution Systems via Probabilistic Graphical Models

Efficient outage location is critical to enhancing the resilience of power distribution systems. However, accurate outage location requires combining massive evidence received from diverse data sources, including smart meter (SM) last gasp signals, customer trouble calls, social media messages, weather data, vegetation information, and physical parameters of the network. This is a computationally complex task due to the high dimensionality of data in distribution grids. In this paper, we propose a multi-source data fusion approach to locate outage events in partially observable distribution systems using Bayesian networks (BNs). A novel aspect of the proposed approach is that it takes multi-source evidence and the complex structure of distribution systems into account using a probabilistic graphical method. Our method can radically reduce the computational complexity of outage location inference in high-dimensional spaces. The graphical structure of the proposed BN is established based on the network’s topology and the causal relationship between random variables, such as the states of branches/customers and evidence. Utilizing this graphical model, accurate outage locations are obtained by leveraging a Gibbs sampling (GS) method, to infer the probabilities of de-energization for all branches. Compared with commonly-used exact inference methods that have exponential complexity in the size of the BN, GS quantifies the target conditional probability distributions in a timely manner. As a result, a case study of several real-world distribution systems is presented to validate the proposed method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

ML for microbiomes

The software provides machine learning analysis and visualization to detect patterns in microbiome data, including topic modeling, probabilistic graphical modeling, conventional machine learning methods, and deep learning. The software is written in python and R, it uses some python and R libraries as well as big open-source libraries like sklearn, networkX, pytorch (python), pgmpy (python), and bnlearn (R). It also has a script to use for MALLET and DTM (open-source packages for topic modeling, written in Java).

Kim, Anastasiia↗

Conin

SAND2025-07645O Conin is a Python library that supports constrained analysis of probabilistic graphical models (PGMs). It enables constrained inference and learning for hidden Markov models, Bayesian networks, dynamic Bayesian networks, and Markov networks. Conin interfaces with the pgmpy library to specify general probabilistic graphical models with a variety of optimization solvers to support learning and inference. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Hart, William [Sandia National Lab. (SNL-CA), Live↗

Graph-learning approach to combine multiresolution seismic velocity models

SUMMARY The resolution of velocity models obtained by tomography varies due to multiple factors and variables, such as the inversion approach, ray coverage, data quality, etc. Combining velocity models with different resolutions can enable more accurate ground motion simulations. Toward this goal, we present a novel methodology to fuse multiresolution seismic velocity maps with probabilistic graphical models (PGMs). The PGMs provide segmentation results, corresponding to various velocity intervals, in seismic velocity models with different resolutions. Further, by considering physical information (such as ray path density), we introduce physics-informed probabilistic graphical models (PIPGMs). These models provide data-driven relations between subdomains with low (LR) and high (HR) resolutions. Transferring (segmented) distribution information from the HR regions enhances the details in the LR regions by solving a maximum likelihood problem with prior knowledge from HR models. When updating areas bordering HR and LR regions, a patch-scanning policy is adopted to consider local patterns and avoid sharp boundaries. To evaluate the efficacy of the proposed PGM fusion method, we tested the fusion approach on both a synthetic checkerboard model and a fault zone structure imaged from the 2019 Ridgecrest, CA, earthquake sequence. The Ridgecrest fault zone image consists of a shallow (top 1 km) high-resolution shear-wave velocity model obtained from ambient noise tomography, which is embedded into the coarser Statewide California Earthquake Center Community Velocity Model version S4.26-M01. The model efficacy is underscored by the deviation between observed and calculated traveltimes along the boundaries between HR and LR regions, 38 per cent less than obtained by conventional Gaussian interpolation. The proposed PGM fusion method can merge any gridded multiresolution velocity model, a valuable tool for computational seismology and ground motion estimation.

Geochemistry & Geophysics↗

GNET2: an R package for constructing gene regulatory networks from transcriptomic data

Abstract Motivation The Gene Network Estimation Tool (GNET) is designed to build gene regulatory networks (GRNs) from transcriptomic gene expression data with a probabilistic graphical model. The data preprocessing, model construction and visualization modules of the original GNET software were developed on different programming platforms, which were inconvenient for users to deploy and use. Results Here, we present GNET2, an improved implementation of GNET as an integrated R package. GNET2 provides more flexibility for parameter initialization and regulatory module construction based on the core iterative modeling process of the original algorithm. The data exchange interface of GNET2 is handled within an R session automatically. Given the growing demand for regulatory network reconstruction from transcriptomic data, GNET2 offers a convenient option for GRN inference on large datasets. Availability and implementation The source code of GNET2 is available at https://github.com/jianlin-cheng/GNET2. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Risk-Informed Condition Evaluation of Solar-centered Energy Generation and Distribution Networks through Bayesian Learning and Inference

We develop a methodology based on Bayesian inference over Probabilistic Graphical Models (PGMs) to understand and quantify risk in solar-centered grids using targeted measurements and learned system behavior. Being non-prescriptive but, rather, able to infer system behavior and, ultimately, address risk queries from data, our machine learning-type paradigm is tailored for diverse topologies and threat scenarios often associated with distributed energy generation and photovoltaic distributed energy resources (PV-DERs) in particular. We describe algorithmic processes for: (i) learning the structure of PGMs that result from attack-prone PV-DER-proliferated distribution systems, (ii) quantifying cause-effect relationships, and (iii) evaluating risk queries based on diverse evidence. The contributions are illustrated on a residential grid subject to output impairment attacks on its PV-DER infrastructure.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Explainable and trustworthy artificial intelligence for correctable modeling in chemical sciences

Data science has primarily focused on big data, but for many physics, chemistry, and engineering applications, data are often small, correlated and, thus, low dimensional, and sourced from both computations and experiments with various levels of noise. Typical statistics and machine learning methods do not work for these cases. Expert knowledge is essential, but a systematic framework for incorporating it into physics-based models under uncertainty is lacking. Here, we develop a mathematical and computational framework for probabilistic artificial intelligence (AI)–based predictive modeling combining data, expert knowledge, multiscale models, and information theory through uncertainty quantification and probabilistic graphical models (PGMs). We apply PGMs to chemistry specifically and develop predictive guarantees for PGMs generally. Our proposed framework, combining AI and uncertainty quantification, provides explainable results leading to correctable and, eventually, trustworthy models. The proposed framework is demonstrated on a microkinetic model of the oxygen reduction reaction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Quantum Markov chain Monte Carlo with digital dissipative dynamics on quantum computers

Modeling the dynamics of a quantum system connected to the environment is critical for advancing our understanding of complex quantum processes, as most quantum processes in nature are affected by an environment. Modeling a macroscopic environment on a quantum simulator may be achieved by coupling independent ancilla qubits that facilitate energy exchange in an appropriate manner with the system and mimic an environment. This approach requires a large, and possibly exponential number of ancillary degrees of freedom which is impractical. In contrast, we develop a digital quantum algorithm that simulates interaction with an environment using a small number of ancilla qubits. By combining periodic modulation of the ancilla energies, or spectral combing, with periodic reset operations, we are able to mimic interaction with a large environment and generate thermal states of interacting many-body systems. We evaluate the algorithm by simulating preparation of thermal states of the transverse Ising model. Our algorithm can also be viewed as a quantum Markov chain Monte Carlo process that allows sampling of the Gibbs distribution of a multivariate model. To illustrate this we evaluate the accuracy of sampling Gibbs distributions of simple probabilistic graphical models using the algorithm.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Artificial Reasoning System for Symptom-Based Conditional Failure Probability Estimation Using Bayesian Network

Advances in nuclear power technologies require enhanced capabilities for operator advice and autonomous control. One of the first tasks in the development of such capabilities is the formulation of symptom-based conditional failure probabilities for structures, systems, and components (SSCs) of interest, for which the primary goal is to aid plant personnel in deducing the probabilistic performance status of the monitored SSCs and in detecting impending faults/failure. The task of conditional failure probability estimation is a bidirectional inference problem and shall be logically tackled by the Bayesian network (BN) approach. As a knowledge-based artificial intelligence tool and a probabilistic graphical model, BN offers the capability of reasoning under uncertainty and graphical representation emulating the physical behavior of the target SSC. This paper provides a systematic overview of the BN technique and the software tools for handling implementation of BN models, along with the associated knowledge representation and reasoning paradigm. Both operational data and expert judgement can be readily incorporated into the knowledge base of a BN model. The challenges with data availability are highlighted, and the general approach to target SSC identification is presented. Our focus is upon failure-prone and risk-important balance of plant assets, especially cases having strong operator involvement. An exemplary case study on the failure of a motor-driven centrifugal pump is also conducted to demonstrate the usefulness and technical feasibility of the proposed artificial reasoning system using an expert system shell.

Zhao, Xingang↗

How to Obtain the Redshift Distribution from Probabilistic Redshift Estimates

Abstract A reliable estimate of the redshift distribution n ( z ) is crucial for using weak gravitational lensing and large-scale structures of galaxy catalogs to study cosmology. Spectroscopic redshifts for the dim and numerous galaxies of next-generation weak-lensing surveys are expected to be unavailable, making photometric redshift (photo- z ) probability density functions (PDFs) the next best alternative for comprehensively encapsulating the nontrivial systematics affecting photo- z point estimation. The established stacked estimator of n ( z ) avoids reducing photo- z PDFs to point estimates but yields a systematically biased estimate of n ( z ) that worsens with a decreasing signal-to-noise ratio, the very regime where photo- z PDFs are most necessary. We introduce Cosmological Hierarchical Inference with Probabilistic Photometric Redshifts ( CHIPPR ), a statistically rigorous probabilistic graphical model of redshift-dependent photometry that correctly propagates the redshift uncertainty information beyond the best-fit estimator of n ( z ) produced by traditional procedures and is provably the only self-consistent way to recover n ( z ) from photo- z PDFs. We present the chippr prototype code, noting that the mathematically justifiable approach incurs computational cost. The CHIPPR approach is applicable to any one-point statistic of any random variable, provided the prior probability density used to produce the posteriors is explicitly known; if the prior is implicit, as may be the case for popular photo- z techniques, then the resulting posterior PDFs cannot be used for scientific inference. We therefore recommend that the photo- z community focus on developing methodologies that enable the recovery of photo- z likelihoods with support over all redshifts, either directly or via a known prior probability density.

79 ASTRONOMY AND ASTROPHYSICS↗

A Graphical Model for Fusing Diverse Microbiome Data

This paper develops a Bayesian graphical model for fusing disparate types of count data. The motivating application is the study of bacterial communities from diverse high-dimensional features, in this case, transcripts, collected from different treatments. In such datasets, there are no explicit correspondences between the communities and each corresponds to different factors, making data fusion challenging. We introduce a flexible multinomial-Gaussian generative model for jointly modeling such count data. This latent variable model jointly characterizes the observed data through a common multivariate Gaussian latent space that parameterizes the set of multinomial probabilities of the transcriptome counts. The covariance matrix of the latent variables induces a covariance matrix of co-dependencies between all the transcripts, effectively fusing multiple data sources. We present a computationally scalable variational Expectation-Maximization (EM) algorithm for inferring the latent variables and the parameters of the model. Here, the inferred latent variables provide a common dimensionality reduction for visualizing the data and the inferred parameters provide a predictive posterior distribution. In addition to simulation studies that demonstrate the variational EM procedure, we apply our model to a bacterial microbiome dataset.

59 BASIC BIOLOGICAL SCIENCES↗

Learning from Crowds by Modeling Common Confusions

Crowdsourcing provides a practical way to obtain large amounts of labeled data at a low cost. However, the annotation quality of annotators varies considerably, which imposes new challenges in learning a high-quality model from the crowdsourced annotations. In this work, we provide a new perspective to decompose annotation noise into common noise and individual noise and differentiate the source of confusion based on instance difficulty and annotator expertise on a per-instance-annotator basis. We realize this new crowdsourcing model by an end-to-end learning solution with two types of noise adaptation layers: one is shared across annotators to capture their commonly shared confusions, and the other one is pertaining to each annotator to realize individual confusion. To recognize the source of noise in each annotation, we use an auxiliary network to choose from the two noise adaptation layers with respect to both instances and annotators. Extensive experiments on both synthesized and real-world benchmarks demonstrate the effectiveness of our proposed common noise adaptation solution.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

235-F GoldSim Fate and Transport Model: Uncertainty Quantification

Building 235-F was configured with two missions in mind: Actinide Billet Line (ABL) and the fabrication of Pu-238 oxide for space program applications. ABL produced Np-237 billets for use in SRS reactors, whereas the design process, fabrication, and examination of Pu-238 oxide powder occurred within the following areas, respectively: Plutonium Experimental Facility (PEF), Plutonium Fuel Form (PuFF), and Old Metallography Lab (OML). By 1990 production ceased and by 2006 de-inventory occurred; however, assays have shown significant holdup remains within ABL and PuFF. As a result, 235-F is a Category 2 nuclear facility, with plans to undergo deactivation and decommission (D and D) via In-Situ Disposal (ISD). The purpose of this project is to ensure United States Environmental Protection Agency (USEPA) groundwater radiation maximum contaminant level (MCL) and dosage standards are met during the D and D of 235-F by quantifying uncertainty through probabilistic modeling and evaluation of various ISD alternatives. GoldSim is a dynamic modeling software package with a graphical, object-oriented interface capable of capturing the influence of complex system input variability on probabilistic system outcomes. A GoldSim stochastic fate and transport model for 235-F was developed and matched with a PORFLOW deterministic model to simulate probabilistic release and flow of radionuclides from ABL and PuFF into the vadose zone, the Upper Three Runs (UTR) Aquifer, and UTR Creek. The GoldSim model was used to probabilistically evaluate four ISD scenarios against USEPA groundwater radiation MCLs and dosage standards. The deterministic 235-F GoldSim fate and transport model continues to be refined to match the results of the PORFLOW deterministic model to ensure the model accurately represents radionuclide movement through the groundwater system. The stochastic variables that are utilized within the GoldSim model are founded on the most current data; a conservative perspective is taken where needed. Alignment with the PORFLOW deterministic model, coupled with input stochastic variability, allows the probabilistic 235-F GoldSim model to capture the conservative breadth of possible outcomes for radionuclide fate within this particular system. This ensures that the USEPA MCLs and dosage limits hold even in the worst case scenarios.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Scalable probabilistic estimates of electric vehicle charging given observed driver behavior

To prepare for rapid growth in global electric vehicle adoption, grid and policy planners depend on detailed forecasts of future charging demand. In this paper we propose a novel holistic, scalable, probabilistic framework to produce large-scale estimates of electric vehicle charging load for long-term planning that capture real drivers’ charging patterns. Our framework captures the uncertainty and stochasticity in charging demand by taking a graphical modeling approach. It has three core elements: driver groups, charging segment choices, and charging session time and energy requirements. The framework uses hierarchical clustering to group drivers by their charging histories, capturing their heterogeneous behaviors and preferences across different segments or types of charging. The framework uses probabilistic mixture models for each driver group’s sessions to identify the unique charging behaviors observed within each segment. We illustrate its application with a large data set from California, profiling the charging patterns and unique driver clusters it identifies. Using the model knobs representing drivers’ battery capacities, behavior, and segment access we present scenarios for California’s charging demand in 2030 with 8 million passenger electric vehicles. Peak charging demand ranged from 3.3 to 8.7 GW across scenarios. Furthermore, each was calculated in under 45 s on a laptop computer.

33 ADVANCED PROPULSION SYSTEMS↗

User’s Manual for RESRAD-BUILD Code V.4: Vol. 1 – Methodology and Models Used in RESRAD-BUILD Code

The RESRAD-BUILD computer code models radionuclide release and transport in indoor environments and performs pathway analyses to evaluate the potential radiological dose and risk incurred by an individual who works or lives in a building contaminated with radioactive material or housing radioactively contaminated furniture or equipment. The code provides four geometries to characterize a radiation source: point, line, area, and volume, in which radionuclides are homogeneously distributed. Radionuclides contained in a source are considered to be released to the indoor air due to various processes including erosion (mechanically or weathering), diffusion (for tritium and radon in a volume source), or emanation (radon in a point, line, or area source). The release can proceed through different time phases with different rates. In RESRAD-BUILD Version 4.0, a dynamic ventilation model is implemented to simulate the fate and transport of source material particles and radionuclides after their releases. This dynamic ventilation model considers (1) air exchange between rooms in the building and between the rooms and the outdoor environment, (2) deposition from air to floor, (3) resuspension from the floor to the air, and (4) periodical vacuuming that reduces the floor deposition. The fate and transport modeling provides estimates of radionuclide concentrations in the source, in the air, and on the floor at different times, which are then integrated over the exposure duration for the calculation of radiation doses and cancer risks. A single run of the RESRAD-BUILD code can model a building with up to 9 rooms, 10 sources, and 10 receptors. The potential radiation dose and cancer risk incurred by each receptor are calculated for seven exposure pathways: (1) external radiation directly from the sources (accounting for shielding), (2) external radiation from radioactive particles deposited on the floors, (3) external radiation from airborne radionuclides, (4) inhalation of airborne radionuclides, (5) inhalation of radon and radon progenies, (6) inadvertent ingestion of radioactive particles directly from the source, and (7) ingestion of radioactive particles deposited on the floors. Various exposure scenarios can be modeled with RESRAD-BUILD, including but are not limited to, office worker, renovation worker, decontamination worker, building visitor, and resident. Both deterministic and probabilistic analyses can be performed to obtain results in both text reports and graphic displays.

61 RADIATION PROTECTION AND DOSIMETRY↗

GDSA framework, a computational framework for complex modeling problems in radioactive waste management

This paper details a computational framework to produce automated, graphical workflows, and how this framework can be deployed to support complex modeling problems like those in nuclear engineering. Key benefits of the framework include: automating previously manual workflows; intuitive construction and communication of workflows through a graphical interface; and automated file transfer and handling for workflows deployed across heterogeneous computing resources. This paper demonstrates the framework's application to probabilistic post-closure performance assessment of systems for deep geologic disposal of nuclear waste. However, the framework is a general capability that can help users running a variety of computational studies.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗