Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graph analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

PlasmoData.jl — A Julia framework for modeling and analyzing complex data as graphs

Datasets encountered in scientific and engineering applications appear in complex formats (e.g., images, multivariate time series, molecules, video, text strings, networks). Graph theory provides a unifying framework to model such datasets and enables the use of powerful tools that can help analyze, visualize, and extract value from data. In this work, we present PlasmoData.jl, an open-source, Julia framework that uses concepts of graph theory to facilitate the modeling and analysis of complex datasets. The core of our framework is a general data modeling abstraction, which we call a DataGraph. We show how the abstraction and software implementation can be used to represent diverse data objects as graphs and to enable the use of tools from topology, graph theory, and machine learning (e.g., graph neural networks) to conduct a variety of tasks. We illustrate the versatility of the framework by using real datasets: (i) an image classification problem using topological data analysis to extract features from the graph model to train machine learning models; (ii) a disease outbreak problem where we model multivariate time series as graphs to detect abnormal events; and (iii) a technology pathway analysis problem where we highlight how we can use graphs to navigate connectivity. Further, our discussion also highlights how PlasmoData.jl leverages native Julia capabilities to enable compact syntax, scalable computations, and interfaces with diverse packages. Overall, we show that the DataGraph abstraction and PlasmoData.jl Julia package are able to model data within graphs and enable useful analysis.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Transmission and Distribution Real-Time Analysis Software for Monitoring and Control: Design and Simulation Testing

The US electric grid is facing operational, stability, and security challenges. Transmission system operators need some measure of visibility into distribution system renewable generation. Distribution system generation needs to support transmission system voltage. The grid is experiencing an expansion in measurement systems. How to take full advantage of this expansion and defend against attacks, both cyber and physical, poses additional challenges. This paper introduces software designed to meet these challenges. At the center of the software is an Integrated System Model (ISM) that spans from transmission to secondary distribution. The ISM is employed in real-time abnormality detection, voltage stability forecasting, and multi-mode control. The software architecture along with selected analysis modules is presented. Testing results are presented for: 1—attacks on utility infrastructure; 2—energy savings from optimal control; 3—distribution system control response during a low voltage transmission system event; 4—cyber-attacks on PV inverters, where physical inverters are used in hardware-in-the-simulation-loop studies. Contributions of this work include real-time analysis that spans from three-phase transmission through secondary distribution; an approach for detecting abnormalities that employs measurements from three independent measurement systems; and a multi-mode distribution system control that responds to cyber-attacks, physical attacks, equipment failures, and transmission system needs.

14 SOLAR ENERGY↗

The public health exposome and pregnancy-related mortality in the United States: a high-dimensional computational analysis

Racial inequities in maternal mortality in the U.S. continue to be stark. The 2015–2018, 4-year total population, county-level, pregnancy-related mortality ratio (PRM; deaths per 100,000 live births; National Center for Health Statistics (NCHS), restricted use mortality file) was linked with the Public Health Exposome (PHE). Using data reduction techniques, 1591 variables were extracted from over 62,000 variables for use in this analysis, providing information on the relationships between PRM and the social, health and health care, natural, and built environments. Graph theoretical algorithms and Bayesian analysis were applied to PHE/PRM linked data to identify latent networks. PHE variables most strongly correlated with total population PRM were years of potential life lost and overall life expectancy. Population-level indicators of PRM were overall poverty, smoking, lack of exercise, heat, and lack of adequate access to food. In this high-dimensional analysis, overall life expectancy, poverty indicators, and health behaviors were found to be the strongest predictors of pregnancy-related mortality. This provides strong evidence that maternal death is part of a broader constellation of both similar and unique health behaviors, social determinants and environmental exposures as other causes of death.

60 APPLIED LIFE SCIENCES↗

A Visual Comparison of Silent Error Propagation

High-performance computing (HPC) systems play a critical role in facilitating scientific discoveries. Their scale and complexity (e.g., the number of computational units and software stack) continue to grow as new systems are expected to process increasingly more data and reduce computing time. However, with more processing elements, the probability that these systems will experience a random bit-flip error that corrupts a program's output also increases, which is often recognized as silent data corruption. Analyzing the resiliency of HPC applications in extreme-scale computing to silent data corruption is crucial but difficult. An HPC application often contains a large number of computation units that need to be tested, and error propagation caused by error corruption is complex and difficult to interpret. Here, to accommodate this challenge, we propose an interactive visualization system that helps HPC researchers understand the resiliency of HPC applications and compare their error propagation. Our system models an application's error propagation to study a program's resiliency by constructing and visualizing its fault tolerance boundary. Coordinating with multiple interactive designs, our system enables domain experts to efficiently explore the complicated spatial and temporal correlation between error propagations. At the end, the system integrated a nonmonotonic error propagation analysis with an adjustable graph propagation visualization to help domain experts examine the details of error propagation and answer such questions as why an error is mitigated or amplified by program execution.

97 MATHEMATICS AND COMPUTING↗

A Systems Approach to Estimating the Uncertainty Limits of X-Ray Radiographic Metrology

Micro- and nanomanufacturing capabilities have rapidly expanded over the past decade to include complex three-dimensional (3D) structure fabrication; however, the metrology required to accurately assess these processes via part inspection and characterization has struggled to keep pace. X-ray computed tomography (CT) is considered an ideal candidate for providing the critically needed metrology on the smallest scales, especially internal features, or inaccessible regions. X-ray CT supporting micro- and nanomanufacturing often push against the poorly understood resolution and variation limits inherent to the machines, which can distort or hide fine structures. In this study, we have developed and experimentally verify a comprehensive analytical uncertainty propagation signal variation flow graph (SVFG) model for X-ray radiography in this work to better understand resolution and image variability limits on the small scale. The SVFG approach captures, quantifies, and predicts variations occurring in the system that limit metrology capabilities, particularly in the micro/nanodomain. This work is the first step to achieving full uncertainty modeling of CT reconstructions and provides insight into improving X-ray attenuation imaging systems. The SVFG methodology framework is applied to generate a complete basis set of functions describing the major sources of variation in radiographs. Five models are identified, covering variation in energy, intensity, length, blur, and position. Radiographic system experiments are defined to measure the parameters required by the SVFGs. Best practices are identified for these measurements. The SVFG models are confirmed via direct measurement of variation to predict variation within 30% on average.

47 OTHER INSTRUMENTATION↗

Topological Interpretability for Deep Learning

With the growing adoption of AI-based systems across everyday life, the need to understand their decision-making mechanisms is correspondingly increasing. The level at which we can trust the statistical inferences made from AI-based decision systems is an increasing concern, especially in high-risk systems such as criminal justice or medical diagnosis, where incorrect inferences may have tragic consequences. Despite their successes in providing solutions to problems involving real-world data, deep learning (DL) models cannot quantify the certainty of their predictions. These models are frequently quite confident, even when their solutions are incorrect. This work presents a method to infer prominent features in two DL classification models trained on clinical and non-clinical text by employing techniques from topological and geometric data analysis. We create a graph of a model's feature space and cluster the inputs into the graph's vertices by the similarity of features and prediction statistics. We then extract subgraphs demonstrating high-predictive accuracy for a given label. These subgraphs contain a wealth of information about features that the DL model has recognized as relevant to its decisions. We infer these features for a given label using a distance metric between probability measures, and demonstrate the stability of our method compared to the LIME and SHAP interpretability methods. This work establishes that we may gain insights into the decision mechanism of a DL model. This method allows us to ascertain if the model is making its decisions based on information germane to the problem or identifies extraneous patterns within the data.

Spannaus, Adam↗

Python HITMIX

SAND2024-10158O Python HITMIX is a software package for computing hitting time moments of vertices in a graph. Hitting time moments can be used to rank the strengths of relationships between vertices in a graph. Examples are provided for computing and using hitting time moments on generic graphs and similarity graphs arising from the analysis of text documents in information retrieval applications. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Dunlavy, Daniel↗

MCNPy

SAND2026-20425O MCNPy runs and analyzes simulations from MCNP, a software that models radiation transport of neutrons and gamma rays. MCNPy uses Python to start MCNP, retrieve event data files, and convert them into graph structures for detailed analysis. It offers visualization tools, including 2D views of particle histories, making complex simulation data easier to interpret for researchers and engineers. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Nowack, Aaron [Sandia National Lab. (SNL-CA), Live↗

Modern chemical graph theory

Abstract Graph theory has a long history in chemistry. Yet as the breadth and variety of chemical data is rapidly changing, so too do graph encoding methods and analyses that yield qualitative and quantitative insights. Using illustrative cases within a basic mathematical framework, we showcase modern chemical graph theory's utility in Chemists' analysis and model development toolkit. The encoding of both experimental and simulation data is discussed at various levels of granularity of information. This is followed by a discussion of the two major classes of graph theoretical analyses: identifying connectivity patterns and partitioning methods. Measures, metrics, descriptors, and topological indices are then introduced with an emphasis upon enhancing interpretability and incorporation into physical models. Challenging data cases are described that include strategies for studying time dependence. Throughout, we incorporate recent advancements in computer science and applied mathematics that are propelling chemical graph theory into new domains of chemical study. This article is categorized under: Molecular and Statistical Mechanics > Molecular Dynamics and Monte‐Carlo Methods Structure and Mechanism > Computational Materials Science Structure and Mechanism > Molecular Structures

Leite, Leonardo S. G.↗

Tackling the Challenges in Scene Graph Generation With Local-to-Global Interactions

In this work, we seek new insights into the underlying challenges of the scene graph generation (SGG) task. Quantitative and qualitative analysis of the visual genome (VG) dataset implies: 1) ambiguity: even if interobject relationship contains the same object (or predicate), they may not be visually or semantically similar; 2) asymmetry: despite the nature of the relationship that embodied the direction, it was not well addressed in previous studies; and 3) higher-order contexts: leveraging the identities of certain graph elements can help generate accurate scene graphs. Motivated by the analysis, we design a novel SGG framework, Local-to-global interaction networks (LOGINs). Locally, interactions extract the essence between three instances of subject, object, and background, while baking direction awareness into the network by explicitly constraining the input order of subject and object. Globally, interactions encode the contexts between every graph component (i.e., nodes and edges). Finally, Attract and Repel loss is utilized to fine-tune the distribution of predicate embeddings. By design, our framework enables predicting the scene graph in a bottom-up manner, leveraging the possible complementariness. To quantify how much LOGIN is aware of relational direction, a new diagnostic task called Bidirectional Relationship Classification (BRC) is also proposed. Overall, experimental results demonstrate that LOGIN can successfully distinguish relational direction than existing methods (in BRC task), while showing state-of-the-art results on the VG benchmark (in SGG task).

97 MATHEMATICS AND COMPUTING↗

An uncertainty-aware strategy for plasma mechanism reduction with directed weighted graphs

In this work, we present a framework for the analysis and reduction of plasma mechanisms by means of weighted directed graphs, in which reactions and species are both treated as nodes. The methodology consists of two distinct analyses. The first, which is qualitative, relies on graph spatializations via force-directed algorithms to discover the predominant global patterns in the chemical model. The second ranks the reactions based on their shortest paths' lengths from/to the species of interest and their relative contributions to the power balance. Further, this quantitative investigation enables a strategy for mechanism reduction that is fully automatized, as it does not require any expert knowledge, highly effective, as it generates reduced mechanisms that are highly accurate while relying on a small number of processes, and easily interpretable, as the algorithm justifies the importance of the retained reactions by outputting their related chemical pathways. Additionally, the work proposes a methodology extension that employs ensembles of graphs to improve the robustness of the reduced mechanism to reaction parameter uncertainties. The approach, here tested for steady-state predictions of a plasma system characterizing negative hydrogen ion sources, is general and can be used in a wide variety of applications outside the particular nuclear fusion context demonstrated in this work.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Reduced-Order Models of Static Power Grids based on Spectral Clustering

For large-scale interconnected power systems that cover large geographical areas, certain electrical studies are required so that appropriate decisions ensure system reliability and low cost. For such studies, it is often neither practical nor necessary to model in detail the entire power system, which is increasingly complex due to a more diverse range of grid assets to choose from in both short and long-term planning. The goal of this paper is to present a methodology to reduce the order of large-scale power networks based on spectral graph theory given that current methods for static network reduction are not scalable. A brief analysis of some spectral clustering properties to determine which graph Laplacian matrix should be used and why is included. The analysis shows that the utilization of the normalized graph Laplacian is more advantageous for clustering purposes. Techniques are proposed to approximate cost functions for the aggregated generators. This is done via linear regression. The reduced-order model obtained with the proposed methodology has an accuracy above 94% and solves the scalability issue commonly present in other reduction methods. If the utilization of the reduced-order model is either constrained to load levels above mid-peak demand, or cost functions of aggregated units are approximated via a piecewise quadratic approach, then the error distribution is in the order of 10^-3. .

Baquedano-Aguilar, Mario D.↗

The future evolution of energy-water-agriculture interconnectivity across the US

Abstract Energy, water, and agricultural resources across the globe are highly interconnected. This interconnectivity poses science challenges, such as understanding and modeling interconnections, as well as practical challenges, such as efficiently managing interdependent resource systems. Using the US as an example, this study seeks to define and explore how interconnectivity evolves over space and time under a range of influences. Concepts from graph theory and input–output analysis are used to visualize and quantify key intersectoral linkages using two new indices: the ‘Interconnectivity Magnitude Index’ and the ‘Interconnectivity Spread Index’. Using the Global Change Analysis Model (GCAM-USA), we explore the future evolution of these indices under four scenarios that explore a range of forces, including socioeconomic and technological change. Analysis is conducted at both national and state level spatial scales from 2015 to 2100. Results from a Reference scenario show that resource interconnectivity in the US is primarily driven by water use amongst different sectors, while changes in interconnectivity are driven by a decoupling of the water and electricity systems, as power plants become more water-efficient over time. High population and GDP growth results in relatively more decoupling of sectors, as a larger share of water and energy is used outside of interconnected sector feedback loops. Lower socioeconomic growth results in the opposite trend. Transitioning to a low-carbon economy increases interconnectivity because of the expansion of purpose-grown biomass, which strengthens the connections between water and energy. The results highlight that while some regions may experience similar sectoral stress projections, the composition of the intersectoral connectivity leading to that sectoral stress may call for distinctly different multi-sector co-management strategies. The methodology we introduce here can be applied in diverse geographical and sectoral contexts to enable better understanding of where, when, and how coupling or decoupling between sectors could evolve and be better managed.

Khan, Zarrar (ORCID:0000000281478553)↗

An ontology-based knowledge graph for representing interactions involving RNA molecules

The "RNA world" represents a novel frontier for the study of fundamental biological processes and human diseases and is paving the way for the development of new drugs tailored to each patient's biomolecular characteristics. Although scientific data about coding and non-coding RNA molecules are constantly produced and available from public repositories, they are scattered across different databases and a centralized, uniform, and semantically consistent representation of the "RNA world" is still lacking. We propose RNA-KG, a knowledge graph (KG) encompassing biological knowledge about RNAs gathered from more than 60 public databases, integrating functional relationships with genes, proteins, and chemicals and ontologically grounded biomedical concepts. To develop RNA-KG, we first identified, pre-processed, and characterized each data source; next, we built a meta-graph that provides an ontological description of the KG by representing all the bio-molecular entities and medical concepts of interest in this domain, as well as the types of interactions connecting them. Finally, we leveraged an instance-based semantically abstracted knowledge model to specify the ontological alignment according to which RNA-KG was generated. RNA-KG can be downloaded in different formats and also queried by a SPARQL endpoint. A thorough topological analysis of the resulting heterogeneous graph provides further insights into the characteristics of the "RNA world". RNA-KG can be both directly explored and visualized, and/or analyzed by applying computational methods to infer bio-medical knowledge from its heterogeneous nodes and edges. The resource can be easily updated with new experimental data, and specific views of the overall KG can be extracted according to the bio-medical problem to be studied.

59 BASIC BIOLOGICAL SCIENCES↗

Open Source Software Prevalence Ingest Tool

The OSSP Ingest Tool accepts user-input organizational information, ingests IT/OT asset lists in Excel format, and ingests the associated CycloneDX SBOM's. It then performs analytics demonstrating the ability to answer the follow research questions: o RQ1. Ability to identify all OSS services running on, and all OSS components present within, an OT device o RQ1a: Ability to differentiate multiple versions of the same OSS component within each OT device. o RQ1b: Ability to differentiate running from not-running OSS components. o RQ1c: Ability to differentiate based on the originator of the component, because a supplier may have modified it after retrieval from the upstream software source. o RQ2. Ability to correlate the identity of a single OSS component across multiple OT devices, mitigating common name variations such as differences in capitalization, '-' vs '_', and so on. o RQ3. Ability to perform subset analysis of OSS components across multiple OT devices o RQ3a: Ability to perform subset analysis across OSS libraries, generating density & distribution graphs to identify commonly-used libraries and outliers. o RQ3b: Ability to perform subset analysis of a single OSS library, generating density & distribution by CI sector, by device type, by device make/model, and/or by firmware version. o RQ3c: Ability to perform subset analysis by grouping OSS libraries according to programming language, then overlay with RQ4b. o RQ3d: Ability to perform subset analysis by OSS upstream source, providing insight into degree of modifications performed by suppliers. o RQ4. Ability to identify dependencies (transitive and direct) of each differentiated OSS library within each OT device, and enable RQ1,2,3 iteratively for dependencies. o RQ1. Ability to identify all OSS services running on, and all OSS components present within, an OT device o RQ1a: Ability to differentiate multiple versions of the same OSS component within each OT device. o RQ1b: Ability Page

Kapadia, Shayna [Lawrence Livermore National Labor↗