Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graph analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

BAMCensus (The Behavior and Advanced Mobility Census Dataset Aggregator) [SWR-25-120]

This software is a high-performance tool developed in Rust for downloading and processing large-scale geospatial datasets, specifically focusing on US Census data. It is designed to address scaling limitations found in existing tools, such as R's [tidycensus](https://walker-data.com/tidycensus/), by providing performant streaming dataset JOIN operations between various US Census datasets (like ACS and LEHD) and their corresponding geometries stored on the TIGER/Lines web server. The tool automates the process of joining these data sources, returning aggregated data to the user based on a specified census GEOID type. The tool automates the process of joining these data sources, returning aggregated data to the user based on a specified census GEOID type. Its primary motivation stems from the need for a high-performance solution to combine spatial datasets with graph traversals within the context of mobility analysis tooling being developed at NREL's Behavior and Advanced Mobility (BAM) group.

Fitzgerald, Robert [National Renewable Energy Labo↗

TriC: Distributed-memory Triangle Counting by Exploiting the Graph Structure

Graph analytics has emerged as an important tool in the analysis of large scale data from diverse application domains such as social networks, cyber security and bioinformatics. Counting the number of triangles in a graph is a fundamental kernel with several applications such as detecting the community structure of a graph or in identifying important vertices in a graph. The ubiquity of massive datasets is driving the need to scale graph analytics on parallel systems. However, numerous challenges exist in efficiently parallelizing graph algorithms, especially on distributed-memory systems. Irregular memory accesses and communication patterns, low computation to communication ratios, and the need for frequent synchronization are some of the leading challenges. In this paper, we present TriC, our distributed-memory implementation of triangle counting in graphs using the Message Passing Interface (MPI), as a submission to the 2020 GraphChallenge competition. Using a set of synthetic and real-world inputs from the challenge, we demonstrate a speedup of up to 90x relative to previous work on 32 processor-cores of a NERSC Cori node. We also provide details from distributed runs with up to8192 processes along with strong scaling results. The observations presented in this work provide an understanding of the system-level bottlenecks at scale that specifically impact sparse-irregular workloads and will therefore benefit other efforts to parallelize graph algorithms.

Halappanavar, Mahantesh↗

Structure–Property Linkage in Alloys Using Graph Neural Network and Explainable Artificial Intelligence

Deep learning tools have recently shown significant potential for accelerating the prediction of microstructure–property linkage in materials. While deep neural networks like convolution neural networks (CNNs) can extract physics information from 3D microstructure images, they often require a large network architecture and substantial training time. In this research, we trained a graph neural network (GNN) using phase field generated microstructures of Ni-Al alloys to predict the evolution of mechanical properties. We found that a single GNN is capable of accurately predicting the strengthening of Ni-Al alloys with microstructures of varying sizes and dimensions, which cannot otherwise be done with a CNN. Additionally, GNN requires significantly less GPU utilization than CNN and offers more interpretable explanation of predictions using saliency analysis as features are manually defined in the graph. We also utilize explainable artificial intelligence tool Bayesian Inference to determine the coefficients in the power law equation that governs coarsening of precipitates. Overall, our work demonstrates the ability of the GNN to accurately and efficiently extract relevant information from material microstructures without having restrictions on microstructure size or dimension and offers an interpretable explanation.

Chemistry↗

Development of Whole System Digital Twins for Advanced Reactors: Leveraging Graph Neural Networks and SAM Simulations

Here, in this work, we introduce a novel method to develop whole system digital twins (DTs) for advanced nuclear reactors. This method treats a complex reactor system as a heterogeneous graph: with the system components as different types of graph nodes and their physical interconnections as edges. Based on the heterogeneous graph, a graph neural network combining graph convolution and temporal node attention is developed as the DT, facilitating a comprehensive understanding of the system's dynamic behavior. By utilizing the System Analysis Module (SAM) code for simulating various operational transients, we develop a graph-based database that trains the DT. This DT is characterized by two primary functions: It can infer the entire system's status using sparse node information, and it can predict the progress of transients based on current and historical system information. Our approach is validated through case studies on the Experimental Breeder Reactor II (EBR-II) system and a generic Fluoride-salt-cooled High-temperature Reactor (gFHR), demonstrating the DT's accuracy in forecasting operational transients. The DT's rapid computation capabilities enhance its potential for supporting advanced reactor operations, offering benefits in intelligent simulation, autonomous control, and anomaly detection, paving the way for improved safety analysis and intelligent component health management for advanced reactor systems and reducing their operations and maintenance cost.

EBR-II↗

Gaps labeling theorem for the bubble-diamond self-similar graphs

Abstract Motivated by the appearance of fractals in several areas of physics, especially in solid state physics and the physics of aperiodic order, and in other sciences, including the quantum information theory, we present a detailed spectral analysis for a new class of fractal-type diamond graphs, referred to as bubble-diamond graphs, and provide a gap-labeling theorem in the sense of Bellissard for the corresponding probabilistic graph Laplacians using the technique of spectral decimation. Labeling the gaps in the Cantor set by the normalized eigenvalue counting function, also known as the integrated density of states, we describe the gap labels as orbits of a second dynamical system that reflects the branching parameter of the bubble construction and the decimation structure. The spectrum of the natural Laplacian on limit graphs is shown generically to be pure point supported on a Cantor set, though one particular graph has a mixture of pure point and singularly continuous components.

Physics↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

Leveraging Topological Events in Tracking Graphs for Understanding Particle Diffusion

Abstract Single particle tracking (SPT) of fluorescent molecules provides significant insights into the diffusion and relative motion of tagged proteins and other structures of interest in biology. However, despite the latest advances in high‐resolution microscopy, individual particles are typically not distinguished from clusters of particles. This lack of resolution obscures potential evidence for how merging and splitting of particles affect their diffusion and any implications on the biological environment. The particle tracks are typically decomposed into individual segments at observed merge and split events, and analysis is performed without knowing the true count of particles in the resulting segments. Here, we address the challenges in analyzing particle tracks in the context of cancer biology. In particular, we study the tracks of KRAS protein, which is implicated in nearly 20% of all human cancers, and whose clustering and aggregation have been linked to the signaling pathway leading to uncontrolled cell growth. We present a new analysis approach for particle tracks by representing them as tracking graphs and using topological events – merging and splitting, to disambiguate the tracks. Using this analysis, we infer a lower bound on the count of particles as they cluster and create conditional distributions of diffusion speeds before and after merge and split events. Using thousands of time‐steps of simulated and in‐vitro SPT data, we demonstrate the efficacy of our method, as it offers the biologists a new, detailed look into the relationship between KRAS clustering and diffusion speeds.

McDonald, T.↗

Grassmannian Diffusion Maps--Based Dimension Reduction and Classification for High-Dimensional Data

This work introduces the Grassmannian diffusion maps (GDMaps), a novel nonlinear dimensionality reduction technique that defines the affinity between points through their representation as low-dimensional subspaces corresponding to points on the Grassmann manifold. Here, the method is designed for applications, such as image recognition and data-based classification of constrained high-dimensional data where each data point itself is a high-dimensional object (i.e., a large matrix) that can be compactly represented in a lower-dimensional subspace. The GDMaps is composed of two stages. The first is a pointwise linear dimensionality reduction wherein each high-dimensional object is mapped onto the Grassmann manifold representing the low-dimensional subspace on which it resides. The second stage is a multipoint nonlinear kernel-based dimension reduction using diffusion maps to identify the subspace structure of the points on the Grassmann manifold. To this end, an appropriate Grassmannian kernel is used to construct the transition matrix of a random walk on a graph connecting points on the Grassmann manifold. Spectral analysis of the transition matrix yields low-dimensional Grassmannian diffusion coordinates embedding the data into a low-dimensional reproducing kernel Hilbert space. Further, a novel data classification/recognition technique is developed based on the construction of an overcomplete dictionary of reduced dimension whose atoms are given by the Grassmannian diffusion coordinates. Three examples are considered. First, a "toy" example shows that the GDMaps can identify an appropriate parametrization of structured points on the unit sphere. The second example demonstrates the ability of the GDMaps to revealing the intrinsic subspace structure of high-dimensional random field data. In the last ex- ample, a face recognition problem is solved considering face images subject to varying illumination conditions, changes in face expressions, and occurrence of occlusions. The technique presented high recognition rates (i.e., 95% in the best case) using a fraction of the data required by conventional methods.

42 ENGINEERING↗

Computational Estimation by Scientific Data Mining with Classical Methods to Automate Learning Strategies of Scientists

Experimental results are often plotted as 2-dimensional graphical plots (aka graphs) in scientific domains depicting dependent versus independent variables to aid visual analysis of processes. Repeatedly performing laboratory experiments consumes significant time and resources, motivating the need for computational estimation. The goals are to estimate the graph obtained in an experiment given its input conditions, and to estimate the conditions that would lead to a desired graph. Existing estimation approaches often do not meet accuracy and efficiency needs of targeted applications. We develop a computational estimation approach called AutoDomainMine that integrates clustering and classification over complex scientific data in a framework so as to automate classical learning methods of scientists. Knowledge discovered thereby from a database of existing experiments serves as the basis for estimation. Challenges include preserving domain semantics in clustering, finding matching strategies in classification, striking a good balance between elaboration and conciseness while displaying estimation results based on needs of targeted users, and deriving objective measures to capture subjective user interests. These and other challenges are addressed in this work. The AutoDomainMine approach is used to build a computational estimation system, rigorously evaluated with real data in Materials Science. Our evaluation confirms that AutoDomainMine provides desired accuracy and efficiency in computational estimation. It is extendable to other science and engineering domains as proved by adaptation of its sub-processes within fields such as Bioinformatics and Nanotechnology.

Computer Science↗

TX$^2$: Transformer eXplainability and eXploration

The Transformer eXplainability and eXploration (Martindale & Stewart, 2021), or TX 2 software package, is a library designed for artificial intelligence researchers to better understand the performance of transformer models (Vaswani et al., 2017) used for sequence classification. The tool is capable of integrating with a trained transformer model and a dataset split into training and testing populations to produce an ipywidget (Project Jupyter Contributors, 2021) dashboard with a number of visualizations to understand model performance with an emphasis on explainability and interpretability. The TX 2 package is primarily intended to integrate into a workflow centered around Jupyter Notebooks (Kluyver et al., 2016), and currently assumes the use of PyTorch (Paszke et al., 2019) and Hugging Face transformers library (Wolf et al., 2020). The dashboard includes visualization and data exploration features to aid researchers, including an interactive UMAP embedding graph (McInnes et al., 2018) to understand classification clusters, a word salience map that can be updated as researchers alter textual entries in near real time, a set of tools to understand word frequency and importance based on the clusters in the UMAP embedding graph, and a set of traditional confusion matrix analysis tools.

97 MATHEMATICS AND COMPUTING↗

Graph-based featurization methods for classifying small molecule compounds

For over a decade, drug-induced liver injury (DILI) has posed significant drawbacks in the synthesis and development of drugs and remains a consequential concern. With finite success within the existing preclinical models, DILI is one of the main causes of drug withdrawal or termination from the market. Particularly, this withdrawal occurs during the late stages of drug development (Kullak-Ublick, 2017). Since DILI is difficult to diagnose and treat, it has become an obstacle in the drug production market that in turn affects clinicians, pharmaceutical companies, and consumers. We propose a method for learning features of DILI-positive drugs based on the graphical relationships and patterns they possess within a network of biological databases. We also train various statistical and machine learning models on these learned features in order to classify the drugs as DILI-positive or negative. Our methods include Random Forest, Neural networks, and logistic regression classification. We utilize labeled DILI-positive and DILI-negative datasets, which were developed by the FDA and the National center for toxicological research, as well as additional literature datasets (Thakkar, 2020) in order to validate our results and assess our featurization and model accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A General Materials Data Science Framework for Quantitative 2D Analysis of Particle Growth from Image Sequences

Abstract Phase transformations are a challenging problem in materials science, which lead to changes in properties and may impact performance of material systems in various applications. We introduce a general framework for the analysis of particle growth kinetics by utilizing concepts from machine learning and graph theory. As a model system, we use image sequences of atomic force microscopy showing the crystallization of an amorphous fluoroelastomer film. To identify crystalline particles in an amorphous matrix and track the temporal evolution of the particle dispersion, we have developed quantitative methods of 2D analysis. 700 image sequences were analyzed using a neural network architecture, achieving 0.97 pixel-wise classification accuracy as a measure of the correctly classified pixels. The growth kinetics of isolated and impinged particles were tracked throughout time using these image sequences. The relationship between image sequences and spatiotemporal graph representations was explored to identify the proximity of crystallites from each other. The framework enables the analysis of all image sequences without the requirement of sampling for specific particles or timesteps for various materials systems.

36 MATERIALS SCIENCE↗

Collection And Analysis Of Telemetry For The Cyote Heuristic

CATCH CLI focuses on gathering telemetry data, storing it in the Neo4j database, querying for Mitre ATT&CK patterns, and creating STIX 2.1 reports. Key Components: Analysis Modules: Analyze data to detect attack patterns. GoSTOTS Collection Engines: Collect telemetry data. These tools can be used together or individually. Analysis modules rely on data from specific engines to identify attack patterns. Source Code Organization: Engines: CATCH/catch/cmd/collection Modules: CATCH/catch/cmd/analysis CGUI Overview CATCH Graphical User Interface (CGUI) offers a graphical shell to execute CATCH CLI, allowing easy editing of: Analysis Modules Database configurations Profiles (collection and device settings) Neo4j Overview Neo4j is a graph database using the Cypher query language, storing data in JSON. It seamlessly integrates with STIX 2.1 data for: Data Submission: CATCH Collection Engines Data Querying: Analysis Modules CATCH modifies STIX 2.1 data for Neo4j submission and reverts it back during querying. STIG Overview Structured Threat Intelligence Graph (STIG) is a tool for creating, editing, querying, analyzing, and visualizing threat intelligence using STIX 2.1 and storing data in Neo4j. Usage Tools can be run: Manually (CLI): Refer to CATCH documentation User Interface: Run ./cgui/CGUI or go run ./cgui/ Additional Information Logging System: Detailed in the config documentation Further Documentation: Available for CATCH and CGUI

Madsen, MichaelJ. [Idaho National Laboratory (INL)↗

Modified Energy Span Analysis of Catalytic Parallel Pathways and Selectivity

Mechanistic modeling provides vital insights into catalytic reactions. To analyze complex reaction networks with parallel pathways, we leverage the graph theory approach of the Energy Span Model (ESM) to develop a modified energy span analysis (MESA). A new method of cycle plots is proposed to perform reaction pathways analysis visually. We demonstrate this method on two published models: one describing carbon monoxide oxidation and the other simulating ethylene conversion to propanal via hydroformylation or ethane via hydrogenation. Fundamental insights explain kinetic observables, such as a reactant’s negative reaction order. General principles are revealed, such as rate-determining surface species being outside the primary reaction flux cycle and pathway selectivity being a purely kinetic property when reaction conditions are not near equilibrium. Lastly, we demonstrate MESA’s consistency with published microkinetic modeling results, highlighting this technique’s extension of the ESM to heterogeneous catalysts using collision theory to describe adsorption steps and concentration effects.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ExTreeM: Scalable Augmented Merge Tree Computation via Extremum Graphs

Over the last decade merge trees have been proven to support a plethora of visualization and analysis tasks since they effectively abstract complex datasets. Here, this paper describes the ExTreeM-Algorithm: A scalable algorithm for the computation of merge trees via extremum graphs. The core idea of ExTreeM is to first derive the extremum graph G of an input scalar field f defined on a cell complex K, and subsequently compute the unaugmented merge tree of f on G instead of K; which are equivalent. Any merge tree algorithm can be carried out significantly faster on G, since K in general contains substantially more cells than G. To further speed up computation, ExTreeM includes a tailored procedure to derive merge trees of extremum graphs. The computation of the fully augmented merge tree, i.e., a merge tree domain segmentation of K, can then be performed in an optional post-processing step. All steps of ExTreeM consist of procedures with high parallel efficiency, and we provide a formal proof of its correctness. Our experiments, performed on publicly available datasets, report a speedup of up to one order of magnitude over the state-of-the-art algorithms included in the TTK and VTK-m software libraries, while also requiring significantly less memory and exhibiting excellent scaling behavior.

97 MATHEMATICS AND COMPUTING↗

People who inject drugs in metropolitan Chicago: A meta-analysis of data from 1997-2017 to inform interventions and computational modeling toward hepatitis C microelimination

Progress toward hepatitis C virus (HCV) elimination in the United States is not on track to meet targets set by the World Health Organization, as the opioid crisis continues to drive both injection drug use and increasing HCV incidence. A pragmatic approach to achieving this is using a microelimination approach of focusing on high-risk populations such as people who inject drugs (PWID). Computational models are useful in understanding the complex interplay of individual, social, and structural level factors that might alter HCV incidence, prevalence, transmission, and treatment uptake to achieve HCV microelimination. However, these models need to be informed with realistic sociodemographic, risk behavior and network estimates on PWID. We conducted a meta-analysis of research studies spanning 20 years of research and interventions with PWID in metropolitan Chicago to produce parameters for a synthetic population for realistic computational models (e.g., agent-based models). We then fit an exponential random graph model (ERGM) using the network estimates from the meta-analysis in order to develop the network component of the synthetic population.

60 APPLIED LIFE SCIENCES↗

Street-level temperature estimation using graph neural networks: Performance, feature embedding and interpretability

Estimating street-level air temperature is a challenging task due to the highly heterogeneous urban surfaces, canyon-like street morphology, and the diverse physical processes in the built environment. Though pioneering studies have embarked on investigations via data-driven approaches, many questions remain to be answered. Here, in this study, we leveraged an innovative framework and redefined the street-level temperature estimation problem using Graph Neural Networks (GNN) with spatial embedding techniques. The results showed that GNN models are more capable and consistent of estimating street-level temperature among tested locations, benefiting from its unique strength in handling extensive data over unstructured graph topology. In addition, we conducted in-depth analysis of feature importance to enhance the model interpretability. Among the urban features analyzed in this study, the time-variant canopy density and meter-level land use data emerge as crucial factors. Our findings highlight GNN 's high potential in capturing the complex dynamics between urban elements and their impacts on microclimate, thus offering valuable insights for comprehensive urban data collection and urban climate modeling in general. Collectively, this study also contributes to urban planning and policy by providing avenues to enhance city resilience against climate change, thereby advancing the agenda for environmental stewardship and urban sustainability.

54 ENVIRONMENTAL SCIENCES↗

A Comparison between Invariant and Equivariant Classical and Quantum Graph Neural Networks

Machine learning algorithms are heavily relied on to understand the vast amounts of data from high-energy particle collisions at the CERN Large Hadron Collider (LHC). The data from such collision events can naturally be represented with graph structures. Therefore, deep geometric methods, such as graph neural networks (GNNs), have been leveraged for various data analysis tasks in high-energy physics. One typical task is jet tagging, where jets are viewed as point clouds with distinct features and edge connections between their constituent particles. The increasing size and complexity of the LHC particle datasets, as well as the computational models used for their analysis, have greatly motivated the development of alternative fast and efficient computational paradigms such as quantum computation. In addition, to enhance the validity and robustness of deep networks, we can leverage the fundamental symmetries present in the data through the use of invariant inputs and equivariant layers. In this paper, we provide a fair and comprehensive comparison of classical graph neural networks (GNNs) and equivariant graph neural networks (EGNNs) and their quantum counterparts: quantum graph neural networks (QGNNs) and equivariant quantum graph neural networks (EQGNN). The four architectures were benchmarked on a binary classification task to classify the parton-level particle initiating the jet. Based on their area under the curve (AUC) scores, the quantum networks were found to outperform the classical networks. However, seeing the computational advantage of quantum networks in practice may have to wait for the further development of quantum technology and its associated application programming interfaces (APIs).

Forestano, Roy T. (ORCID:0000000203552076)↗