Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graph metrics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Reference-free structural variant detection in microbiomes via long-read co-assembly graphs

Motivation: The study of bacterial genome dynamics is vital for understanding the mechanisms underlying microbial adaptation, growth, and their impact on host phenotype. Structural variants (SVs), genomic alterations of 50 base pairs or more, play a pivotal role in driving evolutionary processes and maintaining genomic heterogeneity within bacterial populations. While SV detection in isolate genomes is relatively straightforward, metagenomes present broader challenges due to the absence of clear reference genomes and the presence of mixed strains. In response, our proposed method rhea, forgoes reference genomes and metagenome-assembled genomes (MAGs) by encompassing all metagenomic samples in a series (time or other metric) into a single co-assembly graph. The log fold change in graph coverage between successive samples is then calculated to call SVs that are thriving or declining. Results: We show rhea to outperform existing methods for SV and horizontal gene transfer (HGT) detection in two simulated mock metagenomes, particularly as the simulated reads diverge from reference genomes and an increase in strain diversity is incorporated. We additionally demonstrate use cases for rhea on series metagenomic data of environmental and fermented food microbiomes to detect specific sequence alterations between successive time and temperature samples, suggesting host advantage. Our approach leverages previous work in assembly graph structural and coverage patterns to provide versatility in studying SVs across diverse and poorly characterized microbial communities for more comprehensive insights into microbial gene flux.

59 BASIC BIOLOGICAL SCIENCES↗

Metric Learning for Hyperspectral Image Segmentation

We present a metric learning approach to improve the performance of unsupervised hyperspectral image segmentation. Unsupervised spatial segmentation can assist both user visualization and automatic recognition of surface features. Analysts can use spatially-continuous segments to decrease noise levels and/or localize feature boundaries. However, existing segmentation methods use tasks-agnostic measures of similarity. Here we learn task-specific similarity measures from training data, improving segment fidelity to classes of interest. Multiclass Linear Discriminate Analysis produces a linear transform that optimally separates a labeled set of training classes. The defines a distance metric that generalized to a new scenes, enabling graph-based segmentation that emphasizes key spectral features. We describe tests based on data from the Compact Reconnaissance Imaging Spectrometer (CRISM) in which learned metrics improve segment homogeneity with respect to mineralogical classes.

Compact Reconnaissance Imaging Spectrometer (CRISM↗

Modern chemical graph theory

Abstract Graph theory has a long history in chemistry. Yet as the breadth and variety of chemical data is rapidly changing, so too do graph encoding methods and analyses that yield qualitative and quantitative insights. Using illustrative cases within a basic mathematical framework, we showcase modern chemical graph theory's utility in Chemists' analysis and model development toolkit. The encoding of both experimental and simulation data is discussed at various levels of granularity of information. This is followed by a discussion of the two major classes of graph theoretical analyses: identifying connectivity patterns and partitioning methods. Measures, metrics, descriptors, and topological indices are then introduced with an emphasis upon enhancing interpretability and incorporation into physical models. Challenging data cases are described that include strategies for studying time dependence. Throughout, we incorporate recent advancements in computer science and applied mathematics that are propelling chemical graph theory into new domains of chemical study. This article is categorized under: Molecular and Statistical Mechanics > Molecular Dynamics and Monte‐Carlo Methods Structure and Mechanism > Computational Materials Science Structure and Mechanism > Molecular Structures

Leite, Leonardo S. G.↗

A Transductive Graph Neural Network learning for Grid Resilience Analysis

Power grids are critical infrastructures that require robust resilience analysis to ensure reliable and uninterrupted electricity supply. Traditional simulation-based methods for grid resilience analysis suffer from computational complexity and limited ability to capture the full spectrum of potential disruptions. This paper presents a novel approach to enhance grid resilience by leveraging transductive graph neural network (GNN) learning to identify critical nodes and links. By leveraging the graph structure and system features, GNNs effectively learn resilience metrics and accurately identify critical nodes based on actual grid operational behavior. The efficacy of the proposed approach is demonstrated through case studies on node criticality scoring and critical node/line identification in cascading outage scenarios. The results highlight the advantages of learning-based methods over traditional simulation-based approaches and their potential to revolutionize grid resilience analysis. The contributions of this paper include a graph-based scalable approach for fast cascading analysis, an inductive formulation for training GNN models, and a transfer learning-based approach to scale the model to largescale power systems.

grid resilience, graph neural networks, transducti↗

Efficient Sampling of Complex Interdependent and Multiplex Networks

Efficient sampling of interdependent and multiplex infrastructure networks is critical for effectively applying failure and recovery algorithms in real-world settings, as well as to generate property-preserving reduced-order graph-based ensembles that address topological uncertainties. In this paper, we first explore the performance, i.e. the success in preserving graph properties, of graph sampling algorithms for interdependent and multiplex networks with synthetic and real-world graphs. We simulate sampling algorithms under different parameter settings. These settings include probabilistic graph generators, coupling patterns, and various performance metrics. Our results show that while Random Node and Random Walk sampling algorithms perform best for interdependent networks, Random Edge and Forest Fire sampling algorithms perform best for multiplex networks. Second, we propose and implement a novel similarity-based sampling algorithm for multiplex networks that samples only log(N) number of layers of an N-layer multiplex network while yielding computational savings with performance guarantees. Experimental results show that similarity sampling outperforms complete sampling of all layers while decreasing performance costs from a linear scale to a logarithmic one. Our results also indicate that similarity-based sampling outperforms complete sampling and random selection in nearly all scenarios when tested with real-world data.

Subasi, Omer↗

Automating Log Synthesis and Visualization with Python and Splunk

The goal of this project is to automate log analysis by utilizing Splunk, Bash, and Python together. Simplifying the monitoring and analysis of network traffic was the main goal. In order to accomplish this, a Bash script was created to use 'tcpdump' to automate network sniffing. It also included a 24-hour file rotation mechanism to effectively manage the pcap files that were generated. After that, a Python script was written to read these pcap files and retrieve pertinent data about network traffic. After processing the collected data, Splunk is used to summarize the important metrics and visualize said information with relevant graphs.

99 GENERAL AND MISCELLANEOUS↗

Initial Ada components evaluation

The SAIC has the responsibility for independent test and validation of the SSE. They have been using a mathematical functions library package implemented in Ada to test the SSE IV and V process. The library package consists of elementary mathematical functions and is both machine and accuracy independent. The SSE Ada components evaluation includes code complexity metrics based on Halstead's software science metrics and McCabe's measure of cyclomatic complexity. Halstead's metrics are based on the number of operators and operands on a logical unit of code and are compiled from the number of distinct operators, distinct operands, and total number of occurrences of operators and operands. These metrics give an indication of the physical size of a program in terms of operators and operands and are used diagnostically to point to potential problems. McCabe's Cyclomatic Complexity Metrics (CCM) are compiled from flow charts transformed to equivalent directed graphs. The CCM is a measure of the total number of linearly independent paths through the code's control structure. These metrics were computed for the Ada mathematical functions library using Software Automated Verification and Validation (SAVVAS), the SSE IV and V tool. A table with selected results was shown, indicating that most of these routines are of good quality. Thresholds for the Halstead measures indicate poor quality if the length metric exceeds 260 or difficulty is greater than 190. The McCabe CCM indicated a high quality of software products.

Moebes, Travis↗

Minimization of Measurement Uncertainty in Optical Frequency Domain Reflectometry

Optical frequency domain reflectometry (OFDR) is a technique for interrogating optical fiber sensors to generate relative, quasi-distributed measurements. Although Optical frequency domain reflectometry (OFDR) is increasingly being adopted for aerospace, energy production, and structural monitoring applications, the quantification of uncertainty for OFDR measurements has not been developed beyond sparse empirical relationships. To address this knowledge gap, an uncertainty metric for OFDR measurements was developed. This uncertainty metric was applied to weight the edges between OFDR measurements on directed correlation graphs and analyzed to minimize the cumulative uncertainty. In conclusion, this work is the first to propose an uncertainty metric for OFDR and provides a generalized mathematical framework for optimizing OFDR hardware selection, optical fiber sensor selection, and postprocessing strategy.

42 ENGINEERING↗

Quantifying disorder one atom at a time using an interpretable graph neural network paradigm

Abstract Quantifying the level of atomic disorder within materials is critical to understanding how evolving local structural environments dictate performance and durability. Here, we leverage graph neural networks to define a physically interpretable metric for local disorder, called SODAS. This metric encodes the diversity of the local atomic configurations as a continuous spectrum between the solid and liquid phases, quantified against a distribution of thermal perturbations. We apply this methodology to four prototypical examples with varying levels of disorder: (1) grain boundaries, (2) solid-liquid interfaces, (3) polycrystalline microstructures, and (4) tensile failure/fracture. We also compare SODAS to several commonly used methods. Using elemental aluminum as a case study, we show how our paradigm can track the spatio-temporal evolution of interfaces, incorporating a mathematically defined description of the spatial boundary between order and disorder. We further show how to extract physics-preserved gradients from our continuous disorder fields, which may be used to understand and predict materials performance and failure. Overall, our framework provides a simple and generalizable pathway to quantify the relationship between complex local atomic structure and coarse-grained materials phenomena.

36 MATERIALS SCIENCE↗

Cross-Domain Reasoning for Neuromorphic Model Design

Designing performant neuromorphic models requires reasoning across neuroscience, neuromorphic computing, and machine learning, making it a natural target for cross-domain hypothesis generation. Our primary contribution is a multi-corpus knowledge graph spanning all three domains, which we show substantially increases cross-domain retrieval novelty over single-corpus baselines. We additionally introduce NeuKReAct, an agentic reasoning framework that iteratively retrieves from this graph and synthesizes design hypotheses via a step-by-step blackboard architecture, enabling structured compartmentalization of design decisions. Lastly, we introduce an execution head that translates hypotheses into structured design documents and runnable code. We evaluate novelty using a combinatorial creativity metric that measures cross-domain retrieval distance across the citation graph. Our results confirm that corpus breadth is the dominant driver of novelty. Moreover, we highlight a concrete instance of the novelty-utility tradeoff within NeuKReAct, underscoring a need for joint creativity evaluation, balancing both novelty and utility.

Ramavarapu, Vikram [ORNL] (ORCID:0009000188757213)↗

Graph-Theoretic Approaches to Quantifying Power System Resiliency

Although gaining growing importance, the subject of power system resiliency still lacks a commonly acknowledged metric. As a contribution to solving this complication, in this paper we leverage the concepts of spanning trees and Fiedler value from graph theory to propose two topology-based indices for quantifying the resiliency of power systems. The proposed indices require least information and may be applied to any other flow network, such as water or gas pipeline networks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Visual Understanding of COVID-19 Knowledge Graph for Predictive Analysis

This study aims to effectively analyze and visualize the concept to concept network derived from the COVID-19 Open Research Dataset (CORD-19) dataset, where we have more than 48,000 concepts with more than 300,000 relationships between concepts. In analyzing networks, we focus on finding relationship patterns between the coronavirus disease 2019 (COVID-19) concepts and other concepts. Given the node and edge datasets, we construct directional graphs and calculate all pair shortest paths based on multiple edge weight schemes. However, statistical metrics are not sufficient to identify specific relationships represented in the network. Therefore, we also propose a visual analytics approach to effectively understand the knowledge graph. Our highly interactive visual analytics allows users to effectively analyze the evolving graphs and (COVID-19) concept nodes and other nodes related to the COVID-19 nodes. We envision that this study will pave the path to develop strategies to provide more accurate and scalable predictive analysis on knowledge graphs related to CORD19 and other biomedical knowledge graphs.

Lim, Seung-Hwan↗

Illuminating the pathways to carbon liberation: a systems approach to characterizing the consequential unknowns of carbon transformation and loss from thawing permafrost peatlands (Final Report)

The IsoGenie3 Project delivered new systems-level insights into carbon cycling in thawing permafrost landscapes, with an emphasis on methane and carbon dioxide emissions. From >200 samples from the site collected over a decade, co-analyzed for geochemistry and microbiology, the team recovered ~1,500 assembled microbial genomes and ~1,900 viral population genomes, revealing appreciable genetic novelty - from a new highly abundant bacterial phylum, to novel methane consumers and their activities, to rampant viral novelty. IsoGenie3 linked these organisms to carbon compound transformations (which define the cycling of organic matter in soils, and the loss of the greenhouse gases carbon dioxide and methane), and saw that the microbes at each stage of permafrost thaw had different genetic potential to degrade categories of compounds, expressed that genetic potential differently, and actually transformed carbon compounds into greenhouse gases in different ways. IsoGenie 3 identified that some of the thaw-stage differences were due to plant-microbiome relationships; the plant species across the thaw gradient contributed different carbon compounds into the soil, and hosted distinct microbiota (differing among parts of plants as well as species). Lastly, microbes in the saturated post-thaw conditions appeared likely to contribute to the mobilization and toxification of mercury released during thaw. In parallel with ongoing field sampling and analysis, hypotheses arising from field observations were tested via lab incubation experiments. When communities are taken out of their native habitats, they behave differently, and the team first rigorously quantified the magnitude of this effect on microbiome composition and functional capacity, organic matter composition, and gas production; overall the main system processes were maintained in the lab incubations under the conditions tested. Further, the microbial data could inform geochemical reaction network models of those processes. Then, the team ran experiments with additions of compounds, varying temperature, and “live” vs. “dead” peat (the latter having been gamma irradiated, with a few additional variants to control for methodological artifacts). From these, we (a) determined the importance of plant-derived soluble phenolic compounds in bogs’ extraordinary recalcitrance of organic matter, and carbon gas emissions skewed to carbon dioxide; (b) proposed an abiotic ‘tanning’ mechanism, which could contribute to Sphagnum’s inhibitory effect on anaerobic decomposition through alteration of N availability. IsoGenie3 illuminated longer-term and landscape-scale interactions of permafrost thaw and carbon cycling, advancing knowledge of the drivers of methane dynamics not only across in the permafrost-associated peatland (where hydrology and plant communities dictate microbiomes) but also their interconnected lakes (where sediment carbon quality and resident microbiota are determined by position within lake, and lake features). By leveraging observations of site methane dynamics extending well before this project, the team was able to construct a 44-year portrait of the interplay of permafrost thaw, hydrology, vegetation dynamics, and carbon gas emissions, and the doubling of the fully-thawed fens over this time. From the detailed study of this focal site, IsoGenie3 also aimed to improve model representation of these kinds of sites and processes. To improve predictions of methane transformations, we incorporated acetate and isotope dynamics into the ‘DNDC’ biogeochemistry model. In addition, recovered genomes were grouped into ‘functional groups’, i.e. the genomes that perform a specific function of interest, then used to parameterize maximum growth rate and optimum growth temperature (via signatures in their sequence composition) for the BioCrunch model. The BioCrunch model was then in turn used to test the impact of increasing functional resolution of the microbes, on the carbon gas emissions. Lastly for modeling, the ecosys model was parameterized from the microbial and other data, and used to evaluate drivers of e.g. change in methane emissions. Finally, this project also led to the development of a range of new methods and tools, a new metric of organic matter decomposability, as well as a graph-database solution to multidisciplinary data storage and querying. This project’s ongoing analyses at our focal site also contributed to broader advancements in understanding elements of genetic plasticity and methane metabolism, climate change microbiology and community assembly, global peatland geochemistry and Arctic lakes’ roles in climate feedbacks.

54 ENVIRONMENTAL SCIENCES↗

Topologies on directed graphs

Given a directed graph, a natural topology is defined and relationships between standard topological properties and graph theoretical concepts are studied. In particular, the properties of connectivity and separatedness are investigated. A metric is introduced which is shown to be related to separatedness. The topological notions of continuity and homeomorphism. A class of maps is studied which preserve both graph and topological properties. Applications involving strong maps and contractions are also presented.

Lieberman, R. N.↗

An automated procedure built on MTEX for reconstructing deformation twin hierarchies from electron backscattered diffraction datasets of heavily twinned microstructures

Here we present a set of algorithms built on the MTEX and MATLAB graph toolboxes for automatic reconstruction of deformation twin hierarchies from Electron Backscatter Diffraction (EBSD) datasets with a focus on developing methods for heavily twinned microstructures (twin fractions >0.5). The algorithms address key issues arising at large strains, mainly: missing twin relationships, grouping of heavily deformed grain fragments into families of similar orientation originating from a single initial grain, identification of parent fragments for large twin volume fractions, and classification of families having twin relationships with multiple families. To facilitate the development of these algorithms, large-grained ultra-high purity α-Ti deformed in compression along two directions is investigated. Graphs are utilized to handle non-local geometric merging and to represent relationships throughout the reconstruction process. When determining if a grain fragment is from the undeformed microstructure, the combined metrics of the fragment's orientation volume fraction in the initial texture and the directed graph centrality measure of out-closeness (the number of nodes reached in a graph from a given node) are essential. To address automation in reconstructing the sequence of twinning and relating fragments originating from a single grain in the initial microstructure, the twin family tree is formulated as a minimum spanning tree emanating from the initial grain family. A scheme constructing the distances associated with twin relationship comprising the spanning tree is developed, and a novel quasi-directional Prim spanning tree algorithm is used to determine the twin family tree. The procedure is demonstrated to significantly improve the level of automation in reconstructing twin hierarchies in heavily twinned microstructure compared to other methodologies in literature. The procedure can readily be applied to analyses of twinning in metals, as well as provide an approach for routinely extracting twin statistics at larger deformation levels than previously possible. Significantly, the procedure is demonstrated to be capable of identifying third generation twinning in α-Ti microstructures.

36 MATERIALS SCIENCE↗

NaroNet: Discovery of tumor microenvironment elements from highly multiplexed images

Understanding the spatial interactions between the elements of the tumor microenvironment -i.e. tumor cells. fibroblasts, immune cells- and how these interactions relate to the diagnosis or prognosis of a tumor is one of the goals of computational pathology. We present NaroNet, a deep learning framework that models the multi-scale tumor microenvironment from multiplex-stained cancer tissue images and provides patient-level interpretable predictions using a seamless end-to-end learning pipeline. Trained only with multiplex-stained tissue images and their corresponding patient-level clinical labels, NaroNet unsupervisedly learns which cell phenotypes, cell neighborhoods, and neighborhood interactions have the highest influence to predict the correct label. To this end, NaroNet incorporates several novel and state-of-the-art deep learning techniques, such as patch-level contrastive learning, multi-level graph embeddings, a novel max-sum pooling operation, or a metric that quantifies the relevance that each microenvironment element has in the individual predictions. We validate NaroNet using synthetic data simulating multiplex-immunostained images where a patient label is artificially associated to the -adjustable- probabilistic incidence of different microenvironment elements. We then apply our model to two sets of images of human cancer tissues: 336 seven-color multiplex-immunostained images from 12 high-grade endometrial cancer patients; and 382 35-plex mass cytometry images from 215 breast cancer patients. In both synthetic and real datasets, NaroNet provides outstanding predictions of relevant clinical information while associating those predictions to the presence of specific microenvironment elements.

60 APPLIED LIFE SCIENCES↗

Experimental Observations of the Topology of Convolutional Neural Network Activations

Topological data analysis (TDA) is a branch of computational mathematics, bridging algebraic topology and data science, that provides compact, noise-robust representations of complex structures. Deep neural networks (DNNs) learn millions of parameters associated with a series of transformations defined by the model architecture resulting in high-dimensional, difficult to interpret internal representations of input data. As DNNs become more ubiquitous across multiple sectors of our society, there is increasing recognition that mathematical methods are needed to aid analysts, researchers, and practitioners in understanding and interpreting how these models' internal representations relate to the final classification. In this paper we apply cutting edge techniques from TDA with the goal of gaining insight towards interpretability of convolutional neural networks used for image classification. We use two common TDA approaches to explore several methods for modeling hidden layer activations as high-dimensional point clouds, and provide experimental evidence that these point clouds capture valuable structural information about the model's process. First, we demonstrate that a distance metric based on persistent homology can be used to quantify meaningful differences between layers and discuss these distances in the broader context of existing representational similarity metrics for neural network interpretability. Second, we show that a mapper graph can provide semantic insight as to how these models organize hierarchical class knowledge at each layer. These observations demonstrate that TDA is a useful tool to help deep learning practitioners unlock the hidden structures of their models.

topological data analysis, deep learning↗