Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graph analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

An ontology-based knowledge graph for representing interactions involving RNA molecules

The "RNA world" represents a novel frontier for the study of fundamental biological processes and human diseases and is paving the way for the development of new drugs tailored to each patient's biomolecular characteristics. Although scientific data about coding and non-coding RNA molecules are constantly produced and available from public repositories, they are scattered across different databases and a centralized, uniform, and semantically consistent representation of the "RNA world" is still lacking. We propose RNA-KG, a knowledge graph (KG) encompassing biological knowledge about RNAs gathered from more than 60 public databases, integrating functional relationships with genes, proteins, and chemicals and ontologically grounded biomedical concepts. To develop RNA-KG, we first identified, pre-processed, and characterized each data source; next, we built a meta-graph that provides an ontological description of the KG by representing all the bio-molecular entities and medical concepts of interest in this domain, as well as the types of interactions connecting them. Finally, we leveraged an instance-based semantically abstracted knowledge model to specify the ontological alignment according to which RNA-KG was generated. RNA-KG can be downloaded in different formats and also queried by a SPARQL endpoint. A thorough topological analysis of the resulting heterogeneous graph provides further insights into the characteristics of the "RNA world". RNA-KG can be both directly explored and visualized, and/or analyzed by applying computational methods to infer bio-medical knowledge from its heterogeneous nodes and edges. The resource can be easily updated with new experimental data, and specific views of the overall KG can be extracted according to the bio-medical problem to be studied.

59 BASIC BIOLOGICAL SCIENCES↗

Open Source Software Prevalence Ingest Tool

The OSSP Ingest Tool accepts user-input organizational information, ingests IT/OT asset lists in Excel format, and ingests the associated CycloneDX SBOM's. It then performs analytics demonstrating the ability to answer the follow research questions: o RQ1. Ability to identify all OSS services running on, and all OSS components present within, an OT device o RQ1a: Ability to differentiate multiple versions of the same OSS component within each OT device. o RQ1b: Ability to differentiate running from not-running OSS components. o RQ1c: Ability to differentiate based on the originator of the component, because a supplier may have modified it after retrieval from the upstream software source. o RQ2. Ability to correlate the identity of a single OSS component across multiple OT devices, mitigating common name variations such as differences in capitalization, '-' vs '_', and so on. o RQ3. Ability to perform subset analysis of OSS components across multiple OT devices o RQ3a: Ability to perform subset analysis across OSS libraries, generating density & distribution graphs to identify commonly-used libraries and outliers. o RQ3b: Ability to perform subset analysis of a single OSS library, generating density & distribution by CI sector, by device type, by device make/model, and/or by firmware version. o RQ3c: Ability to perform subset analysis by grouping OSS libraries according to programming language, then overlay with RQ4b. o RQ3d: Ability to perform subset analysis by OSS upstream source, providing insight into degree of modifications performed by suppliers. o RQ4. Ability to identify dependencies (transitive and direct) of each differentiated OSS library within each OT device, and enable RQ1,2,3 iteratively for dependencies. o RQ1. Ability to identify all OSS services running on, and all OSS components present within, an OT device o RQ1a: Ability to differentiate multiple versions of the same OSS component within each OT device. o RQ1b: Ability Page

Kapadia, Shayna [Lawrence Livermore National Labor↗

Temporal Analysis and Scene Change Detection in Multispectral Overhead Imagery

Scene change detection can be a tedious and time consuming process especially when concerning large geographical areas, and the process can be even more cumbersome when analyzing changes in an area over large spans of time. Developing a useful way to help analysts recognize at what points in time significant changes to a scene have occurred can allow them to better focus their efforts in characterizing events. Applications include: Facility monitoring, Construction chronology, Monitoring of vehicle/aircraft activity, Characterization of larger sequences of events. In large areas exceeding hundreds to thousands of square kilometers in size, it can be difficult localizing when scene changes have occurred. Analysts can spend hours going through imagery to try to identify new construction, monitor facility activities, monitor vehicle movement, etc. where the object of interest may only be a few square meters. Our goal is to help cut down this time by giving analysts change maps with hot spots of change, allowing them to focus on regions that have experienced actual change in time frames they're interested in. Additionally, by combining these change maps into layers within a data cube, analysts can examine the change maps from a temporal perspective, allowing events to be characterized over spans of time. By opening the data cube in an imaging software capable of separating the layers, we can analyze the change maps sequentially, allowing us to examine scene changes occurring over time. As an example, we examined overhead imagery from Planet Labs of what appears to be a parking lot on Fort Irwin over the course of a year using ENVI, a geospatial satellite imagery analysis software. Using ENVI, we generate a graph of changes over time, and notice a particular segment near the end of our analysis window where no changes are detected. Examination of the actual satellite imagery reveals that during this time span, the parking lot was empty. This could be due to facility shutdown for maintenance or upgrades, or possibly even total workforce/vehicle fleet movement. Information like this could help analysts better characterize events, as well as to help create clearer timelines in larger sequences of events. Workflow steps: - Collect multiple maps of the same AOI (Area of Interest) during a time span of interest; - Generate change maps from AOI maps; - Generate data cube from change maps. An analyst can use the data cube to help inspect an AOI for activities within a time span of interest. If an event of interest is discovered, the analyst can then refer to the maps corresponding to the appropriate dates and times in the data cube to see precisely what is transpiring. The biggest objective being worked on is improving the change detection methodology employed. We currently use PCA-EM (Principal Component Analysis with Expectation Maximization), but we are currently focusing on implementing IR-MAD (Iteratively Reweighted Multivariate Alteration Detection) to be used in conjunction with PCA-EM in an effort to decrease false positivity and noise in the change maps we generate.

42 ENGINEERING↗

Spatiotemporal Characteristics and Propagation of Summer Extreme Precipitation Events over United States: A Complex Network Analysis

Complex Network (CN) is a graph-theory based depiction of relation shared by various elements of a complex-dynamical system such as the atmosphere. Here we applied the concept of CN to understand the directionality and topological structure of summer extreme precipitation events (SEPEs) over the conterminous United States (CONUS). The SEPEs are calculated based on the 95th percentile daily rainfall at 0.5ox0.5o spatial resolution for CONUS to investigate the multi-dimensional characteristics of precipitation extremes. The derived CN coefficients (e.g., betweenness centrality, clustering coefficient, orientation, and network divergence) reveal important structural and dynamical information about the topology of the SEPEs and improve understanding of the dominant meteorological patterns. The initiation and propagation of SEPEs from the source-zones to the sink-zones are identified. The SEPEs are influenced by topography, dominant wind patterns, and moisture sources in terms of their topological structure and spatial dynamics.

Mondal, Somnath↗

Neuromorphic Graph Algorithms

Graph algorithms enable myriad large-scale applications including cybersecurity, social network analysis, resource allocation, and routing. The scalability of current graph algorithm implementations on conventional computing architectures are hampered by the demise of Moore’s law. We present a theoretical framework for designing and assessing the performance of graph algorithms executing in networks of spiking artificial neurons. Although spiking neural networks (SNNs) are capable of general-purpose computation, few algorithmic results with rigorous asymptotic performance analysis are known. SNNs are exceptionally well-motivated practically, as neuromorphic computing systems with 100 million spiking neurons are available, and systems with a billion neurons are anticipated in the next few years. Beyond massive parallelism and scalability, neuromorphic computing systems offer energy consumption orders of magnitude lower than conventional high-performance computing systems. We employ our framework to design and analyze new spiking algorithms for shortest path and dynamic programming problems. Our neuromorphic algorithms are message-passing algorithms relying critically on data movement for computation. For fair and rigorous comparison with conventional algorithms and architectures, which is challenging but paramount, we develop new models of data-movement in conventional computing architectures. This allows us to prove polynomial-factor advantages, even when we assume a SNN consisting of a simple grid-like network of neurons. To the best of our knowledge, this is one of the first examples of a rigorous asymptotic computational advantage for neuromorphic computing.

97 MATHEMATICS AND COMPUTING↗

Performance and usability enhancements for continuous subgraph matching queries on graph-structured data

A query graph, which includes vertices and edges, represents a query on graph-structured data. The query graph is decomposed into query subgraphs. A network analysis tool performs continuous subgraph matching queries to facilitate analysis of computer network traffic, social media events, or other streams of data represented as a dynamic data graph (graph-structured data). This can help identify emerging trends in the data. Some features of the network analysis tool enhance performance by effectively utilizing distributed computing resources (including processing cores and memory at different nodes of a cluster) to speed up the process of updating the dynamic data graph and detecting matches of query subgraphs. Features of a query graph building tool enhance usability by providing intuitive ways to specify query graphs and their subgraphs. Features of a results visualization tool enhance usability by providing an intuitive way to present the results of continuous subgraph matching queries.

Choudhury, Sutanay↗

A low-latency graph computer to identify metastable particles at the Large Hadron Collider for real-time analysis of potential dark matter signatures

Abstract Image recognition is a pervasive task in many information-processing environments. We present a solution to a difficult pattern recognition problem that lies at the heart of experimental particle physics. Future experiments with very high-intensity beams will produce a spray of thousands of particles in each beam-target or beam-beam collision. Recognizing the trajectories of these particles as they traverse layers of electronic sensors is a massive image recognition task that has never been accomplished in real time. We present a real-time processing solution that is implemented in a commercial field-programmable gate array using high-level synthesis. It is an unsupervised learning algorithm that uses techniques of graph computing. A prime application is the low-latency analysis of dark-matter signatures involving metastable charged particles that manifest as disappearing tracks.

47 OTHER INSTRUMENTATION↗

Benchmarking the PCMCI Causal Discovery Algorithm for Spatiotemporal Systems

Causal discovery algorithms construct hypothesized causal graphs that depict causal dependencies among variables in observational data. While powerful, the accuracy of these algorithms is highly sensitive to the underlying dynamics of the system in ways that have not been fully characterized in the literature. In this report, we benchmark the PCMCI causal discovery algorithm in its application to gridded spatiotemporal systems. Effectively computing grid-level causal graphs on large grids will enable analysis of the causal impacts of transient and mobile spatial phenomena in large systems, such as the Earth’s climate. We evaluate the performance of PCMCI with a set of structural causal models, using simulated spatial vector autoregressive processes in one- and two-dimensions. We develop computational and analytical tools for characterizing these processes and their associated causal graphs. Our findings suggest that direct application of PCMCI is not suitable for the analysis of dynamical spatiotemporal gridded systems, such as climatological data, without significant preprocessing and downscaling of the data. PCMCI requires unrealistic sample sizes to achieve acceptable performance on even modestly sized problems and suffers from a notable curse of dimensionality. This work suggests that, even under generous structural assumptions, significant additional algorithmic improvements are needed before causal discovery algorithms can be reliably applied to grid-level outputs of earth system models.

54 ENVIRONMENTAL SCIENCES↗

Correction to Unveiling the Role of Al 2 O 3 in Preventing Surface Reconstruction During High-Voltage Cycling of Lithium-Ion Batteries

A correction to this article was necessary to replace the original Figure 2 with the actual data in the revised Figure 2 shown here. An error was made while preparing the graphs from the neutron diffraction analysis software, and panel a was accidentally inserted into the panel b–d spots, such that all four panels were replicates. Since the original analysis of the data collected was made correctly, the conclusions and key points related to this section and the whole article remain unchanged. The revised goodness factors are slightly different because the refinements were repeated with small changes to the refinement conditions. The corresponding author has obtained approval from all coauthors in the article prior the submission of the Correction. The authors apologize for any inconveniences caused.

25 ENERGY STORAGE↗

Optimal D-FACTS Placement in Moving Target Defense Against False Data Injection Attacks

Moving target defense (MTD) is a defense strategy to detect stealthy false data injection (FDI) attacks against the power system state estimation using distributed flexible AC transmission system (D-FACTS) devices. However, existing studies neglect to address a fundamental yet critical issue, i.e., the D-FACTS placement, by assuming that all lines are equipped with D-FACTS devices. Here, to tackle this problem, we first derive analytical necessary conditions and requirements on the D-FACTS placement for a complete MTD. Further, we propose sufficient conditions using a graph theory-based topology analysis to ensure that the MTD under the proposed D-FACTS placement has the maximum rank of its composite matrix, which is indicative of the MTD effectiveness. Based on the analytical conditions, we design D-FACTS placement algorithms by using the minimum number of D-FACTS devices to achieve the maximum MTD effectiveness. A novel MTD-based ACOPF model, in which the reactance of D-FACTS lines is introduced as decision variables, is proposed to find a trade-off between the system loss and the MTD effectiveness. Numerical results on 6-bus, IEEE 14-bus, and IEEE 118-bus systems show the efficacy of MTDs using the proposed D-FACTS placement algorithms in maximizing the composite matrix rank and detecting FDI attacks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

DISCOverflow

Procedure for reverse engineering binary files and storing the result in a graph database structure for later analysis. This is achieved through the use of Python, angr, and OrientDB Community Edition.

Beckman, BryanR↗

Simplifying and Visualizing the Ontology of Systems Engineering Models

The credibility of an engineering model is of critical importance in large-scale projects. How concerned should an engineer be when reusing someone else's model when they may not know the author or be familiar with the tools that were used to create it? In this report, the authors advance engineers' capabilities for assessing models through examination of the underlying semantic structure of a model--the ontology. This ontology defines the objects in a model, types of objects, and relationships between them. In this study, two advances in ontology simplification and visualization are discussed and are demonstrated on two systems engineering models. These advances are critical steps toward enabling engineering models to interoperate, as well as assessing models for credibility. For example, results of this research show an 80% reduction in file size and representation size, dramatically improving the throughput of graph algorithms applied to the analysis of these models. Finally, four future problems are outlined in ontology research toward establishing credible models--ontology discovery, ontology matching, ontology alignment, and model assessment.

42 ENGINEERING↗

Knowledge Graph for End-to-End Traceability of an Integrated Human-Earth System Model

Integrated human-Earth system models inform energy-water-land system dynamics and policies, yet their results are difficult to trace through input-data, model structure, scenario configurations, and solved outputs. Because this information is siloed across disconnected artifacts, process-based IAMs have historically lacked a unified, queryable representation. Such lack of traceability prevents researchers from systematically isolating the multi-sector drivers of complex outcomes (such as tracing water-scarcity results back to distant energy-system dynamics) or conducting holistic uncertainty attribution across hundreds of interacting parameters. To address this concern, our work documents the software engineering process of a knowledge graph that unifies these four layers for the Global Change Analysis Model (GCAM-USA_Reference scenario, GCAM v9.1). The graph was built as a relational property graph in DuckDB from the run’s own artifacts: the input-preparation dependency map (gcamdata chunk map), the model’s XML input files, the run configuration, and the results database (BaseX), successfully mapping the model’s declared structure. The resulting graph comprises 204,321 nodes and 1,687,814 edges across 16 node types and 15 edge types, with approximately 16.3 million time-series values stored separately to maintain structural efficiency. To ensure representation fidelity, every edge carries an epistemic-status annotation recording the warrant for the relationship (structural, provenance, dependency, or model-derived), and a machine-readable provenance ledger classifying the origin of every schema element. Evaluation against a fixed five-benchmark suite with locked baselines reports zero structural orphans, zero dangling edge endpoints, and 100% of output-producing technologies traceable to raw input files. Two interactive interfaces present the graph, including a serverless browser application built on DuckDB-Wasm. By establishing the first end-to-end provenance framework for an IAM, this work enables researchers and scientists to systematically audit complex policy scenarios, debug model structures, and trace policy-relevant outputs to their data origins in real time.

Artifical Intelligence↗

Evaluation of Graph Analytics Frameworks Using the GAP Benchmark Suite

The analysis of connected data is an increasingly important application in high-performance computing. Such analyses can reveal fraudulent patterns in financial transactions, optimize telecommunications networks, predict information flow in social networks, etc. However, the landscape of graph analytics is highly diverse. Graph algorithms stress processor architectures differently, and no one graph can represent all topologies. Consequently, no single approach or framework is expected to be optimal for all graph analytics problems. To help make sense of this diverse landscape, we evaluated four approaches to graph analytics: GraphBLAS, Galois, BGL17, GraphIt; and compare them against hand-tuned implementations that take advantage of hardware features on our test platform. Graph- BLAS formulates graph analytics as sparse linear algebra. Galois provides syntactic constructs for data parallelism over irregular data structures. BGL17 is a generic C++ template library for implementing graph algorithms. GraphIt provides a domain- specific language to describe and optimize graph algorithms. We use the GAP Benchmark Suite to establish baseline performance and guide the side-by-side evaluation of each framework. GAP consists of 30 tests: six graph analytics algorithms (breadth- first search, single-source shortest path, PageRank, betweenness centrality, connected components, and triangle counting) run on five graphs, each with different topological characteristics (e.g., high diameter, skewed degree distribution, high average degree). High-performance reference implementations are included for each benchmark algorithm. Because a graph can be loaded into memory a number of ways (e.g., flat file on disk, compressed sparse format, data frames, retrieved from SQL or NoSQL databases), our evaluation focused on computational performance rather than I/O. Our results show the relative strengths of each framework.

Graph algorithms, Benchmarking, shared-memory prog↗

On the integration of molecular dynamics, data science, and experiments for studying solvent effects on catalysis

Computational workflows that combine molecular dynamics (MD) simulations and emerging data-centric (DC) methods can accelerate the screening and analysis of solvent systems experimentally and computationally. Here, MD simulations provide atomic positions and velocities of reactant, solvent, and catalyst materials that can be manipulated into data representations that in turn can be used by DC techniques to conduct predictive modeling, feature extraction, and experimental design. For liquid-phase catalytic applications, emerging DC techniques such as Convolutional and Graph Neural Networks (CNN/GNN), Topological Data Analysis (TDA), and Active Learning (AL) can leverage MD and experimental data to quickly predict solvent effects on reaction outcomes. For instance, in recent studies, 3D solvent environments obtained with MD have been exploited by CNNs to predict experimental reaction rates for homogeneous acid-catalyzed lignocellulosic processes. In this perspective, we discuss basic principles of DC methods and how these can be combined with MD to enable high-throughput screening of solvent selection for diverse catalysis applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

PersGNN: Applying Topological Data Analysis and Geometric Deep Learning to Structure-Based Protein Function Prediction

Understanding protein structure-function relationships is a key challenge in computational biology, with applications across the biotechnology and pharmaceutical industries. While it is known that protein structure directly impacts protein function, many functional prediction tasks use only protein sequence. In this work, we isolate protein structure to make functional annotations for proteins in the Protein Data Bank in order to study the expressiveness of different structure-based prediction schemes. We present PersGNN - an end-to-end trainable deep learning model that combines graph representation learning with topological data analysis to capture a complex set of both local and global structural features. While variations of these techniques have been successfully applied to proteins before, we demonstrate that our hybridized approach, PersGNN, outperforms either method on its own as well as a baseline neural network that learns from the same information. PersGNN achieves a 9.3% boost in area under the precision recall curve (AUPR) compared to the best individual model, as well as high F1 scores across different gene ontology categories, indicating the transferability of this approach.

Swenson, Nicolas↗