Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graph databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Capturing Historic Reliability Performance Through Graph Databases: A Model Based System Engineering Approach

With the goal of improving the performance and reliability of high dependable technological systems such as nuclear power plants, advanced monitoring and health management systems are employed to inform system engineers on observed degradation processes and anomalous behaviors of assets and components. This information is captured in the form of large amount of data which can be heterogenous in nature (e.g., numeric, textual). Such large data availability poses challenges when system engineers are required to parse and analyze them in order to track historic reliability performance of assets and components. This paper tackles directly this challenge by providing means to organize data in the form of a graph: a knowledge graph. The presented approach distinguish itself from current knowledge graph-based methods by the fact that model-based system engineering (MBSE) models are used to “put data into context”. In particular, MBSE models are used as skeleton of a knowledge graph; numeric and textual data elements, once processed, are associated to MBSE model elements. Thus, a knowledge graph captures both system architecture (though MBSE models) and health/performance data. Such feature opens the door to new data analytics methods designed to identify causal relations between observed phenomena.

97 - MATHEMATICS AND COMPUTING↗

Evidence-based Graph Adversary Mapping (EGRAM) [Poster]

Cybersecurity companies such as CrowdStrike, Dragos, Microsoft and Unit 42 categorize Advanced Persistent Threats (APTs) using their own naming schemes. As a result, these APTs are mapped to different malware sources and campaigns, all from differing sources, leading to inconsistent mapping. Inconsistent mapping causes confusion and adds further obscurity around these groups, making it difficult to track and mitigate APT cyberattacks. The Evidence-based Graph Adversary Mapping (EGRAM) tool remediates the mapping challenge by collecting, updating and converting adversary data and their sources into a valid, codified STIX v2.1 bundle which is then stored in a Neo4j graph database. It utilizes graph traversal methods and centrality analysis to generate actionable information as a Structured Threat Intelligence Graph (STIG), based on user queries. EGRAM exists as Python code and a Jupyter Notebook that acts as a searchable, evidence-based, source of intelligence for APT groups’ artifacts and cyber campaigns.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Collection And Analysis Of Telemetry For The Cyote Heuristic

CATCH CLI focuses on gathering telemetry data, storing it in the Neo4j database, querying for Mitre ATT&CK patterns, and creating STIX 2.1 reports. Key Components: Analysis Modules: Analyze data to detect attack patterns. GoSTOTS Collection Engines: Collect telemetry data. These tools can be used together or individually. Analysis modules rely on data from specific engines to identify attack patterns. Source Code Organization: Engines: CATCH/catch/cmd/collection Modules: CATCH/catch/cmd/analysis CGUI Overview CATCH Graphical User Interface (CGUI) offers a graphical shell to execute CATCH CLI, allowing easy editing of: Analysis Modules Database configurations Profiles (collection and device settings) Neo4j Overview Neo4j is a graph database using the Cypher query language, storing data in JSON. It seamlessly integrates with STIX 2.1 data for: Data Submission: CATCH Collection Engines Data Querying: Analysis Modules CATCH modifies STIX 2.1 data for Neo4j submission and reverts it back during querying. STIG Overview Structured Threat Intelligence Graph (STIG) is a tool for creating, editing, querying, analyzing, and visualizing threat intelligence using STIX 2.1 and storing data in Neo4j. Usage Tools can be run: Manually (CLI): Refer to CATCH documentation User Interface: Run ./cgui/CGUI or go run ./cgui/ Additional Information Logging System: Detailed in the config documentation Further Documentation: Available for CATCH and CGUI

Madsen, MichaelJ. [Idaho National Laboratory (INL)↗

Intern Poster: STIG Shouldn't Drop ACID

STIG (Structured Threat Intelligence Graph) is an open-source graph database tool from INL. It’s used to create and process cyber intelligence graphs, which are shared in the cyber threat intelligence community and used to train INL machine learning products like @DisCo. For quality machine learning and critical infrastructure defense, STIG’s database must be ACID: Atomic, Consistent, Isolated, Durable. Various ACID tests were designed and applied to STIG to ensure its behavior follows these properties.

99 - GENERAL AND MISCELLANEOUS↗

DISCOverflow

Procedure for reverse engineering binary files and storing the result in a graph database structure for later analysis. This is achieved through the use of Python, angr, and OrientDB Community Edition.

Beckman, BryanR↗

idaholab/cape2stix

This software allows for the conversion, extraction, and transformation of malware behavior data from "Malware Configuration And Payload Extraction" (CAPEv2) sandbox reports, to Structured Threat Information eXpression (STIX). This allows for further analysis to be performed, sharing of threat data, and transit to a graph database.

Cutshaw, Michael [Idaho National Laboratory (INL),↗

Schema Grapher

PNNL contributions to project site managed by federal agency. Schema Grapher is an application that takes in CSV datasets and converts them to RDF for loading into a graph database. Schema Grapher is responsible for the extract and transform part of an ETL (i.e., extract, transform, load) pipeline.

Avila, Andrew↗

Illuminating the pathways to carbon liberation: a systems approach to characterizing the consequential unknowns of carbon transformation and loss from thawing permafrost peatlands (Final Report)

The IsoGenie3 Project delivered new systems-level insights into carbon cycling in thawing permafrost landscapes, with an emphasis on methane and carbon dioxide emissions. From >200 samples from the site collected over a decade, co-analyzed for geochemistry and microbiology, the team recovered ~1,500 assembled microbial genomes and ~1,900 viral population genomes, revealing appreciable genetic novelty - from a new highly abundant bacterial phylum, to novel methane consumers and their activities, to rampant viral novelty. IsoGenie3 linked these organisms to carbon compound transformations (which define the cycling of organic matter in soils, and the loss of the greenhouse gases carbon dioxide and methane), and saw that the microbes at each stage of permafrost thaw had different genetic potential to degrade categories of compounds, expressed that genetic potential differently, and actually transformed carbon compounds into greenhouse gases in different ways. IsoGenie 3 identified that some of the thaw-stage differences were due to plant-microbiome relationships; the plant species across the thaw gradient contributed different carbon compounds into the soil, and hosted distinct microbiota (differing among parts of plants as well as species). Lastly, microbes in the saturated post-thaw conditions appeared likely to contribute to the mobilization and toxification of mercury released during thaw. In parallel with ongoing field sampling and analysis, hypotheses arising from field observations were tested via lab incubation experiments. When communities are taken out of their native habitats, they behave differently, and the team first rigorously quantified the magnitude of this effect on microbiome composition and functional capacity, organic matter composition, and gas production; overall the main system processes were maintained in the lab incubations under the conditions tested. Further, the microbial data could inform geochemical reaction network models of those processes. Then, the team ran experiments with additions of compounds, varying temperature, and “live” vs. “dead” peat (the latter having been gamma irradiated, with a few additional variants to control for methodological artifacts). From these, we (a) determined the importance of plant-derived soluble phenolic compounds in bogs’ extraordinary recalcitrance of organic matter, and carbon gas emissions skewed to carbon dioxide; (b) proposed an abiotic ‘tanning’ mechanism, which could contribute to Sphagnum’s inhibitory effect on anaerobic decomposition through alteration of N availability. IsoGenie3 illuminated longer-term and landscape-scale interactions of permafrost thaw and carbon cycling, advancing knowledge of the drivers of methane dynamics not only across in the permafrost-associated peatland (where hydrology and plant communities dictate microbiomes) but also their interconnected lakes (where sediment carbon quality and resident microbiota are determined by position within lake, and lake features). By leveraging observations of site methane dynamics extending well before this project, the team was able to construct a 44-year portrait of the interplay of permafrost thaw, hydrology, vegetation dynamics, and carbon gas emissions, and the doubling of the fully-thawed fens over this time. From the detailed study of this focal site, IsoGenie3 also aimed to improve model representation of these kinds of sites and processes. To improve predictions of methane transformations, we incorporated acetate and isotope dynamics into the ‘DNDC’ biogeochemistry model. In addition, recovered genomes were grouped into ‘functional groups’, i.e. the genomes that perform a specific function of interest, then used to parameterize maximum growth rate and optimum growth temperature (via signatures in their sequence composition) for the BioCrunch model. The BioCrunch model was then in turn used to test the impact of increasing functional resolution of the microbes, on the carbon gas emissions. Lastly for modeling, the ecosys model was parameterized from the microbial and other data, and used to evaluate drivers of e.g. change in methane emissions. Finally, this project also led to the development of a range of new methods and tools, a new metric of organic matter decomposability, as well as a graph-database solution to multidisciplinary data storage and querying. This project’s ongoing analyses at our focal site also contributed to broader advancements in understanding elements of genetic plasticity and methane metabolism, climate change microbiology and community assembly, global peatland geochemistry and Arctic lakes’ roles in climate feedbacks.

54 ENVIRONMENTAL SCIENCES↗

Digital Twin for Optimizing Real-time Economy of the Integrated Energy Systems

Economic and safe operation of integrated energy systems (IES) requires real-time optimization (RTO) of the control and actions conducted on each system component. In this regard, digital twins (DTs), which consist of a physical system, a virtual system, and the data communication that occurs between the two, are essential for effective RTO. Through the data warehouse, the virtual system is constantly updated with real-time data from the physical system, and functions as the model in the optimization framework. The reduced-order model of the dynamic process model in the virtual system is used in the optimization framework. The optimization results are then returned, via the data warehouse, as control actions to the physical system. This work demonstrates the software capabilities of DT assets for an IES in the context of preparing a DT for an experimental system comprised of Idaho National Laboratory (INL)’s Thermal Energy Delivery System and battery system. For the virtual demonstration, the DTs encompass (1) a physical system, including the Modelica models of the Thermal Energy Delivery System and the battery system; (2) virtual optimization via the Optimization of Real-Time Capacity Allocation (ORCA) platform; and (3) the open-source data warehouse software DeepLynx. This work assesses the performance of ORCA, which utilizes a reduced-order model built using the Risk Analysis Virtual Environment (RAVEN) and trained on the Modelica models and real-time data pipeline through the graph database hosted in DeepLynx. The proposed optimization workflow will be an RTO model based on DTs and the data they generate.

25 ENERGY STORAGE↗

Session Introduction: Graph Representations and Algorithms in Biomedicine

Connectivity is a fundamental property of biological systems: on the cellular level, proteins interact with each other to form protein-protein interaction networks (PPIs); on the organism level, neurons are arranged in a network; and on a community-level, species can have complex relationships with one another that drive the development and balance of an ecosystem. Graphs, representations of systems consisting of entities as vertices and their connections as edges, are a useful structure to characterize many such systems. Such models can be used to understand biological systems that naturally have a network structure, including PPIs, biological neurons, and ecosystems. In today’s information age, graph representations and algorithms (often in combination with machine learning techniques) are used to organize massive amounts of related data, much of which may be heterogeneous or unstructured, and identify patterns that represent novel biological insights. PSB’s 2023 session “Graph Representations and algorithms in Biomedicine,” encompasses modern developments in graph theory and its applications to various fields of biomedicine. This session includes a wide range of research - knowledge graphs built from text-mined health data, heterogeneous networks using multi-omic databases, and graphs refined to represent uncertainty or improve memory usage.

Chrisman, Brianna S.↗

Acceleration of Graph Neural Network-Based Prediction Models in Chemistry via Co-Design Optimization on Intelligence Processing Units

Atomic structure prediction and associated property calculations are the bedrock of chemical physics. Since high-fidelity ab initio modeling techniques for computing the structure and properties can be prohibitively expensive, this motivates the development of machine-learning (ML) models that make these predictions more efficiently. Training graph neural networks over large atomistic databases introduces unique computational challenges such as the need to process millions of small graphs with variable size and support communication patterns that are distinct from learning over large graphs such as social networks. We demonstrate a novel hardware-software co-design approach to scale up the training of atomistic graph neural networks (GNN) for structure and property prediction. First, to eliminate redundant computation and memory associated with alternative padding techniques and to improve throughput via minimizing communication, we formulate the effective coalescing of the batches of variable-size atomistic graphs as the bin packing problem and introduce a hardware-agnostic algorithm to pack these batches. In addition, we propose hardware-specific optimizations including a planner and vectorization for the gather-scatter operations targeted for Graphcore’s Intelligence Processing Unit (IPU), as well as model-specific optimizations such as merged communication collectives and optimized softplus. Putting these all together, we demonstrate the effectiveness of the proposed co-design approach by providing an implementation of a well-established atomistic GNN on the Graphcore IPUs. We evaluate the training performance on multiple atomistic graph databases with varying degrees of graph counts, sizes and sparsity. Here, we demonstrate that such a co-design approach can reduce the training time of atomistic GNNs and can improve the performance by up to 1.5× compared to the baseline implementation of the model on the IPUs. Additionally, we compare our IPU implementation with a Nvidia GPU-based implementation and show that our atomistic GNN implementation on the IPUs can run 1.8× faster on average compared to the execution time on the GPUs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

SEAFORML (Smart Exploration and Analysis For Optimal and Robust Machine Learning)

The poster discusses data analysis of the WAVgraph database and applied machine learning methods for it. The database is a long-term project that seeks to be a comprehensive repository of information on cyber threats and is updated regularly. It was previously unanalyzed and unexplored. The goal was to learn more about it and its contents in order to have a better understanding and enable better use. The data analysis and discovery enabled further exploration through natural language processing, similarity, and clustering methods. The poster shows some of the insights from the analysis and explains the methods used for the machine learning applications.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Development of Whole System Digital Twins for Advanced Reactors: Leveraging Graph Neural Networks and SAM Simulations

Here, in this work, we introduce a novel method to develop whole system digital twins (DTs) for advanced nuclear reactors. This method treats a complex reactor system as a heterogeneous graph: with the system components as different types of graph nodes and their physical interconnections as edges. Based on the heterogeneous graph, a graph neural network combining graph convolution and temporal node attention is developed as the DT, facilitating a comprehensive understanding of the system's dynamic behavior. By utilizing the System Analysis Module (SAM) code for simulating various operational transients, we develop a graph-based database that trains the DT. This DT is characterized by two primary functions: It can infer the entire system's status using sparse node information, and it can predict the progress of transients based on current and historical system information. Our approach is validated through case studies on the Experimental Breeder Reactor II (EBR-II) system and a generic Fluoride-salt-cooled High-temperature Reactor (gFHR), demonstrating the DT's accuracy in forecasting operational transients. The DT's rapid computation capabilities enhance its potential for supporting advanced reactor operations, offering benefits in intelligent simulation, autonomous control, and anomaly detection, paving the way for improved safety analysis and intelligent component health management for advanced reactor systems and reducing their operations and maintenance cost.

EBR-II↗

Resilience Assessment Framework For Electric Distribution Systems Performance Under Extreme Conditions

The devastating impact of extreme weather-related events is increasingly evident on power grids, especially on distribution grids. The severity of their potential impact calls for 1) developing a suitable resilience assessment framework to capture the system performance and 2) assessing relevant mitigative strategies to lessen the impact of such events. This paper proposes a framework to identify grid vulnerabilities using the energy-at-risk concept to select, disconnect and isolate grid portions due to a resilience event. The proposed framework mainly consists of two steps; i) processing the utility's available infrastructure, i.e., a network model, possible switching combinations, and outage information for those combinations, as a graph-based database, and ii) implementing a novel optimal switching algorithm leveraging database and grid simulated metrics. These switching actions are generated to implement load curtailment in a rolling manner during anticipated grid scarcity conditions. In this study, a test case is created using two taxonomy feeders and is simulated against an extreme temperature event, e.g., a long, relatively cold, and prolonged freeze peak, thereby creating stress on the grid. It is demonstrated that the proposed framework allows utilities to predict the energy-at-risk during such resilience events and design suitable outage management strategies.

Poudel, Shiva↗

DNA parts and gene constructs for plant biodesign

Plant biodesign requires the knowledge of DNA parts (e.g., genes, promoters, terminators), along with their combinations (as gene constructs) linked to engineered traits. DNA parts with validated or predicted functions in plants have been deposited in various online databases. However, these existing databases focus on basic biological functions of individual DNA parts, leaving a gap between basic knowledge and bioengineering applications. To fill this knowledge gap, we have created a user-friendly, open-ended database as a knowledge graph linking DNA parts to gene constructs to traits. This database contains experimentally validated DNA parts and gene constructs documented in peer-reviewed publications. The DNA parts include 1) molecular components with biological functions, such as genes involved in various biological processes (e.g., metabolic and signal transduction pathways) and 2) molecular components with technical functions, such as gene expression, genome engineering and sequence splicing. The gene constructs deposited in this database include both single-gene and multi-gene constructs. This database allows users to submit DNA parts and gene construct compositions linked to engineered traits described in peer-reviewed publications, providing a public digital repository for sharing the biodesign information among the researchers in the fields of plant biotechnology and plant synthetic biology.

plant biodesign synthetic biology gene constructs ↗

System Engineers and Decisions: It?s All about Knowledge

In order to guarantee that a system meets adequate levels of reliability and availability, system performances are continuously monitored and analyzed thanks to the technological advancements driving the Industry 4.0 revolution. An Industry 4.0 approach is typically based on advanced statistical, big data mining, machine learning, and internet-of-things methods designed to detect anomalies in the behavior of system, detect the most likely failure modes, and provide indications to system engineers on when maintenance activities should be performed before system performance are deemed unacceptable (which can be generated by diagnostic and prognostic methods). However, these analyses, which are designed to automatize and increase the efficacy of the system maintenance program, require large amount of data which can come in various forms: numeric, textual, images, sounds etc. Such data constitutes the historic knowledge benchmark to track system performances and support system engineer decisions. Here we claim that data is not sufficient to support this kind of analyses when applied to systems characterized by complex architectures and behaviors. Robust system engineer decisions require the ability to understand the system operational context that lies behind the observed data elements. In this respect, system models are in fact necessary to “put data in context” and capture relationships between data elements. Industry 4.0 methods require in fact contextual knowledge as a basis upon which hypotheses can be generated and assumptions tested. In our view, for complex systems, model-based system engineering (MBSE) models can afford this contextual knowledge, as they are typically used to describe systems architecture and dynamic behaviors. System knowledge is here intended as the blending of collected data and system architecture which takes the form of a “knowledge graph”. A knowledge graph is a database which consists of a large set of nodes (in our case an entity can be either a data or an MBSE element) which are linked to each other. The types of nodes and links follow a pre-defined topology, sometimes also refers as an ontology, that is designed to fit the actual decisions that needs to be performed. We show here how a knowledge graph can be defined to support system engineer maintenance decisions and how the same graph can be built based on system MBSE models and pre-processed data from numeric (through anomaly detections and diagnostic methods) and textual elements (through technical language processing TLP).

97 - MATHEMATICS AND COMPUTING↗

Retrieval Augmented Generation for Robust Cyber Defense

In cybersecurity, the ability to efficiently analyze and respond to vulnerabilities, weaknesses, attack patterns, and threat tactics is critical for effective defense strategies. With the increasing complexity and volume of cybersecurity data, traditional methods of querying and retrieving information are often inadequate. To address this challenge, we implemented Retrieval-Augmented Generation (RAG) systems—CyRAG and GraphCyRAG—that integrate large language models (LLMs) with both structured data from relational databases and knowledge graphs such as Neo4j. CyRAG is designed to handle structured data, focusing on CVE (Common Vulnerabilities and Exposures) and CWE (Common Weakness Enumeration) entities to generate accurate and context-rich responses. In contrast, GraphCyRAG leverages Neo4j knowledge graphs to retrieve interconnected information from CVE, CWE, CAPEC (Common Attack Pattern Enumeration and Classification), and ATT&CK (Adversarial Tactics, Techniques, and Common Knowledge) datasets. By utilizing Neo4j’s graph-based framework, GraphCyRAG enables deeper traversal of relationships between vulnerabilities and attack patterns, providing cybersecurity analysts with more comprehensive insights into potential attack vectors and mitigation strategies. Our preliminary results demonstrate that integrating knowledge graphs with RAG significantly enhances both the accuracy and depth of threat analysis, allowing for the retrieval of dynamic, real-time data and the generation of contextually aware responses. This approach helps analysts uncover hidden relationships between cyber entities, predict exploit paths, and prioritize mitigation efforts effectively. The integration of RAG with cybersecurity knowledge graphs represents a significant advancement in cybersecurity threat intelligence, enabling more informed decision-making and stronger defense strategies.

97 MATHEMATICS AND COMPUTING↗