Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graph databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Automated Generation of Graph-based Cyber Threat Intel

With the advancement of AI technology and tools, specifically in the cybersecurity domain, both cyber defenders and threat actors are continuously adapting the use of these capabilities to expedite their operations. With this phenomenon, threat intelligence that is up to date, refreshable, and has relevant context to a specific threat becomes more and more important as it enables cybersecurity professionals to gain insight into relevant data and relationships to guide their operations. This project enables users to frequently aggregate threat intelligence from various sources, such as vendor vulnerability advisories affecting critical infrastructure, malware reports, and adversary writeups into a centralized, standardized database. The project utilizes the Structured Threat Intelligence eXpression (STIX) for a standardized, shareable threat intelligence data format and Neo4j as a graph database solution to store STIX nodes and relationships. Initial results of the project include datasets of over 8,000 nodes and 20,000 relationships extracted from over 500 data sources that have been released within the past month.

Threat Intelligence↗

Graph-based featurization methods for classifying small molecule compounds

For over a decade, drug-induced liver injury (DILI) has posed significant drawbacks in the synthesis and development of drugs and remains a consequential concern. With finite success within the existing preclinical models, DILI is one of the main causes of drug withdrawal or termination from the market. Particularly, this withdrawal occurs during the late stages of drug development (Kullak-Ublick, 2017). Since DILI is difficult to diagnose and treat, it has become an obstacle in the drug production market that in turn affects clinicians, pharmaceutical companies, and consumers. We propose a method for learning features of DILI-positive drugs based on the graphical relationships and patterns they possess within a network of biological databases. We also train various statistical and machine learning models on these learned features in order to classify the drugs as DILI-positive or negative. Our methods include Random Forest, Neural networks, and logistic regression classification. We utilize labeled DILI-positive and DILI-negative datasets, which were developed by the FDA and the National center for toxicological research, as well as additional literature datasets (Thakkar, 2020) in order to validate our results and assess our featurization and model accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

QLiG: Query Like a Graph For Subgraph Matching

A graph is a natural and flexible modeling approach to represent entities and relationships between them in real-world. A Knowledge Graphs (KG) is a specialized graph with formal and structured representation of facts, relationships, annotated with semantic descriptions. Subgraph matching is one of the fundamental graph problems to identify relationships, interactions and activities of interest within a large graph. A query specification is a collection of abstract components, operations, and constraints to express a pattern. The specification can be implemented in different ways based on underlying data model. Various graph query specifications have been developed over the years and have led to the development of different open-sourced and vendor-specific query languages. Such specification are modeled as an extension of relational algebra used to develop relational query languages such as SQL. Such relational concepts do not inherently support graph queries. There is a need to represent graph queries in terms on graph-based components to expedite query construction by non-database experts. We present a graph-based query approach QLiG (pronounced cleeg), to perform subgraph matching in Labeled Property Graph. We present the query specifications, salient features, and a use case to show functional examples.

Purohit, Sumit↗

Efficient QAOA Optimization using Directed Restarts and Graph Lookup

Variational Quantum Algorithms (VQA) aim to enhance the capabilities of Noisy Intermediate-Scale Quantum (NISQ) devices. These algorithms utilize parameterized circuits and classical optimizers to iteratively execute circuits with varying parameters. However, VQA faces computational overheads due to repeated iterations and random restarts. Prior work suggests using basic sub-graphs to transfer parameters for the input graph, reducing optimizer overheads but limiting applicability to structured regular graphs. In real-world applications, random irregular graphs are common, and existing methods are not scalable or practical for such graphs. This paper presents a framework that aims to improve random irregular graphs in VQA. The framework uses graph similarity and important features like total edge counts, average edge counts, and variance. It follows an iterative process to choose basis sub-graphs from a small database and adjust parameters accordingly. Classical optimizers then utilize these parameters to determine when to restart and perform gradient descent. This approach increases the chances of reaching global maximum points.

Wang, Meng↗

Exaflops Biomedical Knowledge Graph Analytics

We are motivated by newly proposed methods for mining large-scale corpora of scholarly publications (e.g., full biomedical literature), which consists of tens of millions of papers spanning decades of research. In this setting, analysts seek to discover relationships among concepts. They construct graph representations from annotated text databases and then formulate the relationship-mining problem as an all-pairs shortest paths (APSP) and validate connective paths against curated biomedical knowledge graphs (e.g., Spoke). In this context, we present Coast (Exascale Communication-Optimized All-Pairs Shortest Path) and demonstrate 1.004 EF/s on 9,200 Frontier nodes (73,600 GCDs). We develop hyperbolic performance models (HYPERMOD), which guide optimizations and parametric tuning. The proposed Coast algorithm achieved the memory constant parallel efficiency of 99% in the single-precision tropical semiring. Looking forward, Coast will enable the integration of scholarly corpora like PubMed into the Spoke biomedical knowledge graph.

Kannan, Ramakrishnan {ramki}↗

Enabling Efficient Sparse Computations using Linear Algebra Aware Compilers

This project developed the LAPIS compiler framework, built on the Multilevel Intermediate Representation (MLIR), to optimize sparse linear algebra operations and support performance portability across diverse architectures. The main innovation of LAPIS is the Kokkos dialect, which allows for lowering codes from a high productivity language to different architectures in an elegant way. The dialect also allows the conversion of lower-level MLIR code to C++ Kokkos code, facilitating the integration of scientific machine learning (SciML) models into applications. To extend LAPIS for distributed memory architectures, a new partition dialect was created to manage the distribution of sparse tensors and express communication patterns for sparse linear algebra operations. This dialect also supports the distributed execution of operators and includes algorithmic optimizations to minimize communication to improve performance. The project also demonstrates that MLIR can enable effective linear algebra-level optimizations, improving performance on different GPUs for both sparse and dense linear algebra kernels. Key applications of LAPIS include sparse linear algebra and graph kernels, TenSQL, a relational database management solution built on GraphBLAS, and the development of subgraph isomorphism and monomorphism kernels, showcasing performance portability. In summary, the LAPIS framework supports productivity, performance, portability, and distributed memory execution, while also enabling linear algebra-level optimizations that are challenging in traditional programming languages, with successful applications ranging from simple sparse linear algebra to complex graph kernels.

97 MATHEMATICS AND COMPUTING↗

Regularized machine learning on molecular graph model explains systematic error in DFT enthalpies

Abstract A major goal of materials research is the discovery of novel and efficient heterogeneous catalysts for various chemical processes. In such studies, the candidate catalyst material is modeled using tens to thousands of chemical species and elementary reactions. Density Functional Theory (DFT) is widely used to calculate the thermochemistry of these species which might be surface species or gas-phase molecules. The use of an approximate exchange correlation functional in the DFT framework introduces an important source of error in such models. This is especially true in the calculation of gas phase molecules whose thermochemistry is calculated using the same planewave basis set as the rest of the surface mechanism. Unfortunately, the nature and magnitude of these errors is unknown for most practical molecules. Here, we investigate the error in the enthalpy of formation for 1676 gaseous species using two different DFT levels of theory and the ‘ground truth values’ obtained from the NIST database. We featurize molecules using graph theory. We use a regularized algorithm to discover a sparse model of the error and identify important molecular fragments that drive this error. The model is robust to rigorous statistical tests and is used to correct DFT thermochemistry, achieving more than an order of magnitude improvement.

36 MATERIALS SCIENCE↗

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab↗

End-to-end optimization for battery materials and molecules by combining graph neural networks and reinforcement learning

The National Renewable Energy Laboratory (NREL), together with the Colorado School of Mines (CSM) and Colorado State University (CSU), has developed a machine learning-enhanced approach to design new battery materials. Currently, such materials are designed in part via numerous expensive high-fidelity computational simulations that predict the performance of a given composition. Even with computational screening tools, the vast landscape of possible molecular or crystal structures exceeds current and future computational capacity. Improving the efficiency by which new materials can be optimized will therefore disrupt the cost, risk, and time required to bring new energy solutions to the marketplace. Predicting the properties of an organic molecule or periodic crystalline material given its structure has grown increasingly common. These approaches leverage large-scale computational and experimental databases and ML approaches such as graph neural networks. The inverse design problem of finding a material that possesses desired properties is substantially more challenging, since enumerating all valid material structures is not feasible. In this project, we leveraged recent success in reinforcement learning to efficiently navigate this high-dimensional search space. Just as algorithms can find the optimal chess moves from nearly limitless options, we train an approach to evolve a simple starting structure into a complex structure that possess the desired properties. Our solution has been demonstrated by applying it to two related design application tasks for short- and long-term energy storage, respectively: (1) the design of solid-state ion conductors and (2) the design of organic redox-active materials. The project has resulted an open-source software library for material design, documented examples of applying the library to both organic and inorganic material optimization, and peer-reviewed publications detailing the data, computational models, and resulting candidate materials.

25 ENERGY STORAGE↗

End-to-End Optimization for Battery Materials and Molecules by Combining Graph Neural Networks and Reinforcement Learning

The National Renewable Energy Laboratory (NREL), together with the Colorado School of Mines (CSM) and Colorado State University (CSU), has developed a machine learning-enhanced approach to the design of new battery materials. Currently, such materials are designed in part via numerous expensive high-fidelity computational simulations that predict the performance of a given composition. Even with computational screening tools, the vast landscape of possible molecular or crystal structures exceeds current and future computational capacity. Improving the efficiency by which new materials can be optimized will therefore disrupt the cost, risk, and time required to bring new energy solutions to the marketplace. Predicting the properties of an organic molecule or periodic crystalline material given its structure has grown increasingly common. These approaches leverage large-scale computational and experimental databases and ML approaches such as graph neural networks. The inverse design problem of finding a material that possesses desired properties is substantially more challenging, since enumerating all valid material structures is not feasible. In this project, we leveraged recent success in reinforcement learning to efficiently navigate this high-dimensional search space. Just as algorithms can find the optimal chess moves from nearly limitless options, we train an approach to evolve a simple starting structure into a complex structure that possess the desired properties. Our solution has been demonstrated by applying it to two related design application tasks for short- and long-term energy storage, respectively: (1) the design of solid-state ion conductors and (2) the design of organic redox-active materials. The project has resulted an open-source software library for material design, documented examples of applying the library to both organic and inorganic material optimization, and peer-reviewed publications detailing the data, computational models, and resulting candidate materials.

25 ENERGY STORAGE↗

An automated integrated web-based smart tool for open stope design

The Stability Graph is a widely used tool for the design of open stopes in underground mining. Many users of the Stability Graph still apply this design method manually. Although the manual approach has benefits, using multiple graphs and stability number computation charts for each stope surface is time-consuming, even for the experienced mining engineer. Current practice in the use of the method also limits data sharing. This paper presents a StopeSoft web-based tool for open stope stability prediction that is developed on the basis of the Stability Graph method and is available at openstope.com. StopeSoft incorporates flexibility in terms of Stability Graph options and incorporates additional critical factors often overlooked. As a web-based tool, StopeSoft encourages and makes data sharing possible globally, focused on expanding the database and improving the current limitations of the Stability Graph to provide practical, reliable solutions for mining engineers, consultants, and academics. The StopeSoft automated process facilitates the process of open stope stability prediction, saving time and minimizing potential human errors. Statistical treatment of the data accounts for the variability of input parameters to emphasize the probabilistic nature of the Stability Graph method. The probabilistic interpretation of the stability states of stope surfaces eliminates the false feeling of absolute stope performance based on its location on the Stability Graph , as implied by the deterministic approach.

58 GEOSCIENCES↗

LENS: Learning Enabled Network Synthesis

RTRC and UMD have developed novel machine learning based methods under the ARPA-E DIFFERENTIATE program for rapid acceleration of hypothesis generation in complex architecture design spaces involving both discrete choices of component inclusion and interconnection and continuous parametric decisions. The project named Learning Enabled Network Synthesis (LENS) further demonstrated the developed methods on challenging electrical power converter design problems by identifying the most suitable circuit topologies and simultaneously selecting the most appropriate components to achieve optimized design of power converter with improved performances. We demonstrated that LENS could enable exploration of very large design space of circuit topologies and components by addressing the limitations of conventional design process in non-linear, high switching speed, multi-dimensional power converter design and optimization. The key innovation developed in LENS is the seamless integration of statistical learning and logical reasoning techniques and building on the individual strengths of these techniques for rapid hypothesis discovery. The main component of LENS comprises of: 1) Graph Reasoning Engine (GRE) to enforce composition rules that rapidly reject all discrete architectures that are composed incorrectly and generates an adaptive database of feasible designs which can be used by ML modules, 2) Graph Generative Learning module which is a deep neural network based generative model for graph architectures which can enable design space exploration beyond the dataset generated by the GRE, 3) Graph Reduced Order Model (ROM) for graph domains for accelerating computation of output metrics, and 4) Active learning and Rule Discovery module for sample efficient learning and extracting logical rules from the learned ML models which will be integrated in the GRE to enhance the filtering effectiveness. LENS approach can be applied to any design domains where designs can be represented as multi-attribute graphs. The LENS team integrated the various technical innovations listed above into an optimization pipeline and exercised the optimization pipeline on the converter design problem. The LENS project demonstrated that the developed AI/ML technologies can be used to generate novel converter circuits >45x faster than experts on chosen use-cases. This can enable faster design space exploration and identification of new designs which are not considered by experts due to the increasing design space complexity. This has significant potential impact on the public and energy needs of the country. It is currently estimated that 30% of all electrical powers generated passes through power converters. The future estimate is that 80% of all power generated would be passing through converters. LENS fills a critical gap in this space since by accelerating the design process the designers would be able to generate more efficient converters which can lead to significant energy savings for the country.

42 ENGINEERING↗

AI-powered exploration of molecular vibrations, phonons, and spectroscopy

The vibrational dynamics of molecules and solids play a critical role in defining material properties, particularly their thermal behaviors. However, theoretical calculations of these dynamics are often computationally intensive, while experimental approaches can be technically complex and resource-demanding. Recent advancements in data-driven artificial intelligence (AI) methodologies have substantially enhanced the efficiency of these studies. This review explores the latest progress in AI-driven methods for investigating atomic vibrations, emphasizing their role in accelerating computations and enabling rapid predictions of lattice dynamics, phonon behaviors, molecular dynamics, and vibrational spectra. Key developments are discussed, including advancements in databases, structural representations, machine-learning interatomic potentials, graph neural networks, and other emerging approaches. Compared to traditional techniques, AI methods exhibit transformative potential, dramatically improving the efficiency and scope of research in materials science. The review concludes by highlighting the promising future of AI-driven innovations in the study of atomic vibrations.

Han, Bowen [Oak Ridge National Laboratory (ORNL), ↗

Global horizontal spectral irradiance and module spectral response measurements: an open dataset for PV research

This report describes the creation process and final content of a spectral irradiance dataset for Albuquerque, New Mexico accompanied by a set of spectral response measurements for modules deployed at the same location. The spectral irradiance measurements were made using horizontally mounted spectroradiometers; therefore, they represent global horizontal irradiance. The dataset combines non-continuous spectroradiometer and weather measurements from a two-year period into a single calendar year. The data files are accompanied by extensive metadata as well as example calculations and graphs to demonstrate the potential uses of this database. The spectral response measurements were carried out by the National Renewable Energy Laboratory using 12 commercial silicon modules types that are undergoing long-term evaluation at Sandia National Laboratories in Albuquerque.

14 SOLAR ENERGY↗

Knowledge Graph for End-to-End Traceability of an Integrated Human-Earth System Model

Integrated human-Earth system models inform energy-water-land system dynamics and policies, yet their results are difficult to trace through input-data, model structure, scenario configurations, and solved outputs. Because this information is siloed across disconnected artifacts, process-based IAMs have historically lacked a unified, queryable representation. Such lack of traceability prevents researchers from systematically isolating the multi-sector drivers of complex outcomes (such as tracing water-scarcity results back to distant energy-system dynamics) or conducting holistic uncertainty attribution across hundreds of interacting parameters. To address this concern, our work documents the software engineering process of a knowledge graph that unifies these four layers for the Global Change Analysis Model (GCAM-USA_Reference scenario, GCAM v9.1). The graph was built as a relational property graph in DuckDB from the run’s own artifacts: the input-preparation dependency map (gcamdata chunk map), the model’s XML input files, the run configuration, and the results database (BaseX), successfully mapping the model’s declared structure. The resulting graph comprises 204,321 nodes and 1,687,814 edges across 16 node types and 15 edge types, with approximately 16.3 million time-series values stored separately to maintain structural efficiency. To ensure representation fidelity, every edge carries an epistemic-status annotation recording the warrant for the relationship (structural, provenance, dependency, or model-derived), and a machine-readable provenance ledger classifying the origin of every schema element. Evaluation against a fixed five-benchmark suite with locked baselines reports zero structural orphans, zero dangling edge endpoints, and 100% of output-producing technologies traceable to raw input files. Two interactive interfaces present the graph, including a serverless browser application built on DuckDB-Wasm. By establishing the first end-to-end provenance framework for an IAM, this work enables researchers and scientists to systematically audit complex policy scenarios, debug model structures, and trace policy-relevant outputs to their data origins in real time.

Artifical Intelligence↗

Deep-freeze graph training for latent learning

Scientific and engineering advances are primarily driven by multi-tier conceptual constructs and conditional theoretical frameworks. The theories allow predictions of hypothetical system responses, given a set of approximate conditions (ranges of applicability) imposed on latent parameters that cannot be measured directly. Learning to estimate the latent variables (Latent Learning) helps to pinpoint the anticipated range-edge anomalies and improves the confidence in interpretation, interpolation and extrapolation of limited experimental data. Due to high dimensionality and extreme non-linearity of the materials science problems, very large datasets are typically required for conventional data-driven model development. The vital experimental data collection, particularly on microstructural phases, is very challenging, which makes it difficult to compile a high-quality database. Incorporation of the domain knowledge into the computational graph structure, initialization and optimization processes presents a viable mechanism for developing accurate models, with limited datasets. Furthermore, this study successfully utilized the approach to build the Deep Freeze Graph (DeepFreG) by mapping known causality relationships and by digitizing empirical domain knowledge for Latent Learning (LL), with specific applications in materials science.

36 MATERIALS SCIENCE↗

Graphical Gaussian Process Regression Model for Aqueous Solvation Free Energy Prediction of Organic Molecules in Redox Flow Battery

The solvation free energy of organic molecules is a critical parameter in determining emergent properties such as solubility, liquid-phase equilibrium constants, and pKa and redox potentials in an organic redox flow battery. In this work, we present a machine learning (ML) model that can learn and predict the aqueous solvation free energy of an organic molecule using Gaussian process regression method based on a new molecular graph kernel. To investigate the performance of the ML model on electrostatic interaction, the nonpolar interaction contribution of solvent and the conformational entropy of solute in solvation free energy, three data sets with implicit or explicit water solvent models, and contribution of conformational entropy of solute are tested. We demonstrate that our ML model can predict the solvation free energy of molecules at chemical accuracy with a mean absolute error of less than 1 kcal/mol for subsets of the QM9 dataset and the Freesolv database. To solve the general data scarcity problem for a graph-based ML model, we propose a dimension reduction algorithm based on the distance between molecular graphs, which can be used to examine the diversity of the molecular data set. It provides a promising way to build a minimum training set to improve prediction for certain test sets where the space of molecular structures is predetermined.

25 ENERGY STORAGE↗