Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graph databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Enabling Efficient Sparse Computations using Linear Algebra Aware Compilers

This project developed the LAPIS compiler framework, built on the Multilevel Intermediate Representation (MLIR), to optimize sparse linear algebra operations and support performance portability across diverse architectures. The main innovation of LAPIS is the Kokkos dialect, which allows for lowering codes from a high productivity language to different architectures in an elegant way. The dialect also allows the conversion of lower-level MLIR code to C++ Kokkos code, facilitating the integration of scientific machine learning (SciML) models into applications. To extend LAPIS for distributed memory architectures, a new partition dialect was created to manage the distribution of sparse tensors and express communication patterns for sparse linear algebra operations. This dialect also supports the distributed execution of operators and includes algorithmic optimizations to minimize communication to improve performance. The project also demonstrates that MLIR can enable effective linear algebra-level optimizations, improving performance on different GPUs for both sparse and dense linear algebra kernels. Key applications of LAPIS include sparse linear algebra and graph kernels, TenSQL, a relational database management solution built on GraphBLAS, and the development of subgraph isomorphism and monomorphism kernels, showcasing performance portability. In summary, the LAPIS framework supports productivity, performance, portability, and distributed memory execution, while also enabling linear algebra-level optimizations that are challenging in traditional programming languages, with successful applications ranging from simple sparse linear algebra to complex graph kernels.

97 MATHEMATICS AND COMPUTING↗

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab↗

End-to-end optimization for battery materials and molecules by combining graph neural networks and reinforcement learning

The National Renewable Energy Laboratory (NREL), together with the Colorado School of Mines (CSM) and Colorado State University (CSU), has developed a machine learning-enhanced approach to design new battery materials. Currently, such materials are designed in part via numerous expensive high-fidelity computational simulations that predict the performance of a given composition. Even with computational screening tools, the vast landscape of possible molecular or crystal structures exceeds current and future computational capacity. Improving the efficiency by which new materials can be optimized will therefore disrupt the cost, risk, and time required to bring new energy solutions to the marketplace. Predicting the properties of an organic molecule or periodic crystalline material given its structure has grown increasingly common. These approaches leverage large-scale computational and experimental databases and ML approaches such as graph neural networks. The inverse design problem of finding a material that possesses desired properties is substantially more challenging, since enumerating all valid material structures is not feasible. In this project, we leveraged recent success in reinforcement learning to efficiently navigate this high-dimensional search space. Just as algorithms can find the optimal chess moves from nearly limitless options, we train an approach to evolve a simple starting structure into a complex structure that possess the desired properties. Our solution has been demonstrated by applying it to two related design application tasks for short- and long-term energy storage, respectively: (1) the design of solid-state ion conductors and (2) the design of organic redox-active materials. The project has resulted an open-source software library for material design, documented examples of applying the library to both organic and inorganic material optimization, and peer-reviewed publications detailing the data, computational models, and resulting candidate materials.

25 ENERGY STORAGE↗

End-to-End Optimization for Battery Materials and Molecules by Combining Graph Neural Networks and Reinforcement Learning

The National Renewable Energy Laboratory (NREL), together with the Colorado School of Mines (CSM) and Colorado State University (CSU), has developed a machine learning-enhanced approach to the design of new battery materials. Currently, such materials are designed in part via numerous expensive high-fidelity computational simulations that predict the performance of a given composition. Even with computational screening tools, the vast landscape of possible molecular or crystal structures exceeds current and future computational capacity. Improving the efficiency by which new materials can be optimized will therefore disrupt the cost, risk, and time required to bring new energy solutions to the marketplace. Predicting the properties of an organic molecule or periodic crystalline material given its structure has grown increasingly common. These approaches leverage large-scale computational and experimental databases and ML approaches such as graph neural networks. The inverse design problem of finding a material that possesses desired properties is substantially more challenging, since enumerating all valid material structures is not feasible. In this project, we leveraged recent success in reinforcement learning to efficiently navigate this high-dimensional search space. Just as algorithms can find the optimal chess moves from nearly limitless options, we train an approach to evolve a simple starting structure into a complex structure that possess the desired properties. Our solution has been demonstrated by applying it to two related design application tasks for short- and long-term energy storage, respectively: (1) the design of solid-state ion conductors and (2) the design of organic redox-active materials. The project has resulted an open-source software library for material design, documented examples of applying the library to both organic and inorganic material optimization, and peer-reviewed publications detailing the data, computational models, and resulting candidate materials.

25 ENERGY STORAGE↗

Dynamic IT Security Database and Analytics for Launch Control Systems Software

During the Summer 2020 session, I worked with intern Destani S. Van Arsdalen of EGS Software. Together, we co-created a tool to aid the dynamic investigation, updated over time,of the security compliance of LCS COTS and open source software. We originally planned touse spreadsheet software for management and analysis, but through this exploratoryproject, chose to use Python and JSON after receiving feedback on our project’s current anddesired capabilities at that time.At first, the project was solely designed to help on-board new COTS software, based on aquestionnaire that could be filled out for each software package. This, combined with usingthe spreadsheet application’s web-query capabilities to fetch information from the NVD,allowed presentation and analytics cells to automatically populate as elements of themanually-filled questionnaire changed. While this system was promising, we decided tochange technologies for a few reasons. In the spreadsheet, single cells could not hold complexdata like arrays and objects. The automatic population of cells and dynamic updates made itdifficult to manage and add new features. And finally, it had limited extensibility sinceadding new software required significant understanding of how both the spreadsheet wasconstructed, and the more obscure, proprietary scripting languages packaged with it.The pivot to a standard computer science database language of JSON, aided by thescripting capabilities of Python, greatly helped to improve the project’s functionality. First,and most importantly, the script’s import and analysis of database data is easilyreproducible. Additional data analysis can be modularly added without requiringmodification of the script and is capable of routine scheduling. The revised process can besplit into three parts. First, the conversion of LCS asset and software documentation into theJSON hierarchical database format. Second, the merging of this database with the NVD,forming a new data structure, using CPEs of the CVE object as a linking element betweenthem. And third, the automatically performed analytics and analysis of the combined data,in a modular and extensible format, to produce better informed business decisions. The outputted graphs, for example, are automatically generated by the Python script inconnection with the combined database. This allows updated graphs and any analytics to be re-rendered automatically following updates to the LCS’s initial asset documentation. Afinal report can then be programmatically and easily constructed from these sources to allow fully reproducible metrics for heavily evidenced risk management decisions.

it↗

An automated integrated web-based smart tool for open stope design

The Stability Graph is a widely used tool for the design of open stopes in underground mining. Many users of the Stability Graph still apply this design method manually. Although the manual approach has benefits, using multiple graphs and stability number computation charts for each stope surface is time-consuming, even for the experienced mining engineer. Current practice in the use of the method also limits data sharing. This paper presents a StopeSoft web-based tool for open stope stability prediction that is developed on the basis of the Stability Graph method and is available at openstope.com. StopeSoft incorporates flexibility in terms of Stability Graph options and incorporates additional critical factors often overlooked. As a web-based tool, StopeSoft encourages and makes data sharing possible globally, focused on expanding the database and improving the current limitations of the Stability Graph to provide practical, reliable solutions for mining engineers, consultants, and academics. The StopeSoft automated process facilitates the process of open stope stability prediction, saving time and minimizing potential human errors. Statistical treatment of the data accounts for the variability of input parameters to emphasize the probabilistic nature of the Stability Graph method. The probabilistic interpretation of the stability states of stope surfaces eliminates the false feeling of absolute stope performance based on its location on the Stability Graph , as implied by the deterministic approach.

58 GEOSCIENCES↗

A Tool for Automatic Data Distribution for CFD Applications on Structured Grids

Development of HPF versions of NPB and ARC3D has shown that HPF provides an efficient, concise way to express parallelism and to organize data traffic. The use of HPF, as noted in the papers, requires an intimate knowledge of the applications and a detailed analysis of data affinity, data movement, and data granularity. To simplify and accelerate the task of developing HPF versions of existing CFD applications we have designed and implemented ADAPT (Automatic Data Alignment and Placement Tool). ADAPT analyzes a CFD application working on a single structured grid and generates HPF TEMPLATE, (RE)DISTRIBUTION, ALIGNMENT, and INDEPENDENT directives. The directives can be generated on the nest level, subroutine level, application level, or on the application interface level. ADAPT annotates an existing CFD FORTRAN application, performing computations on single or multiple grids. On each grid the application is considered as a sequence of operators, each applied to a set of variables defined in a particular grid domain. ADAPT automatically detects implicit operators (i.e., having data dependences) and explicit operators (without data dependences). For parallelization of an explicit operator ADAPT creates a template for the operator domain, aligns arrays used in the operator with the template, distributes the template, and declares the loops over the distributed dimensions as INDEPENDENT. For parallelization of an implicit operator, the distribution of the operator's domain should be consistent with the operator's dependences. Any dependence between sections distributed on different processors would preclude parallelization if the compiler does not have an ability to pipeline computations. If a data distribution is "orthogonal" to the dependences of an implicit operator, then the loop which implements the operator can be declared as INDEPENDENT. ADAPT starts with an analysis of array index expressions of the loop nests. For each pair of arrays referenced in an assignment statement, it generates an arc in the alignment graph and annotates it with an affinity relation. The template, alignment, and distribution directives for a particular loop nest are then derived from a transitive closure of the affinity relation. A compromise of data distributions in different nests and subroutines is achieved by merging annotated alignment graphs for adjacent nests/stibroutine calls in the nest/call graph of the application in the process called distribution lifting. ADAPT has been implemented as a C++ program running in conjunction with a parallelization tool called CAPTools. ADAPT uses the parse tree, interprocedural analysis and application database generated by CAPTools. It also uses the Directed Graph class, initially implemented in p2d2 (parallel debugger oi distributed programs), and some other classes supporting symbolic computations. ADAPT uses data distribution techniques described. ADAPT was tested with ARC3D and the FT benchmark and has demonstrated a code performance within a factor of 1.5 of handwritten versions.

Frumkin, Michael↗

LENS: Learning Enabled Network Synthesis

RTRC and UMD have developed novel machine learning based methods under the ARPA-E DIFFERENTIATE program for rapid acceleration of hypothesis generation in complex architecture design spaces involving both discrete choices of component inclusion and interconnection and continuous parametric decisions. The project named Learning Enabled Network Synthesis (LENS) further demonstrated the developed methods on challenging electrical power converter design problems by identifying the most suitable circuit topologies and simultaneously selecting the most appropriate components to achieve optimized design of power converter with improved performances. We demonstrated that LENS could enable exploration of very large design space of circuit topologies and components by addressing the limitations of conventional design process in non-linear, high switching speed, multi-dimensional power converter design and optimization. The key innovation developed in LENS is the seamless integration of statistical learning and logical reasoning techniques and building on the individual strengths of these techniques for rapid hypothesis discovery. The main component of LENS comprises of: 1) Graph Reasoning Engine (GRE) to enforce composition rules that rapidly reject all discrete architectures that are composed incorrectly and generates an adaptive database of feasible designs which can be used by ML modules, 2) Graph Generative Learning module which is a deep neural network based generative model for graph architectures which can enable design space exploration beyond the dataset generated by the GRE, 3) Graph Reduced Order Model (ROM) for graph domains for accelerating computation of output metrics, and 4) Active learning and Rule Discovery module for sample efficient learning and extracting logical rules from the learned ML models which will be integrated in the GRE to enhance the filtering effectiveness. LENS approach can be applied to any design domains where designs can be represented as multi-attribute graphs. The LENS team integrated the various technical innovations listed above into an optimization pipeline and exercised the optimization pipeline on the converter design problem. The LENS project demonstrated that the developed AI/ML technologies can be used to generate novel converter circuits >45x faster than experts on chosen use-cases. This can enable faster design space exploration and identification of new designs which are not considered by experts due to the increasing design space complexity. This has significant potential impact on the public and energy needs of the country. It is currently estimated that 30% of all electrical powers generated passes through power converters. The future estimate is that 80% of all power generated would be passing through converters. LENS fills a critical gap in this space since by accelerating the design process the designers would be able to generate more efficient converters which can lead to significant energy savings for the country.

42 ENGINEERING↗

AI-powered exploration of molecular vibrations, phonons, and spectroscopy

The vibrational dynamics of molecules and solids play a critical role in defining material properties, particularly their thermal behaviors. However, theoretical calculations of these dynamics are often computationally intensive, while experimental approaches can be technically complex and resource-demanding. Recent advancements in data-driven artificial intelligence (AI) methodologies have substantially enhanced the efficiency of these studies. This review explores the latest progress in AI-driven methods for investigating atomic vibrations, emphasizing their role in accelerating computations and enabling rapid predictions of lattice dynamics, phonon behaviors, molecular dynamics, and vibrational spectra. Key developments are discussed, including advancements in databases, structural representations, machine-learning interatomic potentials, graph neural networks, and other emerging approaches. Compared to traditional techniques, AI methods exhibit transformative potential, dramatically improving the efficiency and scope of research in materials science. The review concludes by highlighting the promising future of AI-driven innovations in the study of atomic vibrations.

Han, Bowen [Oak Ridge National Laboratory (ORNL), ↗

Global horizontal spectral irradiance and module spectral response measurements: an open dataset for PV research

This report describes the creation process and final content of a spectral irradiance dataset for Albuquerque, New Mexico accompanied by a set of spectral response measurements for modules deployed at the same location. The spectral irradiance measurements were made using horizontally mounted spectroradiometers; therefore, they represent global horizontal irradiance. The dataset combines non-continuous spectroradiometer and weather measurements from a two-year period into a single calendar year. The data files are accompanied by extensive metadata as well as example calculations and graphs to demonstrate the potential uses of this database. The spectral response measurements were carried out by the National Renewable Energy Laboratory using 12 commercial silicon modules types that are undergoing long-term evaluation at Sandia National Laboratories in Albuquerque.

14 SOLAR ENERGY↗

Knowledge Graph for End-to-End Traceability of an Integrated Human-Earth System Model

Integrated human-Earth system models inform energy-water-land system dynamics and policies, yet their results are difficult to trace through input-data, model structure, scenario configurations, and solved outputs. Because this information is siloed across disconnected artifacts, process-based IAMs have historically lacked a unified, queryable representation. Such lack of traceability prevents researchers from systematically isolating the multi-sector drivers of complex outcomes (such as tracing water-scarcity results back to distant energy-system dynamics) or conducting holistic uncertainty attribution across hundreds of interacting parameters. To address this concern, our work documents the software engineering process of a knowledge graph that unifies these four layers for the Global Change Analysis Model (GCAM-USA_Reference scenario, GCAM v9.1). The graph was built as a relational property graph in DuckDB from the run’s own artifacts: the input-preparation dependency map (gcamdata chunk map), the model’s XML input files, the run configuration, and the results database (BaseX), successfully mapping the model’s declared structure. The resulting graph comprises 204,321 nodes and 1,687,814 edges across 16 node types and 15 edge types, with approximately 16.3 million time-series values stored separately to maintain structural efficiency. To ensure representation fidelity, every edge carries an epistemic-status annotation recording the warrant for the relationship (structural, provenance, dependency, or model-derived), and a machine-readable provenance ledger classifying the origin of every schema element. Evaluation against a fixed five-benchmark suite with locked baselines reports zero structural orphans, zero dangling edge endpoints, and 100% of output-producing technologies traceable to raw input files. Two interactive interfaces present the graph, including a serverless browser application built on DuckDB-Wasm. By establishing the first end-to-end provenance framework for an IAM, this work enables researchers and scientists to systematically audit complex policy scenarios, debug model structures, and trace policy-relevant outputs to their data origins in real time.

Artifical Intelligence↗

Deep-freeze graph training for latent learning

Scientific and engineering advances are primarily driven by multi-tier conceptual constructs and conditional theoretical frameworks. The theories allow predictions of hypothetical system responses, given a set of approximate conditions (ranges of applicability) imposed on latent parameters that cannot be measured directly. Learning to estimate the latent variables (Latent Learning) helps to pinpoint the anticipated range-edge anomalies and improves the confidence in interpretation, interpolation and extrapolation of limited experimental data. Due to high dimensionality and extreme non-linearity of the materials science problems, very large datasets are typically required for conventional data-driven model development. The vital experimental data collection, particularly on microstructural phases, is very challenging, which makes it difficult to compile a high-quality database. Incorporation of the domain knowledge into the computational graph structure, initialization and optimization processes presents a viable mechanism for developing accurate models, with limited datasets. Furthermore, this study successfully utilized the approach to build the Deep Freeze Graph (DeepFreG) by mapping known causality relationships and by digitizing empirical domain knowledge for Latent Learning (LL), with specific applications in materials science.

36 MATERIALS SCIENCE↗

Knowledge Network Embedding of Transcriptomic Data From Spaceflown Mice Uncovers Signs and Symptoms Associated With Terrestrial Diseases

There has long been an interest in understanding how the hazards from spaceflight may trigger or exacerbate human diseases. With the goal of advancing our knowledge on physiological changes during space travel, NASA GeneLab provides an open-source repository of multi-omics data from real and simulated spaceflight studies. Alone, this data enables identification of biological changes during spaceflight, but cannot infer how that may impact an astronaut at the phenotypic level. To bridge this gap, SPOKE, a heterogeneous knowledge graph connecting biological and clinical data from over 30 databases, was used in combination with GeneLab transcriptomic data from six studies. This integration identified critical symptoms and physiological changes incurred during spaceflight.

spaceflight↗

Database and deep-learning scalability of anharmonic phonon properties by automated brute-force first-principles calculations

Understanding the anharmonic phonon properties of crystal compounds—such as phonon lifetimes and thermal conductivities—is essential for investigating and optimizing their thermal transport behaviors. These properties also impact optical, electronic, and magnetic characteristics through interactions between phonons and other quasiparticles and fields. In this study, we develop an automated first-principles workflow to calculate anharmonic phonon properties and build a comprehensive database encompassing more than 6500 inorganic compounds. Utilizing this dataset, we train a graph neural network model to predict thermal conductivity values and spectra from structural parameters, demonstrating a scaling law in which prediction accuracy improves with increasing training data size. High-throughput screening with the model enables the identification of materials exhibiting extreme thermal conductivities—both high and low. The resulting database offers valuable insights into the anharmonic behavior of phonons, thereby accelerating the design and development of advanced functional materials.

Ohnishi, Masato [University of Tokyo (Japan); Inst↗

An ontology-based knowledge graph for representing interactions involving RNA molecules

The "RNA world" represents a novel frontier for the study of fundamental biological processes and human diseases and is paving the way for the development of new drugs tailored to each patient's biomolecular characteristics. Although scientific data about coding and non-coding RNA molecules are constantly produced and available from public repositories, they are scattered across different databases and a centralized, uniform, and semantically consistent representation of the "RNA world" is still lacking. We propose RNA-KG, a knowledge graph (KG) encompassing biological knowledge about RNAs gathered from more than 60 public databases, integrating functional relationships with genes, proteins, and chemicals and ontologically grounded biomedical concepts. To develop RNA-KG, we first identified, pre-processed, and characterized each data source; next, we built a meta-graph that provides an ontological description of the KG by representing all the bio-molecular entities and medical concepts of interest in this domain, as well as the types of interactions connecting them. Finally, we leveraged an instance-based semantically abstracted knowledge model to specify the ontological alignment according to which RNA-KG was generated. RNA-KG can be downloaded in different formats and also queried by a SPARQL endpoint. A thorough topological analysis of the resulting heterogeneous graph provides further insights into the characteristics of the "RNA world". RNA-KG can be both directly explored and visualized, and/or analyzed by applying computational methods to infer bio-medical knowledge from its heterogeneous nodes and edges. The resource can be easily updated with new experimental data, and specific views of the overall KG can be extracted according to the bio-medical problem to be studied.

59 BASIC BIOLOGICAL SCIENCES↗

Geospatial Data Platform for All

Spatiotemporal data has evolved in scale due to augmented use in cross-domain applications. Simultaneously, there is substantial growth in the availability of Geographic Information Systems (GIS) data provided by the United States Geological Survey (USGS) along with other federal, state, county, or local agencies through open-data portals and public access APIs. However, data availability does not equate with accessibility. Large-scale analyses and applications require robust, performant data management with co-location of data storage and computing. The insufficiency of data management infrastructure compels researchers to adopt ad hoc project- specific GIS data storage solutions (e.g., copying data to High-Performance computer file systems). As an ad hoc storage strategy does not scale, it hampers cross-domain analyses causing difficulty in data reuse and utilizing existing code bases. Furthermore, GIS data is complex and requires expertise to analyze and manipulate due to its intricate data structures and data-specific projection transformations. Despite the challenges, we recognize that derived GIS data products, e.g., satellite or LIDAR-based images, can be used in downstream applications such as AI by domain, but non-GIS experts. To address the data needs and overcome the challenges, we are working towards a GIS Data Platform focused on efficient data storage, data discovery and access, and an API to enable common workflows. We propose a knowledge-graph (KG) approach for data discovery, whereby datasets are semantically linked to higher- level constructs such as projects and research areas. The semantic data links enable researchers to explore datasets in a top-down approach by specifying relevant and meaningful terms (assists in finding hidden data). An advantage is that the nodes and edges in a knowledge graph create built-in semantic documentation. Deeper spatiotemporal connections between data sources can be encoded via Graph Neural Networks (GNN) (Zhang et al., 2021). The KG approach can be extended to integrate the data itself in a Virtual KG (VKG). Our work will derive inspiration from large-scale VKG efforts that have been undertaken or are currently underway as part of the OpenStreetMap project (Ding et al., 2021). For DOE Data Days, we share the proposed geospatial data platform hybrid (cloud/on-prem) architecture, our work-to-date on storing, retrieving, and transforming LiDAR and raster data relevant to two important NREL use-cases, including the Renewable Energy Potential (reV) Model, and present our proposal for a KG based data discovery engine.

data platform↗

High-Throughput Screening of Li Solid-State Electrolytes With Bond Valence Methods and Graph Neural Networks

Li-based solid-state electrolyte (Li-SSE) materials enable safer, all-solid-state batteries but the computational search for candidates with favorable stability and Li-ion conductivity is challenging due to the size of the search space and the cost of evaluating transport properties with ab initio methods. We present a high-throughput screening approach for Li-SSE materials using a combination of bond-valence methods and graph neural networks. We demonstrate the screening approach with a dataset containing tens of thousands of Li-containing compounds. Furthermore, we combine the machine-learning screening procedure with an isovalent substitution scheme to generate and screen additional Li SSE candidates beyond existing databases. Finally, we discuss relative importances of geometric and bond-valence quantities in the training of graph neural networks, providing insight for future modeling of ionic conductivity in Li-SSE materials.

Materials discovery↗

Screening of Li-Based Solid Electrolytes Using Bond-Valence Methods and Graph Neural Networks

Li-based solid-state electrolyte (Li-SSE) materials enable safer, all-solid-state batteries but the computational search for candidates with favorable stability and Li-ion conductivity is challenging due to the size of the search space and the cost of evaluating transport properties with ab initio methods. We present a high-throughput screening approach for Li-SSE materials using a combination of bond-valence methods and graph neural networks. We demonstrate the screening approach with a dataset containing tens of thousands of Li-containing compounds. Furthermore, we combine the machine-learning screening procedure with an isovalent substitution scheme to generate and screen additional Li SSE candidates beyond existing databases. Finally, we discuss relative importances of geometric and bond-valence quantities in the training of graph neural networks, providing insight for future modeling of ionic conductivity in Li-SSE materials.

Materials discovery↗