Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Addressing the dynamic nature of reference data: a new nucleotide database for robust metagenomic classification

Accurate metagenomic classification relies on comprehensive, up-to-date, and validated reference databases. While the NCBI BLAST Nucleotide (nt) database, encompassing a vast collection of sequences from all domains of life, represents an invaluable resource, its massive size—currently exceeding 10 12 nucleotides—and exponential growth pose significant challenges for researchers seeking to maintain current nt-based indices for metagenomic classification. Recognizing that no current nt-based indices exist for the widely used Centrifuge classifier, and the last public version currently available was released in 2018, we addressed this critical gap by leveraging advanced high-performance computing resources. We present new Centrifuge-compatible nt databases, meticulously constructed using a novel pipeline incorporating different quality control measures, including reference decontamination and filtering. These measures demonstrably reduce spurious classifications, as shown through our reanalysis of published metagenomic data where Plasmodium annotations were dramatically reduced using our decontaminated database, highlighting how database quality can significantly impact research conclusions. Through temporal comparisons, we also reveal how our approach minimizes inconsistencies in taxonomic assignments stemming from asynchronous updates between public sequence and taxonomy databases. These discrepancies are particularly evident in taxa such as Listeria monocytogenes and Naegleria fowleri, where classification accuracy varied significantly across database versions. These new databases, made available as pre-built Centrifuge indexes, respond to the need for an open, robust, nt-based pipeline for taxonomic classification in metagenomics. Applications such as environmental metagenomics, forensics, and clinical metagenomics, which require comprehensive taxonomic coverage, will benefit from this resource. Our work highlights the importance of treating reference databases as dynamic entities, subject to ongoing quality control and validation akin to software development best practices. This approach is crucial for ensuring accuracy and reliability of metagenomic analysis, especially as databases continue to expand in size and complexity.

59 BASIC BIOLOGICAL SCIENCES↗

Carbon Storage Technical Viability Approach (CS TVA) Database

The Carbon Storage Technical Viability Approach (CS TVA) database was developed to support the implementation of the CS TVA Matrix to a national data availability assessment for technically viable carbon storage. This database leverages the efforts of multiple adjacent and overlapping databases by non-redundantly combining the databases into a single database along with additionally providing tags facilitating the CS TVA. The non-redundant aspect of the database permits an accurate assessment of the concentration of available data, aiding in spatial and categorical data gaps analysis relative to the individual CS TVA Matrix Components. Version 2.0 of the database is an expansion of Version 1.0. Version 2.0 was created to include additional data gathered to fill gaps in the existing data set. Downloading the CS TVA v2.0 database will result in two separate databases, the version 1.0 original .gdb, and a second addendum .gdb with the new data gathered, together these two databases make up v2.0. Please see the ReadMe file below for full details, metadata information, use disclaimer, and attributions.

Coal↗

WELLS Database

The Wellbore Exploration and Location Logistic System (WELLS) is a living national wellbore database - created and maintained by the National Energy Technology Laboratory (NETL). This resource contains more than seven million public wellbore records from state, federal, and tribal resources. Sourced from over 65 authoritative, yet disparate resources, the WELLS Database combines and synthesizes well data from oil, gas, underground injection, research, geothermal, geotechnical, groundwater and other types of wells in a single, unified system. This resource can be explored and visualized through the WELLS Interactive Application, also on EDX: https://edx.netl.doe.gov/dataset/wells-interactive-application The WELLS Database (formerly titled CO2-Locate) is an integrated national well dataset, representing open-source wellbore data from disparate state, tribal, and federal entities. The database provides publicly available well header data with key attributes such as well age, depth, and status. The database contains a fully integrated CSV file with all values in numerical columns, such as depth, converted into numbers. This version has a NETL derived API (American Petroleum Institute) number column and has been handled for redundancies, resulting in one record for every unique API number. The database also contains a fully integrated CSV file, where all original data are kept as text values. Additionally, the database includes a shapefile containing key attributes and coordinates from the integrated dataset, reformatted public wells CSV files, and a proprietary well density grid shapefile. Notes for consideration: The WELLS Database will be updated periodically with new datasets and information. A field dictionary with field (i.e., attribute) coverage across acquired public well resources, and the resulting integrated public well datasets are available in the spreadsheet, WELLS_Field_Dictionary.xlsx. Summary layers provided in this database are derived from proprietary layers and do not always contain key features (status, type, true vertical depth, or spud year) and therefore might not be shown when data are queried for those features.

AS↗

Design and requirements of a hydrogen component reliability database (HyCReD)

Hydrogen technologies are expected to play a key role in the decarbonization of several sectors including energy storage and transportation. Rigorous investigation and quantification of the risk and reliability issues associated with hydrogen technologies will be critical to ensuring both their wider adoption and safe, economical operations. Quantitative risk assessment (QRA) is an important tool that has been used to enable the safe deployment of many engineering systems, including hydrogen fueling stations and hydrogen storage systems. However, QRA studies require reliability data which is currently lacking for expanding applications of hydrogen systems. Here, to address this gap, we present a new structure for a hydrogen component reliability database (HyCReD) that can be used to generate reliability data to be used in QRA, reliability, safety studies, maintenance planning, and more. Building on our previous work examining four major hydrogen safety data collection tools (West et al., 2022) [1], our approach in this work was to consult scientific literature on reliability data collection as well as a number of existing reliability engineering databases in the oil & gas, chemical processing, and nuclear power plant sectors. The evaluation of these databases led to identifying best practices to be implemented in a data collection framework for a hydrogen component reliability database. Based on these best practices, a set of 24 requirements for the proposed database are presented, covering its characteristics and the types of data to be collected. We define the structure of the HyCReD database and 25 data elements to be collected, spanning system description, failure, shutdown, or near-miss events, and maintenance events. The data elements are then defined according to international standards used in the safety and reliability practice and potential choice lists are provided for each field. Since this database is being piloted for hydrogen fueling stations, a generic station component hierarchy developed by West (2021) [2] is used to standardize system data. Finally, we demonstrate populating the database with information extracted from five narrative reports on hydrogen fueling station incidents.

08 HYDROGEN↗

TropiRoot 1.0: Database of tropical root characteristics across environments

Tropical ecosystems contain the world's largest biodiversity of vascular plants. Yet, our understanding of tropical functional diversity and its contribution to global diversity patterns is constrained by data availability. This discrepancy underscores an urgent need to bridge data gaps by incorporating comprehensive tropical root data into global datasets. Here, we provide a database of tropical root characteristics. This new database, TropiRoot 1.0, will be instrumental in evaluating an array of hypotheses pertaining to root functional ecology and plant biogeography, both within the tropics and relative to other global biomes. The data compilation was conducted by the TropiRoot Initiative, in partnership with the Fine-Root Ecology Database (FRED) and the Global Root Trait (GRooT) database, Colorado State University (CSU) and the Smithsonian Tropical Research Institute (STRI). Literature search and data extraction were conducted between 2020 and 2024. Literature was identified using Web of Science, Scopus, and complemented using the expert knowledge of members of TropiRoot. To provide broad environmental and geographical distributions, literature searches included root characteristics (traits) across global change drivers, natural gradients, and from different continents. We adopted FRED standardized data columns and streamlined the format to enhance accessibility for data extraction across various user groups. This optimized framework resulted in a smaller, yet comprehensive datasheet. To make the database compatible with other global root trait initiatives, column identification was standardized following the codes provided by FRED. These efforts culminated in data extracted from 104 new sources, resulting in more than 8000 rows of data (either species or community data). Most of the data in TropiRoot 1.0 include root characteristics such as root biomass, morphology, root dynamics, mass fraction, architecture, anatomy, physiology, and root chemistry. This initiative represents a 30% increase in the currently available data for tropical roots in FRED. TropiRoot 1.0 contains root characteristics from 25 different countries, where seven are located in Asia, six in South America, five in Central America and the Caribbean, four in Africa, two in North America, and 1 in Oceania. Due to the volume of data, when ancillary data were available, including soil data, these data were either extracted and included in the database or its availability was recorded in an additional column. Multiple contributors checked the entries for outliers during the collation process to ensure data quality. For text-based observations, we examined all cells to ensure that their content relates to their specific categories. For numerical observations, we ordered each numerical value from least to greatest and plotted the values, checking apparent outliers against the data in their respective sources and correcting or removing incorrect or impossible values. Some data (soil and aboveground) have different columns for the same variable presented in different units, including originally published units, but root characteristics data had units converted to match those reported in FRED. By filling a gap from global databases, TropiRoot 1.0 expands our knowledge of otherwise so far underrepresented regions and our ability to assess global trends. This advancement can be used to improve tropical forest representation in vegetation models. The data are freely available and should be cited when used.

FRED↗

Cross-database comparisons on the greenhouse gas emissions, water consumption, and fossil-fuel use of plastic resin production and their post-use phase impacts

Resin production and post-use pathways have cross-database discrepancies due to different temporal/geographic representation, technologies, and assumptions. These differences can distort comparisons across alternatives and confound efforts to establish standards based on the plastics’ life-cycle impacts. Thus, this study quantifies the degree and identifies the sources of cross-database discrepancies across four LCA databases (GREET, USLCI, Ecoinvent, and GaBi) for five resin production pathways and three post-use phases. For resin production pathways, all resins showed significant cross-database discrepancy in their global warming impacts: the degree of discrepancy was significant to recommend a consistent choice of database to users any LCA comparisons across different products. For post-use phases, landfill datasets had relatively lower degree of cross-database discrepancy than incineration and mechanical recycling. Different metadata characteristics were the sources of some cross-database discrepancies while other parts could be explained by the original differences in the life-cycle inventory sourced from different producers and plants.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

SoK: What does it Mean to Benchmark Database Forensics?

Relational Database Management Systems are the backbone of modern enterprises and public-sector services, and are thus frequent targets of security incidents, insider threats, and thorough regulatory audits. Consequently, databases have become key sources of digital evidence, requiring investigators to reconstruct past activity from audit logs, transaction logs, and backups. Although benchmarking frameworks such as those developed by the Transaction Processing Performance Council (TPC) are widely used to evaluate database performance, they do not capture forensic requirements such as evidentiary completeness, tamper-evidence, chain of custody, or regulatory compliance under GDPR and CCPA. This survey examines the emerging domain of forensic database benchmarking. We gathered prior research on database forensics, secure logging, and tamper-evident data structures; we analyze modern forensic-ready features in commercial and open-source systems (SQL Server Ledger, Oracle Blockchain Tables, PostgreSQL pgAudit, Db2 Audit, Aurora Database Activity Streams, Oracle Real Application Security and IBM Guardium) and assess why existing benchmarks are insufficient. We propose forensic workloads, metrics, and methodologies that incorporate adversarial stressors, deleted-record recovery, and backup analysis. We also identify open research problems and call for a community-driven forensic benchmark suite. The result is an idea for evaluating not only database performance but also forensic soundness, bridging the gap between system engineering, compliance, and digital investigations.

Lenard, Ben↗

The User Guide for the ComPro Database

An extensive database of glass composition and durability data has been compiled at Savannah River National Laboratory to support the development of nuclear waste glasses. This database is referred to as the Glass Composition-Properties Database (ComPro). The ComPro Database, Revision 3, contains 14,134 total rows of data and 125 columns of composition, durability, as defined by the Product Consistency Test, and other fabrication and characterization information, if available, for each glass. Of the 14,134 total rows, 8,484 rows have been classified as “Model” data and 5,650 rows have been classified as “Non-Model” data. An integral supplement to the ComPro database is the User Guide. The User Guide was developed as a tool to aid the End User in a more effective use of the ComPro database. The User Guide provides a road-map of the specific datasets that comprise the ComPro database (both “Model” and “Non-Model” data) as well as a technical basis for the terminology and definitions the End User will encounter. In this report, a general description of the format and information contained in the User Guide is provided. In addition, specific terminology used in the User Guide is also discussed.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

SERDP PFAS Database

The PNNL-PFAS database was constructed to provide quantitative infrared spectra of per- and polyfluoroalkyl substances (PFAS) in the gas phase. These fluorine-rich, organic compounds have been used as non-flammable solvents, cleaning agents, and in the manufacture of a diverse variety of consumer products. Due to their widespread use and relative inertness, PFAS have accumulated in the natural environment, and exposure to these chemicals may be linked to harmful health effects. Development of methods for detecting and remediating PFAS contamination are thus active areas of research. This database supports PFAS investigations where gas-phase infrared signatures may be measured (e.g., during treatment by thermal decomposition). The database currently contains infrared spectra for 15 compounds; these were selected based on the potential of these compounds being thermal degradation products of the manufactured chemicals mentioned above. • Additional compounds will be added to the database in the future. • Additional information about the database may be obtained from [link to future paper about the PFAS database].

PFAS, infrared(IR)spectroscopy, DATABASE, FTIR, sp↗

A new database of building-space-specific internal loads and load schedules for performance based code compliance modeling of commercial buildings

Building-level loads and load profiles prescribed by current modeling rules save modelers time and avoid gaming during whole building performance modeling. However, recent studies show that they sometimes insufficiently capture the entire building performance due to the varied loads and load profiles for different space types. As a solution to this issue, this paper develops a database of building-space-specific loads and load profiles used in code compliance modeling. The existing sets of loads and load profiles are reviewed and the challenges behind using them for specific research topics are discussed. Then, the proposed method to develop the building-space-specific loads and load profiles is introduced. After that, the database for these building-space-specific loads and load profiles is presented. In addition, one case is studied to demonstrate the applications of these loads and load profiles. In this case study, three methods are used to develop building energy models: space-specific (using knowledge of the distribution and location of space types and applying the space-specific data in the developed database), building-level (assuming a lack of knowledge of the space types and using the building-level data in the developed database), and calculated-ratio (assuming knowledge of the distribution of space types but not their locations and calculating weighted average values based on the space-specific data in the developed database). Finally, the energy results simulated by using these three methods are compared, which show building-level methods can produce energy results up to 20% different than the space-specific methods. Finally, this paper discusses the application scope and maintenance of this new database.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Auto-generating databases of Yield Strength and Grain Size using ChemDataExtractor

Abstract The emerging field of material-based data science requires information-rich databases to generate useful results which are currently sparse in the stress engineering domain. To this end, this study uses the’materials-aware’ text-mining toolkit, ChemDataExtractor, to auto-generate databases of yield-strength and grain-size values by extracting such information from the literature. The precision of the extracted data is 83.0% for yield strength and 78.8% for grain size. The automatically-extracted data were organised into four databases: a Yield Strength, Grain Size, Engineering-Ready Yield Strength and Combined database. For further validation of the databases, the Combined database was used to plot the Hall-Petch relationship for, the alloy, AZ31, and similar results to the literature were found, demonstrating how one can make use of these automatically-extracted datasets.

36 MATERIALS SCIENCE↗

A Global Building Occupant Behavior Database

This paper introduces a database of 34 field-measured building occupant behavior datasets collected from 15 countries and 39 institutions across 10 climatic zones covering various building types in both commercial and residential sectors. This is a comprehensive global database about building occupant behavior. The database includes occupancy patterns (i.e., presence and people count) and occupant behaviors (i.e., interactions with devices, equipment, and technical systems in buildings). Brick schema models were developed to represent sensor and room metadata information. The database is publicly available, and a website was created for the public to access, query, and download specific datasets or the whole database interactively. The database can help to advance the knowledge and understanding of realistic occupancy patterns and human-building interactions with building systems (e.g., light switching, set-point changes on thermostats, fans on/off, etc.) and envelopes (e.g., window opening/closing). With these more realistic inputs of occupants’ schedules and their interactions with buildings and systems, building designers, energy modelers, and consultants can improve the accuracy of building energy simulation and building load forecasting.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A database of thermally activated delayed fluorescent molecules auto-generated from scientific literature with ChemDataExtractor

A database of thermally activated delayed fluorescent (TADF) molecules was automatically generated from the scientific literature. It consists of 25,482 data records with an overall precision of 82%. Among these, 5,349 records have chemical names in the form of SMILES strings which are represented with 91% accuracy; these are grouped in a subsidiary database. Each data record contains one of the following four properties: maximum emission wavelength (λ EM ), photoluminescence quantum yield (PLQY), singlet-triplet energy splitting (ΔE ST ), and delayed lifetime (τ D ). The databases were created through text mining using ChemDataExtractor, a chemistry-aware natural-language-processing toolkit, which has been adapted for TADF research. The text-mined corpus consisted of 2,733 papers from the Royal Society of Chemistry and Elsevier. To the best of our knowledge, these databases are the first databases that have been auto-generated for TADF molecules from existing publications. The databases have been publicly released for experimental and computational applications in the TADF research field.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Database of Stress-Strain Properties Auto-generated from the Scientific Literature using ChemDataExtractor

Abstract There has been an ongoing need for information-rich databases in the mechanical-engineering domain to aid in data-driven materials science. To address the lack of suitable property databases, this study employs the latest version of the chemistry-aware natural-language-processing (NLP) toolkit, ChemDataExtractor, to automatically curate a comprehensive materials database of key stress-strain properties. The database contains information about materials and their cognate properties: ultimate tensile strength, yield strength, fracture strength, Young’s modulus, and ductility values. 720,308 data records were extracted from the scientific literature and organized into machine-readable databases formats. The extracted data have an overall precision, recall and F-score of 82.03%, 92.13% and 86.79%, respectively. The resulting database has been made publicly available, aiming to facilitate data-driven research and accelerate advancements within the mechanical-engineering domain.

Kumar, Pankaj↗

Database of virus genomes from ultra-deep sequencing of wastewater

Researchers at University of Missouri have conducted ultra-deep RNA sequencing of viral concentrates from wastewater (1 billion Illumina reads per sample). The resulting dataset spans 321 samples collected weekly from 11 cities between 2023-2025. As part of a tri-lab collaboration, scientists at LLNL and LANL cleaned, assembled, and annotated this metagenomic data, identifying nearly 200,000 viral genomes. Careful data curation resulted in a database containing 21,015 high-quality, near-complete viral genomes from wastewater. This database contains viruses predicted to infect a range of hosts including bacteria (most common viruses), plants (most abundant viruses), and vertebrates (rarest viruses). There are also numerous novel viruses that could not be well identified and whose host(s) are unknown. Just 7% of all genomes in the wastewater virus database had genus-level matches in the public NCBI database, and 17% matched to a recently created metagenomic virus database at that level (metaVR). The database will provide baseline information about viruses in wastewater that may be used to additional identify novel viruses during ongoing monitoring

Allen, Jonathan [Lawrence Livermore National Labor↗

An Overview of the Molten Salt Thermal Properties Database–Thermophysical, Version 2.1.1 (MSTDB-TP v.2.1.1)

The current status of the Molten Salt Thermodynamic Database- Thermal Physical (MSTDB-TP) is reported. Building off of MSTDB-TP and through the direction of the Roadmap for thermal property measurement of Molten Salt Reactor systems, MSTDB-TP 2.1.1 has now been released containing a total 448 salt entries (data from 140+independent studies) compared to the original commit of 62, containing thermophysical properties including density, viscosity, thermal conductivity, and heat capacity. Along with the new database release, further advancement of Saline, a C++application programming interface (API), and a graphical user interface (GUI) has facilitated increased user/developer interaction with the database. Furthermore, estimation techniques, first principles calculations (Ab-Initio) and interpolation/extrapolation methods (Redlich-Kister/Muggianu), have shown great promise in the future of thermophysical property determination for filling out compositional spaces and investigating experimentally difficult salts (hazardous/expensive). This report describes the database composition, development and advancement of database tools, and the strategy of advancing and implementing estimation data into future iterations of the database for MSTDB-TP.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Facilitating Data Collection of Maintenance Events to Populate the Hydrogen Component Reliability Database (HyCReD)

The Hydrogen Component Reliability Database (HyCReD) is a collaborative project between the National Renewable Energy Laboratory, the University of Maryland, and hydrogen stakeholders to improve safety and reliability for hydrogen facilities by implementing component reliability data taxonomies that support hydrogen infrastructure failure rate analysis. The project aims to quantify failure rates of hydrogen components through high-quality data collection and analysis on root causes and maintenance needed. HyCReD provides a common database for cataloging hydrogen component failures which exists for reliability research in many other mature industries [2]. The database fills a gap for the hydrogen community by providing a scientifically rigorous approach to quantitative risk assessment (QRA), prognostic health management (PHM), and reliability-centered maintenance (RCM) analysis. High level results will be aggregated and anonymized to protect company sensitive information; detailed results will be used to help address issues of hydrogen components. These advanced analytics will support accelerated deployment of hydrogen infrastructure by enabling better: design and safety of projects (safety codes and standards development), infrastructure reliability and cost (component failure rates, maintenance protocols), and component R&D needs (robust supply chain). A key to a successful HyCReD implementation is facilitating the ease of reporting and data quality in the database that can be used for analysis. Maintenance data was a previously identified gap in initial efforts to populate and validate the database taxonomies [3]. Collection of maintenance data will be instrumental in identifying failure modes and rates, identifying incipient component failures or reduced performance, cataloging best practices for maintenance routines and methods for prognostic health management, and quantifying the risk and effect of different failure modes. Several key priorities are identified for streamlined data collection to achieve quality and detailed failure data: Applicability, Ease of Use, Accessibility, and Information Security. The HyCReD team has now begun deployment of the database to several companies and groups that have signed non-disclosure agreements to facilitate the data collection of failures in industry hydrogen refueling station infrastructure. This paper will provide an update into the process of HyCReD deployment including the development of a coding guide for facility personnel to reference and ensure data quality and consistency from one station to another as well as implementation of contextually dependent data fields of system taxonomy and formatted entries to provide ease of use. The goal is to communicate the lessons learned from the roll-out to technicians and engineers in the field, and the addition of need for high level of security to protect all stakeholders.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Evaluation of BISON metallic fuel performance modeling against experimental measurements within FIPD and IMIS databases

Simulations were conducted using the BISON fuel performance code on an automated process to read initial and operating conditions from two databases—the Fuels Irradiation and Physics Database (FIPD) and Integral Fast Reactor Materials Information System (IMIS) database. These databases contain metallic fuel data from the Experimental Breeder Reactor-II (EBR-II) and the Fast Flux Test Facility (FFTF). The work demonstrates use of an integrated framework to access EBR-II fuel pin data for evaluating fuel performance models contained within BISON to predict fuel performance of next-generation metallic fuel systems. Between IMIS and FIPD, there is enough information to conduct 1,977 unique EBR-II metallic fuel pin histories from 29 different experiments, and 338 pins from FFTF MFF-3 and MFF-5 with varying levels of details between the two databases. Each of these fuel performance histories includes a high-resolution power history, flux history, coolant channel flow rates, and coolant channel temperatures, and new model developments in BISON since the initial demonstration of this integrated framework. Fission gas release (FGR), cumulative damage fraction, fuel axial swelling, FCCI wastage thickness, cladding profilometry, and burnup were all simulated in BISON and compared to post-irradiation examination (PIE) results to evaluate BISON fuel performance modeling. Implementation of new fuel performance models into a generic BISON input file coupled with IMIS and FIPD yielded results with a better representation of physics than the initial evaluation of the integrated framework. Cladding profilometry, FGR, and fuel axial swelling were found to be in good agreement with PIE measurements for most of the pins simulated. The chosen mechanical contact solver was found to significantly impact the axial fuel swelling and cladding strain predictions when used in conjunction with the U-Pu-Zr hot-pressing model since it bound the fuel to prevent further swelling and increased hydrostatic stresses. This work suggests that fuel performance modeling in BISON under steady-state conditions represents the PIE data well and should be reassessed when new PIE data become available in IMIS and FIPD databases and when improved physical models to better capture fuel performance are added to BISON.

Paaren, Kyle M.↗