Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “databases)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Existing Hydropower Assets (EHA) Capacity Plant Database, 2005-2024

Existing Hydropower Asset (EHA) Annual Capacity is a geospatial point-level dataset containing annual capacity over the years (2005-2024) and key characteristics of operational U.S. hydropower plants with 1 megawatt or greater of nameplate capacity. EIA form 860 and EHA are the primary sources of the derived data.

Johnson, Megan [ORNL] (ORCID:0000000290141741)↗

Existing Hydropower Assets (EHA) Capacity Factor Plant Database, 2005-2024

Existing Hydropower Asset (EHA) Annual Capacity Factor is a geospatial point-level dataset containing annual capacity factors over the years (2005-2024) and key characteristics of operational U.S. hydropower plants with 1 megawatt or greater of nameplate capacity. EIA form 860 and EHA are the primary sources of the derived data. Pumped storage and hybrid plants are excluded.

Johnson, Megan [ORNL] (ORCID:0000000290141741)↗

Comprehensive Database of Environmental Mitigations Extracted from FERC-Licensed Hydropower Projects Using Artificial Intelligence Techniques, 1998-2023

This dataset provides a comprehensive inventory of environmental mitigation measures required by Federal Energy Regulatory Commission (FERC) licensed hydropower facilities from 461 licenses that were issued from 1998 to 2023. These licenses constitute 446 of the 1015 FERC projects that were active at the end of 2023. 17,612 mentions of environmental mitigations were identified and categorized in 128 unique categories. Mitigations were identified using a Natural Language Processing (NLP) approach, specifically with a Bidirectional Encoder Representations from Transformer (BERT) model. Model-derived results were then reviewed and updated by a subject matter expert as needed. This dataset introduces important enhancements to previous efforts to inventory environmental mitigations, such as including associated license text for each mitigation, tracking the number of instances a mitigation was identified within a license, and providing improved location information. These enhancements significantly expand the dataset's utility, offering greater analytical capabilities and ensuring reproducibility. The dataset is downloadable as a zip file containing the metadata and dataset files.

Ruggles, Thomas [Oak Ridge National Laboratory (OR↗

Optimizing metaproteomics database construction: lessons from a study of the vaginal microbiome

Metaproteomics, a method for untargeted, high-throughput identification of proteins in complex samples, provides functional information about microbial communities and can tie functions to specific taxa. Metaproteomics often generates less data than other omics techniques, but analytical workflows can be improved to increase usable data in metaproteomic outputs. Identification of peptides in the metaproteomic analysis is performed by comparing mass spectra of sample peptides to a reference database of protein sequences. Although these protein databases are an integral part of the metaproteomic analysis, few studies have explored how database composition impacts peptide identification. Here, we used cervicovaginal lavage (CVL) samples from a study of bacterial vaginosis (BV) to compare the performance of databases built using six different strategies. We evaluated broad versus sample-matched databases, as well as databases populated with proteins translated from metagenomic sequencing of the same samples versus sequences from public repositories. Smaller sample-matched databases performed significantly better, driven by the statistical constraints on large databases. Additionally, large databases attributed up to 34% of significant bacterial hits to taxa absent from the sample, as determined orthogonally by 16S rRNA gene sequencing. We also tested a set of hybrid databases which included bacterial proteins from NCBI RefSeq and translated bacterial genes from the samples. These hybrid databases had the best overall performance, identifying 1,068 unique human and 1,418 unique bacterial proteins, ~30% more than a database populated with proteins from typical vaginal bacteria and fungi. Our findings can help guide the optimal identification of proteins while maintaining statistical power for reaching biological conclusions.

59 BASIC BIOLOGICAL SCIENCES↗

Addressing the dynamic nature of reference data: a new nucleotide database for robust metagenomic classification

Accurate metagenomic classification relies on comprehensive, up-to-date, and validated reference databases. While the NCBI BLAST Nucleotide (nt) database, encompassing a vast collection of sequences from all domains of life, represents an invaluable resource, its massive size—currently exceeding 10 12 nucleotides—and exponential growth pose significant challenges for researchers seeking to maintain current nt-based indices for metagenomic classification. Recognizing that no current nt-based indices exist for the widely used Centrifuge classifier, and the last public version currently available was released in 2018, we addressed this critical gap by leveraging advanced high-performance computing resources. We present new Centrifuge-compatible nt databases, meticulously constructed using a novel pipeline incorporating different quality control measures, including reference decontamination and filtering. These measures demonstrably reduce spurious classifications, as shown through our reanalysis of published metagenomic data where Plasmodium annotations were dramatically reduced using our decontaminated database, highlighting how database quality can significantly impact research conclusions. Through temporal comparisons, we also reveal how our approach minimizes inconsistencies in taxonomic assignments stemming from asynchronous updates between public sequence and taxonomy databases. These discrepancies are particularly evident in taxa such as Listeria monocytogenes and Naegleria fowleri, where classification accuracy varied significantly across database versions. These new databases, made available as pre-built Centrifuge indexes, respond to the need for an open, robust, nt-based pipeline for taxonomic classification in metagenomics. Applications such as environmental metagenomics, forensics, and clinical metagenomics, which require comprehensive taxonomic coverage, will benefit from this resource. Our work highlights the importance of treating reference databases as dynamic entities, subject to ongoing quality control and validation akin to software development best practices. This approach is crucial for ensuring accuracy and reliability of metagenomic analysis, especially as databases continue to expand in size and complexity.

59 BASIC BIOLOGICAL SCIENCES↗

Carbon Storage Technical Viability Approach (CS TVA) Database

The Carbon Storage Technical Viability Approach (CS TVA) database was developed to support the implementation of the CS TVA Matrix to a national data availability assessment for technically viable carbon storage. This database leverages the efforts of multiple adjacent and overlapping databases by non-redundantly combining the databases into a single database along with additionally providing tags facilitating the CS TVA. The non-redundant aspect of the database permits an accurate assessment of the concentration of available data, aiding in spatial and categorical data gaps analysis relative to the individual CS TVA Matrix Components. Version 2.0 of the database is an expansion of Version 1.0. Version 2.0 was created to include additional data gathered to fill gaps in the existing data set. Downloading the CS TVA v2.0 database will result in two separate databases, the version 1.0 original .gdb, and a second addendum .gdb with the new data gathered, together these two databases make up v2.0. Please see the ReadMe file below for full details, metadata information, use disclaimer, and attributions.

Coal↗

WELLS Database

The Wellbore Exploration and Location Logistic System (WELLS) is a living national wellbore database - created and maintained by the National Energy Technology Laboratory (NETL). This resource contains more than seven million public wellbore records from state, federal, and tribal resources. Sourced from over 65 authoritative, yet disparate resources, the WELLS Database combines and synthesizes well data from oil, gas, underground injection, research, geothermal, geotechnical, groundwater and other types of wells in a single, unified system. This resource can be explored and visualized through the WELLS Interactive Application, also on EDX: https://edx.netl.doe.gov/dataset/wells-interactive-application The WELLS Database (formerly titled CO2-Locate) is an integrated national well dataset, representing open-source wellbore data from disparate state, tribal, and federal entities. The database provides publicly available well header data with key attributes such as well age, depth, and status. The database contains a fully integrated CSV file with all values in numerical columns, such as depth, converted into numbers. This version has a NETL derived API (American Petroleum Institute) number column and has been handled for redundancies, resulting in one record for every unique API number. The database also contains a fully integrated CSV file, where all original data are kept as text values. Additionally, the database includes a shapefile containing key attributes and coordinates from the integrated dataset, reformatted public wells CSV files, and a proprietary well density grid shapefile. Notes for consideration: The WELLS Database will be updated periodically with new datasets and information. A field dictionary with field (i.e., attribute) coverage across acquired public well resources, and the resulting integrated public well datasets are available in the spreadsheet, WELLS_Field_Dictionary.xlsx. Summary layers provided in this database are derived from proprietary layers and do not always contain key features (status, type, true vertical depth, or spud year) and therefore might not be shown when data are queried for those features.

AS↗

Design and requirements of a hydrogen component reliability database (HyCReD)

Hydrogen technologies are expected to play a key role in the decarbonization of several sectors including energy storage and transportation. Rigorous investigation and quantification of the risk and reliability issues associated with hydrogen technologies will be critical to ensuring both their wider adoption and safe, economical operations. Quantitative risk assessment (QRA) is an important tool that has been used to enable the safe deployment of many engineering systems, including hydrogen fueling stations and hydrogen storage systems. However, QRA studies require reliability data which is currently lacking for expanding applications of hydrogen systems. Here, to address this gap, we present a new structure for a hydrogen component reliability database (HyCReD) that can be used to generate reliability data to be used in QRA, reliability, safety studies, maintenance planning, and more. Building on our previous work examining four major hydrogen safety data collection tools (West et al., 2022) [1], our approach in this work was to consult scientific literature on reliability data collection as well as a number of existing reliability engineering databases in the oil & gas, chemical processing, and nuclear power plant sectors. The evaluation of these databases led to identifying best practices to be implemented in a data collection framework for a hydrogen component reliability database. Based on these best practices, a set of 24 requirements for the proposed database are presented, covering its characteristics and the types of data to be collected. We define the structure of the HyCReD database and 25 data elements to be collected, spanning system description, failure, shutdown, or near-miss events, and maintenance events. The data elements are then defined according to international standards used in the safety and reliability practice and potential choice lists are provided for each field. Since this database is being piloted for hydrogen fueling stations, a generic station component hierarchy developed by West (2021) [2] is used to standardize system data. Finally, we demonstrate populating the database with information extracted from five narrative reports on hydrogen fueling station incidents.

08 HYDROGEN↗

TropiRoot 1.0: Database of tropical root characteristics across environments

Tropical ecosystems contain the world's largest biodiversity of vascular plants. Yet, our understanding of tropical functional diversity and its contribution to global diversity patterns is constrained by data availability. This discrepancy underscores an urgent need to bridge data gaps by incorporating comprehensive tropical root data into global datasets. Here, we provide a database of tropical root characteristics. This new database, TropiRoot 1.0, will be instrumental in evaluating an array of hypotheses pertaining to root functional ecology and plant biogeography, both within the tropics and relative to other global biomes. The data compilation was conducted by the TropiRoot Initiative, in partnership with the Fine-Root Ecology Database (FRED) and the Global Root Trait (GRooT) database, Colorado State University (CSU) and the Smithsonian Tropical Research Institute (STRI). Literature search and data extraction were conducted between 2020 and 2024. Literature was identified using Web of Science, Scopus, and complemented using the expert knowledge of members of TropiRoot. To provide broad environmental and geographical distributions, literature searches included root characteristics (traits) across global change drivers, natural gradients, and from different continents. We adopted FRED standardized data columns and streamlined the format to enhance accessibility for data extraction across various user groups. This optimized framework resulted in a smaller, yet comprehensive datasheet. To make the database compatible with other global root trait initiatives, column identification was standardized following the codes provided by FRED. These efforts culminated in data extracted from 104 new sources, resulting in more than 8000 rows of data (either species or community data). Most of the data in TropiRoot 1.0 include root characteristics such as root biomass, morphology, root dynamics, mass fraction, architecture, anatomy, physiology, and root chemistry. This initiative represents a 30% increase in the currently available data for tropical roots in FRED. TropiRoot 1.0 contains root characteristics from 25 different countries, where seven are located in Asia, six in South America, five in Central America and the Caribbean, four in Africa, two in North America, and 1 in Oceania. Due to the volume of data, when ancillary data were available, including soil data, these data were either extracted and included in the database or its availability was recorded in an additional column. Multiple contributors checked the entries for outliers during the collation process to ensure data quality. For text-based observations, we examined all cells to ensure that their content relates to their specific categories. For numerical observations, we ordered each numerical value from least to greatest and plotted the values, checking apparent outliers against the data in their respective sources and correcting or removing incorrect or impossible values. Some data (soil and aboveground) have different columns for the same variable presented in different units, including originally published units, but root characteristics data had units converted to match those reported in FRED. By filling a gap from global databases, TropiRoot 1.0 expands our knowledge of otherwise so far underrepresented regions and our ability to assess global trends. This advancement can be used to improve tropical forest representation in vegetation models. The data are freely available and should be cited when used.

FRED↗

Cross-database comparisons on the greenhouse gas emissions, water consumption, and fossil-fuel use of plastic resin production and their post-use phase impacts

Resin production and post-use pathways have cross-database discrepancies due to different temporal/geographic representation, technologies, and assumptions. These differences can distort comparisons across alternatives and confound efforts to establish standards based on the plastics’ life-cycle impacts. Thus, this study quantifies the degree and identifies the sources of cross-database discrepancies across four LCA databases (GREET, USLCI, Ecoinvent, and GaBi) for five resin production pathways and three post-use phases. For resin production pathways, all resins showed significant cross-database discrepancy in their global warming impacts: the degree of discrepancy was significant to recommend a consistent choice of database to users any LCA comparisons across different products. For post-use phases, landfill datasets had relatively lower degree of cross-database discrepancy than incineration and mechanical recycling. Different metadata characteristics were the sources of some cross-database discrepancies while other parts could be explained by the original differences in the life-cycle inventory sourced from different producers and plants.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

SoK: What does it Mean to Benchmark Database Forensics?

Relational Database Management Systems are the backbone of modern enterprises and public-sector services, and are thus frequent targets of security incidents, insider threats, and thorough regulatory audits. Consequently, databases have become key sources of digital evidence, requiring investigators to reconstruct past activity from audit logs, transaction logs, and backups. Although benchmarking frameworks such as those developed by the Transaction Processing Performance Council (TPC) are widely used to evaluate database performance, they do not capture forensic requirements such as evidentiary completeness, tamper-evidence, chain of custody, or regulatory compliance under GDPR and CCPA. This survey examines the emerging domain of forensic database benchmarking. We gathered prior research on database forensics, secure logging, and tamper-evident data structures; we analyze modern forensic-ready features in commercial and open-source systems (SQL Server Ledger, Oracle Blockchain Tables, PostgreSQL pgAudit, Db2 Audit, Aurora Database Activity Streams, Oracle Real Application Security and IBM Guardium) and assess why existing benchmarks are insufficient. We propose forensic workloads, metrics, and methodologies that incorporate adversarial stressors, deleted-record recovery, and backup analysis. We also identify open research problems and call for a community-driven forensic benchmark suite. The result is an idea for evaluating not only database performance but also forensic soundness, bridging the gap between system engineering, compliance, and digital investigations.

Lenard, Ben↗

efam: an e xpanded, metaproteome-supported HMM profile database of viral protein fam ilies

Viruses infect, reprogram and kill microbes, leading to profound ecosystem consequences, from elemental cycling in oceans and soils to microbiome-modulated diseases in plants and animals. Although metagenomic datasets are increasingly available, identifying viruses in them is challenging due to poor representation and annotation of viral sequences in databases. Here, we establish efam, an expanded collection of Hidden Markov Model (HMM) profiles that represent viral protein families conservatively identified from the Global Ocean Virome 2.0 dataset. This resulted in 240 311 HMM profiles, each with at least 2 protein sequences, making efam >7-fold larger than the next largest, pan-ecosystem viral HMM profile database. Adjusting the criteria for viral contig confidence from ‘conservative’ to ‘eXtremely Conservative’ resulted in 37 841 HMM profiles in our efam-XC database. To assess the value of this resource, we integrated efam-XC into VirSorter viral discovery software to discover viruses from less-studied, ecologically distinct oxygen minimum zone (OMZ) marine habitats. This expanded database led to an increase in viruses recovered from every tested OMZ virome by ~24% on average (up to ~42%) and especially improved the recovery of often-missed shorter contigs (<5 kb). Additionally, to help elucidate lesser-known viral protein functions, we annotated the profiles using multiple databases from the DRAM pipeline and virion-associated metaproteomic data, which doubled the number of annotations obtainable by standard, single-database annotation approaches. Together, these marine resources (efam and efam-XC) are provided as searchable, compressed HMM databases that will be updated bi-annually to help maximize viral sequence discovery and study from any ecosystem.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The User Guide for the ComPro Database

An extensive database of glass composition and durability data has been compiled at Savannah River National Laboratory to support the development of nuclear waste glasses. This database is referred to as the Glass Composition-Properties Database (ComPro). The ComPro Database, Revision 3, contains 14,134 total rows of data and 125 columns of composition, durability, as defined by the Product Consistency Test, and other fabrication and characterization information, if available, for each glass. Of the 14,134 total rows, 8,484 rows have been classified as “Model” data and 5,650 rows have been classified as “Non-Model” data. An integral supplement to the ComPro database is the User Guide. The User Guide was developed as a tool to aid the End User in a more effective use of the ComPro database. The User Guide provides a road-map of the specific datasets that comprise the ComPro database (both “Model” and “Non-Model” data) as well as a technical basis for the terminology and definitions the End User will encounter. In this report, a general description of the format and information contained in the User Guide is provided. In addition, specific terminology used in the User Guide is also discussed.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

SERDP PFAS Database

The PNNL-PFAS database was constructed to provide quantitative infrared spectra of per- and polyfluoroalkyl substances (PFAS) in the gas phase. These fluorine-rich, organic compounds have been used as non-flammable solvents, cleaning agents, and in the manufacture of a diverse variety of consumer products. Due to their widespread use and relative inertness, PFAS have accumulated in the natural environment, and exposure to these chemicals may be linked to harmful health effects. Development of methods for detecting and remediating PFAS contamination are thus active areas of research. This database supports PFAS investigations where gas-phase infrared signatures may be measured (e.g., during treatment by thermal decomposition). The database currently contains infrared spectra for 15 compounds; these were selected based on the potential of these compounds being thermal degradation products of the manufactured chemicals mentioned above. • Additional compounds will be added to the database in the future. • Additional information about the database may be obtained from [link to future paper about the PFAS database].

PFAS, infrared(IR)spectroscopy, DATABASE, FTIR, sp↗

A new database of building-space-specific internal loads and load schedules for performance based code compliance modeling of commercial buildings

Building-level loads and load profiles prescribed by current modeling rules save modelers time and avoid gaming during whole building performance modeling. However, recent studies show that they sometimes insufficiently capture the entire building performance due to the varied loads and load profiles for different space types. As a solution to this issue, this paper develops a database of building-space-specific loads and load profiles used in code compliance modeling. The existing sets of loads and load profiles are reviewed and the challenges behind using them for specific research topics are discussed. Then, the proposed method to develop the building-space-specific loads and load profiles is introduced. After that, the database for these building-space-specific loads and load profiles is presented. In addition, one case is studied to demonstrate the applications of these loads and load profiles. In this case study, three methods are used to develop building energy models: space-specific (using knowledge of the distribution and location of space types and applying the space-specific data in the developed database), building-level (assuming a lack of knowledge of the space types and using the building-level data in the developed database), and calculated-ratio (assuming knowledge of the distribution of space types but not their locations and calculating weighted average values based on the space-specific data in the developed database). Finally, the energy results simulated by using these three methods are compared, which show building-level methods can produce energy results up to 20% different than the space-specific methods. Finally, this paper discusses the application scope and maintenance of this new database.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Auto-generating databases of Yield Strength and Grain Size using ChemDataExtractor

Abstract The emerging field of material-based data science requires information-rich databases to generate useful results which are currently sparse in the stress engineering domain. To this end, this study uses the’materials-aware’ text-mining toolkit, ChemDataExtractor, to auto-generate databases of yield-strength and grain-size values by extracting such information from the literature. The precision of the extracted data is 83.0% for yield strength and 78.8% for grain size. The automatically-extracted data were organised into four databases: a Yield Strength, Grain Size, Engineering-Ready Yield Strength and Combined database. For further validation of the databases, the Combined database was used to plot the Hall-Petch relationship for, the alloy, AZ31, and similar results to the literature were found, demonstrating how one can make use of these automatically-extracted datasets.

36 MATERIALS SCIENCE↗

A Global Building Occupant Behavior Database

This paper introduces a database of 34 field-measured building occupant behavior datasets collected from 15 countries and 39 institutions across 10 climatic zones covering various building types in both commercial and residential sectors. This is a comprehensive global database about building occupant behavior. The database includes occupancy patterns (i.e., presence and people count) and occupant behaviors (i.e., interactions with devices, equipment, and technical systems in buildings). Brick schema models were developed to represent sensor and room metadata information. The database is publicly available, and a website was created for the public to access, query, and download specific datasets or the whole database interactively. The database can help to advance the knowledge and understanding of realistic occupancy patterns and human-building interactions with building systems (e.g., light switching, set-point changes on thermostats, fans on/off, etc.) and envelopes (e.g., window opening/closing). With these more realistic inputs of occupants’ schedules and their interactions with buildings and systems, building designers, energy modelers, and consultants can improve the accuracy of building energy simulation and building load forecasting.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗