Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Existing Hydropower Assets (EHA) Annual Gross Generation Plant Database, 2003-2024

Existing Hydropower Asset (EHA) Annual Gross Generation is a geospatial point-level dataset containing annual gross generation over time (2003-2024) and key characteristics of operational U.S. pumped storage and hybrid plants with 1 megawatt or greater of nameplate capacity. EIA 923 and EHA are the primary sources of the derived data. Hydropower units are excluded.

Johnson, Megan [ORNL] (ORCID:0000000290141741)↗

Existing Hydropower Assets (EHA) Annual Net Generation Plant Database, 2003-2024

Existing Hydropower Asset (EHA) Annual Net Generation is a geospatial point-level dataset containing annual net generation over time (2003-2024) and key characteristics of operational U.S. hydropower plants with 1 megawatt or greater of nameplate capacity. EIA 923 and EHA are the primary sources of the derived data. Pumped storage and hybrid plants are excluded.

Johnson, Megan [ORNL] (ORCID:0000000290141741)↗

BEAST DB: Grand-Canonical Database of Electrocatalyst Properties

We present BEAST DB, an open-source database comprised of ab initio electrochemical data computed using grand-canonical density functional theory in implicit solvent at consistent calculation parameters. The database contains over 20,000 surface calculations and covers a broad set of heterogeneous catalyst materials and electrochemical reactions. Calculations were performed at self-consistent fixed potential as well as constant charge to facilitate comparisons to the computational hydrogen electrode. This article presents common use cases of the database to rationalize trends in catalyst activity, screen catalyst material spaces, understand elementary mechanistic steps, analyze the electronic structure, and train machine learning models to predict higher fidelity properties. Users can interact graphically with the database by querying for individual calculations to gain a granular understanding of reaction steps or by querying for an entire reaction pathway on a given material using an interactive reaction pathway tool. BEAST DB will be periodically updated, with planned future updates to include advanced electronic structure data, surface speciation studies, and greater reaction coverage.

database↗

A decay database of coincident γ–γ and γ–X -ray branching ratios for in-field spectroscopy applications

Current fieldable spectroscopy techniques often use single detector systems heavily impacted by interferences from intense background radiation fields. These effects result in low-confidence measurements that can lead to misinterpretation of the collected spectrum. To help improve interpretation of the fission products and short-lived radionuclides produced in a composite sample, a coincidence-database is being developed in support of a robust portable and X-ray coincidence detector system concurrently under development at the Pacific Northwest National Laboratory for in-field deployment. Hitherto, no database exists containing coincident γ–γ and γ–X-ray branching-ratio intensities on an absolute scale that will greatly enhance isotopic identification for in-field applications. As part of this project, software has been developed to parse all radioactive-decay data sets from the Evaluated Nuclear Structure Data File (ENSDF) archive to enable translation into a more useful JavaScript Object Notation (JSON) formats that more readily supports query-based data manipulation. The coincident database described in this work is the first of its kind and contains coincidence γ–γ and γ–X-ray intensities and their corresponding uncertainties, together with auxiliary metadata associated with each decay data set. The new JSON format provides a convenient and portable means of data storage that can be imported into analysis frameworks with relatively low overhead allowing for meaningful comparison with measured data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

NMR Database of Lignin and Cell Wall Model Compounds

This database was designed to provide a coherent, single source of NMR data of lignin and other plant cell wall model compounds. The database exists as an Adobe pdf cross-platform file for viewing and printing. This is the latest public version of the Database, version 2024/08 updated from the 2009 version.

cell wall↗

Integration of the NCRC Database and Other INL Databases

The Nuclear Computational Resource Center provides a portal by which industry professionals, educational staff, students, national laboratory employees, and others may request access to certain engineering software tools. As the tools provided through the Nuclear Computational Resource Center portal are not open-source and freely available, a set of approvals are necessary before access is granted. All code recipients must be associated with an institution that has a license with Idaho National Laboratory for the code requested. Information about these licenses is controlled by Idaho National Laboratory’s Technology Deployment organization and housed in a Technology Deployment database. Those requesting code access who are not citizens of the United States must also have a security plan, mandated by Idaho National Laboratory policy. Security plans are managed by the International Access Program and are stored in an International Access Program database known as IFacts. Granting access to software thus depends on information stored in the Technology Deployment database and IFacts. In the past, no connection between the Nuclear Computational Resource Center portal and these databases existed, making checking the status of license agreements and security plans time consuming and error prone. This report demonstrates that the Nuclear Computational Resource Center portal now connects to both the Technology Deployment database and IFacts, greatly improving the ease of use of the Nuclear Computational Resource Center system for administrators, which leads to a better overall experience for those requesting code access.

99 GENERAL AND MISCELLANEOUS↗

Scalability Testing Approach for Internet of Things for Manufacturing SQL and NoSQL Database Latency and Throughput

The proliferation of low-cost sensors and industrial data solutions has continued to push the frontier of manufacturing technology. Machine learning and other advanced statistical techniques stand to provide tremendous advantages in production capabilities, optimization, monitoring, and efficiency. The tremendous volume of data gathered continues to grow, and the methods for storing the data are critical underpinnings for advancing manufacturing technology. This work aims to investigate the ramifications and design tradeoffs within a decoupled architecture of two prominent database management systems (DBMS): sql and NoSQL. A representative comparison is carried out with Amazon Web Services (AWS) DynamoDB and AWS Aurora MySQL. The technologies and accompanying design constraints are investigated, and a side-by-side comparison is carried out through high-fidelity industrial data simulated load tests using metrics from a major US manufacturer. The results support the use of simulated client load testing for comparing the latency of database management systems as a system scales up from the prototype stage into production. As a result of complex query support, MySQL is favored for higher-order insights, while NoSQL can reduce system latency for known access patterns at the expense of integrated query flexibility. Here, by reviewing this work, a manufacturer can observe that the use of high-fidelity load testing can reveal tradeoffs in IoTfM write/ingestion performance in terms of latency that are not observable through prototype-scale testing of commercially available cloud DB solutions.

AWS↗

Reservoir Sediment Management and Monitoring Database

Overview This dataset compiles dam sediment management and monitoring information from surveys, case studies, and journal articles. Additionally, features described by the National Inventory of Dams (i.e., presence of sluice gates) are included to indicate known infrastructure features that may address sediment releases. The location and description of records from downstream monitoring gages are catalogued in order to help with tracking conditions over time (e.g., before and after management actions, as operations change, etc.). The data help address national scale understanding of challenges and solutions related to the accumulation of sediment behind a dam as well as downstream passage. Sediment trapping causes problems as it reduces storage capacity, disrupts dam and reservoir function, impedes access for recreation, alters water quality/habitat conditions, and contributes to riverbank and coastal erosion within the reservoir. Data compilation from a variety of sources is a first step towards assessing system-wide efficacy of management solutions. This dataset was developed under the Water Power Technologies Office funded effort which began as a Seedling on Reservoir Sedimentation Data, and was supported by the Reservoir Sedimentation Modeling Framework and Data Analysis project. These projects have addressed challenges in describing sediment transport, trapping, and management at dams throughout the US. Methodology An outer join on dams/reservoirs with surveys and survey reports (documented in the RESSED database, USBR or USACE databases, project websites, etc.) with the National Inventory of Dams, based on the NIDID to determine dams with documented management and/or sluice gates. Additional dams with documented management activity were identified through review of technical articles from the past 25 years in Journal of Hydrology, Journal of Water Resources Planning and Management, Geomorphology, Journal of Hydraulic Engineering, Water, Journal of Cleaner Production, International Journal of Sediment Research, Nature Scientific Reports, Earth Surface Processes and Landforms, and Environmental Science and Pollution Research. Individual records were created for each survey or management activity documented. To evaluate downstream sediment monitoring records, the nhdPlusTools and dataRetrieval packages in R were used to find gages within 10km of each dam in the management database. Length of record and location of matched gages were retrieved for those parameters relevant to sediment concentration or total sediment discharge.

Hansen, Carly [ORNL] (ORCID:0000000193280838)↗

Produced Water DNA Database (PW-DNA): Utilizing KBase to generate an environmental specific curated molecular database

The deep subsurface is estimated to host the majority of Earth’s microbial biomass yet remains one of the most challenging environments to access and study. One common approach to investigate these microbial communities is through the analysis of produced water from subsurface reservoirs, where researchers can assess water and gas chemistry along with molecular (DNA/RNA) sequence data. Advances in high-throughput sequencing have greatly expanded our understanding of these environments and their biotechnological potential. However, further progress requires large-scale, integrative meta-analyses across diverse datasets. To address this need, we developed the Produced Water-DNA (PW-DNA) Database, a curated, publicly available resource that consolidates microbial DNA/RNA sequences, geochemical data, and relevant metadata from in situ hydrocarbon environments such as coal beds, oil reservoirs, and natural gas systems. The PW-DNA database delivers three core benefits to the research community: (1) it improves data sharing by linking environmental microbial datasets with corresponding geochemical parameters, enabling more robust filtering and analysis; (2) it connects with complementary research databases to promote broader dissemination and interoperability; and (3) it supports technological innovation by serving as a resource for identifying microbial trends and exploring genetic potential. While individual studies have highlighted basin-specific microbial communities and functional redundancy in biogeochemical cycling, a comprehensive, system-wide perspective is needed to better understand connectivity and novelty across subsurface ecosystems. By designing the PW-DNA in the KBase platform, we provide a reproducible, visual framework for integrating large-scale genomic and geochemical data, enabling researchers to perform more informed analyses and experimental design. Ultimately, this resource enhances the ability to identify, characterize, and interpret microbial functions across diverse subsurface environments, thereby accelerating discovery in subsurface microbiology and biotechnology.

59 BASIC BIOLOGICAL SCIENCES↗

HTGR Validation: NEUP Survey and Database Database Development for HTGR Thermal-Fluid Experiments

This set of slides is aimed to improve access to the High-Temperature Gas-cooled Reactor (HTGR) validation data and optimize the return on the significant investment made by DOE. Supported by the Advanced Reactor Technologies (ART) Gas-Cooled Reactor (GCR) program. Slides include information from FY2009 to FY2023, there are in total 35 DOE NEUP projects focusing on the thermal-fluid experiments related with High-Temperature Gas-cooled Reactor (HTGR), producing a large amount of high-quality validation data. NEUP Survey and Database Development for HTGR Thermal-Fluid Experiments is distributed at universities and has not been disseminated to the HTGR community. More collaborating with university PIs and refining the HTGR phenomena summary chart continuously is expected in the future.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Existing Hydropower Assets (EHA) Capacity Plant Database, 2005-2024

Existing Hydropower Asset (EHA) Annual Capacity is a geospatial point-level dataset containing annual capacity over the years (2005-2024) and key characteristics of operational U.S. hydropower plants with 1 megawatt or greater of nameplate capacity. EIA form 860 and EHA are the primary sources of the derived data.

Johnson, Megan [ORNL] (ORCID:0000000290141741)↗

Existing Hydropower Assets (EHA) Capacity Factor Plant Database, 2005-2024

Existing Hydropower Asset (EHA) Annual Capacity Factor is a geospatial point-level dataset containing annual capacity factors over the years (2005-2024) and key characteristics of operational U.S. hydropower plants with 1 megawatt or greater of nameplate capacity. EIA form 860 and EHA are the primary sources of the derived data. Pumped storage and hybrid plants are excluded.

Johnson, Megan [ORNL] (ORCID:0000000290141741)↗

Comprehensive Database of Environmental Mitigations Extracted from FERC-Licensed Hydropower Projects Using Artificial Intelligence Techniques, 1998-2023

This dataset provides a comprehensive inventory of environmental mitigation measures required by Federal Energy Regulatory Commission (FERC) licensed hydropower facilities from 461 licenses that were issued from 1998 to 2023. These licenses constitute 446 of the 1015 FERC projects that were active at the end of 2023. 17,612 mentions of environmental mitigations were identified and categorized in 128 unique categories. Mitigations were identified using a Natural Language Processing (NLP) approach, specifically with a Bidirectional Encoder Representations from Transformer (BERT) model. Model-derived results were then reviewed and updated by a subject matter expert as needed. This dataset introduces important enhancements to previous efforts to inventory environmental mitigations, such as including associated license text for each mitigation, tracking the number of instances a mitigation was identified within a license, and providing improved location information. These enhancements significantly expand the dataset's utility, offering greater analytical capabilities and ensuring reproducibility. The dataset is downloadable as a zip file containing the metadata and dataset files.

Ruggles, Thomas [Oak Ridge National Laboratory (OR↗

Optimizing metaproteomics database construction: lessons from a study of the vaginal microbiome

Metaproteomics, a method for untargeted, high-throughput identification of proteins in complex samples, provides functional information about microbial communities and can tie functions to specific taxa. Metaproteomics often generates less data than other omics techniques, but analytical workflows can be improved to increase usable data in metaproteomic outputs. Identification of peptides in the metaproteomic analysis is performed by comparing mass spectra of sample peptides to a reference database of protein sequences. Although these protein databases are an integral part of the metaproteomic analysis, few studies have explored how database composition impacts peptide identification. Here, we used cervicovaginal lavage (CVL) samples from a study of bacterial vaginosis (BV) to compare the performance of databases built using six different strategies. We evaluated broad versus sample-matched databases, as well as databases populated with proteins translated from metagenomic sequencing of the same samples versus sequences from public repositories. Smaller sample-matched databases performed significantly better, driven by the statistical constraints on large databases. Additionally, large databases attributed up to 34% of significant bacterial hits to taxa absent from the sample, as determined orthogonally by 16S rRNA gene sequencing. We also tested a set of hybrid databases which included bacterial proteins from NCBI RefSeq and translated bacterial genes from the samples. These hybrid databases had the best overall performance, identifying 1,068 unique human and 1,418 unique bacterial proteins, ~30% more than a database populated with proteins from typical vaginal bacteria and fungi. Our findings can help guide the optimal identification of proteins while maintaining statistical power for reaching biological conclusions.

59 BASIC BIOLOGICAL SCIENCES↗