Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

An infrared spectral database for gas-phase quantitation of volatile per- and polyfluoroalkyl substances (PFAS)

We report the construction of a database of vetted infrared spectra specifically targeting volatile fluorocarbon gases that may be emitted during thermal treatment of per- and polyfluoroalkyl substances (PFAS) to assist understanding of treatment processes and improve quantification. To populate this database, protocols derived from the Pacific Northwest National Laboratory (PNNL) infrared spectral database are used, curtailing the species selection for this data set. Each spectrum in the database is a weighted average derived from 10 or more individual measurements at different partial pressures (static method) or flow rates (dissemination method) to yield good fidelity of both strong and weak infrared signatures, with each composite spectrum ranging from = 6500 cm-1 to = 600 cm-1 with an apodized resolution of 0.112 cm-1. This resolution was chosen to fully resolve all spectral features, recognizing that atmospheric pressure broadening results in nearly all ro-vibrational lines having linewidths = 0.1 cm-1. As an example case, application of the database is demonstrated via identification and quantification of dominant 1H-perfluoroheptane and perfluorohept-1-ene fluorocarbon products resulting from thermal decomposition of perfluorooctanoate (PFOA) below 450 °C.

Infrared, Gas-phase spectra, FTIR, Spectral databa↗

The Natural Products Magnetic Resonance Database (NP-MRD) for 2025

The Natural Products Magnetic Resonance Database or NP-MRD (https://np-mrd.org) is a comprehensive, freely accessible, web-based resource for the deposition, distribution, extraction and retrieval of nuclear magnetic resonance (NMR) data on natural products. The NP-MRD was initially established to support compound de-replication and data dissemination for the natural products community. However, that community has now grown to include many users from the metabolomics, microbiomics, foodomics and nutrition science fields. Indeed, since its launch in 2021, the NP-MRD has expanded enormously in size, scope and popularity. The current version of NP-MRD now contains nearly 7X more compounds (281,859 vs. 40,908) and 7X more NMR spectra (5.1 million vs. 817,000) than the first release. More specifically, an additional 4.6 million predicted spectra and another 11,000 spectra simulated from experimental chemical shifts were deposited into the database. Likewise, the number of NMR raw spectral data depositions has grown from a 165 spectra per year to more than 10,000 per year. As a result of this expansion, the number of monthly webpage views has grown from 55 to 20,000 and the number of monthly visitors has increased from 7 to 2500. To address this growth and to better support the expanding needs of its diverse community of users, many additional improvements to the NP-MRD have been made. These include significant enhancements to the data submission process, important improvements to the visualization and display of NMR spectra, notable updates to the database’s spectral search utilities and useful additions to support better NMR spectral analysis/prediction. Significant efforts have also been undertaken to remediate and update many of NP-MRD’s database entries. This manuscript describes these database improvements and expansion efforts, along with how they have been implemented and what future upgrades to the NP-MRD are planned.

Artifical Intelligence↗

SERDP PFAS 2.0 - An infrared spectral database for gas-phase quantitation of volatile per- and polyfluoroalkyl substances (PFAS)

We report the construction of a database of vetted infrared spectra specifically targeting volatile fluorocarbon gases that may be emitted during thermal treatment of per- and polyfluoroalkyl substances (PFAS) to assist understanding of treatment processes and improve quantification. To populate this database, protocols derived from the Pacific Northwest National Laboratory (PNNL) infrared spectral database are used, curtailing the species selection for this data set. Each spectrum in the database is a weighted average derived from 10 or more individual measurements at different partial pressures (static method) or flow rates (dissemination method) to yield good fidelity of both strong and weak infrared signatures, with each composite spectrum ranging from = 6500 cm-1 to = 600 cm-1 with an apodized resolution of 0.112 cm-1. This resolution was chosen to fully resolve all spectral features, recognizing that atmospheric pressure broadening results in nearly all ro-vibrational lines having linewidths = 0.1 cm-1. As an example case, application of the database is demonstrated via identification and quantification of dominant 1H-perfluoroheptane and perfluorohept-1-ene fluorocarbon products resulting from thermal decomposition of perfluorooctanoate (PFOA) below 450 °C.

Infrared, Gas-phase spectra, FTIR, Spectral databa↗

G2Aero Database of Airfoils - Curated Airfoils

This dataset contains a curated set of 19,164 airfoil shapes from various applications and the data-driven design space of separable shape tensors (PGA space), which can be used as a parameter space for machine-learning applications focused on airfoil shapes. We constructed the airfoil dataset in two main stages. First, we identified 13 baseline airfoils from the NREL 5MW and IEA 15MW reference wind turbines. We reparameterized these shapes using least-squares fits of 8-order CST parametrizations, which involve 18 coefficients. By uniformly perturbing all 18 CST coefficients by +/-20% around each baseline airfoil, we generated 1,000 unique airfoils. Each airfoil was sampled with 1,001 shape landmarks whose x-coordinates followed a cosine distribution along the chord. This process resulted in a total of 13,000 airfoil shapes, each with 1,001 landmarks. In the second phase, we gathered additional airfoils from the extensive BigFoil database, which consolidates data from sources such as the University of Illinois Urbana-Champaign (UIUC) airfoil database, the JavaFoil database, the NACA-TR-824 database, and others. We undertook a thorough pre-processing step to filter out shapes with sparse, noisy, or incomplete data. We also removed airfoils with sharp leading edge and those exceeding our threshold for trailing edge thickness. Additionally, we thinned out the collection of NACA airfoils-- parametric sweeps of NACA airfoils with increasing thickness and camber present in BigFoil database-- by selecting every fourth step in the parameter sweeps. Finally, we regularized the airfoils by reparametrizing them with an 8-order CST parametrization (with 1,001 shape landmarks with x coordinated following cosine distribution along the chord) and removing airfoils with high reconstruction errors. This data pre-processing resulted in a set of 6,164 airfoils. In total, our curated airfoil dataset comprises 19,164 airfoils, each with 1,001 landmarks, and is stored in the curated_airfoils.npz file. Using this curated airfoil dataset, we utilized the separable shape tensors framework to develop a data-driven parameterization of airfoils based on principal geodesic analysis (PGA) of separable shape tensors. This PGA space is provided in PGAspace.npz file.

airfoils↗

A database of refractive indices and dielectric constants auto-generated using ChemDataExtractor

The ability to auto-generate databases of optical properties holds great potential for advancing optical research, especially with regards to the data-driven discovery of optical materials. An optical property database of refractive indices and dielectric constants is presented, which comprises a total of 49,076 refractive index and 60,804 dielectric constant data records on 11,054 unique chemicals. The database was auto-generated using the state-of-the-art natural language processing software, ChemDataExtractor, using a corpus of 388,461 scientific papers. The data repository offers a representative overview of the information on linear optical properties that resides in scientific papers from the past 30 years. Public availability of these data will enable a quick search for the optical property of certain materials. The large size of this repository will accelerate data-driven research on the design and prediction of optical materials and their properties. To the best of our knowledge, this is the first auto-generated database of optical properties from a large number of scientific papers. We provide a web interface to aid the use of our database.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Perovskite- and Dye-Sensitized Solar-Cell Device Databases Auto-generated Using ChemDataExtractor

The number of scientific publications reporting cutting-edge third-generation photovoltaic devices is increasing rapidly, owing to the pressing need to develop renewable-energy technologies that address the climate-change crisis. Consequently, the field could benefit from a central repository where photovoltaic-performance metrics, such as the power-conversion efficiency (η) are recorded. We present two automatically generated databases that contain photovoltaic properties and device material data for dye-sensitized solar cells (DSCs) and perovskite solar cells (PSCs), totalling 660,881 data entries representing 57,678 photovoltaic devices. The databases were generated by applying the text-mining toolkit ChemDataExtractor on a corpus of 25,720 articles. A multi-faceted evaluation, incorporating manual and automatic methods, was applied to ensure that the data contained therein were of the highest quality, with precision metrics ranging from 73.1% to 95.8%. The DSC database contains 475,045 entries representing 41,680 devices, and the PSC database contains 185,836 entries representing 15,818 devices. The databases are available in MongoDB and JSON formats, which can be queried in Python, R, Java and MATLAB for data-driven photovoltaic materials discovery.

14 SOLAR ENERGY↗

DUNE Database Development

The DUNE experiment will produce vast amounts of metadata, which describe the data coming from the read-out of the primary DUNE detectors. Various databases will make up the overall DB architecture for this metadata. ProtoDUNE at CERN is the largest existing prototype for DUNE and serves as a testing ground for - among other things - possible database solutions for DUNE. The subset of all metadata that is accessed during offline data reconstruction and analysis is referred to as ‘conditions data’ and it is stored in a dedicated database. As offline data reconstruction and analysis will be deployed on HTC and HPC resources, conditions data is expected to be accessed at very high rates. It is therefore crucial to store it in a granularity that matches the expected access patterns allowing for extensive caching. This requires a good understanding of the sources and use cases of conditions data. This contribution will briefly summarize the database architecture deployed at ProtoDUNE and explain the various sources of conditions data. We will present how the conditions data is retrieved and streamed from the databases and how it is handled to match expected access patterns.

Vizcaya Hernandez, Ana Paula↗

Mapping Inquiry Tool (MapIT) Database

The Mapping Inquiry Tool (MapIT) database consists of a geodatabase and data catalog of geologic, geophysical, structural, hydrologic, and contextual data, based on the data types to support geologic carbon storage activities and other subsurface energy systems resource assessments. The database was aggregated from publicly available data across the USA from state and federal entities. The database is structured by categories including rock unit geology, boundaries, national CS datasets, geophysical data, faults and structural data, infrastructure, surface hydrology, groundwater, and more. The data described in the data catalog is also available in the Mapping Inquiry Tool (https://edx.netl.doe.gov/dataset/mapping-inquiry-tool). Version 3 of the geodatabase and data catalog have been updated as of 5/17/2024. The database was published with a limited number of layers. The Catalog V3 contains many more resources than the geodatabase, documenting all layers that will be included in MapIT, and includes links to the original sources of the data. Within the catalog, in the final column, there is information about if the file is included in the geodatabase or not. Use the links provided in the catalog to download data directly from the original source if not included in the geodatabase. Four resources are included in this submission: 1. Geodatabase 2. ReadMe file 3. Catalog of data layers and additional data resources 4. Web link to a resource describing the motivation and reviewing the content of the geodatabase - DOE NETL Carbon Storage Site Mapping Inquiry Tool Database

carbon storage↗

PRO-X Research Reactor Database Status

The PRO-X Research Reactor Database is a functioning tool that can be used to quickly view attributes of past, current, and future research reactors. It contains all data found in the IAEA Research Reactor Database and entries generated from the M3 fuel conversion program. Additionally, it can be downloaded to any computer with the ability to run Microsoft ACCESS. Several research reactor attributes visible to the user have a limited amount of information, much of which is available and vetted for entry into the database. However, the necessary resources have not been allocated to add this information to the data tables. A change management system has also been developed to track changes to the database. This would need to be incorporated into the database program prior to data table modification in order to ensure proper management of the changes.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

An Overview of the Molten Salt Thermal Properties Database--Thermophysical, Version 3.1 (MSTDB-TP v.3.1)

This report presents the current status of the Molten Salt Thermal Properties Database–Thermophysical (MSTDB-TP). Information regarding version 3.1 is provided herein, which contains 820 individual salt entries (data from 180+ independent studies); the thermophysical properties contained in the database include density, viscosity, thermal conductivity, and heat capacity. The major updates to the database include a significant expansion of pseudobinary and higher-order chloride salt mixtures, many of which bearing actinides, and an incorporation of more recent literature data (i.e., that within the past 5 years). Also, modifications have been made to the pure compound data in the database as a consequence of an external quality assessment of duplicate datasets. The user-facing API for the MSTDB-TP, Saline, has been updated to include viscosity estimation capabilities based on the Redlich–Kister formalism; this is an advancement with respect to the existing density estimation capabilities. The graphical user interface was also updated to include a density estimation capability, backed by Saline. Finally, additional preliminary efforts to include surface tension into the database, as well as an investigation on formalisms that would be appropriate for thermal conductivity estimation, are reported herein.

36 MATERIALS SCIENCE↗

Biosphere Futures: a database of social-ecological scenarios

T. Biosphere Futures (https://biospherefutures.net/) is a new online database to collect and discover scenario studies from across the world, with a specific focus on scenarios that explicitly incorporate interdependencies between humans and their supporting ecosystems. It provides access to a globally diverse collection of case studies that includes most ecosystems and regions, enabling exploration of the multifaceted ways in which the future might unfold. Together, the case studies illuminate the diversity and plurality of people’s expectations and aspirations for the future. The objective of Biosphere Futures is to promote the use of scenarios for sustainable development of the biosphere and to foster a community of practice around social-ecological scenarios. We do so by facilitating the assessment, synthesis, and comparative analysis of scenario case studies, pointing to relevant resources, and by helping practitioners and researchers to disseminate and showcase their own work. This article begins by outlining the rationale behind the creation of the database, followed by an introduction to its functionality and the criteria employed for selecting case studies. Subsequently, we present a synthesis of the first 100 case studies included in the scenarios database, highlighting emerging patterns and identifying potential avenues for further research. Finally, given that broader utilization and contributions to the database will enhance the achievement of Biosphere Futures’ objectives, we invite the creators of social-ecological scenarios to contribute additional case studies. By expanding the database’s breadth and depth, we can collectively foster a more nuanced understanding of the possible trajectories of our biosphere and enable better decision making for sustainable development.

54 ENVIRONMENTAL SCIENCES↗

North American Lithium-Ion Battery Supply Chain Database Development - Phase II

Lithium-ion batteries (LIBs) are used in a wide range of applications, including cell phones, laptops, power tools, electric vehicles, and grid storage, and are essential for economic growth and addressing climate change. However, the significant demand for LIBs has led to supply chain issues for the United States, as China dominates the processing of battery materials and battery production. To address this concern, NAATBatt International, a trade association of North American battery companies, supported the National Renewable Energy Laboratory in developing a database of companies that mine, process, manufacture, reuse, and recycle batteries in North America. The purpose of this database was to identify strengths and gaps in the supply chain, so that private-government partnerships could develop strategies to create a competitive LIB supply chain in the US. NREL published the first version of this database in 2021 and the second version in 2022. The database includes companies that have a manufacturing facility in North America and are engaged in materials, cells, packs, end-of-life management, as well as those involved in LIB battery modeling, distribution, service and repair, and R&D. In this presentation, we will discuss our approach to collecting data and categorizing various segments and products. We will also provide a summary of the data and present various maps to illustrate the distribution of companies in the database.

ADVANCED PROPULSION SYSTEMS,ENERGY STORAGE↗

DUNE Database Development

The DUNE experiment will produce vast amounts of metadata, which describe the data coming from the read-out of the primary DUNE detectors. Various databases will make up the overall DB architecture for this metadata. ProtoDUNE at CERN is the largest existing prototype for DUNE and serves as a testing ground for - among other things - possible database solutions for DUNE. The subset of all metadata that is accessed during offline data reconstruction and analysis is referred to as ‘conditions data’ and it is stored in a dedicated database. As offline data reconstruction and analysis will be deployed on HTC and HPC resources, conditions data is expected to be accessed at very high rates. It is therefore crucial to store it in a granularity that matches the expected access patterns allowing for extensive caching. This requires a good understanding of the sources and use cases of conditions data. This contribution will briefly summarize the database architecture deployed at ProtoDUNE and explain the various sources of conditions data. We will present how the conditions data is retrieved and streamed from the databases and how it is handled to match expected access patterns.

43 PARTICLE ACCELERATORS↗

BRE‐X Emissions Database for End‐of‐Life Scenarios of Selective Building Construction Materials to Enable Circular Economy in Construction

In the United States, construction and demolition debris predominately end up in landfills with minimal end‐of‐life Re‐X (recover, recycle, reuse, etc.) scenarios, resulting in large environmental impacts and lost opportunities for material recovery. Except for concrete and metals, which seem to have a few well‐defined end‐of‐life pathways, there seems to be a lack of well‐documented end‐of‐life scenarios for other construction materials, let alone their emissions data. Hence, there is a need for documented end‐of‐life Re‐X scenarios and end‐of‐life data of more building materials to motivate widespread use of Re‐X strategies in building design. This paper outlines the efforts of the National Renewable Energy Laboratory, Carbon Leadership Forum, Building Transparency, and Skidmore, Owings & Merrill to (a) create an open‐access BRE‐X (Building Re‐X) end‐of‐life emissions database consisting of greenhouse gas emissions data associated with various end‐of‐life scenarios for a select list of high‐impact building construction materials, and (b) integrate the BRE‐X end‐of‐life emissions database with CAD/BIM/LCA tools for evaluating various end‐of‐life scenarios. The paper also presents a few existing life cycle inventory databases that contain sparse amounts of end‐of‐life data for a few construction materials and their limitations in terms of scaling and data consolidation. Finally, a sample of how the collected data can be ingested into whole‐building LCA tools using open data formats and a public access link to the BRE‐X end‐of‐life emissions database is also included.

36 MATERIALS SCIENCE↗

Large language model-driven database for thermoelectric materials

Thermoelectric materials have the ability to convert waste heat into electricity, offering a valuable solution for energy harvesting. However, their widespread use is hindered by low conversion efficiency, the reliance on expensive rare earth elements, and the environmental and regulatory concerns associated with lead-based materials. A fast and cost-effective way to identify highly efficient thermoelectric materials is through data-driven methods. These approaches rely on robust and comprehensive datasets to train models. Although there are several databases on thermoelectric materials, there is still a need to collect and integrate experimental data from peer-reviewed research articles to capture diverse compositions and properties of materials. Here, in this work, we developed a comprehensive database of 7,123 thermoelectric compounds, containing key information such as chemical composition, structural detail, seebeck coefficient, electrical and thermal conductivity, power factor, and figure of merit (ZT). We used the GPTArticleExtractor workflow, powered by large language models (LLM), to extract and curate data automatically from the scientific literature published in Elsevier journals. This process enabled the creation of a structured database that addresses the challenges of manual data collection. The open access database could stimulate data-driven research and advance thermoelectric material analysis and discovery.

Database↗

Reflections on one million compounds in the open quantum materials database (OQMD)

Abstract Density functional theory (DFT) has been widely applied in modern materials discovery and many materials databases, including the open quantum materials database (OQMD), contain large collections of calculated DFT properties of experimentally known crystal structures and hypothetical predicted compounds. Since the beginning of the OQMD in late 2010, over one million compounds have now been calculated and stored in the database, which is constantly used by worldwide researchers in advancing materials studies. The growth of the OQMD depends on project-based high-throughput DFT calculations, including structure-based projects, property-based projects, and most recently, machine-learning-based projects. Another major goal of the OQMD is to ensure the openness of its materials data to the public and the OQMD developers are constantly working with other materials databases to reach a universal querying protocol in support of the FAIR data principles.

08 HYDROGEN↗

Data from TropiRoot 1.0 database: tropical root characteristics across environments

TropiRoot 1.0 is a new tropical root database with root characteristics across environment gradients. It has data extracted from 104 new sources, resulting in more than 8000 rows of data (either species or community data). Most of the data in TropiRoot 1.0 includes root characteristics such as root biomass, morphology, root dynamics, mass fraction, architecture, anatomy, physiology and root chemistry. This initiative represents an approximately 30% increase in the currently available data for tropical roots in the Fine Root Ecology Database (FRED). TropiRoot 1.0, contains root characteristics from 25 different countries where seven are located in Asia, six in South America, five in Central America and the Caribbean, four in Africa, two in North America, and 1 in Oceania. Due to the volume of data, when ancillary data was available, including soil data, these data was either extracted and included in the database or their availability was recorded in an additional column. Multiple contributors checked the entries for outliers during the collation process to ensure data quality. For text-based observations, we examined all cells to ensure that their content relates to their specific categories. For numerical observations, we ordered each numerical value from least to greatest and plotted the values, checking apparent outliers against the data in their respective sources and correcting or removing incorrect or impossible values. Some data (soil and aboveground) have different columns for the same variable presented in different units, including originally published units, but root characteristics data had units converted to match the ones reported in FRED. By filling a gap from global databases, TropiRoot 1.0 expands our knowledge of otherwise so far underrepresented regions, and our ability to assess global trends. This advancement can be used to improve tropical forest representation in vegetation models.

54 ENVIRONMENTAL SCIENCES↗

Advances in Metallic Fuel Database Development and Data Qualification

The Fuels Irradiation and Physics Database (FIPD [1]) is a comprehensive repository of data and documents related to Uranium-Zirconium based metallic fuel test pins. This database stores operational conditions of these pins, calculated using a suite of Argonne National Laboratory analysis codes developed during the Integral Fast Reactor (IFR) program. Key calculated data include axial distributions of power, temperature, fluence, burnup, and isotopic densities. Additionally, the FIPD holds post-irradiation examination (PIE) data such as fission gas release, gas chemistry measurements, and axial distributions derived from profilometry, gamma scanning, and neutron radiography. Complementing these data is an extensive archive of documents related to various pins and experiments. These include raw PIE records, design details, safety analyses, and operational reports. More detail about FIPD can be found in ref. [2]. The database development is an ongoing effort covering metallic fuel experiments from the Experimental Breeder Reactor II (EBR-II) and the Fast Flux Test Facility (FFTF). The recent improvements to the database and the data QA status are summarized in this paper.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗