Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Large language model-driven database for thermoelectric materials

Thermoelectric materials have the ability to convert waste heat into electricity, offering a valuable solution for energy harvesting. However, their widespread use is hindered by low conversion efficiency, the reliance on expensive rare earth elements, and the environmental and regulatory concerns associated with lead-based materials. A fast and cost-effective way to identify highly efficient thermoelectric materials is through data-driven methods. These approaches rely on robust and comprehensive datasets to train models. Although there are several databases on thermoelectric materials, there is still a need to collect and integrate experimental data from peer-reviewed research articles to capture diverse compositions and properties of materials. Here, in this work, we developed a comprehensive database of 7,123 thermoelectric compounds, containing key information such as chemical composition, structural detail, seebeck coefficient, electrical and thermal conductivity, power factor, and figure of merit (ZT). We used the GPTArticleExtractor workflow, powered by large language models (LLM), to extract and curate data automatically from the scientific literature published in Elsevier journals. This process enabled the creation of a structured database that addresses the challenges of manual data collection. The open access database could stimulate data-driven research and advance thermoelectric material analysis and discovery.

Database↗

An uncertainty-focused database approach to extract spatiotemporal trends from qualitative and discontinuous lake-status histories

Changes in lake status are often interpreted as palaeoclimate indicators due to their dependence on precipitation and evaporation. The Global Lake Status Database (GLSDB) has since long provided a standardised synopsis of qualitative lake status over the last 30,000 14C years. Potential sources of uncertainty however are not recorded in the GLSDB. Here we present an updated and improved relational-database framework that incorporates uncertainty in both chronology and the interpretation of palaeoenvironmental data. The database uses peer-reviewed palaeolimnological studies to produce a consensus on qualitative lake-status histories, whose chronologies are revised and standardized through the recalibration of radiocarbon dates and the application of Bayesian age-depth modelling for stratigraphic archives. Quantitative information on absolute water-level elevation is preserved if available from geomorphological sources. We also propose a new probabilistic analytical framework that accounts for these uncertainties to reconstruct synoptic, integrated environmental signals. The process is based on a Monte Carlo algorithm that iteratively samples individual lake-status histories within the limits of their uncertainties to produce many possible scenarios. We then use Recursively-Subtracted Empirical Orthogonal Function analysis to extract dominant patterns of lake-status variability from these scenarios. As a proof of concept, we apply this framework to 67 sites in eastern and southern Africa whose lake-status histories cover part of the late Pleistocene and/or Holocene. We show that, despite the sometimes large temporal and interpretation uncertainties, and the inclusion of highly discontinuous lake-status time series, identifying the major known millennial-scale climatic phases during the last 20,000 years is possible. Our framework was also able to identify an antiphased response between the lake basins in eastern and interior southern Africa to these changes. Here, we propose that our new database and methodology framework serves as a template for efficient lake-status data synthesis, encourages the incorporation of lake-status data in palaeoclimate syntheses, and expands the possibilities for the use of such data in the evaluation of climate models.

58 GEOSCIENCES↗

Identifying Abandoned Well Sites Using Database Records and Aeromagnetic Surveys

Oil and natural gas are primary sources of energy in the United States. Improved drilling and extracting techniques have led to a renewed interest in historic oil and gas fields, but limited records of legacy wells make new drilling efforts more difficult, as abandoned wells may provide conduits for liquids and gases to migrate into groundwater reservoirs or the atmosphere. Well finding using aeromagnetic surveys pinpoints the location of steel-cased wells, detecting both active and abandoned wells, including buried casings lacking aboveground markers. Here, we present six aeromagnetic surveys conducted in Pennsylvania and Wyoming as case studies, comparing the magnetic points to locations known in databases. In all study sites, more magnetic points were detected than recorded in databases. Based on differences between theoretical database well counts and the actual number of wells detected in surveys, we estimated the total number of wells in Pennsylvania to be 395 000–466 000 and 181 000–182 000 in Wyoming. Extrapolating to the national level, in this work, we estimate the average number of wells in the continental United States is 6.04 ± 19.97 million wells with 1.16 ± 3.84 million of those designated as abandoned wells, within the range of previous abandoned well count estimations. Although aeromagnetic surveys are limited to detecting steel-cased wells and do not differentiate sites based on well status, this study nevertheless demonstrates the utility of aeromagnetic surveys in well finding efforts across the country and shows limitations in database records of oil and natural gas wells.

54 ENVIRONMENTAL SCIENCES↗

Reflections on one million compounds in the open quantum materials database (OQMD)

Abstract Density functional theory (DFT) has been widely applied in modern materials discovery and many materials databases, including the open quantum materials database (OQMD), contain large collections of calculated DFT properties of experimentally known crystal structures and hypothetical predicted compounds. Since the beginning of the OQMD in late 2010, over one million compounds have now been calculated and stored in the database, which is constantly used by worldwide researchers in advancing materials studies. The growth of the OQMD depends on project-based high-throughput DFT calculations, including structure-based projects, property-based projects, and most recently, machine-learning-based projects. Another major goal of the OQMD is to ensure the openness of its materials data to the public and the OQMD developers are constantly working with other materials databases to reach a universal querying protocol in support of the FAIR data principles.

08 HYDROGEN↗

The updated ITPA global H-mode confinement database: description and analysis

The multi-machine ITPA Global H-mode Confinement Database has been upgraded with new data from JET with the ITER-like wall and ASDEX Upgrade with the full tungsten wall. This paper describes the new database and presents results of regression analysis to estimate the global energy confinement scaling in H-mode plasmas using a standard power law. Various subsets of the database are considered, focusing on type of wall and divertor materials, confinement regime (all H-modes, ELMy H or ELM-free) and ITER-like constraints. Apart from ordinary least squares, two other, robust regression techniques are applied, which take into account uncertainty on all variables. Regression on data from individual devices shows that, generally, the confinement dependence on density and the power degradation are weakest in the fully metallic devices. Using the multi-machine scalings, predictions are made of the confinement time in a standard ELMy H-mode scenario in ITER. The uncertainty on the scaling parameters is discussed with a view to practically useful error bars on the parameters and predictions. One of the derived scalings for ELMy H-modes on an ITER-like subset is studied in particular and compared to the IPB98(y,2) confinement scaling in engineering and dimensionless form. Transformation of this new scaling from engineering variables to dimensionless quantities is shown to result in large error bars on the dimensionless scaling. Regression analysis in the space of dimensionless variables is therefore proposed as an alternative, yielding acceptable estimates for the dimensionless scaling. The new scaling, which is dimensionally correct within the uncertainties, suggests that some dependencies of confinement in the multi- machine database can be reconciled with parameter scans in individual devices. This includes vanishingly small dependence of confinement on line-averaged density and normalized plasma pressure (β), as well as a noticeable, positive dependence on effective atomic mass and plasma triangularity. Extrapolation of this scaling to ITER yields a somewhat lower confinement time compared to the IPB98(y, 2) prediction, possibly related to the considerably weaker dependence on major radius in the new scaling (slightly above linear). Further studies are needed to compare more flexible regression models with the power law used here. In addition, data from more devices concerning possible ‘hidden variables’ could help to determine their influence on confinement, while adding data in sparsely populated areas of the parameter space may contribute to further disentangling some of the global confinement dependencies in tokamak plasmas.

Database↗

Data from TropiRoot 1.0 database: tropical root characteristics across environments

TropiRoot 1.0 is a new tropical root database with root characteristics across environment gradients. It has data extracted from 104 new sources, resulting in more than 8000 rows of data (either species or community data). Most of the data in TropiRoot 1.0 includes root characteristics such as root biomass, morphology, root dynamics, mass fraction, architecture, anatomy, physiology and root chemistry. This initiative represents an approximately 30% increase in the currently available data for tropical roots in the Fine Root Ecology Database (FRED). TropiRoot 1.0, contains root characteristics from 25 different countries where seven are located in Asia, six in South America, five in Central America and the Caribbean, four in Africa, two in North America, and 1 in Oceania. Due to the volume of data, when ancillary data was available, including soil data, these data was either extracted and included in the database or their availability was recorded in an additional column. Multiple contributors checked the entries for outliers during the collation process to ensure data quality. For text-based observations, we examined all cells to ensure that their content relates to their specific categories. For numerical observations, we ordered each numerical value from least to greatest and plotted the values, checking apparent outliers against the data in their respective sources and correcting or removing incorrect or impossible values. Some data (soil and aboveground) have different columns for the same variable presented in different units, including originally published units, but root characteristics data had units converted to match the ones reported in FRED. By filling a gap from global databases, TropiRoot 1.0 expands our knowledge of otherwise so far underrepresented regions, and our ability to assess global trends. This advancement can be used to improve tropical forest representation in vegetation models.

54 ENVIRONMENTAL SCIENCES↗

MIMIC II: a massive temporal ICU patient database to support research in intelligent patient monitoring

Development and evaluation of Intensive Care Unit (ICU) decision-support systems would be greatly facilitated by the availability of a large-scale ICU patient database. Following our previous efforts with the MIMIC (Multi-parameter Intelligent Monitoring for Intensive Care) Database, we have leveraged advances in networking and storage technologies to develop a far more massive temporal database, MIMIC II. MIMIC II is an ongoing effort: data is continuously and prospectively archived from all ICU patients in our hospital. MIMIC II now consists of over 800 ICU patient records including over 120 gigabytes of data and is growing. A customized archiving system was used to store continuously up to four waveforms and 30 different parameters from ICU patient monitors. An integrated user-friendly relational database was developed for browsing of patients' clinical information (lab results, fluid balance, medications, nurses' progress notes). Based upon its unprecedented size and scope, MIMIC II will prove to be an important resource for intelligent patient monitoring research, and will support efforts in medical data mining and knowledge-discovery.

NASA Discipline Cardiopulmonary↗

Experimental Database with Baseline CFD Solutions: 2-D and Axisymmetric Hypersonic Shock-Wave/Turbulent-Boundary-Layer Interactions

A database compilation of hypersonic shock-wave/turbulent boundary layer experiments is provided. The experiments selected for the database are either 2D or axisymmetric, and include both compression corner and impinging type SWTBL interactions. The strength of the interactions range from attached to incipient separation to fully separated flows. The experiments were chosen based on criterion to ensure quality of the datasets, to be relevant to NASA's missions and to be useful for validation and uncertainty assessment of CFD Navier-Stokes predictive methods, both now and in the future. An emphasis on datasets selected was on surface pressures and surface heating throughout the interaction, but include some wall shear stress distributions and flowfield profiles. Included, for selected cases, are example CFD grids and setup information, along with surface pressure and wall heating results from simulations using current NASA real-gas Navier-Stokes codes by which future CFD investigators can compare and evaluate physics modeling improvements and validation and uncertainty assessments of future CFD code developments. The experimental database is presented tabulated in the Appendices describing each experiment. The database is also provided in computer-readable ASCII files located on a companion DVD.

Database↗

ThermoBase: A Database of the Phylogeny and Physiology of Thermophilic and Hyperthermophilic Organisms

Thermophiles and hyperthermophiles are those organisms which grow at high temperature (> 40°C). The unusual properties of these organisms have received interest in multiple fields of biological research, and have found applications in biotechnology, especially in industrial processes. However, there are few listings of thermophilic and hyperthermophilic organisms and their relevant environmental and physiological data. Such repositories can be used to standardize definitions of thermophile and hyperthermophile limits and tolerances and would mitigate the need for extracting organism data from diverse literature sources across multiple, sometimes loosely related, research fields. Therefore, we have developed ThermoBase, a web-based and freely available database which currently houses comprehensive descriptions for 1238 thermophilic or hyperthermophilic organisms. ThermoBase reports taxonomic, metabolic, environmental, experimental, and physiological information in addition to literature resources. This includes parameters such as coupling ions for chemiosmosis, optimal pH and range, optimal temperature and range, optimal pressure, and optimal salinity. The database interface allows for search features and sorting of parameters. As such, it is the goal of ThermoBase to facilitate and expedite hypothesis generation, literature research, and understanding relating to thermophiles and hyperthermophiles within the scientific community in an accessible and centralized repository. ThermoBase is freely available online at the Astrobiology Habitable Environments Database (AHED; https://ahed.nasa.gov), at the Database Center for Life Science (TogoDB; http://togodb.org/db/thermobase), and in the S1 File.

Thermophiles↗

Comparison of Entry Descent and Landing Aerodynamic Databases with Uncertainty Quantification Developed Using Machine Learning Techniques

When developing the aerodynamic databases for use in trajectory simulations, it is important to develop a system of metrics to qualify which aerodynamic models are best to use. Since aerodynamics are just one input into trajectory simulations, the results of these simulations do not reflect on the quality of the aerodynamic database used. This means that aerodynamic database comparisons must be done offline. While traditional metrics that focus on mean/nominal predictions are a good first step, more robust estimates of the prediction interval become important as more focused uncertainty models are developed. We explore the limitations of evaluating aerodynamic models based purely on nominal-centered response surfaces. Before elaborating and evaluating metrics based on distributed models, the value of evaluating prediction interval and confidence interval are discussed to conclude that prediction intervals are more relevant to the use of trajectory analysis. Several metrics to evaluate the prediction interval are introduced with a focus on the standard calibration metric. Finally, we compare candidate models using both mean and distributed metrics. A finalized candidate model developed using state of the art machine learning methods is compared to a baseline model developed using traditional aerodynamic database modeling techniques.

Aerodynamic Database↗

Improved Aerothermal Reliability Analysis Enabled by the Mars Sample Return Earth Entry System Aerothermal Database

The Mars Sample Return (MSR) campaign is a series of missions designed to retrieve Martian rock and soil samples for detailed study on Earth. The campaign is split into three primary phases: sample collection with the Mars2020 rover, retrieval with the Sample Return Lander (SRL) and Mars Ascent Vehicle (MAV), and then return to Earth with the Earth Return Orbiter (ERO) and Capture, Containment, and Return System (CCRS) [1]. The final sequence in the Earth return phase is the delivery and entry of the Earth Entry System (EES) sample return capsule. Due to unprecedented planetary protection concerns, the sample return capsule is subject to strict reliability requirements. To this end, the MSR-EES aerothermal team has implemented a flexible aerothermal database architecture capable of integration with state-of-the-art trajectory codes to provide a more rigorous aerothermal reliability analysis. The EES database enables the generation of environments at any location on the heatshield and can incorporate trajectory uncertainties to both statistically quantify aerothermal environments for arcjet testing and produce material response boundary conditions to rigorously select thermal protection system (TPS) sizing environments. This poster will not discuss the fundamental modeling assumptions included in the database, and will instead focus on the downstream reliability analyses that can be performed with a database of this architecture.

MSR-EES↗

Advances in Metallic Fuel Database Development and Data Qualification

The Fuels Irradiation and Physics Database (FIPD [1]) is a comprehensive repository of data and documents related to Uranium-Zirconium based metallic fuel test pins. This database stores operational conditions of these pins, calculated using a suite of Argonne National Laboratory analysis codes developed during the Integral Fast Reactor (IFR) program. Key calculated data include axial distributions of power, temperature, fluence, burnup, and isotopic densities. Additionally, the FIPD holds post-irradiation examination (PIE) data such as fission gas release, gas chemistry measurements, and axial distributions derived from profilometry, gamma scanning, and neutron radiography. Complementing these data is an extensive archive of documents related to various pins and experiments. These include raw PIE records, design details, safety analyses, and operational reports. More detail about FIPD can be found in ref. [2]. The database development is an ongoing effort covering metallic fuel experiments from the Experimental Breeder Reactor II (EBR-II) and the Fast Flux Test Facility (FFTF). The recent improvements to the database and the data QA status are summarized in this paper.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A database of ultrastable MOFs reassembled from stable fragments with machine learning models

High-throughput screening of hypothetical metal-organic framework (MOF) databases can uncover new materials, but their stability in real-world applications is often unknown. We leverage community knowledge and machine learning (ML) models to identify MOFs that are thermally stable and stable upon activation. We separate these MOFs into their building blocks and recombine them to make a new hypothetical MOF database of over 50,000 structures with orders of magnitude more (1) connectivity nets and (2) inorganic building blocks than were present in prior databases. Further, this database shows a 10-fold enrichment of ultrastable MOF structures that are stable upon activation and more than 1 standard deviation more thermally stable than the average experimentally characterized MOF. For nearly 10,000 ultrastable MOFs, we compute elastic moduli to confirm that these materials have good mechanical stability, and we report methane deliverable capacities. We identify privileged metal nodes in ultrastable MOFs that optimize gas storage and mechanical stability simultaneously.

36 MATERIALS SCIENCE↗

FIPD: The SFR metallic fuels irradiation & physics database

The DOE Advanced Reactor Technology (ART) program has supported efforts to recover and preserve metallic fuel data generated throughout the US sodium-cooled fast reactor (SFR) program. Those efforts have been focused on establishing databases of the experimental data that were mainly generated during the Integral Fast Reactor (IFR) program including data generated at Experimental Breeder Reactor-II (EBR-II), Fast Flux Test Facility (FFTF), and Transient Reactor Test Facility (TREAT) reactors, as well as out of pile data. The data is essential for future licensing activities of metallic fuel based advanced fast reactors. This paper describes the development of the SFR Metallic Fuels Irradiation & Physics Database (FIPD) and covers the scientific knowledge available in the database. Furthermore, the architecture of the FIPD is described by showing the available reactor operation data, fabrication, and post-irradiation examination (PIE) data, and other documents. The applications of the fuel database are also discussed.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Initial demonstration of automated fuel performance modeling with 1977 EBR-II metallic fuel pins using BISON code with FIPD and IMIS databases

Using the BISON fuel performance code, simulations were conducted using an automated process to read initial and operating conditions from the Fuels Irradiation and Physics Database (FIPD) and Integral Fast Reactor materials information system (IMIS) database, which contains metallic fuel data from the Experimental Breeder Reactor-II (EBR-II). This work demonstrates use of an integrated framework to access the vast majority of EBR-II experimental fuel pin data to support rapid development of fuel performance models for next-generation metallic fuel systems. With this capability, validation for fuel qualification can be performed rapidly. Between IMIS and FIPD, there is enough information to conduct 1977 unique EBR-II metallic fuel pin histories from 24 different experiments, at varying levels of detail between the two databases. Each of these histories includes a high-resolution power history, flux history, coolant channel flow rates, and coolant channel temperatures. Fission gas release (FGR), cumulative damage fraction (CDF), fuel axial swelling, cladding profilometry, and burnup were all simulated in BISON. The results were compared to post-irradiation examination (PIE) results for the initial demonstration of automated BISON modeling. BISON simulations conducted with IMIS and FIPD were in rough agreement with PIE measurements and calculations. Cladding profilometry, FGR, and fuel axial swelling were found to be in rough agreement with PIE measurements, depending on the physics used within the BISON input files. Here, the mechanical contact solver chosen was found to significantly impact axial fuel swelling and cladding strain predictions. CDF values were assessed to see whether pin failure may have been predicted (CDF ≥ 1). This work suggests that continued development of an automated tool for BISON should focus on inclusion of the Fast Flux Test Facility (FFTF) experimental data for a larger database for metallic fuel, improved physical models to better capture fuel performance, such as fuel-cladding interactions, and a more detailed comparison with available PIE data to further the BISON model development.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

The northeast materials database for magnetic materials

The discovery of magnetic materials with high operating temperature ranges and optimized performance is essential for advanced applications. Current data-driven approaches are limited by the lack of accurate, comprehensive, and feature-rich databases. This study aims to address this challenge by using Large Language Models (LLMs) to create a comprehensive, experiment-based, magnetic materials database named the Northeast Materials Database (NEMAD), which consists of 67,573 magnetic materials entries (www.nemad.org). The database incorporates chemical composition, magnetic phase transition temperatures, structural details, and magnetic properties. Enabled by NEMAD, we trained machine learning models to classify materials and predict transition temperatures. Our classification model achieved an accuracy of 90% in categorizing materials as ferromagnetic (FM), antiferromagnetic (AFM), and non-magnetic (NM). The regression models predict Curie (Néel) temperature with a coefficient of determination (R 2 ) of 0.87 (0.83) and a mean absolute error (MAE) of 56K (38K). These models identified 25 (13) FM (AFM) candidates with a predicted Curie (Néel) temperature above 500K (100K) from the Materials Project. This work shows the feasibility of combining LLMs for automated data extraction and machine learning models to accelerate the discovery of magnetic materials.

Ferromagnetism↗

A database of battery materials auto-generated using ChemDataExtractor

A database of battery materials is presented which comprises a total of 292,313 data records, with 214,617 unique chemical-property data relations between 17,354 unique chemicals and up to five material properties: capacity, voltage, conductivity, Coulombic efficiency and energy. 117,403 data are multivariate on a property where it is the dependent variable in part of a data series. The database was auto-generated by mining text from 229,061 academic papers using the chemistry-aware natural language processing toolkit, ChemDataExtractor version 1.5, which was modified for the specific domain of batteries. The collected data can be used as a representative overview of battery material information that is contained within text of scientific papers. Public availability of these data will also enable battery materials design and prediction via data-science methods. To the best of our knowledge, this is the first auto-generated database of battery materials extracted from a relatively large number of scientific papers. We also provide a Graphical User Interface (GUI) to aid the use of this database.

25 ENERGY STORAGE↗

A thermoelectric materials database auto-generated from the scientific literature using ChemDataExtractor

An auto-generated thermoelectric-materials database is presented, containing 22,805 data records, automatically generated from the scientific literature, spanning 10,641 unique extracted chemical names. Each record contains a chemical entity and one of the seminal thermoelectric properties: thermoelectric figure of merit, ZT; thermal conductivity, κ; Seebeck coefficient, S; electrical conductivity, σ; power factor, PF; each linked to their corresponding recorded temperature, T. The database was auto-generated using the automatic sentence-parsing capabilities of the chemistry-aware, natural language processing toolkit, ChemDataExtractor 2.0, adapted for application in the thermoelectric-materials domain, following a rule-based sentence-simplification step. Data were mined from the text of 60,843 scientific papers that were sourced from three scientific publishers: Elsevier, the Royal Society of Chemistry, and Springer. To the best of our knowledge, this is the first automatically-generated database of thermoelectric materials and their properties from existing literature. The database was evaluated to have a precision of 82.25% and has been made publicly available to facilitate the application of data science in the thermoelectric-materials domain, for analysis, design, and prediction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗