Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

A thermoelectric materials database auto-generated from the scientific literature using ChemDataExtractor

An auto-generated thermoelectric-materials database is presented, containing 22,805 data records, automatically generated from the scientific literature, spanning 10,641 unique extracted chemical names. Each record contains a chemical entity and one of the seminal thermoelectric properties: thermoelectric figure of merit, ZT; thermal conductivity, κ; Seebeck coefficient, S; electrical conductivity, σ; power factor, PF; each linked to their corresponding recorded temperature, T. The database was auto-generated using the automatic sentence-parsing capabilities of the chemistry-aware, natural language processing toolkit, ChemDataExtractor 2.0, adapted for application in the thermoelectric-materials domain, following a rule-based sentence-simplification step. Data were mined from the text of 60,843 scientific papers that were sourced from three scientific publishers: Elsevier, the Royal Society of Chemistry, and Springer. To the best of our knowledge, this is the first automatically-generated database of thermoelectric materials and their properties from existing literature. The database was evaluated to have a precision of 82.25% and has been made publicly available to facilitate the application of data science in the thermoelectric-materials domain, for analysis, design, and prediction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Simulated sulfur K-edge X-ray absorption spectroscopy database of lithium thiophosphate solid electrolytes

X-ray absorption spectroscopy (XAS) is a premier technique for materials characterization, providing key information about the local chemical environment of the absorber atom. In this work, we develop a database of sulfur K-edge XAS spectra of crystalline and amorphous lithium thiophosphate materials based on the atomic structures reported in Chem. Mater., 34, 6702 (2022). The XAS database is based on simulations using the excited electron and core-hole pseudopotential approach implemented in the Vienna Ab initio Simulation Package. Our database contains 2681 S K-edge XAS spectra for 66 crystalline and glassy structure models, making it the largest collection of first-principles computational XAS spectra for glass/ceramic lithium thiophosphates to date. This database can be used to correlate S spectral features with distinct S species based on their local coordination and short-range ordering in sulfide-based solid electrolytes. The data is openly distributed via the Materials Cloud, allowing researchers to access it for free and use it for further analysis, such as spectral fingerprinting, matching with experiments, and developing machine learning models.

36 MATERIALS SCIENCE↗

Outcomes of WPEC SG47 on "Use of Shielding Integral Benchmark Archive and Database for Nuclear Data Validation"

The Working Party on International Nuclear Data Evaluation Co-operation Subgroup 47 (WPECSG47) entitled "Use of Shielding Integral Benchmark Archive and Database for Nuclear Data Validation" was organised from 2019 and 2022 with the objectives to promote more systematic and wider use of shielding benchmark experiments in nuclear data (ND) and transport code validation and development, to provide feedback on the Shielding Integral Benchmark Archive and Database (SINBAD), and to promote its further development in coordination with the Expert Group on Physics of Reactor Systems (EGPRS). Altogether 9 meetings, the large majority (8) held remotely, were organised during the past 3 years to discuss the experience on the use of SINBAD, evaluation of new benchmarks and improvements to be contributed to the database which was severely neglected and lacking maintenance over the past ← 10+ years. Several proposals for new or updated benchmark evaluation were presented and discussed, such as FNG copper, LLNL pulsed spheres, CIAE iron sphere, KFK 1977 gamma measurements, Rez Fe sphere, ASPIS, ORNL Oxygen broomstick, TIARA and others. Complementing the database with new features was also discussed, for example providing the nuclear data sensitivity profiles more systematically would facilitate and better guide the use of data. Information on the geometry, (radiation source) and materials available in CAD format is expected to allow an easier and less error prone reference for computational model preparation and a potential input to CAD based workflows. Inputs for various transport codes and other benchmark data from participants have been shared via the NEA GitLab which could hopefully in the future evolve and form a bases for critically checked and validated benchmark data. Future development of SINBAD will be monitored by EGPRS and the newly created SINBAD Task Force.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

WPEC SG50: Developing an Automatically Readable, Comprehensive and Curated Experimental Nuclear Reaction Database

The Organisation for Economic Co-operation and Development (OECD) Nuclear Energy Agency (NEA) Working Party on International Nuclear Data Evaluation Co-operation (WPEC) subgroup (SG) 50 was formed in 2020 to develop an automatically readable, comprehensive and curated experimental nuclear reaction database. This database is called MEDUSAL (Machine-readable Experimental Data User Application & Library), and will draw from EXFOR. The EXFOR database preserves experimental nuclear reaction data true to its original documentation and information from the authors of the data. MEDUSAL will deviate from EXFOR by storing additional information from users of the data for their fields of work (evaluation, model development, validation, etc.). This includes expert judgment on the data sets, identification of data points as outliers, renormalization of the data to the newest monitor reactions, and estimations of missing uncertainty sources. The format for MEDUSAL is being developed to enable easy automatic parsing of large amounts of data. Here, we will summarize the use cases, high-level requirements and first steps towards developing the database MEDUSAL and the API to access it.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Quantum computing without quantum computers: Database search and data processing using classical wave superposition

Quantum computers are proven to be more efficient at solving a specific class of problems compared to traditional digital computers. Superposition of states and quantum entanglement are the two key ingredients that make quantum computing so powerful. However, not all quantum algorithms require quantum entanglement (e.g., search through an unsorted database). Is it possible to utilize classical wave superposition to speed up database searching as much as by using quantum computers? There were several attempts to mimic quantum computers using classical waves. It was concluded that the use of classical wave superposition comes with the cost of an exponential increase in resources. In this work, we consider the feasibility of building classical wave-based devices able to provide fundamental speedup over digital counterparts without the exponential overhead. We present experimental data on database searching through a magnetic database using spin wave superposition. The results demonstrate the same speedup as expected for quantum computers. Also, we present examples of numerical modeling demonstrating classical wave interference for period finding. This approach may not compete with quantum computers with efficiency but outperform classical digital computers. We argue that classical wave-based devices can perform some of the quantum algorithms with the same efficiency as quantum computers as long as quantum entanglement is not required.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A database and meta-analysis on the performance of exploding pusher implosions conducted at OMEGA

A database of 222 exploding pusher implosions conducted at the OMEGA Laser Facility is presented. The dataset consists of glass-shell capsules filled with varying pressures of D 2 , T 2 , and 3 He, which were imploded using square laser pulses with intensities ranging from 1 to 1 × 10 15 W/cm 2 . The database includes measurements of bang times, ion temperatures, and yields from the DD, D 3 He, and DT fusion reactions. A semi-analytic exploding pusher model is introduced, which effectively captures the observed trends in the data. This model predicts that the measurements scale according to a power-law relation based on the initial capsule and laser conditions. A generalized power-law scaling relation is directly fit to each dataset, providing a useful interpolation of the entire database. Overall, the database provides a valuable resource to estimating bang times, temperatures, and yields for the design of future experiments. Additionally, it provides a diverse set of data for validating more advanced implosion physics models.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Empirical scaling of the L–H threshold power for metal wall tokamaks using a multi-device database

The empirical scaling for the H-mode power threshold in tokamaks has been revisited using a database with threshold data from machines with a metallic first wall as part of International Tokamak Physics Activity (ITPA) task TC-26. The database contains discharges from ASDEX Upgrade (AUG) (W), JET (Be/W) and Alcator C-Mod (Mo). This was motivated by reports that in like-for-like discharges the power threshold was reduced by approximately 30% after the change from carbon based to metallic first wall materials on AUG (Ryter et al 2013 Nucl. Fusion 53 113003) and JET (Maggi et al 2014 Nucl. Fusion 54 023007). The database contains L–H transition data for all hydrogen isotopes and mixtures, including T and DT from the recent JET campaigns. Compared to the ITPA 2008 scaling (Martin et al 2008 J. Phys.: Conf. Ser. 123 012033), the metal wall scaling has a smaller magnetic field exponent but a larger density exponent. We present an additional parameter to capture the strong dependence of the L–H power threshold (approx. factor 2) on the magnetic configuration in the divertor on JET. The scaling recovers the approximate inverse isotope mass scaling of the threshold power. Alternative scalings involving the plasma current and poloidal magnetic field are explored. Despite the reduction in threshold observed earlier, the scalings based on the metal wall database do not necessarily extrapolate to a lower threshold for ITER compared to the ITPA 2008 scaling, especially at high density. The divertor configuration effect induces the largest uncertainty in the extrapolation.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Prevalence of gp160 polymorphisms known to be related to decreased susceptibility to temsavir in different subtypes of HIV-1 in the Los Alamos National Laboratory HIV Sequence Database

Fostemsavir, a prodrug of the gp120-directed attachment inhibitor temsavir, is indicated for use in heavily treatment-experienced individuals with MDR HIV-1. Reduced susceptibility to temsavir in the clinic maps to discrete changes at amino acid positions in gp160: S375, M426, M434 and M475.To query the Los Alamos National Laboratory (LANL) HIV Sequence Database for the prevalence of polymorphisms at gp160 positions of interest. Full-length gp160 sequences (N = 7560) were queried for amino acid polymorphisms relative to the subtype B consensus at positions of interest; frequencies were reported for all sequences and among subtypes/circulating recombinant forms (CRFs) with ≥10 isolates in the database. Among 239 subtypes in the database, the 5 most prevalent were B (n = 2651, 35.1%), C (n = 1626, 21.5%), CRF01_AE (n = 674, 8.9%), A1 (n = 273, 3.6%) and CRF02_AG (n = 199, 2.6%). Among all 7560 sequences, the most prevalent amino acids at positions of interest (S 375 , 73.5%; M 426 , 82.1%; M 434 , 88.2%; M 475 , 89.9%) were the same as the subtype B consensus. Specific polymorphisms with the potential to decrease temsavir susceptibility (S 375 H/I/M/N/T/Y, M 426 L/P, M 434 I/K and M475I) were found in <10% of isolates of subtypes D, G, A6, BC, F1, CRF07_BC, CRF08_BC, 02A, CRF06_cpx, F2, 02G and 02B. S 375 H and M 475 I were predominant among CRF01_AE (S375H, 99.3%; M 475 I, 76.3%; consistent with previously reported low temsavir susceptibility of this CRF) and 01B (S 375 H, 71.7%; M 475 I, 49.5%). Analysis of the LANL HIV Sequence Database found a low prevalence of gp160 amino acid polymorphisms with the potential to reduce temsavir susceptibility overall and among most of the common subtypes.

59 BASIC BIOLOGICAL SCIENCES↗

The ModelSEED Biochemistry Database for the integration of metabolic annotations and the reconstruction, comparison and analysis of metabolic models for plants, fungi and microbes

Abstract For over 10 years, ModelSEED has been a primary resource for the construction of draft genome-scale metabolic models based on annotated microbial or plant genomes. Now being released, the biochemistry database serves as the foundation of biochemical data underlying ModelSEED and KBase. The biochemistry database embodies several properties that, taken together, distinguish it from other published biochemistry resources by: (i) including compartmentalization, transport reactions, charged molecules and proton balancing on reactions; (ii) being extensible by the user community, with all data stored in GitHub; and (iii) design as a biochemical ‘Rosetta Stone’ to facilitate comparison and integration of annotations from many different tools and databases. The database was constructed by combining chemical data from many resources, applying standard transformations, identifying redundancies and computing thermodynamic properties. The ModelSEED biochemistry is continually tested using flux balance analysis to ensure the biochemical network is modeling-ready and capable of simulating diverse phenotypes. Ontologies can be designed to aid in comparing and reconciling metabolic reconstructions that differ in how they represent various metabolic pathways. ModelSEED now includes 33,978 compounds and 36,645 reactions, available as a set of extensible files on GitHub, and available to search at https://modelseed.org and KBase.

59 BASIC BIOLOGICAL SCIENCES↗

eQuilibrator 3.0: a database solution for thermodynamic constant estimation

Abstract eQuilibrator (equilibrator.weizmann.ac.il) is a database of biochemical equilibrium constants and Gibbs free energies, originally designed as a web-based interface. While the website now counts around 1,000 distinct monthly users, its design could not accommodate larger compound databases and it lacked a scalable Application Programming Interface (API) for integration into other tools developed by the systems biology community. Here, we report on the recent updates to the database as well as the addition of a new Python-based interface to eQuilibrator that adds many new features such as a 100-fold larger compound database, the ability to add novel compounds, improvements in speed and memory use, and correction for Mg2+ ion concentrations. Moreover, the new interface can compute the covariance matrix of the uncertainty between estimates, for which we show the advantages and describe the application in metabolic modelling. We foresee that these improvements will make thermodynamic modelling more accessible and facilitate the integration of eQuilibrator into other software platforms.

59 BASIC BIOLOGICAL SCIENCES↗

VISTA Enhancer browser: an updated database of tissue-specific developmental enhancers

Regulatory elements (enhancers) are major drivers of gene expression in mammals and harbor many genetic variants associated with human diseases. Here, we present an updated VISTA Enhancer Browser (https://enhancer.lbl.gov), a database of transgenic enhancer assays conducted in developing mouse embryos in vivo. Since the original publication in 2007, the database grew nearly 20-fold from 250 to over 4500 experiments and currently harbors over 23 500 images. The updated database provides structured information on experiments conducted at different stages of embryonic development, including enhancer activities of human pathogenic and synthetic variants and sequences derived from a variety of species. In addition to manually curated results of thousands of individual experiments, the new database also features hundreds of manually curated comparisons between alleles. The VISTA Enhancer Browser provides a crucial resource for study of human genetic variation, gene regulation and developmental biology.

59 BASIC BIOLOGICAL SCIENCES↗

Enzyme Engineering Database (EnzEngDB): a platform for sharing and interpreting sequence–function relationships across protein engineering campaigns

The discovery and engineering of new enzymes is important across the bioeconomy, with diverse applications from foods to pharmaceuticals, sensors to agriculture. However, enzyme engineering, in particular machine learning-guided engineering, is hampered by a lack of data. Currently there exists no database designed to capture and interpret datasets created in this domain, nor are there easy analysis and visualisation tools. We developed the Enzyme Engineering Database to provide a centralized resource and an online analysis tool to consolidate sequence-function data from enzyme engineering campaigns, thereby making three contributions: (i) a database into which researchers can deposit public data, (ii) visualisation and analysis tools for protein engineers to analyse their own data or compare enzyme variants to other engineering campaigns, and (iii) a gold-standard dataset for benchmarking automated extraction along with the first large language model extraction pipeline specific for enzyme engineering campaigns. The Enzyme Engineering Database is accessible at http://enzengdb.org/.

Long, Yueming [California Institute of Technology ↗

Tallo: A global tree allometry and crown architecture database

Abstract Data capturing multiple axes of tree size and shape, such as a tree's stem diameter, height and crown size, underpin a wide range of ecological research—from developing and testing theory on forest structure and dynamics, to estimating forest carbon stocks and their uncertainties, and integrating remote sensing imagery into forest monitoring programmes. However, these data can be surprisingly hard to come by, particularly for certain regions of the world and for specific taxonomic groups, posing a real barrier to progress in these fields. To overcome this challenge, we developed the Tallo database, a collection of 498,838 georeferenced and taxonomically standardized records of individual trees for which stem diameter, height and/or crown radius have been measured. These data were collected at 61,856 globally distributed sites, spanning all major forested and non‐forested biomes. The majority of trees in the database are identified to species (88%), and collectively Tallo includes data for 5163 species distributed across 1453 genera and 187 plant families. The database is publicly archived under a CC‐BY 4.0 licence and can be access from: https://doi.org/10.5281/zenodo.6637599 . To demonstrate its value, here we present three case studies that highlight how the Tallo database can be used to address a range of theoretical and applied questions in ecology—from testing the predictions of metabolic scaling theory, to exploring the limits of tree allometric plasticity along environmental gradients and modelling global variation in maximum attainable tree height. In doing so, we provide a key resource for field ecologists, remote sensing researchers and the modelling community working together to better understand the role that trees play in regulating the terrestrial carbon cycle.

Jucker, Tommaso↗

SQL and NoSQL Databases for Cyber Physical Production Systems in Internet of Things for Manufacturing (IoTfM)

Abstract In this paper, the design and performance differences between Relational Database Management Systems (RDBMS) and NoSQL Database Systems are examined, with attention to their applicability for real-world Internet of Things for manufacturing (IoTfM) data. While previous work has extensively compared SQL and NoSQL for both generalized and IoT uses, this work specifically examines the tradeoffs and performance differences for manufacturing applications by using a high-fidelity data set collected from a large US manufacturing firm. Growing an IoT system beyond the pilot stage requires scalable data storage; this work seeks to determine the impact of selected database systems on data write performance at scale. Payload size and message frequency were used as the primary characteristics to maintain model fidelity in simulated clients. As the number of simulated asset clients grow, the data write latency was calculated to determine how both database systems’ performance were affected. To isolate the RDBMS and NoSQL differences, a cloud environment was created using Amazon Web Services (AWS) with two identical data ingestion pipelines: writing data to an RDMBS (1) using AWS Aurora MySQL, and (2) using AWS DynamoDB NoSQL. The findings may provide guidance for further experimentation in large-scale manufacturing IoT implementations.

Gamero, David↗

The Pan-Arctic Vegetation Cover (PAVC) database v1.1

The Pan-Arctic Vegetation Cover (PAVC) database contains synthesized field-data observations of vegetation cover from 978 Arctic Alaska plots with observations from 2010 to 2021. The cover datasets contain plot data at both the plant functional type (PFT) and species-level resolution, with standardized PFT definitions and species names. We synthesized publicly available point-intercept and visual estimate plots from the Arctic Vegetation Archive of Alaska, the Alaska Vegetation Plots Database, the North Slope Science Catalog, and the National Ecological Observatory Network; as well as previously unpublished data from the Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic).Users will find four synthesized datasets, 4 associated data descriptor (dd) files, and 1 metadata file in the PAVC database:synthesized_species_fcover.csv contains fractional cover (fcover) for unique accepted species names, where names include vegetation identified at the family, genus, species, subspecies, and variety levels, as well as general functional types across all 5 data sources. The synthesized_species_fcover_dd.csv accompanies this dataset with header information.synthesized_pft_fcover.csv contains fcover for the following PFTs: non-vascular plants with lichen and bryophyte subcategories, trees with deciduous and evergreen subcategories, shrubs with deciduous and evergreen subcategories, graminoids (grasses), and forbs (herbaceous flowering plants) measured as total cover. Litter and “other” cover are also included as total cover. Additional “types” include water and bare ground, which were measured as top cover. The synthesized_pft_fcover_dd.csv accompanies this dataset with header information.species_pft_checklist.csv is a lookup table containing the translation from a dataset species name to an accepted species name and to a PFT. This table can be used to clarify our species to PFT adjudications, and to aid users in assigning their own PFTs. Any issues found in this checklist should be reported in the Issues tab of our github.survey_unit_information.csv contains auxiliary information about the plots synthesized in this database. It contains useful information for filtering plots of interest based on temporal, geospatial, and contextual information about the plot surveys.flmd.csv contains metadata information about each file in the database.This research was performed as a part of the NGEE Arctic project. The NGEE Arctic project was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Energy-Community-Geo-Database

Fundamental to understanding evolving infrastructure and technology is understanding the drivers for energy communities that affect where and how infrastructure is placed and used. An Energy Community Geo Database aligned to carbon storage systems will aid stakeholders in decision making around key future infrastructure (e.g., CCS injection locations) placement while helping contextualize these analyses against past and present-day social and environmental attributes. Much of the data required to support analyses are available from authoritative, largely governmental sources, but at present are disparate and require time-consuming acquisition, integration, and upkeep for currency. The CCS-EC-GB v2.0 integrates datasets from various federal agencies and authoritative sources to aid stakeholders and decision makers of the social and environmental factors that might impact the viability of CCS and energy related project implementation. There are 5 categories in the CCS-EC-GB v2 database. Most of the layers within each category have been updated in this version. As compared to the old database, there is 1 new category in the CCS-EC-GB v2 database: infrastructure. This is an evolving project and application will be updated periodically with new datasets and information. Notes for consideration: This database/web map will be updated with additional as new data and information becomes available and has been processed, reviewed, and approved for release. Additional state and federal entity data are planned to be integrated and included in future revisions. Summary layers provided in this application are derived from proprietary layers and do not always contain key features (age, status, or TVD) and therefore might not be shown when data are queried for those features.

CCS,Carbon capture and storage,Carbon storage,U.S.↗

Prospective Seal Unit Spatial Extent Database for U.S. Sedimentary Basins

The Prospective Seal Unit Spatial Extent Database for U.S. Sedimentary Basins contains a series of spatial datasets representing spatial extents of publicly available data for caprock and seal rock units within the Appalachian Basin, Denver-Julesburg Basin, Great Valley Basin (Sacramento and San Joaquin Basins), Illinois Basin, Michigan Basin, San Juan Basin, U.S. Gulf Coast Basin, and Williston Basin. The database is designed to support carbon storage feasibility and resources assessment for carbon transport and storage (CTS) projects while displaying the spatial extent of prospective seal units and provide a guide to the original data source. This database leverages publicly available data resources from authoritative sources (e.g. U.S. Geological Survey, State Geologic Surveys, and published reports), and aims to help guide users to understand the seal unit's spatial coverage and data gaps from the regional to sub-basin/field scale. The database is organized by seal unit/formation, including the spatial extent for data found to be available for the seal unit. The various datasets represented include spatial extents of the lithologic formation, depth to top structural contour maps, and thickness/isopach maps. Included in this submission are the following resources: 1. Geodatabase/Dataset: “prospective-seal-unit-extents-2025.gdb” 2. ReadMe: “readme-prospective-seal-unit-spatial-extent-dataset-2025.pdf” 3. Data Catalog: “prospective-seal-unit-spatial-extents-data-catalog-2025.xlsx” 4. Data Sources Key: “data-source.csv” Please see NETL disclaimers here: https://netl.doe.gov/home/disclaimer

Basin↗

M4SF-20LL010301042: Updated Thermodynamic Database for Use in Generic Disposal System Assessment

This progress report (Level 4 Milestone Number M4SF-20LL010301042) summarizes research conducted at Lawrence Livermore National Laboratory (LLNL) within the Argillite Activity Number SF-20LL010301041. LLNL is leading efforts in the development of thermodynamic databases in support of the Spent Fuel and Waste Science Technology (SFWST) program. Thermodynamic models provide the basis for understanding the stability of solid phases and speciation of aqueous species and modeling the evolution of repository conditions. The LLNL effort is being performed in coordination with other US database development efforts. The effort includes a review of available thermochemical databases and a path forward for database integration. International coordination with the NEA-TDB is supported through crystalline international work package.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗