Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A thermochemical database from high-throughput first-principles calculations and its application to analyzing phase evolution in AM-fabricated IN718

A comprehensive thermochemical database is constructed based on high–throughput first-principles phonon calculations of over 3000 atomic structures in limited concentrations in Ni, Fe, and Co alloys involving a total of 26 elements including Al, B, C, Cr, Cu, Hf, La, Mn, Mo, N, Nb, O, P, Re, Ru, S, Si, Ta, Ti, V, W, Y, and Zr, providing thermochemical data largely unavailable from existing experiments. Here, the database can be employed to predict the equilibrium phase compositions and fractions directly from first-principles by minimizing the chemical potential of a multicomponent system with a fixed overall chemical composition and a fixed temperature. It is applied to the additively manufactured nickel-based IN718 superalloy to analyze the phase evolution with temperature. IN718 is known for its great performance in tensile, fatigue, creep, and rupture strength, combined with easy fabrication and corrosion resistance. In particular, we successfully predicted the formation of L1 0 -FeNi, γ’-Ni 3 (Fe,Al), α-Cr, δ-Ni 3 (Nb,Mo), γ”-Ni 3 Nb, and η-Ni 3 Ti at low temperatures (below 680 K), γ’-Ni 3 Al, δ-Ni 3 Nb, γ”-Ni 3 Nb, α-Cr, and γ-Ni(Fe,Cr,Mo) at intermediate temperatures (between 680 and 1140 K), and δ-Ni 3 Nb and γ-Ni(Fe,Cr,Mo) at high temperatures (above 1140 K) in IN718. These predictions are validated by EDS mapping of compositional distributions and corresponding identifications of phase distributions. The database is expected to be a valuable source for future thermodynamic analysis and microstructure prediction of alloys involving the 26 elements.

36 MATERIALS SCIENCE↗

SPEARS: A Database-Invariant Spectral modeling API

The Spectral Physics Environment for Advanced Remote Sensing (SPEARS) application programming interface (API) is a Python-based, line-by-line, local thermal equilibrium (LTE) spectral modeling code which is optimized for simultaneously synthesizing optical spectra from any combination of fundamental spectroscopic databases. In this article, we contribute two novel spectral modeling techniques to the scientific literature. First we describe how SPEARS integrates a physics-based collisional model for calculating pressure broadening in the absence of available broadening coefficients. With this collisional model implementation, a generalized approach to fundamental spectroscopic databases can be achieved across multiple databases. We also detail our adaptive grid mesh algorithm developed to make the code scalable for simulating large spectral bandwidths at high spectral fidelity using intuitive grid parameters. Here, we present comparisons to other modeling tools, experiments, and provide a discussion on the SPEARS user interface.

47 OTHER INSTRUMENTATION↗

MOFX-DB: An Online Database of Computational Adsorption Data for Nanoporous Materials

Machine learning and data mining coupled with molecular modeling have become powerful tools for materials discovery. Metal-organic frameworks (MOFs) are a rich area for this due to their modular construction and numerous applications. Here, we make data from several previous large-scale studies in MOFs and zeolites from our groups (and new data for N 2 and Ar adsorption in MOFs) easily accessible in one place. The database includes over 3 million simulated adsorption data points for H 2 , CH 4 , CO 2 , Xe, Kr, Ar, and N 2 in over 160 000 MOFs and zeolites, textural properties like pore sizes and surface areas, and the structure file for each material. We include metadata about the Monte Carlo simulations to enable reproducibility. The database is searchable by MOF properties, and the data are stored in a standardized JSON format that that is interoperable with the NIST adsorption database. We also identify several MOFs that meet high performance targets for multiple applications, such as high storage capacity for both hydrogen and methane or high CO 2 capacity plus good Xe/Kr selectivity. Here, by providing this data publicly, we hope to facilitate machine learning studies on these materials, leading to new insights on adsorption in MOFs and zeolites.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The ab initio non-crystalline structure database: empowering machine learning to decode diffusivity

Non-crystalline materials exhibit unique properties that make them suitable for various applications in science and technology, ranging from optical and electronic devices and solid-state batteries to protective coatings. However, data-driven exploration and design of non-crystalline materials is hampered by the absence of a comprehensive database covering a broad chemical space. In this work, we present the largest computed non-crystalline structure database to date, generated from systematic and accurate ab initio molecular dynamics (AIMD) calculations. We also show how the database can be used in simple machine-learning models to connect properties to composition and structure, here specifically targeting ionic conductivity. These models predict the Li-ion diffusivity with speed and accuracy, offering a cost-effective alternative to expensive density functional theory (DFT) calculations. Furthermore, the process of computational quenching non-crystalline structures provides a unique sampling of out-of-equilibrium structures, energies, and force landscape, and we anticipate that the corresponding trajectories will inform future work in universal machine learning potentials, impacting design beyond that of non-crystalline materials. In addition, combining diffusion trajectories from our dataset with models that predict liquidus viscosity and melting temperature could be utilized to develop models for predicting glass-forming ability.

36 MATERIALS SCIENCE↗

A database of synthetic inelastic neutron scattering spectra from molecules and crystals

Abstract Inelastic neutron scattering (INS) is a powerful tool to study the vibrational dynamics in a material. The analysis and interpretation of the INS spectra, however, are often nontrivial. Unlike diffraction, for which one can quickly calculate the scattering pattern from the structure, the calculation of INS spectra from the structure involves multiple steps requiring significant experience and computational resources. To overcome this barrier, a database of INS spectra consisting of commonly seen materials will be a valuable reference, and it will also lay the foundation of advanced data-driven analysis and interpretation of INS spectra. Here we report such a database compiled for over 20,000 organic molecules and over 10,000 inorganic crystals. The INS spectra are obtained from a streamlined workflow, and the synthetic INS spectra are also verified by available experimental data. The database is expected to greatly facilitate INS data analysis, and it can also enable the utilization of advanced analytics such as data mining and machine learning. Notice: This manuscript has been authored by UT-Battelle, LLC under Contract No. DE-AC05-00OR22725 with the U.S. Department of Energy. The United States Government retains and the publisher, by accepting the article for publication, acknowledges that the United States Government retains a non-exclusive, paid-up, irrevocable, world-wide license to publish or reproduce the published form of this manuscript, or allow others to do so, for United States Government purposes. The Department of Energy will provide public access to these results of federally sponsored research in accordance with the DOE Public Access Plan ( http://energy.gov/downloads/doe-public-access-plan ).

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A global database of soil microbial phospholipid fatty acids and enzyme activities

Abstract Soil microbes drive ecosystem function and play a critical role in how ecosystems respond to global change. Research surrounding soil microbial communities has rapidly increased in recent decades, and substantial data relating to phospholipid fatty acids (PLFAs) and potential enzyme activity have been collected and analysed. However, studies have mostly been restricted to local and regional scales, and their accuracy and usefulness are limited by the extent of accessible data. Here we aim to improve data availability by collating a global database of soil PLFA and potential enzyme activity measurements from 12,258 georeferenced samples located across all continents, 5.1% of which have not previously been published. The database contains data relating to 113 PLFAs and 26 enzyme activities, and includes metadata such as sampling date, sample depth, and soil pH, total carbon, and total nitrogen. This database will help researchers in conducting both global- and local-scale studies to better understand soil microbial biomass and function.

Science & Technology - Other Topics↗

Database work for the new cross section standards evaluation

An effort is now underway to produce a new evaluation of the neutron standards. It is important to maintain experimental programs to increase the quality and extend the database for the neutron cross section standards in order to improve evaluations of them that will be used to convert cross section measurements made relative to those standards. Measurements have been made for most of the standard cross sections since the last evaluation of the standards. The improved database includes the cross sections for the H(n,n), 6Li(n,t), 10B(n,αγ), 10B(n,α), C(n,n), Au(n,γ), 235U(n,f) and 238U(n,f) standard reactions and ratios among them. The database also includes the 238U(n,γ) and 239Pu(n,f) cross sections in addition to the standard cross sections. Those data were included since there are many ratio measurements of those cross sections with the standards and absolute data are available for them.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Geant4 Monte-Carlo (GEMC) A database-driven simulation program

GEMC[1] is an application that harnesses the power of databases to execute Geant4 Monte-Carlo simulations. The databases (MYSQL, CSQL, TEXT) define the geometry, materials, digitization algorithms, readout electronics and output formats. Implemented in C++, GEMC also boasts a user-friendly Python API that facilitates detector construction and database population. GEMC can handle real-life scenarios such as geometry variations and the run number-dependent calibration constants and digitization parameters. This abstract provides an overview of GEMC, accompanied by examples that showcase its versatility. We delve into the practical application of GEMC within the the CLAS12 experimental program at Jefferson Lab.

Ungaro, Maurizio↗

NENCI-2021. I. A large benchmark database of non-equilibrium non-covalent interactions emphasizing close intermolecular contacts

In this work, we present NENCI-2021, a benchmark database of ~8000 Non-Equilibirum Non-Covalent Interaction energies for a large and diverse selection of intermolecular complexes of biological and chemical relevance. To meet the growing demand for large and high-quality quantum mechanical data in the chemical sciences, NENCI-2021 starts with the 101 molecular dimers in the widely used S66 and S101 databases and extends the scope of these works by (i) including 40 cation–π and anion–π complexes, a fundamentally important class of non-covalent interactions that are found throughout nature and pose a substantial challenge to theory, and (ii) systematically sampling all 141 intermolecular potential energy surfaces (PESs) by simultaneously varying the intermolecular distance and intermolecular angle in each dimer. Designed with an emphasis on close contacts, the complexes in NENCI-2021 were generated by sampling seven intermolecular distances along each PES (ranging from 0.7× to 1.1× the equilibrium separation) and nine intermolecular angles per distance (five for each ion–π complex), yielding an extensive database of 7763 benchmark intermolecular interaction energies (E int ) obtained at the coupled-cluster with singles, doubles, and perturbative triples/complete basis set [CCSD(T)/CBS] level of theory. The E int values in NENCI-2021 span a total of 225.3 kcal/mol, ranging from -38.5 to +186.8 kcal/mol, with a mean (median) E int value of -1.06 kcal/mol (-2.39 kcal/mol). In addition, a wide range of intermolecular atom-pair distances are also present in NENCI-2021, where close intermolecular contacts involving atoms that are located within the so-called van der Waals envelope are prevalent—these interactions, in particular, pose an enormous challenge for molecular modeling and are observed in many important chemical and biological systems. A detailed symmetry-adapted perturbation theory (SAPT)- based energy decomposition analysis also confirms the diverse and comprehensive nature of the intermolecular binding motifs present in NENCI-2021, which now includes a significant number of primarily induction-bound dimers (e.g., cation–π complexes). NENCI-2021 thus spans all regions of the SAPT ternary diagram, thereby warranting a new four-category classification scheme that includes complexes primarily bound by electrostatics (3499), induction (700), dispersion (1372), or mixtures thereof (2192). A critical error analysis performed on a representative set of intermolecular complexes in NENCI-2021 demonstrates that the E int values provided herein have an average error of ±0.1 kcal/mol, even for complexes with strongly repulsive E int values, and maximum errors of ±0.2–0.3 kcal/mol (i.e., ~±1.0 kJ/mol) for the most challenging cases. For these reasons, we expect that NENCI-2021 will play an important role in the testing, training, and development of next-generation classical and polarizable force fields, density functional theory approximations, wavefunction theory methods, and machine learning based intra- and inter-molecular potentials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

High gas throughput SOLPS-ITER simulations extending the ITER database to strong detachment

Abstract SOLPS-ITER simulations performed for Q DT = 10, P SOL = 100 MW burning plasmas on ITER extend the existing database to high values of separatrix averaged neon impurity concentration (⟨ c Ne ⟩ ≈ 6%) and divertor neutral pressure (⟨ p div ⟩ > 25 Pa) in order to determine the heat flux mitigation capability of these scenarios and whether strongly detached states are accessible. In the existing database of ITER simulations, the level of detachment was limited to cases where the integral ion flux to the outer target was greater than 80% of the value at rollover, with the impurity radiation localized near the target. With the possibility of narrow heat flux channels and increased deposited power due to tile shaping, it is important to explore operation at a higher degree of detachment. Two series of simulations were explored to extend the database of SOLPS simulations. By increasing the deuterium and neon puff rates proportionally, the peak divertor energy flux ( q ⊥,max ) is decreased from 5 to 3 MW m −2 while ⟨ p div ⟩ increased from 11 to 27 Pa. By increasing only the neon puff, q ⊥, max can be reduced to <1MW m −2 while ⟨ p div ⟩ is maintained at ∼ 11 Pa. As the neon puff level is increased, the position of the impurity radiation peak is shifted towards the X-point. At the highest neon puff levels with steady-state solutions, the electron temperature is reduced below 1 eV across 50 cm of each divertor target. The new cases extend previously observed tight relationships in power and momentum loss factors to low electron temperature improving their utility for highly detached regimes.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

HIResist: a database of HIV-1 resistance to broadly neutralizing antibodies

Changing the course of the human immunodeficiency virus type I (HIV-1) pandemic is a high public health priority with approximately 39 million people currently living with HIV-1 (PLWH) and about 1.5 million new infections annually worldwide. Broadly neutralizing antibodies (bnAbs) typically target highly conserved sites on the HIV-1 envelope glycoproteins (Envs), which mediate viral entry, and block the infection of diverse HIV-1 strains. But different mechanisms of HIV-1 resistance to bnAbs prevent robust application of bnAbs for therapeutic and preventive interventions. Here we report the development of a new database that provides data and computational tools to aid the discovery of resistant features and may assist in analysis of HIV-1 resistance to bnAbs. Bioinformatic tools allow identification of specific patterns in Env sequences of resistant strains and development of strategies to elucidate the mechanisms of HIV-1 escape; comparison of resistant and sensitive HIV-1 strains for each bnAb; identification of resistance and sensitivity signatures associated with specific bnAbs or groups of bnAbs; and visualization of antibody pairs on cross-sensitivity plots. The database has been designed with a particular focus on user-friendly and interactive interface. Our database is a valuable resource for the scientific community and provides opportunities to investigate patterns of HIV-1 resistance and to develop new approaches aimed to overcome HIV-1 resistance to bnAbs.

60 APPLIED LIFE SCIENCES↗

PYK-SubstitutionOME: an integrated database containing allosteric coupling, ligand affinity and mutational, structural, pathological, bioinformatic and computational information about pyruvate kinase isozymes

Interpreting changes in patient genomes, understanding how viruses evolve and engineering novel protein function all depend on accurately predicting the functional outcomes that arise from amino acid substitutions. To that end, the development of first-generation prediction algorithms was guided by historic experimental datasets. However, these datasets were heavily biased toward substitutions at positions that have not changed much throughout evolution (i.e. conserved). Although newer datasets include substitutions at positions that span a range of evolutionary conservation scores, these data are largely derived from assays that agglomerate multiple aspects of function. To facilitate predictions from the foundational chemical properties of proteins, large substitution databases with biochemical characterizations of function are needed. We report here a database derived from mutational, biochemical, bioinformatic, structural, pathological and computational studies of a highly studied protein family—pyruvate kinase (PYK). A centerpiece of this database is the biochemical characterization—including quantitative evaluation of allosteric regulation—of the changes that accompany substitutions at positions that sample the full conservation range observed in the PYK family. We have used these data to facilitate critical advances in the foundational studies of allosteric regulation and protein evolution and as rigorous benchmarks for testing protein predictions. We trust that the collected dataset will be useful for the broader scientific community in the further development of prediction algorithms.

59 BASIC BIOLOGICAL SCIENCES↗

IMG/PR: a database of plasmids from genomes and metagenomes with rich annotations and metadata

Plasmids are mobile genetic elements found in many clades of Archaea and Bacteria. They drive horizontal gene transfer, impacting ecological and evolutionary processes within microbial communities, and hold substantial importance in human health and biotechnology. To support plasmid research and provide scientists with data of an unprecedented diversity of plasmid sequences, we introduce the IMG/PR database, a new resource encompassing 699 973 plasmid sequences derived from genomes, metagenomes and metatranscriptomes. IMG/PR is the first database to provide data of plasmid that were systematically identified from diverse microbiome samples. IMG/PR plasmids are associated with rich metadata that includes geographical and ecosystem information, host taxonomy, similarity to other plasmids, functional annotation, presence of genes involved in conjugation and antibiotic resistance. The database offers diverse methods for exploring its extensive plasmid collection, enabling users to navigate plasmids through metadata-centric queries, plasmid comparisons and BLAST searches. The web interface for IMG/PR is accessible at https://img.jgi.doe.gov/pr. Plasmid metadata and sequences can be downloaded from https://genome.jgi.doe.gov/portal/IMG_PR.

59 BASIC BIOLOGICAL SCIENCES↗

CABO-16S—a Combined Archaea, Bacteria, Organelle 16S rRNA database framework for amplicon analysis of prokaryotes and eukaryotes in environmental samples

Abstract Identification of both prokaryotic and eukaryotic microorganisms in environmental samples is currently challenged by the need for additional sequencing to obtain separate 16S and 18S ribosomal RNA (rRNA) amplicons or the constraints imposed by “universal” primers. Organellar 16S rRNA sequences are amplified and sequenced along with prokaryote 16S rRNA and provide an alternative method to identify eukaryotic microorganisms. CABO-16S combines bacterial and archaeal sequences from the SILVA database with 16S rRNA sequences of plastids and other organelles from the PR2 database to enable identification of all 16S rRNA sequences. Comparison of CABO-16S with SILVA 138.2 results in equivalent taxonomic classification of mock communities and increased classification of diverse environmental samples. In particular, identification of phototrophic eukaryotes in shallow seagrass environments, marine waters, and lake waters was increased. The CABO-16S framework allows users to add custom sequences for further classification of underrepresented clades and can be easily updated with future releases of reference databases. Addition of sequences obtained from Sanger sequencing of methane seep sediments and curated sequences of the polyphyletic SEEP-SRB1 clade resulted in differentiation of syntrophic and non-syntrophic SEEP-SRB1 in hydrothermal vent sediments. CABO-16S highlights the benefit of combining and amending existing training sets when studying microorganisms in diverse environments.

Eitel, Eryn M. (ORCID:0009000723919297)↗

Open database for GPD analyses

This article summarizes the main ideas behind creating an open database proposed for use in the exploration of generalized parton distributions (GPDs). This lightweight database is well suited for GPD phenomenology and is designed to store both experimental and lattice-QCD data. It can also aid in benchmarking GPD-related developments, such as GPD models. The database utilizes a new data format based on the YAML serialization language, enabling the storage of essential information for modern analyses, such as replica values. It includes interfaces for both Python and C++, allowing straightforward integration with analysis codes.

Burkert, V. D. [Thomas Jefferson National Accelera↗

Relationship and distribution of Salmonella enterica serovar I 4,[5],12:i:- strain sequences in the NCBI Pathogen Detection database

Background: Of the > 2600 Salmonella serovars, Salmonella enterica serovar I 4,[5],12:i:- (serovar I 4,[5],12:i:-) has emerged as one of the most common causes of human salmonellosis and the most frequent multidrug-resistant (MDR; resistance to ≥3 antimicrobial classes) nontyphoidal Salmonella serovar in the U.S. Serovar I 4,[5],12:i:- isolates have been described globally with resistance to ampicillin, streptomycin, sulfisoxazole, and tetracycline (R-type ASSuT) and an integrative and conjugative element with multi-metal tolerance named Salmonella Genomic Island 4 (SGI-4). Results: We analyzed 13,612 serovar I 4,[5],12:i:- strain sequences available in the NCBI Pathogen Detection database to determine global distribution, animal sources, presence of SGI-4, occurrence of R-type ASSuT, frequency of antimicrobial resistance (AMR), and potential transmission clusters. Genome sequences for serovar I 4,[5],12:i:- strains represented 30 countries from 5 continents (North America, Europe, Asia, Oceania, and South America), but sequences from the United States (59%) and the United Kingdom (28%) were dominant. The metal tolerance island SGI-4 and the R-type ASSuT were present in 71 and 55% of serovar I 4,[5],12:i:- strain sequences, respectively. Sixty-five percent of strain sequences were MDR which correlates to serovar I 4,[5],12:i:- being the most frequent MDR serovar. The distribution of serovar I 4,[5],12:i:- strain sequences in the NCBI Pathogen Detection database suggests that swine-associated strain sequences were the most frequent food-animal source and were significantly more likely to contain the metal tolerance island SGI-4 and genes for MDR compared to all other animal-associated isolate sequences. Conclusions: Our study illustrates how analysis of genomic sequences from the NCBI Pathogen Detection database can be utilized to identify the prevalence of genetic features such as antimicrobial resistance, metal tolerance, and virulence genes that may be responsible for the successful emergence of bacterial foodborne pathogens.

59 BASIC BIOLOGICAL SCIENCES↗

Downloadable Dynamometer Database (D3): Public Test Data on Advanced-Technology Vehicles

Access to high-quality, independent vehicle test data is critical to advancing energy-efficient transportation research. The Downloadable Dynamometer Database (D3) is a public repository of dynamometer test data on advanced-technology vehicles, generated at the Advanced Mobility Technology Laboratory (AMTL) at Argonne National Laboratory and hosted by the Transportation and Power Systems Division. The database has been made available to support researchers, students, and professionals engaged in energy-efficient vehicle research, development, and education. A wide range of vehicle categories has been tested (i.e., alternative fuel vehicles, conventional gasoline and diesel vehicles, all-electric vehicles, hybrid electric vehicles, and plug-in hybrid electric vehicles), as well as various drive cycles and test conditions documented in the accompanying D3 user presentation. Stakeholders can select a vehicle type, identify a vehicle of interest, and download the associated test data for use in their own analyses. Data downloaded from D3 must be accompanied by the required attribution: "This data is from the Downloadable Dynamometer Database and was generated at the Advanced Mobility Technology Laboratory (AMTL) at Argonne National Laboratory." These data are critical to vehicle modeling, validation, technology assessment, and educational use.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

The Global Spectra-Trait Initiative: A database of paired leaf spectroscopy and functional traits associated with leaf photosynthetic capacity (v1.0.0)

The Global Spectra-Trait Initiative (GSTI) aims to generate generalizable spectra trait models using reflectance data to predict leaf traits associated with the photosynthesis capacity of leaves. It comprises a synthesized dataset of leaf trait data, input datasets and code. Leaf traits include the maximum carboxylation rate of rubisco (Vcmax), the maximum electron transport rate (Jmax), the dark respiration, as well as the prediction of leaf nitrogen, leaf mass per area (LMA), and leaf water content (LWC). The dataset comprises >7500 paired observations from around 400 species from a broad range of biomes. This dataset comprises a zip file of the GSTI GitHub repository (https://github.com/plantphys/gsti), the synthesized database (.csv) and database metadata files. This dataset was updated on 2025-12-12 with minor edits to mirror the accepted manuscript version and GitHub release (Version 1.0.0 (ESSD accepted version)). Edits included minor changes to the project documentation on GitHub and removal of 12 duplicate entries from the database.

54 ENVIRONMENTAL SCIENCES↗