Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Large-database cross-verification and validation of tokamak transport models using baselines for comparison

State-of-the-art 1D transport solvers ASTRA and TRANSP are verified, then validated across a large database of semi-randomly selected, time-dependent DIII-D discharges. Various empirical models are provided as baselines to contextualize the validation figures of merit using statistical hypothesis tests. For predicting plasma temperature profiles, no statistically significant advantage is found for the ASTRA and TRANSP simulators over a baseline empirical (two-parameter) model. For predicting stored energy, a significant advantage is found for the simulators over a baseline empirical model based on confinement time scaling. Uncertainty in the results due to diagnostic and profile fitting uncertainties is approximated and determined to be insignificant due in part to the large quantity of discharges employed in the study. Advantages are discussed for validation methodologies like this one that employ (1) large databases and (2) baselines for comparison that are specific to the intended use-case of the model.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Simulating charged defects at database scale

Point defects have a strong influence on the physical properties of materials, often dominating the electronic and optical behavior in semiconductors and insulators. The simulation and analysis of point defects is, therefore, crucial for understanding the growth and operation of materials, especially for optoelectronics applications. In this work, we present a general-purpose Python framework for the analysis of point defects in crystalline materials as well as a generalized workflow for their treatment with high-throughput simulations. The distinguishing feature of our approach is an emphasis on a unique, unit cell, structure-only, definition of point defects which decouples the defect definition, and the specific supercell representation used to simulate the defect. This allows the results of first-principles calculations to be aggregated into a database without extensive provenance information and is a crucial step in building a persistent database of point defects that can grow over time, a key component toward realizing the idea of a “defect genome” that can yield more complex relationships governing the behavior of defects in materials. We demonstrate several examples of the approach for three technologically relevant materials and highlight current pitfalls that must be considered when employing these methodologies as well as their potential solutions.

36 MATERIALS SCIENCE↗

United States Nuclear Power Reactor Used Nuclear Fuel Database and Applications

The Unified Database (UDB) within STANDARDS serves as the foundational data infrastructure for managing the United States' spent nuclear fuel inventory of 315,111 discharged assemblies totaling 91,036 metric tons of heavy metal. The database organizes this complex inventory through over 200 interconnected tables structured into eight primary attribute categories, supporting integrated analyses across storage, transportation, and disposal domains. Data enters the UDB through the GC-859 Nuclear Fuel Data Survey, which transitioned to web-based collection in 2023, improving data quality through real-time validation. The UDB enables automated generation of input files for nuclear safety analyses, reducing preparation time from weeks to hours while maintaining traceability. Applications include national inventory reporting, Certificate of Compliance assessments, and facility optimization. The three-tier distribution model balances accessibility with security requirements for federal agencies, national laboratories, and research organizations. The UDB provides essential data infrastructure as spent fuel management transitions from site-specific to integrated national campaigns.

Stefanovic, Peter↗

Second release of the CoRe database of binary neutron star merger waveforms

Abstract We present the second data release of gravitational waveforms from binary neutron star (BNS) merger simulations performed by the Computational Relativity ( CoRe ) collaboration. The current database consists of 254 different BNS configurations and a total of 590 individual numerical-relativity simulations using various grid resolutions. The released waveform data contain the strain and the Weyl curvature multipoles up to ℓ = m = 4 . They span a significant portion of the mass, mass-ratio, spin and eccentricity parameter space and include targeted configurations to the events GW170817 and GW190425. CoRe simulations are performed with 18 different equations of state, seven of which are finite temperature models, and three of which account for non-hadronic degrees of freedom. About half of the released data are computed with high-order hydrodynamics schemes for tens of orbits to merger; the other half is computed with advanced microphysics. We showcase a standard waveform error analysis and discuss the accuracy of the database in terms of faithfulness. We present ready-to-use fitting formulas for equation of state-insensitive relations at merger (e.g. merger frequency), luminosity peak, and post-merger spectrum.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Assessing the numerical stability of physics models to equilibrium variation through database comparisons on DIII-D

High fidelity kinetic equilibria are crucial for tokamak modeling and analysis. Manual workflows for constructing kinetic equilibria are time consuming and subject to user error, motivating development of automated equilibrium reconstruction tools to provide accurate and consistent reconstructions for downstream physics analysis. These automated tools also provide access to kinetic equilibria at large database scales, which enables the quantification of general uncertainties arising from equilibrium reconstruction techniques. In this paper, we compare a large database of DIII-D kinetic equilibria generated manually by physics experts to equilibria from automated kinetic reconstruction tools, assessing the impact of reconstruction method on equilibrium parameters and resulting magnetohydrodynamic stability calculations. We find agreement among scalar parameters, whereas profile quantities, such as the bootstrap current, show larger disagreements. We analyze ideal kink and classical tearing stability with DCON and STRIDE respectively, finding that the kink stability calculation is generally more robust than the tearing index Δ' calculation. We find that in 90% of cases, both kink stability classifications are unchanged between the manual expert and automated kinetic equilibria.

CAKE↗

Physics insights from a large-scale 2D UEDGE simulation database for detachment control in KSTAR

A large-scale database of two-dimensional UEDGE simulations has been developed to study detachment physics in KSTAR and to support surrogate models for control applications. Nearly 70 000 steady-state solutions were generated, systematically scanning upstream density, input power, plasma current, impurity fraction, and anomalous transport coefficients, with magnetic and electric drifts across the magnetic field included. The database identifies robust detachment indicators, with strike-point electron temperature at detachment onset consistently T e,target ~ 3-4 eV, largely insensitive to upstream conditions. Scaling relations reveal weaker impurity sensitivity than one-dimensional models and show that heat flux widths follow Eich’s scaling only for uniform, low D and χ. Distinctive in–out divertor asymmetries are observed in KSTAR, differing qualitatively from DIII-D. Complementary time-dependent simulations quantify plasma response to gas puffing, with delays of 5 - 15 ms at the outer strike point and ∼ 40 ms for the low-magnetic-field-side radiation front. These dynamics are well captured by first-order-plus-dead-time models and are consistent with experimentally observed detachment-control behavior in KSTAR (Gupta et al 2025 Plasma Phys. Control. Fusion (submitted)).

Physics - Plasma physics↗

pKPDB: a protein data bank extension database of p Ka and pI theoretical values

Abstract Summary pKa values of ionizable residues and isoelectric points of proteins provide valuable local and global insights about their structure and function. These properties can be estimated with reasonably good accuracy using Poisson–Boltzmann and Monte Carlo calculations at a considerable computational cost (from some minutes to several hours). pKPDB is a database of over 12 M theoretical pKa values calculated over 120k protein structures deposited in the Protein Data Bank. By providing precomputed pKa and pI values, users can retrieve results instantaneously for their protein(s) of interest while also saving countless hours and resources that would be spent on repeated calculations. Furthermore, there is an ever-growing imbalance between experimental pKa and pI values and the number of resolved structures. This database will complement the experimental and computational data already available and can also provide crucial information regarding buried residues that are under-represented in experimental measurements. Availability and implementation Gzipped csv files containing p Ka and isoelectric point values can be downloaded from https://pypka.org/pKPDB. To query a single PDB code please use the PypKa free server at https://pypka.org. The pKPDB source code can be found at https://github.com/mms-fcul/pKPDB. Supplementary information Supplementary data are available at Bioinformatics online.

Reis, Pedro B. P. S. (ORCID:0000000335636239)↗

Standardized naming of microbiome samples in Genomes OnLine Database

The power of next-generation sequencing has resulted in an explosive growth in the number of projects aiming to understand the metagenomic diversity of complex microbial environments. The interdisciplinary nature of this microbiome research community, along with the absence of reporting standards for microbiome data and samples, poses a significant challenge for follow-up studies. Commonly used names of metagenomes and metatranscriptomes in public databases currently lack the essential information necessary to accurately describe and classify the underlying samples, which makes a comparative analysis difficult to conduct and often results in misclassified sequences in data repositories. The Genomes OnLine Database (GOLD) (https:// gold.jgi.doe.gov/) at the Department of Energy Joint Genome Institute has been at the forefront of addressing this challenge by developing a standardized nomenclature system for naming microbiome samples. GOLD, currently in its twenty-fifth anniversary, continues to enrich the research community with hundreds of thousands of metagenomes and metatranscriptomes with well-curated and easy-to-understand names. Through this manuscript, we describe the overall naming process that can be easily adopted by researchers worldwide. Additionally, we propose the use of this naming system as a best practice for the scientific community to facilitate better interoperability and reusability of microbiome data.

59 BASIC BIOLOGICAL SCIENCES↗

The impact of curation errors in the PDBBind Database on machine learning predictions of protein–protein binding affinity

The PDBBind database has been widely utilized for the computational prediction of protein–protein binding affinities. While the accuracy of the PDBBind-curated equilibrium dissociation constants (K D ) has been reported for the protein–ligand subset of the PDBBind database, the curation accuracy has not been reported for the protein–protein subset. Here, we present a detailed manual analysis for the subset of PDBBind records with PubMed Central Open Access primary publications and find that ~19% of these records had K D values that were not supported by their primary publications. The impact of these putative curation errors on the machine learning-based prediction of K D from experimental protein–protein 3D structures was evaluated and correcting the curation errors improved the Pearson correlation coefficient between measured and random forest-predicted log 10 (K D ) values by ~8 percentage points. This finding underscores the importance of dataset accuracy for computational modelling and highlights the need for more stringent curation processes when extracting information from the scientific literature.

59 BASIC BIOLOGICAL SCIENCES↗

TBDB: a database of structurally annotated T-box riboswitch:tRNA pairs

Abstract T-box riboswitches constitute a large family of tRNA-binding leader sequences that play a central role in gene regulation in many gram-positive bacteria. Accurate inference of the tRNA binding to T-box riboswitches is critical to predict their cis-regulatory activity. However, there is no central repository of information on the tRNA binding specificities of T-box riboswitches, and de novo prediction of binding specificities requires advanced knowledge of computational tools to annotate riboswitch secondary structure features. Here, we present the T-box Riboswitch Annotation Database (TBDB, https://tbdb.io), an open-access database with a collection of 23,535 T-box riboswitch sequences, spanning the major phyla of 3,632 bacterial species. Among structural predictions, the TBDB also identifies specifier sequences, cognate tRNA binding partners, and downstream regulatory targets. To our knowledge, the TBDB presents the largest collection of feature, sequence, and structural annotations carried out on this important family of regulatory RNA.

59 BASIC BIOLOGICAL SCIENCES↗

NP-MRD: the Natural Products Magnetic Resonance Database

The Natural Products Magnetic Resonance Database (NP-MRD) is a comprehensive, freely available electronic resource for the deposition, distribution, searching and retrieval of nuclear magnetic resonance (NMR) data on natural products, metabolites and other biologically derived chemicals. NMR spectroscopy has long been viewed as the ‘gold standard’ for the structure determination of novel natural products and novel metabolites. NMR is also widely used in natural product dereplication and the characterization of biofluid mixtures (metabolomics). All of these NMR applications require large collections of high quality, well-annotated, referential NMR spectra of pure compounds. Unfortunately, referential NMR spectral collections for natural products are quite limited. It is because of the critical need for dedicated, open access natural product NMR resources that the NP-MRD was funded by the National Institute of Health (NIH). Since its launch in 2020, the NP-MRD has grown quickly to become the world's largest repository for NMR data on natural products and other biological substances. It currently contains both structural and NMR data for nearly 41,000 natural product compounds from >7400 different living species. All structural, spectroscopic and descriptive data in the NP-MRD is interactively viewable, searchable and fully downloadable in multiple formats. Extensive hyperlinks to other databases of relevance are also provided. The NP-MRD also supports community deposition of NMR assignments and NMR spectra (1D and 2D) of natural products and related meta-data. The deposition system performs extensive data enrichment, automated data format conversion and spectral/assignment evaluation.

59 BASIC BIOLOGICAL SCIENCES↗

NMPFamsDB: a database of novel protein families from microbial metagenomes and metatranscriptomes

Abstract The Novel Metagenome Protein Families Database (NMPFamsDB) is a database of metagenome- and metatranscriptome-derived protein families, whose members have no hits to proteins of reference genomes or Pfam domains. Each protein family is accompanied by multiple sequence alignments, Hidden Markov Models, taxonomic information, ecosystem and geolocation metadata, sequence and structure predictions, as well as 3D structure models predicted with AlphaFold2. In its current version, NMPFamsDB hosts over 100 000 protein families, each with at least 100 members. The reported protein families significantly expand (more than double) the number of known protein sequence clusters from reference genomes and reveal new insights into their habitat distribution, origins, functions and taxonomy. We expect NMPFamsDB to be a valuable resource for microbial proteome-wide analyses and for further discovery and characterization of novel functions. NMPFamsDB is publicly available in http://www.nmpfamsdb.org/ or https://bib.fleming.gr/NMPFamsDB.

59 BASIC BIOLOGICAL SCIENCES↗

Genomes OnLine Database (GOLD) v.10: new features and updates

The Genomes OnLine Database (GOLD; https://gold.jgi.doe.gov/) at the Department of Energy Joint Genome Institute is a comprehensive online metadata repository designed to catalog and manage information related to (meta)genomic sequence projects. GOLD provides a centralized platform where researchers can access a wide array of metadata from its four organization levels namely Study, Organism/Biosample, Sequencing Project and Analysis Project. GOLD continues to serve as a valuable resource and has seen significant growth and expansion since its inception in 1997. With its expanded role as a collaborative platform, it not only actively imports data from other primary repositories like National Center for Biotechnology Information but also supports contributions from researchers worldwide. This collaborative approach has enriched the database with diverse datasets, creating a more integrated resource to enhance scientific insights. As genomic research becomes increasingly integral to various scientific disciplines, more researchers and institutions are turning to GOLD for their metadata needs. To meet this growing demand, GOLD has expanded by adding diverse metadata fields, intuitive features, advanced search capabilities and enhanced data visualization tools, making it easier for users to find and interpret relevant information. This manuscript provides an update and highlights the new features introduced over the last 2 years.

59 BASIC BIOLOGICAL SCIENCES↗

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N↗

Expansion of the tmRNA sequence database and new tools for search and visualization

Abstract Transfer–messenger RNA (tmRNA) contributes essential tRNA-like and mRNA-like functions during the process of trans-translation, a mechanism of quality control for the translating bacterial ribosome. Proper tmRNA identification benefits the study of trans-translation and also the study of genomic islands, which frequently use the tmRNA gene as an integration site. Automated tmRNA gene identification tools are available, but manual inspection is still important for eliminating false positives. We have increased our database of precisely mapped tmRNA sequences over 50-fold to 97 179 unique sequences. Group I introns had previously been found integrated within a single subsite within the TψC-loop; they have now been identified at four distinct subsites, suggesting multiple founding events of invasion of tmRNA genes by group I introns, all in the same vicinity. tmRNA genes were found in metagenomic archaeal genomes, perhaps a result of misbinning of bacterial sequences during genome assembly. With the expanded database, we have produced new covariance models for improved tmRNA sequence search and new secondary structure visualization tools.

59 BASIC BIOLOGICAL SCIENCES↗

First-principles equation of state database for warm dense matter computation

We put together a first-principles equation of state (FPEOS) database for matter at extreme conditions by combining results from path integral Monte Carlo and density functional molecular dynamics simulations of the elements H, He, B, C, N, O, Ne, Na, Mg, Al, and Si as well as the compounds LiF , B 4 C , BN , CH 4 , CH 2 , C 2 H 3 , CH , C 2 H , MgO , and MgSiO 3 . For all these materials, we provide the pressure and internal energy over a density-temperature range from ~ 0.5 to 50 g cm - 3 and from ~ 10 4 to 10 9 K, which are based on ~ 5000 different first-principles simulations. We compute isobars, adiabats, and shock Hugoniot curves in the regime of L - and K -shell ionization. Invoking the linear mixing approximation, we study the properties of mixtures at high density and temperature. Furthermore, we derive the Hugoniot curves for water and alumina as well as for carbon-oxygen, helium-neon, and CH-silicon mixtures. We predict the maximal shock compression ratios of H 2 O , H 2 O 2 , Al 2 O 3 , CO , and CO 2 to be 4.61, 4.64, 4.64, 4.89, and 4.83, respectively. Finally we use the FPEOS database to determine the points of maximum shock compression for all available binary mixtures. We identify mixtures that reach higher shock compression ratios than their end members. We discuss trends common to all mixtures in pressure-temperature and particle-shock velocity spaces. In the Supplemental Material, we provide all FPEOS tables as well as computer codes for interpolation, Hugoniot calculations, and plots of various thermodynamic functions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Statistics on magnetic properties of Co compounds: A database-driven method for discovering Co-based ferromagnets

The search for new ferromagnetic compounds is often targeted at known structure families, particularly those containing iron, cobalt, and manganese. Here, we propose a method to expand this search to lesser-known structure types, using a database of experimental Curie and Néel temperatures. This study, in particular, illustrates the case of compounds containing the element cobalt, and we demonstrate how the use of such a database can lead to the discovery of ferromagnetic materials that had previously been overlooked. We report statistics from a literature survey of the magnetic properties of Co-based compounds with more than 33 at. % Co. We classify more than 13 000 compounds by structure type, cobalt content, and magnetic ground state. From these data, compounds TaCo 2 Ga, La 6 Co 13 Bi, and Nd 2 Co 3 were identified as potential ferromagnets, and we confirm their ferromagnetic ordering theoretically via first-principles calculations and experimentally via synthesis and characterization measurements. In addition, the analysis is focused on the collection of data trends and discovery of ferromagnetic materials with easy-axis magnetic anisotropy. Both known ferromagnetic materials with unknown magnetic anisotropy, and unstudied compounds were considered. From the subset of known ferromagnets with unknown anisotropy, the compound Co 2 Mg was synthesized, characterized, and determined to have easy-axis magnetic anisotropy at room temperature.

36 MATERIALS SCIENCE↗

Resources, Training, and Education Under the Heliostat Consortium: Industry Gap Analysis and Building a Resource Database

Concentrating solar power is not a widely deployed or known technology area, and the heliostat workforce community in the United States is currently small, with knowledge and expertise not widely available. The resource, training, and education (RTE) topic within the Heliostat Consortium (HelioCon) was established to address this. RTE encompasses resources, practices, and programs to ensure that (1) newcomers to the heliostat development community have an adequate knowledge base and training to conduct R&D efforts, (2) outsiders to the field are provided with resources and opportunities to join the workforce, and (3) the workforce community is a productive, healthy, and fulfilling environment for all workers. In the first year of the project, a roadmap study was conducted, in which the major gaps in RTE were identified by consulting experts in the industry, with the top gap being the lack of public accessibility to concentrating solar-thermal power (CSP) knowledge. Here, to address this, the HelioCon team has been developing a centralized web-based resource database, containing a reference library, educational videos, lists of components suppliers and software/metrology tools, a power tower plant database, and information on existing standards/guidelines.

14 SOLAR ENERGY↗