Engineering PapersSearch

SEARCH · Engineering Papers

Results for “database improvement”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A global database of soil microbial phospholipid fatty acids and enzyme activities

Abstract Soil microbes drive ecosystem function and play a critical role in how ecosystems respond to global change. Research surrounding soil microbial communities has rapidly increased in recent decades, and substantial data relating to phospholipid fatty acids (PLFAs) and potential enzyme activity have been collected and analysed. However, studies have mostly been restricted to local and regional scales, and their accuracy and usefulness are limited by the extent of accessible data. Here we aim to improve data availability by collating a global database of soil PLFA and potential enzyme activity measurements from 12,258 georeferenced samples located across all continents, 5.1% of which have not previously been published. The database contains data relating to 113 PLFAs and 26 enzyme activities, and includes metadata such as sampling date, sample depth, and soil pH, total carbon, and total nitrogen. This database will help researchers in conducting both global- and local-scale studies to better understand soil microbial biomass and function.

Science & Technology - Other Topics

Particle Physics Division Lifting Fixture Database Restructure

The Particle Physics Division (PPD) at Fermilab needed to modernize its outdated lifting fixture database. Many fixtures lacked identification or had incomplete information. During a summer internship, I cross-referenced existing data, photographed and measured fixtures using standardized tools, and updated their locations and details in a modern Excel format. Some fixtures were identified as belonging to other departments, like the Applied Physics and Superconducting Technology Directorate (APS-TD). This effort significantly improved the accuracy and utility of the PPD database, streamlined other departmental inventories, and potentially saved hundreds of thousands of dollars. The work concluded with a clear, organized record of all verified lifting fixtures.

Creedon, Carroll [Unlisted, US, IL; Fermilab]

Produced Water DNA Database (PW-DNA): Utilizing KBase to generate an environmental specific curated molecular database

The deep subsurface is estimated to host the majority of Earth’s microbial biomass yet remains one of the most challenging environments to access and study. One common approach to investigate these microbial communities is through the analysis of produced water from subsurface reservoirs, where researchers can assess water and gas chemistry along with molecular (DNA/RNA) sequence data. Advances in high-throughput sequencing have greatly expanded our understanding of these environments and their biotechnological potential. However, further progress requires large-scale, integrative meta-analyses across diverse datasets. To address this need, we developed the Produced Water-DNA (PW-DNA) Database, a curated, publicly available resource that consolidates microbial DNA/RNA sequences, geochemical data, and relevant metadata from in situ hydrocarbon environments such as coal beds, oil reservoirs, and natural gas systems. The PW-DNA database delivers three core benefits to the research community: (1) it improves data sharing by linking environmental microbial datasets with corresponding geochemical parameters, enabling more robust filtering and analysis; (2) it connects with complementary research databases to promote broader dissemination and interoperability; and (3) it supports technological innovation by serving as a resource for identifying microbial trends and exploring genetic potential. While individual studies have highlighted basin-specific microbial communities and functional redundancy in biogeochemical cycling, a comprehensive, system-wide perspective is needed to better understand connectivity and novelty across subsurface ecosystems. By designing the PW-DNA in the KBase platform, we provide a reproducible, visual framework for integrating large-scale genomic and geochemical data, enabling researchers to perform more informed analyses and experimental design. Ultimately, this resource enhances the ability to identify, characterize, and interpret microbial functions across diverse subsurface environments, thereby accelerating discovery in subsurface microbiology and biotechnology.

59 BASIC BIOLOGICAL SCIENCES

Updating the Building Science Advisor (BSA): A Tool to Assist in the Design of Durable Building Envelopes

Predicting the moisture durability of building envelope components remains challenging due to multiple influencing factors, including material selection, assembly positioning, local climate conditions, air tightness, interior environment, and construction quality. Building codes increasingly emphasize energy efficiency through enhanced insulation and tighter envelopes but offer limited guidance on moisture durability considerations. Consequently, builders face uncertainty, particularly as new materials and assemblies enter the market.The Building Science Advisor (BSA) is a free, web-based expert system developed to address these challenges by providing actionable insights into the moisture durability and energy efficiency of both new and retrofit wall designs. Recently updated, we are now providing version 3.0 of the tool. BSA features significant user interface improvements, enhancing navigation and user interaction through a refreshed, intuitive design. Additionally, the tool incorporates a newly developed database containing pre-simulated wall assembly cases, significantly reducing response times and improving the accuracy of moisture durability assessments. Furthermore, the updated BSA includes moisture content as a performance criterion, providing users with a more comprehensive understanding of moisture-related durability risks. These enhancements enable rapid, reliable assessments tailored to specific climate zones and local building practices. BSA continues to offer targeted guidance on wall retrofit scenarios and delivers access to an expanded library of location-specific building science resources.This paper describes these key updates, highlighting the enhanced features, expanded capabilities, and overall improvements to user experience and educational content. The paper includes a demonstration that illustrates how the revised BSA effectively supports practitioners in designing durable, energy-efficient building envelope assemblies.

Salonvaara, Mikael [ORNL] (ORCID:0000000318991554)

Capturing Historic Reliability Performance Through Graph Databases: A Model Based System Engineering Approach

With the goal of improving the performance and reliability of high dependable technological systems such as nuclear power plants, advanced monitoring and health management systems are employed to inform system engineers on observed degradation processes and anomalous behaviors of assets and components. This information is captured in the form of large amount of data which can be heterogenous in nature (e.g., numeric, textual). Such large data availability poses challenges when system engineers are required to parse and analyze them in order to track historic reliability performance of assets and components. This paper tackles directly this challenge by providing means to organize data in the form of a graph: a knowledge graph. The presented approach distinguish itself from current knowledge graph-based methods by the fact that model-based system engineering (MBSE) models are used to “put data into context”. In particular, MBSE models are used as skeleton of a knowledge graph; numeric and textual data elements, once processed, are associated to MBSE model elements. Thus, a knowledge graph captures both system architecture (though MBSE models) and health/performance data. Such feature opens the door to new data analytics methods designed to identify causal relations between observed phenomena.

97 - MATHEMATICS AND COMPUTING

Expansion of the tmRNA sequence database and new tools for search and visualization

Abstract Transfer–messenger RNA (tmRNA) contributes essential tRNA-like and mRNA-like functions during the process of trans-translation, a mechanism of quality control for the translating bacterial ribosome. Proper tmRNA identification benefits the study of trans-translation and also the study of genomic islands, which frequently use the tmRNA gene as an integration site. Automated tmRNA gene identification tools are available, but manual inspection is still important for eliminating false positives. We have increased our database of precisely mapped tmRNA sequences over 50-fold to 97 179 unique sequences. Group I introns had previously been found integrated within a single subsite within the TψC-loop; they have now been identified at four distinct subsites, suggesting multiple founding events of invasion of tmRNA genes by group I introns, all in the same vicinity. tmRNA genes were found in metagenomic archaeal genomes, perhaps a result of misbinning of bacterial sequences during genome assembly. With the expanded database, we have produced new covariance models for improved tmRNA sequence search and new secondary structure visualization tools.

59 BASIC BIOLOGICAL SCIENCES

Investigating the Impacts of Aqueous-Phase Processing on Organic Aerosol Chemical Climatology Using ARM and ASR Observations

This project improved understanding of how atmospheric aerosol particles form and evolve, with a focus on the role of water-driven (aqueous-phase) chemical reactions in the atmosphere. These processes occur in clouds, fog, and humid air and can significantly change the composition and properties of airborne particles, known as aerosols, which influence air quality and climate. By combining field measurements, laboratory experiments, and advanced analytical techniques, the project identified key chemical signatures that allow scientists to distinguish particles formed through aqueous processes from those formed in the gas phase. Observations from multiple environments, including wildfire smoke, urban regions, and cloud-influenced areas, show that aqueous chemistry is an important pathway for particle formation and aging. The project also developed new measurement approaches using uncrewed aerial systems (UAS) to capture how aerosol composition varies with altitude, providing critical insights into how particles interact with clouds. In addition, new data analysis frameworks were created to better interpret long-term aerosol measurements and improve characterization of particle sources and transformations. These results have been integrated into a global database of aerosol measurements and used to support atmospheric modeling efforts. Overall, the project provides important tools and knowledge to improve predictions of how aerosols affect climate and air quality, particularly through their interactions with radiation and clouds.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Towards a liana plant functional type for vegetation models

Lianas (woody climbers) are crucial components of tropical forests and they have been increasingly recognized to have profound effects on tropical forest carbon dynamics. Despite their importance, lianas' representation in vegetation models remains limited, partly due to the complexity of liana-tree dynamics and the diversity in liana life history strategies. This paper provides a comprehensive review of advances and challenges for mechanistically representing lianas in forest ecosystem models and a proposed path towards effectively representing lianas in these models. Defining a liana plant functional type is a significant challenge because of the high morphological and physiological diversity amongst liana species, and because of their structural association with trees. Here, we identify critical liana traits that likely should contribute to establishing a liana plant functional type, along with key processes to properly represent lianas in ecosystem models. Subsequently, we discuss a variety of possible liana implementation strategies with their associated strengths, limitations, computational costs and data requirements. A fundamental redesign of the tree-centric demographic vegetation models seems appropriate to accommodate the unique growth and competition strategies of lianas. We illustrate the potential of such models with a single-site case study where we disentangle putative mechanisms of liana increasing abundance. Furthermore, we underscore the critical need for comprehensive liana demographic and functional data (including long-term, physiological, and pantropical observations) for the qualitative implementation and evaluation in the proposed modeling efforts. Currently, there is a scarcity of liana data and the data that do exist have a neotropical bias. We finally introduce a new liana functional trait database that can centralize existing liana trait data, incentivize improved data gathering and thus facilitate model development and scientific analyses.

54 ENVIRONMENTAL SCIENCES

DUNE Rucio Server Scalabiilty Studies

The DUNE collaboration has an ongoing production effort to simulate the full detectors and to analyze the various prototypes that are currently running. Rucio is used to manage the 40PB of files made to date. When 500 or more jobs were sending output to Rucio simultaneously via Rucio upload, we observed timeouts, unhandled exceptions, and Rucio server restarts due to slow performance. In collaboration with the core Rucio team we did a full review of the Rucio upload code and identified several optimizations that can be made. We also have deployed the Ingress load balancer in front of our Rucio servers and added a database connection pooling utility. These changes led to significant improvement both in reliability and scalability, yet we anticipate even better performance will eventually be required. We describe in this paper the initial state of the system, the various debugging processes that were used, and our plans to further improve scalability.

Calcutt, J. [Brookhaven Natl. Lab.]

OmniXAS: A universal deep-learning framework for materials x-ray absorption spectra

X-ray absorption spectroscopy (XAS) is a powerful characterization technique for probing the local chemical environment of absorbing atoms. However, analyzing XAS data presents significant challenges, often requiring extensive, computationally intensive simulations, as well as significant domain expertise. These limitations hinder the development of fast, robust XAS analysis pipelines that are essential in high-throughput studies and for autonomous experimentation. Here, we address these challenges with OmniXAS, a framework that contains a suite of transfer learning approaches for XAS prediction, each uniquely contributing to improved accuracy and efficiency, as demonstrated on the K-edge spectra database covering eight 3⁢d transition metals (Ti–Cu). The OmniXAS framework is built upon three distinct strategies. First, we use M3GNet [Nat. Comput. Sci. 2, 718 (2022)] to derive latent representations of the local chemical environment of absorption sites as input for XAS prediction, achieving significant improvements over conventional featurization techniques. Second, we employ a hierarchical transfer learning strategy, training a universal multitask model across elements before fine-tuning for element-specific predictions. Models based on this cascaded approach after elementwise fine-tuning outperform element-specific models by up to 69%. Third, we implement cross-fidelity transfer learning, adapting a universal model to predict spectra generated by simulation of a different fidelity with a much higher computational cost. This approach improves prediction accuracy by up to 11% over models trained on the target fidelity alone. Our approach significantly boosts the throughput of XAS modeling by orders of magnitude as compared to first-principles simulations and is extendable to XAS prediction for a broader range of elements. The proposed transfer learning framework is generalizable to enhance deep-learning models that target other properties in materials research.

36 MATERIALS SCIENCE

Core Model Proposal #410: Updates to Socioeconomic and Macroeconomic Data, Processing Structure, and Visualization

This Core Model Proposal (CMP) comprehensively restructures and updates the macroeconomic and socioeconomic modules in gcamdata. It includes visualizations of key data inputs, accounting identities, and data flows in the context of GCAM-Macro-KLEM. Major improvements include: (1) updating the Penn World Table (PWT) to version 10 and incorporating a new source, the Global Macro Database (GMD); (2) updating the SSP socioeconomics database from version 3.0.1 to 3.2; (3) introducing SSP-specific differentiation of employment and labor force data; (4) improving data integration between national accounts (from PWT, GMD, and GTAP) and GDP/population data from external sources; and (5) general data cleaning and structural refinements. We document the data sources and key assumptions used throughout the processing. These updates establish the foundation for the forthcoming KLEAM version of GCAM-macro.

97 MATHEMATICS AND COMPUTING

Polarization observables 𝑇 and 𝐹 in the 𝛾⁢𝑝→𝜋 0 ⁢𝑝 reaction at CLAS

Pion photoproduction in the $\vec{γ}$$\vec{p}$ → π 0 p reaction has been measured in the FROzen Spin Target (FROST) experiment at the Thomas Jefferson National Accelerator Facility. In this experiment, circularly polarized photons with energies up to 3.082 GeV impinged on a transversely polarized frozen-spin target. Final-state protons were detected in the Continuous Electron Beam Accelerator Facility (CEBAF) Large Acceptance Spectrometer. The polarization observables T and F have been extracted for W from 1445 to 2525 MeV, of which the energy range is much broader, and the precision is better than the existing measurements in higher W ranges. The data generally agree with predictions of present partial-wave analyses but also show marked differences for higher W ranges. By incorporating the present data into the databases, the Scattering Analysis Interactive Data (SAID) fits have been improved with relatively small χ 2 and significant changes in the parameters of the (1910)1/2 + and N(2190)7/2 − have been found with the JüBo model.

particle production

Evaluation of New Additions to OLI Software in Predicting Mercuric and Mercurous Species in Liquid Waste Operations

Speciation of mercury during the pretreatment steps of tank waste processing is critical to successful mercury removal prior to vitrification during Liquid Waste Operations (LWO) at SRS. OLI software has been used to predict mercury speciation and activity throughout LWO. The OLI software operates based on a thermodynamic framework called the Mixed Solvent Electrolyte (MSE) framework. The MSE framework allows prediction in theoretically infinitely dilute to concentrated mixtures (e.g., purely solute solutions). Before modification to the MSE framework databanks, certain critical mercury species were missing in the MSE databank, and some thermodynamic data needed to be updated for the OLI software to accurately predict mercury chemical species in SRS waste tanks. To better reflect streams across LWO, new mercury species were integrated into the MSE database. To evaluate the changes to the OLI MSE framework per the Technical Task Request (TTR) and the Task Technical and Quality Assurance Plan (TTQAP), waste stream compositions from Tanks 38, 43, and Tank 50 decontaminated salt solution (DSS) were used as model inputs. Models were developed and executed using both the old and new databases. Compositional analyses from caustic Tank 50 DSS and caustic Tanks 38 and 43 were used as the input streams. These streams represent the most comprehensive chemical data sets where both mercury and tank constituents were measured together. Results for Tank 50 DSS predict HgO as the predominant species in both databases. Both methyl and dimethyl Hg species are present when the new database is ‘on’ and are not predicted with the new database turned ‘off’. The new database predicts a greater amount of HgO and a greater fraction of it in the solid phase. Pourbaix diagrams (potential vs. pH) generated for each Tank 50 DSS were identical regardless of which database was used. Elemental Hg and HgO were predicted in the water stable region under basic conditions. Tanks 38 and 43 follow similar trends as the Tank 50 DSS models. Unlike Tanks 38 and 50 DSS, the Tank 43 Pourbaix plot shows a region of stability for an aqueous HgOHCO3 - species between approximately pH 7-11. In all streams, when MeHg+ is included in the inputs, the new database predicts aqueous MeHgOH as the dominant species. If elemental or dimethyl mercury is in the waste stream, the new database model predicts they are unchanged and remain in those states and quantities. Additionally, the total mercury values are reported for both the measured input data and the OLI output data for all considered tanks. The summary indicates that the percentage error between the measured and calculated values is less than 1% in all cases The reconciliations and generation of the Pourbaix diagrams for Tank 50 DSS took approximately ten times longer with the new database ‘on’. In addition, over the course of that time, models with the new database ‘on’ were more likely to crash or display an error. Some modest performance improvements were noted when modeling with an i7 processor versus an i5. An example error is found in Appendix A. Furthermore, Appendix B provides V&V for two chemical systems analyzed with the OLI software, results were satisfactory. It is recommended to utilize the new databases (i.e., HCO.ddb and SR-Hg.ddb) in future Savannah River Mission Completion applications of OLI to represent pseudo steady-state. Furthermore, the integration and utilization of the new databases (i.e., HCO.ddb and SR-Hg.ddb) in modeling applications (e.g., Aspen) is also recommended.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

NEAR: Neural Embeddings for Amino acid Relationships

Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.

59 BASIC BIOLOGICAL SCIENCES

A Play-Based Exploration of CO 2 Storage in the Illinois Basin (Final Technical Report)

This report documents the results of "A Play-Based Exploration of CO₂ Storage in the Illinois Basin" (DE-FE0032366), a project funded by the U.S. Department of Energy Office of Fossil Energy and Carbon Management and conducted by the Illinois State Geological Survey (ISGS) at the University of Illinois Urbana-Champaign in partnership with Visage Energy. The project adapted play-based exploration (PBE), a systematic basin-scale evaluation methodology from the petroleum industry, to screen areas of Illinois for commercial geologic carbon storage (GCS) in Cambro-Ordovician strata. The traditional play concept was expanded to encompass three play element groups, subsurface geologic factors, surface features, and societal factors, yielding 24 play elements with defensible suitability criteria applied through a five-tier classification scheme. An integrated geospatial database was assembled from ISGS, MGSC, MRCI, NATCARB, and public data sources, supported by significant data-improvement work including correction of legacy well locations, digitization of more than 1,300 well construction records using the DOE CATALOG team's OGRRE tool, compilation of a statewide 2D seismic database, and production of a refined fault and fold geodatabase.

58 GEOSCIENCES

Toward a Generalizable Prediction Model of Molten Salt Mixture Density with Chemistry‐Informed Transfer Learning

Optimally designing applications of molten salts requires knowledge of their thermophysical properties over a wide range of temperatures and compositions. There exist significant gaps in existing databases and this data can be challenging to experimentally measure due to high temperatures, salt corrosivity, and salt hygroscopicity. Existing databases have been used to create Redlich–Kister (RK) models for mixture density showing improved accuracy with respect to ideal mixing assumptions, but these models require subcomponent data measurements for each new system, therefore lacking generality. In order to address generalizability and data sparsity, a transfer learning procedure is proposed to train deep neural networks (DNNs) using a combination of semi‐empirical relationships (RK), data from the thermophysical arm of the molten salt thermal properties database and universal ab initio properties of component mixtures taken from the joint automated repository for various integrated simulations (JARVIS) classical force‐field inspired descriptors database to predict density in molten salts. Herein, it is shown that DNNs predict molten salt density with an r 2 over 0.99 and a mean absolute percentage error under 1%, outperforming alternative methods.

inorganic materials

Utilizing Ontology Structures To Curate the DOE-NETL Carbon Storage Open Database

The specialized ontology for the Carbon Storage Open Database will enable more rapid assignment of appropriate symbology standards for visualization improvements, optimize topical and spatial tagging within keywords, and improve flexibility for utilization in existing data repositories such as EDX. This effort also aims to establish a foundation for utilization of ontologies for organization of other data related to geologic carbon storage in the future.

Martin, Abigail

Methane fluxes in tidal marshes of the conterminous United States

Abstract Methane (CH 4 ) is a potent greenhouse gas (GHG) with atmospheric concentrations that have nearly tripled since pre‐industrial times. Wetlands account for a large share of global CH 4 emissions, yet the magnitude and factors controlling CH 4 fluxes in tidal wetlands remain uncertain. We synthesized CH 4 flux data from 100 chamber and 9 eddy covariance (EC) sites across tidal marshes in the conterminous United States to assess controlling factors and improve predictions of CH 4 emissions. This effort included creating an open‐source database of chamber‐based GHG fluxes ( https://doi.org/10.25573/serc.14227085 ). Annual fluxes across chamber and EC sites averaged 26 ± 53 g CH 4 m −2 year −1 , with a median of 3.9 g CH 4 m −2 year −1 , and only 25% of sites exceeding 18 g CH 4 m −2 year −1 . The highest fluxes were observed at fresh‐oligohaline sites with daily maximum temperature normals (MATmax) above 25.6°C. These were followed by frequently inundated low and mid‐fresh‐oligohaline marshes with MATmax ≤25.6°C, and mesohaline sites with MATmax >19°C. Quantile regressions of paired chamber CH 4 flux and porewater biogeochemistry revealed that the 90th percentile of fluxes fell below 5 ± 3 nmol m −2 s −1 at sulfate concentrations >4.7 ± 0.6 mM, porewater salinity >21 ± 2 psu, or surface water salinity >15 ± 3 psu. Across sites, salinity was the dominant predictor of annual CH 4 fluxes, while within sites, temperature, gross primary productivity (GPP), and tidal height controlled variability at diel and seasonal scales. At the diel scale, GPP preceded temperature in importance for predicting CH 4 flux changes, while the opposite was observed at the seasonal scale. Water levels influenced the timing and pathway of diel CH 4 fluxes, with pulsed releases of stored CH 4 at low to rising tide. This study provides data and methods to improve tidal marsh CH 4 emission estimates, support blue carbon assessments, and refine national and global GHG inventories.

54 ENVIRONMENTAL SCIENCES