Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

The HTAP_v3 emission mosaic: merging regional and global monthly emissions (2000–2018) to support air quality modelling and policies

This study, performed under the umbrella of the Task Force on Hemispheric Transport of Air Pollution (TF-HTAP), responds to the global and regional atmospheric modelling community's need of a mosaic emission inventory of air pollutants that conforms to specific requirements: global coverage, long time series, spatially distributed emissions with high time resolution, and a high sectoral resolution. The mosaic approach of integrating official regional emission inventories based on locally reported data, with a global inventory based on a globally consistent methodology, allows modellers to perform simulations of high scientific quality while also ensuring that the results remain relevant to policymakers. HTAP_v3, an ad hoc global mosaic of anthropogenic inventories, has been developed by integrating official inventories over specific areas (North America, Europe, Asia including Japan and South Korea) with the independent Emissions Database for Global Atmospheric Research (EDGAR) inventory for the remaining world regions. The results are spatially and temporally distributed emissions of SO 2 , NO $x$ , CO, non-methane volatile organic compounds (NMVOCs), NH 3 , PM 10 , PM 2.5 , black carbon (BC), and organic carbon (OC), with a spatial resolution of 0.1º × 0.1º and time intervals of months and years, covering the period 2000–2018. The emissions are further disaggregated into 16 anthropogenic emitting sectors. This paper describes the methodology applied to develop such an emission mosaic, reports on source allocation, differences among existing inventories, and best practices for the mosaic compilation. One of the key strengths of the HTAP_v3 emission mosaic is its temporal coverage, enabling the analysis of emission trends over the past 2 decades. The development of a global emission mosaic over such long time series represents a unique product for global air quality modelling and for better-informed policymaking, reflecting the community effort expended by the TF-HTAP to disentangle the complexity of transboundary transport of air pollution.

54 ENVIRONMENTAL SCIENCES↗

The HTAP_v3.2 emission mosaic: merging regional and global monthly emissions (2000–2020) to support air quality modelling and policies

This study, performed under the umbrella of the Task Force on Hemispheric Transport of Air Pollution (TF-HTAP), responds to the need of the global and regional atmospheric modelling community of having a mosaic emission inventory of air pollutants that conforms to specific requirements: global coverage, long time series, spatially distributed emissions with high time resolution, and a high sectoral resolution. The mosaic approach of integrating official regional emission inventories based on locally reported data, with a global inventory based on a globally consistent methodology, allows modellers to perform simulations of a high scientific quality while also ensuring that the results remain relevant to policymakers. HTAP_v3.2, an ad-hoc global mosaic of anthropogenic inventories, is an update to the HTAP_v3 global mosaic inventory and has been developed by integrating official inventories over specific areas (North America, Europe, Asia including China, Japan and Korea) with the independent Emissions Database for Global Atmospheric Research (EDGAR) inventory for the remaining world regions. The results are spatially and temporally distributed emissions of SO 2 , NO x , CO, NMVOC, NH 3 , PM 10 , PM 2.5 , Black Carbon (BC), and Organic Carbon (OC), with a spatial resolution of 0.1 × 0.1° and time intervals of months and years covering the period 2000–2020 (https://doi.org/10.5281/zenodo.17086684, Crippa, 2025, https://edgar.jrc.ec.europa.eu/dataset_htap_v32, last access: 27 October 2025). The emissions are further disaggregated to 16 anthropogenic emitting sectors. This paper describes the methodology applied to develop such an emission mosaic, reports on source allocation, differences among existing inventories, and best practices for the mosaic compilation. One of the key strengths of the HTAP_v3.2 emission mosaic is its temporal coverage, enabling the analysis of emission trends over the past two decades. The development of a global emission mosaic over such long time series represents a unique product for global air quality modelling and for better-informed policy making, reflecting the community effort expended by the TF-HTAP to disentangle the complexity of transboundary transport of air pollution.

Guizzardi, Diego [European Commission, Ispra (Ital↗

A North Atlantic synthetic tropical cyclone tracks, intensity, and rainfall dataset

Tropical Cyclones (TCs) cause significant socio-economic damages to the US and Caribbean coastal regions annually, making it important to understand TC risk at the local-to-regional scales where their impacts are most prominent. However, the short length of the observed record and the substantial computational expense associated with high-resolution climate models make it difficult to assess TC risk using either approach. To overcome these challenges, we developed a database of synthetic TCs using the Risk Analysis Framework for Tropical Cyclones (RAFT). The database includes 50,000 synthetic TC tracks, along-track intensities and storm-induced precipitation. TC tracks generated in RAFT are in reasonable agreement with observations for spatial distribution of TC tracks and basin-scale distributions of TC translation speeds, lifetime maximum intensities and intensification rates. Also, spatial variations in coastal frequency and precipitation for landfalling TCs are well-reproduced in RAFT. In summary, the synthetic TC database based on RAFT provides a reasonable pathway for robust assessment of TC wind and rainfall risk for the US coastal regions and other areas affected by Atlantic TCs.

Xu, Wenwei↗

The Natural Products Magnetic Resonance Database (NP-MRD) for 2025

The Natural Products Magnetic Resonance Database or NP-MRD (https://np-mrd.org) is a comprehensive, freely accessible, web-based resource for the deposition, distribution, extraction and retrieval of nuclear magnetic resonance (NMR) data on natural products. The NP-MRD was initially established to support compound de-replication and data dissemination for the natural products community. However, that community has now grown to include many users from the metabolomics, microbiomics, foodomics and nutrition science fields. Indeed, since its launch in 2021, the NP-MRD has expanded enormously in size, scope and popularity. The current version of NP-MRD now contains nearly 7X more compounds (281,859 vs. 40,908) and 7X more NMR spectra (5.1 million vs. 817,000) than the first release. More specifically, an additional 4.6 million predicted spectra and another 11,000 spectra simulated from experimental chemical shifts were deposited into the database. Likewise, the number of NMR raw spectral data depositions has grown from a 165 spectra per year to more than 10,000 per year. As a result of this expansion, the number of monthly webpage views has grown from 55 to 20,000 and the number of monthly visitors has increased from 7 to 2500. To address this growth and to better support the expanding needs of its diverse community of users, many additional improvements to the NP-MRD have been made. These include significant enhancements to the data submission process, important improvements to the visualization and display of NMR spectra, notable updates to the database’s spectral search utilities and useful additions to support better NMR spectral analysis/prediction. Significant efforts have also been undertaken to remediate and update many of NP-MRD’s database entries. This manuscript describes these database improvements and expansion efforts, along with how they have been implemented and what future upgrades to the NP-MRD are planned.

Artifical Intelligence↗

FY 2025 End of Year Report: Seismic Monitoring of Underground Vibration Sources using Distributed Acoustic Sensing (DAS) and Seismometers

This end-of-year report summarizes progress on using seismic monitoring to detect, associate, and locate anomalous vibration signals that may indicate potential containment breaches. The work focused on four key tasks: 1. Developing a database of continuous waveforms and ground-truth event data from multiple sensing modalities. 2. Refining and implementing detection and association algorithms to generate a catalog of anomalous underground activities. 3. Testing and improving distributed acoustic sensing amplitude-based geolocation methods to build an event location catalog. 4. Testing and refining seismic array polarization-based geolocation methods to build an event location catalog. This report provides a brief recap of results from the FY25 midyear report (Tasks 1 and 2) and presents new findings from geolocation methods (Tasks 3 and 4).

58 GEOSCIENCES↗

Unravelling the diversity of magnetotactic bacteria through analysis of open genomic databases

Magnetotactic bacteria (MTB) are prokaryotes that possess genes for the synthesis of membrane-bounded crystals of magnetite or greigite, called magnetosomes. Despite over half a century of studying MTB, only about 60 genomes have been sequenced. Most belong to Proteobacteria, with a minority affiliated with the Nitrospirae, Omnitrophica, Planctomycetes, and Latescibacteria. Due to the scanty information available regarding MTB phylogenetic diversity, little is known about their ecology, evolution and about the magnetosome biomineralization process. This study presents a large-scale search of magnetosome biomineralization genes and reveals 38 new MTB genomes. Several of these genomes were detected in the phyla Elusimicrobia, Candidatus Hydrogenedentes, and Nitrospinae, where magnetotactic representatives have not previously been reported. Analysis of the obtained putative magnetosome biomineralization genes revealed a monophyletic origin capable of putative greigite magnetosome synthesis. The ecological distributions of the reconstructed MTB genomes were also analyzed and several patterns were identified. These data suggest that open databases are an excellent source for obtaining new information of interest.

59 BASIC BIOLOGICAL SCIENCES↗

Synthesizing realistic sand assemblies with denoising diffusion in latent space

Abstract The shapes and morphological features of grains in sand assemblies have far‐reaching implications in many engineering applications, such as geotechnical engineering, computer animations, petroleum engineering, and concentrated solar power. Yet, our understanding of the influence of grain geometries on macroscopic response is often only qualitative, due to the limited availability of high‐quality 3D grain geometry data. In this paper, we introduce a denoising diffusion algorithm that uses a set of point clouds collected from the surface of individual sand grains to generate grains in the latent space. By employing a point cloud autoencoder, the three‐dimensional point cloud structures of sand grains are first encoded into a lower‐dimensional latent space. A generative denoising diffusion probabilistic model is trained to produce synthetic sand that maximizes the log‐likelihood of the generated samples belonging to the original data distribution measured by a Kullback‐Leibler divergence. Numerical experiments suggest that the proposed method is capable of generating realistic grains with morphology, shapes and sizes consistent with the training data inferred from an F50 sand database. We then use a rigid contact dynamic simulator to pour the synthetic sand in a confined volume to form granular assemblies in a static equilibrium state with targeted distribution properties. To ensure third‐party validation, 50,000 synthetic sand grains and the 1542 real synchrotron microcomputed tomography (SMT) scans of the F50 sand, as well as the granular assemblies composed of synthetic sand grains are made available in an open‐source repository.

Vlassis, Nikolaos N.↗

Private Tabular Survey Data Products through Synthetic Microdata Generation

We propose two synthetic microdata approaches to generate private tabular survey data products for public release. We adapt a pseudo posterior mechanism that downweights by-record likelihood contributions with weights ∈[0,1] based on their identification disclosure risks to producing tabular products for survey data. Our method applied to an observed survey database achieves an asymptotic global probabilistic differential privacy guarantee. Our two approaches synthesize the observed sample distribution of the outcome and survey weights, jointly, such that both quantities together possess a privacy guarantee. The privacy-protected outcome and survey weights are used to construct tabular cell estimates (where the cell inclusion indicators are treated as known and public) and associated standard errors to correct for survey sampling bias. Through a real data application to the Survey of Doctorate Recipients public use file and simulation studies motivated by the application, we demonstrate that our two microdata synthesis approaches to construct tabular products provide superior utility preservation as compared to the additive noise approach of the Laplace Mechanism. Moreover, our approaches allow the release of microdata to the public, enabling additional analyses at no extra privacy cost.

Mathematical Methods In Social Sciences↗

cuTS: Scaling Subgraph Isomorphism on Distributed Multi-GPUSystems Using Trie Based Data Structure

Subgraph isomorphism is a pattern-matching algorithm widely used in many domains such as chem-informatics, bioinformatics, databases, and social network analysis. It is computationally expensive and is a proven NP-hard problem. The massive parallelism offered by the GPU hardware is well suited for solving the subgraph isomorphism. However, current GPU implementations are far from the achievable performance. Moreover, the enormous memory requirement of current approaches limits the problem size that can be handled. This work analyzes the fundamental challenges associated with processing the subgraph isomorphism on GPUs and develops an efficient GPU hardware-aware implementation. We also develop a new GPU-friendly trie-based data structure to drastically reduce the intermediate storage space requirement. Hence, our approach runs larger benchmarks than the competitors. We also develop the first distributed sub-graph isomorphism algorithm for GPUs. Our experimental evaluation section demonstrates the efficacy of our approach by comparing the execution time and number of cases that we can handle against the state-of-the-art GPU implementations.

Xiang, Lizhi↗

Distributed Generation Market Demand (dGen) model

The Distributed Generation Market Demand (dGen) model simulates customer adoption of distributed energy resources (DERs) for residential, commercial, and industrial entities in the United States or other countries through 2050. The dGen model can be used for identifying the sectors, locations, and customers for whom adopting DERs would have a high economic value, for generating forecasts as an input to estimate distribution hosting capacity analysis, integrated resource planning, and load forecasting, and for understanding the economic or policy conditions in which DER adoption becomes viable, and for illustrating sensitivity to market and policy changes such as retail electricity rate structures, net energy metering, and technology costs.

Array↗

A hydrogeophysical framework to assess infiltration during a simulated ecosystem-scale flooding experiment

This study presents a framework to quantify changes in soil saturation in response to flooding caused by extreme hydrologic perturbation on coastal ecosystems at the interfaces and transition between terrestrial and aquatic systems. Subsurface heterogeneity limits the use of in situ measurements to quantify subsurface flow during flooding due to the spatial discontinuity in the measured data. While geophysical methods, including time-lapse electrical resistivity imaging (ERI), are increasingly used to monitor soil hydrological processes, their abilities to parameterize flow models have been underutilized. This study combines background ERI, ground penetrating radar (GPR), time-lapse ERI, soil characterization, and a numerical flow model developed using an Advanced Terrestrial Simulator (ATS) code to quantify the infiltration pathway and describe the hydrological dynamics during a simulated flooding experiment. We assessed the use of two conceptual models developed using [1] ERI and GPR data that described the stratigraphic distribution, and time-lapse ERI that mapped permeability contrast, and [2] information from a national soil database for capturing changes in saturation. Combining the ERI and GPR results with soil core data revealed the stratigraphic heterogeneity at the site with a silty clay layer from 1 to 2 m between an overlying loamy topsoil and an underlying saturated silty sand. This silty clay layer could restrict deep infiltration. The time-lapse ERI showed up to a 35% decrease in resistivity, which correlated with soil moisture data (R 2 value > 0.53) and revealed preferential infiltration zones used to inform the flow model. Numerical simulation results from both the geophysics- and soil database-informed models quantified changes in soil saturation with calculated soil moistures that agreed with field data. The geophysics-informed model captured more of the system’s variability, reflective of shallow subsurface heterogeneities. The framework presented will serve as a precursor for a robust ecohydrological model that can describe the impacts of extreme events induced by climate change on coastal ecosystems.

54 ENVIRONMENTAL SCIENCES↗

Macroscopic trends of linear tearing stability in cylindrical current profiles

Abstract The likelihood of realising tokamak power-plants will be greatly improved by the discovery of high-gain equilibria that resist the formation of small islands and hence avoid the disruptive neoclassical tearing mode. We propose a series of studies to understand how simple tokamak design can leverage aspects of tearing onset physics to maximise passive resistance to island formation. Here we investigate the variation that current profiles can bring about in preventing tearing onset through the cylindrical linear tearing stability parameter Δ ′ . A database of 159148 realistic pilot-plant current profiles was generated with Monte Carlo sampling, and the distribution of Δ ′ values was linked with interpretable profile characteristics. In agreement with prior theoretical and experimental studies, Δ ′ was found to be strongly correlated with the existence and steepness of a local toroidal current well or hill, with the former destabilising and the latter stabilising. In the absence of these two cases, the remaining Δ ′ values were linearly bounded by the toroidal current gradient at the rational surface.

Physics↗

mkite

mkite is a distributed computing platform for materials simulation. mkite is built with the server-client pattern, decoupling production databases from client runners. When used in combination with message brokers, mkite enables any available client to perform calculations without prior hardware specification on the server side. Furthermore, the software enables the creation of complex workflows with multiple inputs and branches, facilitating the exploration of combinatorial chemical spaces. The mkite suite provides recipes and tools to interact with package such as VASP, but is extensible to any other simulation package. Finally, mkite helps keeping the provenance of calculations in a SQL database. A complete description of the software is available at https://arxiv.org/abs/2301.08841.

Schwalbe Koda, Daniel↗

KOMPASS-II: Compaction of Crushed salt for Safe Containment – Phase 2

Long-term stable sealing elements are a basic component in the safety concept for a possible repository for heat-emitting radioactive waste in rock salt. The sealing elements will be part of the closure concept for drifts and shafts. They will be made from a welldefinied crushed salt in employ a specific manufacturing process. The use of crushed salt as geotechnical barrier as required by the German Site Selection Act from 2017 /STA 17/ represents a paradigm change in the safety function of crushed salt, since this material was formerly only considered as stabilizing backfill for the host rock. The demonstration of the long-term stability and impermeability of crushed salt is crucial for its use as a geotechnical barrier. The KOMPASS-II project, is a follow-up of the KOMPASS-I project and continues the work with focus on improving the understanding of the thermal-hydraulic-mechanical (THM) coupled processes in crushed salt compaction with the objective to enhance the scientific competence for using crushed salt for the long-term isolation of high-level nuclear waste within rock salt repositories. The project strives for an adequate characterization of the compaction process and the essential influencing parameters, as well as a robust and reliable long-term prognosis using validated constitutive models. For this purpose, experimental studies on long-term compaction tests are combined with microstructural investigations and numerical modeling. The long-term compaction tests in this project focused on the effect of mean stress, deviatoric stress and temperature on the compaction behavior of crushed salt. A laboratory benchmark was performed identifying a variability in compaction behavior. Microstructural investigations were executed with the objective to characterize the influence of pre-compaction procedure, humidity content and grain size/grain size distribution on the overall compaction process of crushed salt with respect to the deformation mechanisms. The created database was used for benchmark calculations aiming for improvement and optimization of a large number of constitutive models available for crushed salt. The models were calibrated, and the improvement process was made visible applying the virtual demonstrator.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

ELM2.1-XGBfire1.0: improving wildfire prediction by integrating a machine learning fire model in a land surface model

Wildfires have shown increasing trends in both frequency and severity across the contiguous United States (CONUS). However, process-based fire models have difficulties in accurately simulating the burned area over the CONUS due to a simplification of the physical process and cannot capture the interplay among fire, ignition, climate, and human activities. The deficiency of burned area simulation deteriorates the description of fire impact on energy balance, water budget, and carbon fluxes in the Earth system models (ESMs). Alternatively, fire models based on machine learning (ML), which capture statistical relationships between the burned area and environmental factors, have shown promising burned area predictions and corresponding fire impact simulation. We develop a hybrid framework (ELM2.1-XGBFire1.0) that integrates an eXtreme Gradient Boosting (XGBoost) wildfire model with the Energy Exascale Earth System Model (E3SM) land model (ELM) version 2.1. A Fortran–C–Python deep learning bridge is adapted to support online communication between ELM and the ML fire model. Specifically, the burned area predicted by the ML-based wildfire model is directly passed to ELM to adjust the carbon pool and vegetation dynamics after disturbance, which are then used as predictors in the ML-based fire model in the next time step. Evaluated against the historical burned area from Global Fire Emissions Database 5 from 2001–2019, the ELM2.1-XGBFire1.0 outperforms process-based fire models in terms of spatial distribution and seasonal variations. The ELM2.1-XGBFire1.0 has proven to be a new tool for studying vegetation–fire interactions and, more importantly, enables seamless exploration of climate–fire feedback, working as an active component of E3SM.

54 ENVIRONMENTAL SCIENCES↗

FIRED (Fire Events Delineation): An Open, Flexible Algorithm and Database of US Fire Events Derived from the MODIS Burned Area Product (2001–2019)

Harnessing the fire data revolution, i.e., the abundance of information from satellites, government records, social media, and human health sources, now requires complex and challenging data integration approaches. Defining fire events is key to that effort. In order to understand the spatial and temporal characteristics of fire, or the classic fire regime concept, we need to critically define fire events from remote sensing data. Events, fundamentally a geographic concept with delineated spatial and temporal boundaries around a specific phenomenon that is homogenous in some property, are key to understanding fire regimes and more importantly how they are changing. Here, we describe Fire Events Delineation (FIRED), an event-delineation algorithm, that has been used to derive fire events (N = 51,871) from the MODIS MCD64 burned area product for the coterminous US (CONUS) from January 2001 to May 2019. The optimized spatial and temporal parameters to cluster burned area pixels into events were an 11-day window and a 5-pixel (2315 m) distance, when optimized against 13,741 wildfire perimeters in the CONUS from the Monitoring Trends in Burn Severity record. The linear relationship between the size of individual FIRED and Monitoring Trends in Burn Severity (MTBS) events for the CONUS was strong (R2 = 0.92 for all events). Importantly, this algorithm is open-source and flexible, allowing the end user to modify the spatio-temporal threshold or even the underlying algorithm approach as they see fit. We expect the optimized criteria to vary across regions, based on regional distributions of fire event size and rate of spread. We describe the derived metrics provided in a new national database and how they can be used to better understand US fire regimes. The open, flexible FIRED algorithm could be utilized to derive events in any satellite product. We hope that this open science effort will help catalyze a community-driven, data-integration effort (termed OneFire) to build a more complete picture of fire.

54 ENVIRONMENTAL SCIENCES↗

Profiling the BLAST bioinformatics application for load balancing on high-performance computing clusters

Abstract Background The Basic Local Alignment Search Tool (BLAST) is a suite of commonly used algorithms for identifying matches between biological sequences. The user supplies a database file and query file of sequences for BLAST to find identical sequences between the two. The typical millions of database and query sequences make BLAST computationally challenging but also well suited for parallelization on high-performance computing clusters. The efficacy of parallelization depends on the data partitioning, where the optimal data partitioning relies on an accurate performance model. In previous studies, a BLAST job was sped up by 27 times by partitioning the database and query among thousands of processor nodes. However, the optimality of the partitioning method was not studied. Unlike BLAST performance models proposed in the literature that usually have problem size and hardware configuration as the only variables, the execution time of a BLAST job is a function of database size, query size, and hardware capability. In this work, the nucleotide BLAST application BLASTN was profiled using three methods: shell-level profiling with the Unix “time” command, code-level profiling with the built-in “profiler” module, and system-level profiling with the Unix “gprof” program. The runtimes were measured for six node types, using six different database files and 15 query files, on a heterogeneous HPC cluster with 500+ nodes. The empirical measurement data were fitted with quadratic functions to develop performance models that were used to guide the data parallelization for BLASTN jobs. Results Profiling results showed that BLASTN contains more than 34,500 different functions, but a single function, RunMTBySplitDB, takes 99.12% of the total runtime. Among its 53 child functions, five core functions were identified to make up 92.12% of the overall BLASTN runtime. Based on the performance models, static load balancing algorithms can be applied to the BLASTN input data to minimize the runtime of the longest job on an HPC cluster. Four test cases being run on homogeneous and heterogeneous clusters were tested. Experiment results showed that the runtime can be reduced by 81% on a homogeneous cluster and by 20% on a heterogeneous cluster by re-distributing the workload. Discussion Optimal data partitioning can improve BLASTN’s overall runtime 5.4-fold in comparison with dividing the database and query into the same number of fragments. The proposed methodology can be used in the other applications in the BLAST+ suite or any other application as long as source code is available.

59 BASIC BIOLOGICAL SCIENCES↗