Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

National Community Solar Partnership (NCSP+) Resource List

NLR compiles and makes quarterly updates to this dataset of publications related to expanding access to solar through its support for the National Community Solar Partnership (NCSP+). The purpose of the list is to help NCSP+ partners and others to find relevant resources on community solar, low- and moderate-income residential rooftop solar + storage, community-serving commercial solar projects, microgrids, and distributed solar + storage aggregations such as virtual power plants serving low-income communities. Resource types include journal articles, reports, fact sheets, presentations, videos, datasets, modeling tools, webinars, and other formats. A web-based version of the database is available on the NCSP+ Data Hub at: https://openei.org/wiki/NCSP_Hub/Resources .

14 SOLAR ENERGY↗

Automatic and rapid calibration of urban building energy models by learning from energy performance database

Urban building energy modeling (UBEM) is attracting increasing attention in the energy modeling filed. Unlike modeling a single building using detailed building systems information, UBEM generally uses limited high-level building stock data to infer default assumptions about building characteristics and operations. Additionally, this practice inherently brings uncertainty to UBEM. This study introduced a novel method of automatic and rapid calibration of UBEM based on the annual electricity and natural gas energy use data by learning the correlations between crucial model input parameters and the building energy use from the reference building models. A case study was presented to calibrate 72 large office buildings built before 1978 in San Francisco. Seventeen model parameters were selected and Monte Carlo sampling was used to create 1000 samples that reasonably represent the parameter space. Then 1000 simulations were performed for the reference building model to create an energy performance database. The results showed that by learning from the energy performance database, it took less than four simulation runs on average to calibrate a building model. After the calibration, the distributions of each parameter were obtained to replace their single predefined default values. For example, the default lighting power density of 21.39 W/m 2 was calibrated to be 7.50 W/m 2 on average. The case study successfully demonstrated the effectiveness of the novel calibration method for UBEM in the mild climate. The method will be further tested in future for other climate zones and other building types.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Simulation of PV Variability as a Function of PV Generation and Plant Size

The deployment of photovoltaic (PV) systems continues to show significant expansion; however, this growth has brought added attention to issues around the variability of the solar resource. Both spatial and temporal variability exist. Temporal scales can range from the sub-second to multiyear, whereas spatial scales can range from a few meters to tens of kilometers. There are multiple methods described in the literature to quantify PV variability at various spatial and temporal scales. This study focuses on short-term temporal variability and uses similar approaches with the addition of PV plant size a parameter to quantify variability. The method employed here incorporates the normalization of clear-and cloudy-sky conditions and PV plant size to quantify nominal variability metrics. The distribution and fluctuations of these metrics provide relevant information that is useful for system operations. The National Solar Radiation Database (NSRDB) is used to simulate PV variability as a function of PV generation and plant size. Hypothetical but realistic system information at 33 locations is used to model PV generation by feeding NSRDB solar irradiance data to the National Renewable Energy Laboratory’s System Advisor Model (SAM). Over the selected region, it is found that the aggregated ramp rates for the 1-minute data are associated with standard deviations ranging from 0.002–0.055 on a daily basis; however, hourly intervals induce higher aggregated ramp rates than the other timescales. Even though minute-to-minute variations are significant for the 1-minute time-scale, the standard deviation aggregated into a daily metric is smaller because of the cancellation of values.

irradiance↗

Convergence in simulating global soil organic carbon by structurally different models after data assimilation

Abstract Current biogeochemical models produce carbon–climate feedback projections with large uncertainties, often attributed to their structural differences when simulating soil organic carbon (SOC) dynamics worldwide. However, choices of model parameter values that quantify the strength and represent properties of different soil carbon cycle processes could also contribute to model simulation uncertainties. Here, we demonstrate the critical role of using common observational data in reducing model uncertainty in estimates of global SOC storage. Two structurally different models featuring distinctive carbon pools, decomposition kinetics, and carbon transfer pathways simulate opposite global SOC distributions with their customary parameter values yet converge to similar results after being informed by the same global SOC database using a data assimilation approach. The converged spatial SOC simulations result from similar simulations in key model components such as carbon transfer efficiency, baseline decomposition rate, and environmental effects on carbon fluxes by these two models after data assimilation. Moreover, data assimilation results suggest equally effective simulations of SOC using models following either first‐order or Michaelis–Menten kinetics at the global scale. Nevertheless, a wider range of data with high‐quality control and assurance are needed to further constrain SOC dynamics simulations and reduce unconstrained parameters. New sets of data, such as microbial genomics‐function relationships, may also suggest novel structures to account for in future model development. Overall, our results highlight the importance of observational data in informing model development and constraining model predictions.

54 ENVIRONMENTAL SCIENCES↗

Simulation of PV Variability as a Function of PV Generation and Plant Size: Preprint

The deployment of photovoltaic (PV) systems continues to show significant expansion; however, this growth has brought added attention to issues around the variability of the solar resource. Both spatial and temporal variability exist. Temporal scales can range from the sub-second to multiyear, whereas spatial scales can range from a few meters to tens of kilometers. There are multiple methods described in the literature to quantify PV variability at various spatial and temporal scales. This study focuses on short-term temporal variability and uses similar approaches with the addition of PV plant size a parameter to quantify variability. The method employed here incorporates the normalization of clear- and cloudy-sky conditions and PV plant size to quantify nominal variability metrics. The distribution and fluctuations of these metrics provide relevant information that is useful for system operations. The National Solar Radiation Database (NSRDB) is used to simulate PV variability as a function of PV generation and plant size. Hypothetical but realistic system information at 33 locations is used to model PV generation by feeding NSRDB solar irradiance data to the National Renewable Energy Laboratory’s System Advisor Model (SAM). Over the selected region, it is found that the aggregated ramp rates for the 1-minute data are associated with standard deviations ranging from 0.002–0.055 on a daily basis; however, hourly intervals induce higher aggregated ramp rates than the other timescales. Even though minute-to-minute variations are significant for the 1-minute timescale, the standard deviation aggregated into a daily metric is smaller because of the cancellation of values.

41 EE - Solar Energy Technologies Office (EE-4S)↗

CASTLE: Conflict Analysis Strategy Testing Laboratory Environment v.1.0.0

SAND2024-01743O The Conflict Analysis Strategy Testing Laboratory Environment (CASTLE) is a software framework that enables and simplifies building a novel, turn-based strategy game in which it can define its own rules, maps, pieces, and interactions. The software is for novice to experienced programmers with some knowledge of Unity3D, a tool used in game production. CASTLE includes a library of common game mechanics used for strategic wargames and traditional board games, such as cards, tokens, dice, and grid maps. It follows design principles popularized by the video game industry and uses singletons for managing portions of the code. CASTLE builds on Unity's component-based design and can respond to engine events during execution. Among the numerous user-friendly features: Build games quickly and cost-effectively Network in real-time and apply data to new games developed on the framework Host multiple participants online Connect rule- or machine learning-based agents to a CASTLE game to serve as opponents or to simulate games Collect data collection from players and in-game behaviors Create a survey to gather demographics or opinions from players Store data locally or save it to an external database through Representational State Transfer (REST) functions CASTLE, which was prototyped using Microsoft Azure, is also designed for easily distributing online games using popular cloud services. The multiplayer functionality includes an agent interface, allowing developers to construct AI players that can substitute for humans in any of the games. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Fabian, Nathan↗

Discrepancy quantification between experimental and simulated data of CO 2 adsorption isotherm using hierarchical Bayesian estimation

Here, to quantitatively analyze the inconsistencies commonly observed between experimental and simulated adsorption isotherms, parameter estimation of adsorption isotherm models was conducted by hierarchical Bayesian estimation with parameter uncertainties being quantified as probability distributions. The estimation method was implemented using Markov Chain Monte Carlo (MCMC) to analyze multiple data sets obtained from different sources, including a publicly available database. To describe the discrepancies of experimental and simulated adsorption data, the simulation data was set as the reference to which experimental measurements were compared. We applied the proposed approach to analyze CO 2 adsorption isotherms that are measured and simulated on zeolite 13X and MIL-101(Cr). In these case studies, the discrepancy of CO 2 adsorption isotherm was successfully quantified between experimental measurements and predictions given by molecular simulations using Grand Canonical Monte Carlo (GCMC), where uncertainties were quantified as probability distributions. Furthermore, experimental data sets that agree well with the GCMC simulation have been identified, providing insights into experimental and measurement methods as well as choosing the right assumptions in the molecular simulation.

42 ENGINEERING↗

G2Aero Database of Airfoils - Curated Airfoils

This dataset contains a curated set of 19,164 airfoil shapes from various applications and the data-driven design space of separable shape tensors (PGA space), which can be used as a parameter space for machine-learning applications focused on airfoil shapes. We constructed the airfoil dataset in two main stages. First, we identified 13 baseline airfoils from the NREL 5MW and IEA 15MW reference wind turbines. We reparameterized these shapes using least-squares fits of 8-order CST parametrizations, which involve 18 coefficients. By uniformly perturbing all 18 CST coefficients by +/-20% around each baseline airfoil, we generated 1,000 unique airfoils. Each airfoil was sampled with 1,001 shape landmarks whose x-coordinates followed a cosine distribution along the chord. This process resulted in a total of 13,000 airfoil shapes, each with 1,001 landmarks. In the second phase, we gathered additional airfoils from the extensive BigFoil database, which consolidates data from sources such as the University of Illinois Urbana-Champaign (UIUC) airfoil database, the JavaFoil database, the NACA-TR-824 database, and others. We undertook a thorough pre-processing step to filter out shapes with sparse, noisy, or incomplete data. We also removed airfoils with sharp leading edge and those exceeding our threshold for trailing edge thickness. Additionally, we thinned out the collection of NACA airfoils-- parametric sweeps of NACA airfoils with increasing thickness and camber present in BigFoil database-- by selecting every fourth step in the parameter sweeps. Finally, we regularized the airfoils by reparametrizing them with an 8-order CST parametrization (with 1,001 shape landmarks with x coordinated following cosine distribution along the chord) and removing airfoils with high reconstruction errors. This data pre-processing resulted in a set of 6,164 airfoils. In total, our curated airfoil dataset comprises 19,164 airfoils, each with 1,001 landmarks, and is stored in the curated_airfoils.npz file. Using this curated airfoil dataset, we utilized the separable shape tensors framework to develop a data-driven parameterization of airfoils based on principal geodesic analysis (PGA) of separable shape tensors. This PGA space is provided in PGAspace.npz file.

airfoils↗

Seismic monitoring of underground vibration: database of seismic data and ground truth

Seismic waves provide valuable insights into underground activities, serving as an essential tool for monitoring anomalies that could signal containment breaches in geological repositories. We aim to test and refine underground detection and geolocation techniques to identify anomalous vibration signals indicative of potential breaches, thereby strengthening georepository safeguards. We evaluate the effectiveness of two distinct, low-maintenance sensing technologies (surface geophones and underground distributed acoustic sensing (DAS) fiber optic cable) leveraging existing datasets. This report details the experimental designs, instrumentation, and data characteristics for both seismic and DAS arrays. Additionally, we provide a ground truth database documenting relevant operational activities for each experiment.

58 GEOSCIENCES↗

Correlated $\ n-γ$ angular distributions from the $\ Q$ = 4.4398 MeV 12 C ($\ n, n' γ$) reaction for incident neutron energies from 6.5 MeV to 16.5 MeV

Neutron scattering cross sections and angular distributions represent one of the most glaring sources of uncertainty in calculations of nuclear systems. Even simple nuclei like 12 C show indications of errors in nuclear databases for scattering reactions. Measurements of inelastic neutron scattering have historically measured either the scattered neutrons or the nuclear deexcitation $\ γ$ emission. Only a very small number of experiments attempted correlated measurements of both the neutron and $\ γ$ data simultaneously, even though these $\ n-γ$ correlations could be essential for understanding particle transport in nuclear systems. In this work we describe a measurement of the $\ n, γ$, and correlated $\ n-γ$ angular distributions from the $\ Q$ = 4.4398 MeV 12 C ($\ n, n'γ$) reaction in a single experiment using an EJ-309 liquid scintillator detector array with wide angular coverage, and with a continuous incident neutron energy range from 6.5 to 16.5 MeV. We also provide a thorough covariance description of these results, including normalization of the probability distribution. While the measured n distributions agree well with the relatively large number of available literature measurements, there are comparatively very few measurements of the γ distributions from this reaction. However, our data support the presence of a nonzero α 4 Legendre polynomial component of the γ angular distribution suggested in past measurements, which is currently not incorporated in the ENDF/B-VIII.0 library despite the use of these same literature data for evaluation of the 12 C ($\ n, n'γ$) cross section. The correlated $\ n-γ$ distribution measurements are limited to three measurements at incident neutron energies near 14 MeV. Our results do not generally agree with any of these literature measurements. We observe clear indications of significant changes in the $\ n$ distribution for specific $\ γ$-detection angles and vice versa especially near thresholds for other reaction channels, which shows the potential for significant bias in experiments that, for example, tag on inelastic scattering using a single or small number of $\ γ$ -detection angles and could impact particle transport calculations.

6 ≤ A ≤ 19↗

COVID19 Disease Map, a computational knowledge repository of virus–host interaction mechanisms

We need to effectively combine the knowledge from surging literature with complex datasets to propose mechanistic models of SARS-CoV-2 infection, improving data interpretation and predicting key targets of intervention. Here, we describe a large-scale community effort to build an open access, interoperable and computable repository of COVID-19 molecular mechanisms. The COVID-19 Disease Map (C19DMap) is a graphical, interactive representation of disease-relevant molecular mechanisms linking many knowledge sources. Notably, it is a computational resource for graph-based analyses and disease modelling. To this end, we established a framework of tools, platforms and guidelines necessary for a multifaceted community of biocurators, domain experts, bioinformaticians and computational biologists. The diagrams of the C19DMap, curated from the literature, are integrated with relevant interaction and text mining databases. We demonstrate the application of network analysis and modelling approaches by concrete examples to highlight new testable hypotheses. This framework helps to find signatures of SARS-CoV-2 predisposition, treatment response or prioritisation of drug candidates. Such an approach may help deal with new waves of COVID-19 or similar pandemics in the long-term perspective.

59 BASIC BIOLOGICAL SCIENCES↗

Corrosion of rebar in concrete. Part II: Literature survey and statistical analysis of existing data on chloride threshold

The literature is reviewed with respect to defining, measuring, and interpreting the chloride threshold (CT). CT was defined with a strong fundamental basis in the Point Defect Model. which also dictates the nature of the statistical analysis of surveyed CT values. A database of CT for various primary and secondary independent variables has been established. Statistical analyses reveal that CT is lognormally distributed, in accordance with the PDM, with most probable values being 0.85, 0.69, and 1.19 for %total Cl/cem, %free Cl/cem, and [Cl - ]/[OH - ], respectively. The breakdown and corrosion potentials follow normal distributions with means of 0.056 and -0.322 V SCE , respectively.

36 MATERIALS SCIENCE↗

Improving North American Wildfire Prediction by Integrating a Machine-Learning Fire Model in a Land Surface Model

Wildfires have shown increasing trends in both frequency and severity across the Contiguous United States (CONUS). However, process-based fire models have difficulties in accurately simulating the burned area over the CONUS due to a simplification of the physical process and cannot capture the interplay among fire, ignition, climate, and human activities. The deficiency of burned area simulation deteriorates the description of fire impact on energy balance, water budget, and carbon fluxes in the Earth System Models (ESMs). Alternatively, machine learning (ML) based fire models, which capture statistical relationships between the burned area and environmental factors, have shown promising burned area predictions and corresponding fire impact simulation. We develop a hybrid framework (ML4Fire-XGB) that integrates a pretrained eXtreme Gradient Boosting (XGBoost) wildfire model with the Energy Exascale Earth System Model (E3SM) land model (ELM) version 2.1. A Fortran-C-Python deep learning bridge is adapted to support online communication between ELM and the ML fire model. Specifically, the burned area predicted by the ML-based wildfire model is directly passed to ELM to adjust the carbon pool and vegetation dynamics after disturbance, which are then used as predictors in the ML-based fire model in the next time step. Evaluated against the historical burned area from Global Fire Emissions Database 5 from 2001-2020, the ML4Fire-XGB model outperforms process-based fire models in terms of spatial distribution and seasonal variations. Sensitivity analysis confirms that the ML4Fire-XGB well captures the responses of the burned area to rising temperatures. The ML4Fire-XGB model has proved to be a new tool for studying vegetation-fire interactions, and more importantly, enables seamless exploration of climate-fire feedback, working as an active component in E3SM.

54 ENVIRONMENTAL SCIENCES↗

Simulated wildfire burned area over the CONUS during 2001-2020

Wildfires have shown increasing trends in both frequency and severity across the Contiguous United States (CONUS). However, process-based fire models have difficulties in accurately simulating the burned area over the CONUS due to a simplification of the physical process and cannot capture the interplay among fire, ignition, climate, and human activities. The deficiency of burned area simulation deteriorates the description of fire impact on energy balance, water budget, and carbon fluxes in the Earth System Models (ESMs). Alternatively, machine learning (ML) based fire models, which capture statistical relationships between the burned area and environmental factors, have shown promising burned area predictions and corresponding fire impact simulation. We develop a hybrid framework (ML4Fire-XGB) that integrates a pretrained eXtreme Gradient Boosting (XGBoost) wildfire model with the Energy Exascale Earth System Model (E3SM) land model (ELM). A Fortran-C-Python deep learning bridge is adapted to support online communication between ELM and the ML fire model. Specifically, the burned area predicted by the ML-based wildfire model is directly passed to ELM to adjust the carbon pool and vegetation dynamics after disturbance, which are then used as predictors in the ML-based fire model in the next time step. Evaluated against the historical burned area from Global Fire Emissions Database 5 from 2001-2020, the ML4Fire-XGB model outperforms process-based fire models in terms of spatial distribution and seasonal variations. Sensitivity analysis confirms that the ML4Fire-XGB well captures the responses of the burned area to rising temperatures. The ML4Fire-XGB model has proved to be a new tool for studying vegetation-fire interactions, and more importantly, enables seamless exploration of climate-fire feedback, working as an active component in E3SM.

Liu, Ye↗

Computational Fluid Dynamics (CFD) Simulations of Taylor Bubbles in Vertical and Inclined Pipes with Upward and Downward Liquid Flow

Summary Two-phase flow is a common occurrence in pipes of oil and gas developments. Current predictive tools are based on the mechanistic two-fluid model, which requires the use of closure relations to predict integral flow parameters such as liquid holdup (or void fraction) and pressure gradient. However, these closure relations carry the highest uncertainties in the model. In particular, significant discrepancies have been found between experimental data and closure relations for the Taylor bubble velocity in slug flow, which has been determined to strongly affect the mechanistic model predictions (Lizarraga-García 2016). In this work, we study the behavior of Taylor bubbles in vertical and inclined pipes with upward and downward flow using a validated 3D computational fluid dynamics (CFD) approach with level set method implemented in a commercial code. A total of 56 cases are simulated, covering a wide range of fluid properties, pipe diameters, and inclination angles: Eo ∈ [10, 700]; Mo ∈ [1×10–6, 5×103]; ReSL ∈ [–40, 10]; θ ∈ [5°, 90°]. For bubbles in vertical upward flows, the simulated distribution parameter, C0, is successfully compared with an existing model. However, the C0 values of downward and inclined slug flows where the bubble becomes asymmetric are shown to be significantly different from their respective vertical upward flow values, and no current model exists for the fluids simulated here. The main contributions of this work are (1) the relatively large 3D numerical database generated for this type of flow, (2) the study of the asymmetric nature of inclined and some vertical downward slug flows, and (3) the analysis of its impact on the distribution parameter, C0.

Engineering↗

Seeing values for LSST strategy simulations

The opsim4 operations simulation program for the LSST astronomical survey uses a database of seeing values covering the range of times to besimulated. Idescribethe creation of such a database using Dual Image Motion Monitor(DIMM)datacollected at Cerro Pachon from 2004-03-17 to 2019-10-07. In times during which the data overlap, I compare the distribution of DIMM seeing values to the seeing measured in DECamimages,takenatasite 10kmaway. Becauseinstrumentalproblemsinthe DIMMmay indicate unreliablemeasurements,cutsonimagequality(asindicatedby the measured Strehlratio)wereexplored. TheDIMMhassignificantgaps,soImodel thedata(withandwithoutcutsonStrehlratio)andgenerateartificialdatainthegaps according to the model. The model consists of a sinusoidal variation with a period of one year, an autoregressive (AR1) model for variations in mean seeing from one night to the next, and another AR1 model for variations on a 5 minute timescale. I create four databases according to thisprocedure, twobasedonDIMMdatastarting 2006-01-01 (with and without a Strehl ratio cut), and two starting 2009-01-01. I then run opsim simulations using each, and an otherwise identical simulation using the default seeing database, and explore the differences

Neilsen, Eric H. [Fermilab]↗

Extraction of the neutron F 2 structure function from inclusive proton and deuteron deep-inelastic scattering data

The available world deep-inelastic scattering (DIS) data on proton and deuteron structure functions F 2 p , F 2 d , and their ratios are leveraged to extract the free neutron F 2 n structure function, the F 2 n / F 2 p ratio, and associated uncertainties using the latest nuclear effect calculations in the deuteron. Special attention is devoted to the normalization of the proton and deuteron experimental datasets and to the treatment of correlated systematic errors, as well as the quantification of procedural and theoretical uncertainties. The extracted F 2 n dataset is utilized to evaluate the Q 2 dependence of the Gottfried sum rule and the nonsinglet F 2 p − F 2 n moments. To facilitate replication of our study, as well as for general applications, we provide a comprehensive DIS database including all recent Jefferson Lab 6 GeV measurements, the extracted F n 2 , a modified CTEQ-JLab global parton distribution function fit named CJ15nlo_mod, and grids with calculated proton, neutron, and deuteron DIS structure functions. Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Scaling Resolution of Gigapixel Whole Slide Images Using Spatial Decomposition on Convolutional Neural Networks

Gigapixel images are prevalent in scientific domains ranging from remote sensing, and satellite imagery to microscopy, etc. However, training a deep learning model at the natural resolution of those images has been a challenge in terms of both, overcoming the resource limit (e.g. HBM memory constraints), as well as scaling up to a large number of GPUs. In this paper, we trained Residual neural Networks (ResNet) on 22,528 x 22,528-pixel size images using a distributed spatial decomposition method on 2,304 GPUs on the Summit Supercomputer. We applied our method on a Whole Slide Imaging (WSI) dataset from The Cancer Genome Atlas (TCGA) database. WSI images can be in the size of 100,000 x 100,000 pixels or even larger, and in this work we studied the effect of image resolution on a classification task, while achieving state-of-the-art AUC scores. Moreover, our approach doesn't need pixel-level labels, since we're avoiding patching from the WSI images completely, while adding the capability of training arbitrary large-size images. This is achieved through a distributed spatial decomposition method, by leveraging the non-block fat-tree interconnect network of the Summit architecture, which enabled GPU-to-GPU direct communication. Finally, detailed performance analysis results are shown, as well as a comparison with a data-parallel approach when possible.

Tsaris, Aristeidis (aris)↗