Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

mkite

mkite is a distributed computing platform for materials simulation. mkite is built with the server-client pattern, decoupling production databases from client runners. When used in combination with message brokers, mkite enables any available client to perform calculations without prior hardware specification on the server side. Furthermore, the software enables the creation of complex workflows with multiple inputs and branches, facilitating the exploration of combinatorial chemical spaces. The mkite suite provides recipes and tools to interact with package such as VASP, but is extensible to any other simulation package. Finally, mkite helps keeping the provenance of calculations in a SQL database. A complete description of the software is available at https://arxiv.org/abs/2301.08841.

Schwalbe Koda, Daniel↗

KOMPASS-II: Compaction of Crushed salt for Safe Containment – Phase 2

Long-term stable sealing elements are a basic component in the safety concept for a possible repository for heat-emitting radioactive waste in rock salt. The sealing elements will be part of the closure concept for drifts and shafts. They will be made from a welldefinied crushed salt in employ a specific manufacturing process. The use of crushed salt as geotechnical barrier as required by the German Site Selection Act from 2017 /STA 17/ represents a paradigm change in the safety function of crushed salt, since this material was formerly only considered as stabilizing backfill for the host rock. The demonstration of the long-term stability and impermeability of crushed salt is crucial for its use as a geotechnical barrier. The KOMPASS-II project, is a follow-up of the KOMPASS-I project and continues the work with focus on improving the understanding of the thermal-hydraulic-mechanical (THM) coupled processes in crushed salt compaction with the objective to enhance the scientific competence for using crushed salt for the long-term isolation of high-level nuclear waste within rock salt repositories. The project strives for an adequate characterization of the compaction process and the essential influencing parameters, as well as a robust and reliable long-term prognosis using validated constitutive models. For this purpose, experimental studies on long-term compaction tests are combined with microstructural investigations and numerical modeling. The long-term compaction tests in this project focused on the effect of mean stress, deviatoric stress and temperature on the compaction behavior of crushed salt. A laboratory benchmark was performed identifying a variability in compaction behavior. Microstructural investigations were executed with the objective to characterize the influence of pre-compaction procedure, humidity content and grain size/grain size distribution on the overall compaction process of crushed salt with respect to the deformation mechanisms. The created database was used for benchmark calculations aiming for improvement and optimization of a large number of constitutive models available for crushed salt. The models were calibrated, and the improvement process was made visible applying the virtual demonstrator.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

ELM2.1-XGBfire1.0: improving wildfire prediction by integrating a machine learning fire model in a land surface model

Wildfires have shown increasing trends in both frequency and severity across the contiguous United States (CONUS). However, process-based fire models have difficulties in accurately simulating the burned area over the CONUS due to a simplification of the physical process and cannot capture the interplay among fire, ignition, climate, and human activities. The deficiency of burned area simulation deteriorates the description of fire impact on energy balance, water budget, and carbon fluxes in the Earth system models (ESMs). Alternatively, fire models based on machine learning (ML), which capture statistical relationships between the burned area and environmental factors, have shown promising burned area predictions and corresponding fire impact simulation. We develop a hybrid framework (ELM2.1-XGBFire1.0) that integrates an eXtreme Gradient Boosting (XGBoost) wildfire model with the Energy Exascale Earth System Model (E3SM) land model (ELM) version 2.1. A Fortran–C–Python deep learning bridge is adapted to support online communication between ELM and the ML fire model. Specifically, the burned area predicted by the ML-based wildfire model is directly passed to ELM to adjust the carbon pool and vegetation dynamics after disturbance, which are then used as predictors in the ML-based fire model in the next time step. Evaluated against the historical burned area from Global Fire Emissions Database 5 from 2001–2019, the ELM2.1-XGBFire1.0 outperforms process-based fire models in terms of spatial distribution and seasonal variations. The ELM2.1-XGBFire1.0 has proven to be a new tool for studying vegetation–fire interactions and, more importantly, enables seamless exploration of climate–fire feedback, working as an active component of E3SM.

54 ENVIRONMENTAL SCIENCES↗

FIRED (Fire Events Delineation): An Open, Flexible Algorithm and Database of US Fire Events Derived from the MODIS Burned Area Product (2001–2019)

Harnessing the fire data revolution, i.e., the abundance of information from satellites, government records, social media, and human health sources, now requires complex and challenging data integration approaches. Defining fire events is key to that effort. In order to understand the spatial and temporal characteristics of fire, or the classic fire regime concept, we need to critically define fire events from remote sensing data. Events, fundamentally a geographic concept with delineated spatial and temporal boundaries around a specific phenomenon that is homogenous in some property, are key to understanding fire regimes and more importantly how they are changing. Here, we describe Fire Events Delineation (FIRED), an event-delineation algorithm, that has been used to derive fire events (N = 51,871) from the MODIS MCD64 burned area product for the coterminous US (CONUS) from January 2001 to May 2019. The optimized spatial and temporal parameters to cluster burned area pixels into events were an 11-day window and a 5-pixel (2315 m) distance, when optimized against 13,741 wildfire perimeters in the CONUS from the Monitoring Trends in Burn Severity record. The linear relationship between the size of individual FIRED and Monitoring Trends in Burn Severity (MTBS) events for the CONUS was strong (R2 = 0.92 for all events). Importantly, this algorithm is open-source and flexible, allowing the end user to modify the spatio-temporal threshold or even the underlying algorithm approach as they see fit. We expect the optimized criteria to vary across regions, based on regional distributions of fire event size and rate of spread. We describe the derived metrics provided in a new national database and how they can be used to better understand US fire regimes. The open, flexible FIRED algorithm could be utilized to derive events in any satellite product. We hope that this open science effort will help catalyze a community-driven, data-integration effort (termed OneFire) to build a more complete picture of fire.

54 ENVIRONMENTAL SCIENCES↗

Profiling the BLAST bioinformatics application for load balancing on high-performance computing clusters

Abstract Background The Basic Local Alignment Search Tool (BLAST) is a suite of commonly used algorithms for identifying matches between biological sequences. The user supplies a database file and query file of sequences for BLAST to find identical sequences between the two. The typical millions of database and query sequences make BLAST computationally challenging but also well suited for parallelization on high-performance computing clusters. The efficacy of parallelization depends on the data partitioning, where the optimal data partitioning relies on an accurate performance model. In previous studies, a BLAST job was sped up by 27 times by partitioning the database and query among thousands of processor nodes. However, the optimality of the partitioning method was not studied. Unlike BLAST performance models proposed in the literature that usually have problem size and hardware configuration as the only variables, the execution time of a BLAST job is a function of database size, query size, and hardware capability. In this work, the nucleotide BLAST application BLASTN was profiled using three methods: shell-level profiling with the Unix “time” command, code-level profiling with the built-in “profiler” module, and system-level profiling with the Unix “gprof” program. The runtimes were measured for six node types, using six different database files and 15 query files, on a heterogeneous HPC cluster with 500+ nodes. The empirical measurement data were fitted with quadratic functions to develop performance models that were used to guide the data parallelization for BLASTN jobs. Results Profiling results showed that BLASTN contains more than 34,500 different functions, but a single function, RunMTBySplitDB, takes 99.12% of the total runtime. Among its 53 child functions, five core functions were identified to make up 92.12% of the overall BLASTN runtime. Based on the performance models, static load balancing algorithms can be applied to the BLASTN input data to minimize the runtime of the longest job on an HPC cluster. Four test cases being run on homogeneous and heterogeneous clusters were tested. Experiment results showed that the runtime can be reduced by 81% on a homogeneous cluster and by 20% on a heterogeneous cluster by re-distributing the workload. Discussion Optimal data partitioning can improve BLASTN’s overall runtime 5.4-fold in comparison with dividing the database and query into the same number of fragments. The proposed methodology can be used in the other applications in the BLAST+ suite or any other application as long as source code is available.

59 BASIC BIOLOGICAL SCIENCES↗

A North Atlantic synthetic tropical cyclone track, intensity, and rainfall dataset

Tropical Cyclones (TCs) cause significant socio-economic damages to the US and Caribbean coastal regions annually, making it important to understand TC risk at the local-to-regional scales. However, the short length of the observed record and the substantial computational expense associated with high-resolution climate models make it difficult to assess TC risk using either approach. To overcome these challenges, we developed a database of synthetic TCs using the Risk Analysis Framework for Tropical Cyclones (RAFT). The database includes 40,000 synthetic TC tracks, along-track intensities and storm-induced precipitation. TC tracks generated in RAFT are in reasonable agreement with the observed spatial distribution of TC tracks and basin-scale TC statistics. Specifically, along the coast, spatial variations in TC crossing probability and extreme winds upon landfall are well-reproduced by RAFT with R-squared values of 0.81 and 0.73, respectively. In summary, the synthetic TC database constructed with RAFT provides a reasonable pathway for the robust assessment of North Atlantic TC wind and rainfall risks.

54 ENVIRONMENTAL SCIENCES↗

The Third Data Release of the KODIAQ Survey

We present and make publicly available the third data release (DR3) of the Keck Observatory Database of Ionized Absorption toward Quasars (KODIAQ) survey. KODIAQ DR3 consists of a fully reduced sample of 727 quasars at 0.1 < z {sub em} < 6.4 observed with the Echellette Sepctrograph and Imager at moderate resolution (4000 ≤ R ≤ 10,000). DR3 contains 872 spectra available in flux calibrated form, representing a sum total exposure time of ∼2.8 megaseconds. These coadded spectra arise from a total of 2753 individual exposures of quasars taken from the Keck Observatory Archive (KOA) in raw form and uniformly processed using a data reduction package made available through the XIDL distribution. DR3 is publicly available to the community, housed as a higher level science product at the KOA and in the igmspec database.

74 ATOMIC AND MOLECULAR PHYSICS↗

Single Gaussian process method for arbitrary tokamak regimes with a statistical analysis

Abstract Gaussian process regression is a Bayesian method for inferring profiles based on input data. The technique is increasing in popularity in the fusion community due to its many advantages over traditional fitting techniques including intrinsic uncertainty quantification and robustness to over-fitting. This work investigates the use of a new method, the change-point method, for handling the varying length scales found in different tokamak regimes. The use of the Student’s t-distribution for the Bayesian likelihood probability is also investigated and shown to be advantageous in providing good fits in profiles with many outliers. To compare different methods, synthetic data generated from analytic profiles is used to create a database enabling a quantitative statistical comparison of which methods perform the best. Using a full Bayesian approach with the change-point method, Matérn kernel for the prior probability, and Student’s t-distribution for the likelihood is shown to give the best results.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Low-mode nonuniformity in direct-drive ICF implosions due to laser smoothing techniques employed on OMEGA

For successful laser-direct-drive inertial confinement fusion implosions, the laser irradiation must be highly uniform over the target surface. On OMEGA, multiple laser beams are used to illuminate targets quasi-uniformly. High-mode-number nonuniformities due to laser speckle on each individual beam are reduced by splitting each beam into two orthogonal polarizations (i.e., polarization smoothing, or PS) and a range of wavelengths (i.e., smoothing by spectral dispersion) that are dispersed at the target plane. However, cross-beam energy transfer (CBET) is sensitive to both the polarizations and wavelengths of the interacting beams, so the interplay between CBET and the laser-smoothing schemes results in unique intensity variation across each beam profile, which is a systematic source of low-mode drive nonuniformity on OMEGA. Here, we model these effects and find that the predicted ℓ = 1 mode in the laser-absorption distribution is consistent with the systematic core-flow direction that has been determined from the OMEGA implosion database. We also observe good agreement with the measured core-flow directions for two specific sets of implosions (one with PS, the other without PS) when we also account for the measured beam mispointing and the beam power imbalance.

Crossed beam scattering↗

OpenSAMPL: An Open Source Library for Timing and Synchronization Measurements and Analytics

Today's power grid operators are implementing timing and synchronization solutions that provide resilience to Global Navigation Satellite System (GNSS) vulnerabilities. These vendor-specific solutions often come with additional software applications that are designed to monitor that vendor's synchronization performance data. However, resilient timing architectures often resulting in multi-vendor solutions, including approaches that blend terrestrial clocks with space-based subscription services. In such an environment, collecting, analyzing, and visualizing data from a variety of sources within a single platform was heretofore not possible. To address this need, the US Department of Energy's Center for Alternative Synchronization and Timing (CAST) developed OpenSAMPL, the Open Synchronized Analytics and Monitoring Platform, an open-source Python framework for processing, loading, and observing clock measurement data from distributed devices. OpenSAMPL enables the ingestion of diverse clock-probe sources into a scalable time-series database and applies robust analytics. OpenSAMPL currently supports two vendor data pipelines, and will be extended to more in the near future, enabling seamless monitoring of a variety of timing and synchronization devices in a common environment.

Grant, Josh [ORNL] (ORCID:0000000163475060)↗

Hydropower Resilience Database for Assessing Microgrid Formation Capability and Enhancing Power Grid Resilience

In the face of increasing frequency of extreme events, enhancing resilience and reliability of energy infrastructure demands innovative solutions. Hydropower, with its inherent generation flexibility and grid-forming capability, offers significant potential for enhancing power grid resilience through the establishment of microgrids. To enable informed decision-making and strategic planning, we present the Hydropower Resilience Database (HRD) that integrates data from various sources such as Oakridge National Laboratory (ORNL) HydroSource, National Inventory of Dams, and US Western Grid database. By integrating relevant information on hydropower plant characteristics, including dams, reservoirs, and connected electrical grid, this resource enables hydropower plant owners, utilities, and community stakeholders to identify and evaluate the feasibility of using hydropower resources in microgrids to support nearby communities and critical infrastructure in various challenging scenarios. Using HRD, a set of metrics are evaluated to quantify the capability of hydropower plants to support essential microgrid functions. An interactive tool is developed using ArcGIS, enabling the visualization and analysis using HRD. Our research contributes to strategic planning efforts for fortifying energy infrastructure, ensuring reliable power supply during disruptions, and advancing the development of robust and resilient energy systems in regions susceptible to grid vulnerabilities.

13 HYDRO ENERGY↗

Hydropower Resilience Database for Assessing Microgrid Formation Capability and Enhancing Power Grid Resilience

In the face of increasing frequency of extreme events, enhancing resilience and reliability of energy infrastructure demands innovative solutions. Hydropower, with its inherent generation flexibility and grid-forming capability, offers significant potential for enhancing power grid resilience through the establishment of microgrids. To enable informed decision-making and strategic planning, we present the Hydropower Resilience Database (HRD) that integrates data from various sources such as Oakridge National Laboratory (ORNL) HydroSource, National Inventory of Dams, and US Western Grid database. By integrating relevant information on hydropower plant characteristics, including dams, reservoirs, and connected electrical grid, this resource enables hydropower plant owners, utilities, and community stakeholders to identify and evaluate the feasibility of using hydropower resources in microgrids to support nearby communities and critical infrastructure in various challenging scenarios. Using HRD, a set of metrics are evaluated to quantify the capability of hydropower plants to support essential microgrid functions. An interactive tool is developed using ArcGIS, enabling the visualization and analysis using HRD. Our research contributes to strategic planning efforts for fortifying energy infrastructure, ensuring reliable power supply during disruptions, and advancing the development of robust and resilient energy systems in regions susceptible to grid vulnerabilities.

13 HYDRO ENERGY↗

Advances in Metallic Fuel Database Development and Data Qualification

The Fuels Irradiation and Physics Database (FIPD [1]) is a comprehensive repository of data and documents related to Uranium-Zirconium based metallic fuel test pins. This database stores operational conditions of these pins, calculated using a suite of Argonne National Laboratory analysis codes developed during the Integral Fast Reactor (IFR) program. Key calculated data include axial distributions of power, temperature, fluence, burnup, and isotopic densities. Additionally, the FIPD holds post-irradiation examination (PIE) data such as fission gas release, gas chemistry measurements, and axial distributions derived from profilometry, gamma scanning, and neutron radiography. Complementing these data is an extensive archive of documents related to various pins and experiments. These include raw PIE records, design details, safety analyses, and operational reports. More detail about FIPD can be found in ref. [2]. The database development is an ongoing effort covering metallic fuel experiments from the Experimental Breeder Reactor II (EBR-II) and the Fast Flux Test Facility (FFTF). The recent improvements to the database and the data QA status are summarized in this paper.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Aerodynamic Sensitivities over Separable Shape Tensors

Here, we present a comprehensive aerodynamic sensitivity analysis of airfoil parameterization informed by separable shape tensors. This parameterization approach uniquely benefits the design process by isolating various well-studied shape characteristics, such as airfoil thickness, and providing a well-regulated low-dimensional parameter domain for aerodynamic designs. Exploring the aerodynamic sensitivities of this novel parameterization can provide valuable insights for more robust designs and future manufacturing efforts. We construct a data-driven parameter space of airfoils using principal geodesic analysis of separable shape tensors informed by a curated database containing almost 20,000 suitable engineering airfoils. Analyzing the shape reconstruction error and the maximum mean discrepancy between joint distributions of aerodynamic quantities, we study the dimensionality of the learned parameter space. This simple numerical experiment demonstrates a dramatic dimension reduction that retains design effectiveness and promotes regularity of the shape representations. Finally, we generate new airfoils and use the HAM2D Reynolds-averaged Navier–Stokes solver to predict lift, drag, and moment coefficients. We compute multiple sensitivity metrics to quantify and assert the consistency of parameter influence on the aerodynamic quantities. We also explore low-dimensional polynomial ridge approximations to motivate physical intuitions and offer explanations of the approximated sensitivities.

17 WIND ENERGY↗

Enabling Efficient Sparse Computations using Linear Algebra Aware Compilers

This project developed the LAPIS compiler framework, built on the Multilevel Intermediate Representation (MLIR), to optimize sparse linear algebra operations and support performance portability across diverse architectures. The main innovation of LAPIS is the Kokkos dialect, which allows for lowering codes from a high productivity language to different architectures in an elegant way. The dialect also allows the conversion of lower-level MLIR code to C++ Kokkos code, facilitating the integration of scientific machine learning (SciML) models into applications. To extend LAPIS for distributed memory architectures, a new partition dialect was created to manage the distribution of sparse tensors and express communication patterns for sparse linear algebra operations. This dialect also supports the distributed execution of operators and includes algorithmic optimizations to minimize communication to improve performance. The project also demonstrates that MLIR can enable effective linear algebra-level optimizations, improving performance on different GPUs for both sparse and dense linear algebra kernels. Key applications of LAPIS include sparse linear algebra and graph kernels, TenSQL, a relational database management solution built on GraphBLAS, and the development of subgraph isomorphism and monomorphism kernels, showcasing performance portability. In summary, the LAPIS framework supports productivity, performance, portability, and distributed memory execution, while also enabling linear algebra-level optimizations that are challenging in traditional programming languages, with successful applications ranging from simple sparse linear algebra to complex graph kernels.

97 MATHEMATICS AND COMPUTING↗

National Community Solar Partnership (NCSP+) Resource List

NLR compiles and makes quarterly updates to this dataset of publications related to expanding access to solar through its support for the National Community Solar Partnership (NCSP+). The purpose of the list is to help NCSP+ partners and others to find relevant resources on community solar, low- and moderate-income residential rooftop solar + storage, community-serving commercial solar projects, microgrids, and distributed solar + storage aggregations such as virtual power plants serving low-income communities. Resource types include journal articles, reports, fact sheets, presentations, videos, datasets, modeling tools, webinars, and other formats. A web-based version of the database is available on the NCSP+ Data Hub at: https://openei.org/wiki/NCSP_Hub/Resources .

14 SOLAR ENERGY↗

Automatic and rapid calibration of urban building energy models by learning from energy performance database

Urban building energy modeling (UBEM) is attracting increasing attention in the energy modeling filed. Unlike modeling a single building using detailed building systems information, UBEM generally uses limited high-level building stock data to infer default assumptions about building characteristics and operations. Additionally, this practice inherently brings uncertainty to UBEM. This study introduced a novel method of automatic and rapid calibration of UBEM based on the annual electricity and natural gas energy use data by learning the correlations between crucial model input parameters and the building energy use from the reference building models. A case study was presented to calibrate 72 large office buildings built before 1978 in San Francisco. Seventeen model parameters were selected and Monte Carlo sampling was used to create 1000 samples that reasonably represent the parameter space. Then 1000 simulations were performed for the reference building model to create an energy performance database. The results showed that by learning from the energy performance database, it took less than four simulation runs on average to calibrate a building model. After the calibration, the distributions of each parameter were obtained to replace their single predefined default values. For example, the default lighting power density of 21.39 W/m 2 was calibrated to be 7.50 W/m 2 on average. The case study successfully demonstrated the effectiveness of the novel calibration method for UBEM in the mild climate. The method will be further tested in future for other climate zones and other building types.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Simulation of PV Variability as a Function of PV Generation and Plant Size

The deployment of photovoltaic (PV) systems continues to show significant expansion; however, this growth has brought added attention to issues around the variability of the solar resource. Both spatial and temporal variability exist. Temporal scales can range from the sub-second to multiyear, whereas spatial scales can range from a few meters to tens of kilometers. There are multiple methods described in the literature to quantify PV variability at various spatial and temporal scales. This study focuses on short-term temporal variability and uses similar approaches with the addition of PV plant size a parameter to quantify variability. The method employed here incorporates the normalization of clear-and cloudy-sky conditions and PV plant size to quantify nominal variability metrics. The distribution and fluctuations of these metrics provide relevant information that is useful for system operations. The National Solar Radiation Database (NSRDB) is used to simulate PV variability as a function of PV generation and plant size. Hypothetical but realistic system information at 33 locations is used to model PV generation by feeding NSRDB solar irradiance data to the National Renewable Energy Laboratory’s System Advisor Model (SAM). Over the selected region, it is found that the aggregated ramp rates for the 1-minute data are associated with standard deviations ranging from 0.002–0.055 on a daily basis; however, hourly intervals induce higher aggregated ramp rates than the other timescales. Even though minute-to-minute variations are significant for the 1-minute time-scale, the standard deviation aggregated into a daily metric is smaller because of the cancellation of values.

irradiance↗