Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

dGen (Distributed Generation Market Demand) Model Data: Alpha Release

Open sourced data needed to run the basic alpha release version of the dGen model. Includes a pre-generated agent file of 100,000 agents in pickle file format along with the base schema and table data in parquet format that are needed to create a postgreSQL database for the model to interact with.

14 SOLAR ENERGY↗

The nth-plant scenario for blended feedstock conversion and preprocessing nationwide: biorefineries and depots

The sustainability of the biofuel industry depends on the development of a mature conversion technology on a national level that can take advantage of the economies of scale: the nth-plant. Defining the future location and supply logistics of conversion plants is imperative to ultimately transform the nation’s renewable biomass resources into cost-competitive, high-performance feedstock for production of biofuels and bioproducts. Since the US has put restrictions on production levels of conventional biofuels from edible resources, the nation needs to plan for the widespread accessibility and development of the cellulosic biofuel scenario. Conventional feedstock supply systems will be unable to handle cellulosic biomass nationwide, making it essential to expand the industry with an advanced feedstock supply system incorporating a distributed network of preprocessing depots and conversion plants, or biorefineries. Current studies are mostly limited to designing supply systems for specific regions of the country. We developed a national database with potential locations for depots and biorefineries to meet the nation’s target demand of cellulosic biofuel. Blended feedstock with switchgrass and corn stover (harvested by either a two- or three-pass method) are considered in a Mixed Integer Linear Programming model to deliver on-spec biomass that considers both, a desired quantity and quality at the biorefinery. A total delivered feedstock cost that is less than $79.07/dt (2016$) is evaluated for years 2022, 2030, and 2040. In 2022, 124 depots and 59 biorefineries could be supplied with 42.8 million dt of corn stover and switchgrass. In 2030 and 2040, the total accessible biomass could increase to 215% and 393% respectively when compared to 2022. However, an $8/dry tons reduction in targeted delivery cost could reduce total accessible biomass by 67%. Kansas, Nebraska, South Dakota and Texas were identified as potential states with a strong biofuel economy given that they had six or more biorefineries located in all scenarios. In some scenarios, Colorado, Alabama, Georgia, Minnesota, Mississippi and South Carolina would greatly benefit from a depot network as these could only deliver to a biorefinery in a nearby state. To elaborate the impact of a nationwide consideration, the findings were compared with existing literature for different US regions. We also present results for biorefinery capacities that are double, triple, and quadruple in size.

09 BIOMASS FUELS↗

Assessing Photovoltaic Capacity Factor Variability Using Long-Term Satellite Derived Solar Resource Data Under Brazilian Climate

Accurate estimation of photovoltaic (PV) energy yield and its variability is essential for reducing financial risk and supporting reliable system planning for rapidly expanding PV markets. In Brazil, high solar adoption and increasing levels of distributed energy resources are beginning to introduce operational challenges such as curtailment and evolving grid requirements. Understanding how natural variability in solar resource propagates into PV system performance is therefore increasingly important for both project design and grid integration. Modern PV yield assessments commonly rely on multi-year meteorological datasets and probabilistic exceedance metrics (e.g., P50/P90) to quantify energy yield uncertainty for project financing. However, the implications of long-term solar resource variability for PV system design choices and high-adoption grid conditions remain less well characterized for rapidly expanding markets such as Brazil. In particular, understanding how weather-driven variability propagates into PV production distributions and capacity factor expectations is important for evaluating curtailment exposure, deployment strategies, and storage requirements in regions experiencing rapid growth of distributed and utility-scale PV. Seasonal and interannual variability in atmospheric conditions can produce substantial fluctuations in monthly PV energy production, which propagate into uncertainty in annual energy yield and capacity factor expectations. Characterizing this variability using long-term meteorological datasets allows probabilistic estimation of PV system performance and provides improved insight into the range of expected PV energy outcomes. This study explores the use of long-term satellite-derived meteorological data from the National Solar Radiation Database (NSRDB) to evaluate the variability of photovoltaic system performance across multiple locations in Brazil. Using a 27-year dataset (1998-2024), PV system simulations are performed to characterize the distribution of annual and seasonal capacity factors and energy yield outcomes, while propagating key sources of meteorological variability and model uncertainty through the PV modeling chain. The analysis also investigates the sensitivity of PV performance outcomes to key system design assumptions within the PV modeling chain, including tracking configuration and system sizing parameters. The resulting probabilistic performance characterization provides insight into how weather-driven variability influences PV production expectations and capacity factor distributions. These results provide a foundation for evaluating how weather-driven variability interacts with high PV adoption and potential storage or curtailment mitigation strategies.

14 SOLAR ENERGY↗

The Ice Particle and Aggregate Simulator (IPAS). Part II: Analysis of a Database of Theoretical Aggregates for Microphysical Parameterization

Abstract Bulk ice-microphysical models parameterize the dynamic evolution of ice particles from advection, collection, and sedimentation through a cloud layer to the surface. Frozen hydrometeors can grow to acquire a multitude of shapes and sizes, which influence the distribution of mass within cloud systems. Aggregates, defined herein as the collection of ice particles, have a variety of formations based on initial ice particle size, shape, falling orientation, and the number of particles that collect. This work focuses on using the Ice Particle and Aggregate Simulator (IPAS) as a statistical tool to repetitively collect ice crystals of identical properties to derive bulk aggregate characteristics. A database of 9 744 000 aggregates is generated with resulting properties analyzed. After 150 single ice crystals (monomers) collect, the most extreme aggregate aspect ratio calculations asymptote toward and ϕ ca ≈ 0.50 for aggregates composed of quasi-horizontally oriented and randomly oriented monomers, respectively. The results presented are largely consistent with both a previous theoretical study and estimates derived from ground-based observations from two different geographic locations. Particle falling orientation highly influences newly formed aggregate aspect ratios from the collection of particles with extreme aspect ratios; quasi-horizontally oriented particles can produce aggregate aspect ratios an order of magnitude more extreme than randomly oriented particles but can also produce near-spherical aggregates as the number of monomers comprising the aggregate reach approximately 100. Finally, a majority of collections result in aggregates that are closer to prolate than oblate spheroids.

54 ENVIRONMENTAL SCIENCES↗

Accurate calculation of many-body energies in water clusters using a classical geometry-dependent induction model

We incorporate geometry-dependent distributed multipole and polarizability surfaces into an induction model that is used to describe the 3- and 4-body terms of the interaction between water molecules. The expansion is carried out up to hexadecapole with the multipoles distributed on the atom sites. Dipole-dipole, dipole-quadrupole, and quadrupole-quadrupole distributed polarizabilities are used to represent the response of the multipoles to an electric field. We compare the model against two large databases consisting of 43,844 3-body terms and 3,603 4-body terms obtained from high level ab initio calculations previously used to fit the MB-pol and q-AQUA interaction potentials. The classical induction model with no adjustable parameters reproduces the ab-initio 3- and 4-body terms contained in these teo Databases with a Root-Mean-Square-Error (RMSE) of 0.104/0.058 and a Mean-Absolute-Error (MAE) of 0.054/0.026 kcal/mol, respectively, results that are on a par with those obtained 1 by fitting the same data using tens of thousands of Permutationally Invariant Polynomials (PIPs). This demonstrates the accuracy of this physically motivated model in describing the 3- and 4-body terms in the interactions between water molecules with no adjustable parameters. The triple-dipole-dispersion energy was included in the 3-body energy and was found to be small but not quite negligible. The model represents a practical, efficient and transferable approach for obtaining accurate non-additive interactions for multi-component systems without the need of performing tens of thousands of high level electronic structure calculations and fitting them with tens of thousands of PIPs.

molecular analysis water, aqueous↗

Global variation in the fraction of leaf nitrogen allocated to photosynthesis

Plants invest a considerable amount of leaf nitrogen in the photosynthetic enzyme ribulose-1,5-bisphosphate carboxylase-oxygenase (RuBisCO), forming a strong coupling of nitrogen and photosynthetic capacity. Variability in the nitrogen-photosynthesis relationship indicates different nitrogen use strategies of plants (i.e., the fraction nitrogen allocated to RuBisCO; fLNR), however, the reason for this remains unclear as widely different nitrogen use strategies are adopted in photosynthesis models. Here, we use a comprehensive database of in situ observations, a remote sensing product of leaf chlorophyll and ancillary climate and soil data, to examine the global distribution in fLNR using a random forest model. We find global fLNR is 18.2 ± 6.2%, with its variation largely driven by negative dependence on leaf mass per area and positive dependence on leaf phosphorus. Some climate and soil factors (i.e., light, atmospheric dryness, soil pH, and sand) have considerable positive influences on fLNR regionally. This study provides insight into the nitrogen-photosynthesis relationship of plants globally and an improved understanding of the global distribution of photosynthetic potential.

54 ENVIRONMENTAL SCIENCES↗

A robust synthetic data generation framework for machine learning in high-resolution transmission electron microscopy (HRTEM)

Machine learning techniques are attractive options for developing highly-accurate analysis tools for nanomaterials characterization, including high-resolution transmission electron microscopy (HRTEM). However, successfully implementing such machine learning tools can be difficult due to the challenges in procuring sufficiently large, high-quality training datasets from experiments. In this work, we introduce Construction Zone, a Python package for rapid generation of complex nanoscale atomic structures which enables fast, systematic sampling of realistic nanomaterial structures and can be used as a random structure generator for large, diverse synthetic datasets. Using Construction Zone, we develop an end-to-end machine learning workflow for training neural network models to analyze experimental atomic resolution HRTEM images on the task of nanoparticle image segmentation purely with simulated databases. Further, we study the data curation process to understand how various aspects of the curated simulated data—including simulation fidelity, the distribution of atomic structures, and the distribution of imaging conditions—affect model performance across three benchmark experimental HRTEM image datasets. Using our workflow, we are able to achieve state-of-the-art segmentation performance on these experimental benchmarks and, further, we discuss robust strategies for consistently achieving high performance with machine learning in experimental settings using purely synthetic data. Construction Zone and its documentation are available at https://github.com/lerandc/construction_zone.

36 MATERIALS SCIENCE↗

Modernizing GlideinWMS Factory Monitoring with Prometheus & Grafana

Large-scale scientific experiments like CMS and DUNE rely on the distributed workload management system GlideinWMS to efficiently utilize computing resources across heterogeneous computing environments. GlideinWMS currently records Factory statistics using Round Robin Databases (RRDBs), XML, and JSON files, and these statistics are displayed via custom monitoring Web pages, thereby limiting integration with modern observability platforms. This project investigates the use of Prometheus-based instrumentation to expose Factory metrics using OpenTelemetry principles. Factory statistics related to Glidein submission and job execution are exported as Prometheus metrics through the Prometheus Python Client Library and are served via an HTTP metrics endpoint. The collected metrics are inspected using the Prometheus web-based interface and are visualized through Grafana dashboards within the Landscape monitoring infrastructure at Fermilab. This project significantly streamlines the integration of modern monitoring technologies into GlideinWMS and establishes a framework for extending observability across additional system components.

Appiah, Gideon [Grambling State U.]↗

The Coupled Influence of Thermal Physiology and Biotic Interactions on the Distribution and Density of Ant Species along an Elevational Gradient

A fundamental tenet of biogeography is that abiotic and biotic factors interact to shape the distributions of species and the organization of communities, with interactions being more important in benign environments, and environmental filtering more important in stressful environments. This pattern is often inferred using large databases or phylogenetic signal, but physiological mechanisms underlying such patterns are rarely examined. We focused on 18 ant species at 29 sites along an extensive elevational gradient, coupling experimental data on critical thermal limits, null model analyses, and observational data of density and abundance to elucidate factors governing species’ elevational range limits. Thermal tolerance data showed that environmental conditions were likely to be more important in colder, more stressful environments, where physiology was the most important constraint on the distribution and density of ant species. Conversely, the evidence for species interactions was strongest in warmer, more benign conditions, as indicated by our observational data and null model analyses. Our results provide a strong test that biotic interactions drive the distributions and density of species in warm climates, but that environmental filtering predominates at colder, high-elevation sites. Such a pattern suggests that the responses of species to climate change are likely to be context-dependent and more specifically, geographically-dependent.

54 ENVIRONMENTAL SCIENCES↗

Online multimedia retrieval on CPU–GPU platforms with adaptive work partition

Nearest neighbors search is a core operation found in several online multimedia services. These services have to handle very large databases, while, at the same time, they must minimize the query response times observed by users. This is specially complex because those services deal with fluctuating query workloads (rates). Consequently, they must adapt at run-time to minimize the response times as the load varies. In this paper, we address the aforementioned challenges with a distributed memory parallelization of the product quantization nearest neighbor search, also known as IVFADC, for hybrid CPU–GPU machines. Overall, our parallel IVFADC implements an out-of-GPU memory execution scheme to use the GPU for databases in which the index does not fit in its memory, which is crucial for searching in very large databases. The careful use of CPU and GPU with work stealing led to an average response time reduction of 2.4 as compared to using the GPU only. Also, our approach to adapt the system to fluctuating loads, called Dynamic Query Processing Policy (DQPP), attained a response time reduction of up to 5 vs. the best static (BS) policy for moderate loads. The system has attained high query processing rates and near-linear scalability in all experiments. We have evaluated our system on a machine with up to 256 NVIDIA V100 GPUs processing a database of 256 billion SIFT features vectors.

97 MATHEMATICS AND COMPUTING↗

Structural Flexibility of Metal Chelate Complexes and Its Relation to Supramolecular Chemistry

In this study, crystal structures of metal chelate and related complexes in the Cambridge Structural Database have been analyzed, with respect to their use as components in supramolecular metal-organic compounds. In β-diketonate complexes, the distribution of angles between ligands is relatively broad; other ligands, such as 2,2'-bipyridine, yield significantly narrower distributions. According to the principle of structure correlation, these distributions reflect the ease of distorting the various families of complexes. A comparison through density functional theory calculations also indicates that angular distortions require significantly less energy for M(β-diketonate) 3 than for M(2,2'-bipyridine) 3 . The differences are likely to affect the construction of supramolecular systems from different combinations of metals and ligands, including the likelihood that the desired structures will be obtained.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Metrics and extrapolation of resonant magnetic perturbation thresholds for ELM suppression

This large database study of resonant magnetic perturbation (RMP) edge localized mode (ELM) suppression thresholds in the AUG, DIII-D, EAST, and KSTAR tokamaks details the key strengths and weaknesses of RMP metrics. The RMP ELM suppression database used for this work contains plasma information at the time of transition from ELMing to ELM suppressed states where a clear experimental threshold is identified. The experimental threshold distributions are compared for five metrics: (1) the island overlap width, (2) pedestal top Chirikov overlap, (3) peeling edge displacement, (4) pedestal top resonant drive, and (5) edge dominant mode overlap. The distributions, the regularity of the dependence on RMP coil currents, and the sensitivities of a given metric to equilibrium reconstruction details are compared. The overlap metric proves to be a good compromise between including the appropriate plasma response physics and maintaining a numerical robustness. This quantity does not exhibit clear power-law scalings for projection, but machine learning can assist in predicting thresholds within the existing parameter ranges and providing uncertainty quantification of those predictions. Two new first-principles models, one utilizing a threshold from the non-linear Modified Rutherford equation evaluated at the pedestal top and one utilizing the SLAYER code to calculate the linear tearing threshold from torque balance, offer possible paths to extrapolation beyond the existing database parameter space.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Aboveground biomass density models for NASA’s Global Ecosystem Dynamics Investigation (GEDI) lidar mission

NASA's Global Ecosystem Dynamics Investigation (GEDI) is collecting spaceborne full waveform lidar data with a primary science goal of producing accurate estimates of forest aboveground biomass density (AGBD). This paper presents the development of the models used to create GEDI's footprint-level (~25 m) AGBD (GEDI04_A) product, including a description of the datasets used and the procedure for final model selection. The data used to fit our models are from a compilation of globally distributed spatially and temporally coincident field and airborne lidar datasets, whereby we simulated GEDI-like waveforms from airborne lidar to build a calibration database. We used this database to expand the geographic extent of past waveform lidar studies, and divided the globe into four broad strata by Plant Functional Type (PFT) and six geographic regions. GEDI's waveform-to-biomass models take the form of parametric Ordinary Least Squares (OLS) models with simulated Relative Height (RH) metrics as predictor variables. From an exhaustive set of candidate models, we selected the best input predictor variables, and data transformations for each geographic stratum in the GEDI domain to produce a set of comprehensive predictive footprint-level models. We found that model selection frequently favored combinations of RH metrics at the 98th, 90th, 50th, and 10th height above ground-level percentiles (RH98, RH90, RH50, and RH10, respectively), but that inclusion of lower RH metrics (e.g. RH10) did not markedly improve model performance. Second, forced inclusion of RH98 in all models was important and did not degrade model performance, and the best performing models were parsimonious, typically having only 1-3 predictors. Third, stratification by geographic domain (PFT, geographic region) improved model performance in comparison to global models without stratification. Fourth, for the vast majority of strata, the best performing models were fit using square root transformation of field AGBD and/or height metrics. There was considerable variability in model performance across geographic strata, and areas with sparse training data and/or high AGBD values had the poorest performance. These models are used to produce global predictions of AGBD, but will be improved in the future as more and better training data become available.

54 ENVIRONMENTAL SCIENCES↗

Remote Sensing Improves Multi‐Hazard Flooding and Extreme Heat Detection by Fivefold Over Current Estimates

The co‐occurrence of multiple hazards is of growing concern globally as the frequency and magnitude of extreme climate events increases. Despite studies examining the spatial distribution of such events, there has been little work in examining if all relevant life threatening and damaging hazards are captured in existing hazard databases and by common hazard metrics. For example, local/regional flash flooding events are seldom captured by optical satellite instruments and are subsequently excluded from global hazard databases. Similarly, the heat hazard definitions most frequently used in multi‐hazard studies inherently fail to capture events that are life‐threatening but climatologically within an expected range. Our goal is to determine the potential for increasing multi‐hazard event detection capabilities by inferring additional hazard footprints from widely accessible satellite data. We use daily precipitation and temperature satellite data to develop an open‐source framework that infers additional hazard footprints that are not included in traditional methods. With the state of Texas as our study area, we detected 2.5 times as many flood hazards, equivalent to $320 million in property and crop damages. Furthermore, our expanded heat hazard definition increases the impacted area by 56.6%, equivalent to 91.5 million km 2 over an 18 year period. Increasing hazard detection capabilities and expanding existing definitions of hazards using daily satellite data increases the temporal and spatial resolutions at which multi‐hazard events are detected. Having more complete data sets of all relevant hazard extents improves our ability to track global trends and more accurately determine the magnitude of hazard exposure inequities.

equity↗

Artificial neural network prediction of self-diffusion in pure compounds over multiple phase regimes

Artificial neural networks (ANNs) were developed to accurately predict the self-diffusion constants for pure components in liquid, gas and super critical phases. The ANNs were tested on an experimental database of 6625 self-diffusion constants for 118 different chemical compounds. The presence of multiple phases results in a heavy skew in the distribution of diffusion constants and multiple approaches were used to address this challenge. First, an ANN was developed with the raw diffusion values to assess what the main drawbacks of this direct method were. The first approach for improving the predictions involved taking the log 10 of diffusion to provide a more uniform distribution and reduce the range of target output values used to develop the ANN. The second approach involved developing individual ANNs for each phase using the raw diffusion values. Results show that the log transformation leads to a model with the best self-diffusion constant predictions and an overall average absolute deviation (AAD) of 6.56%. The resultant ANN is a generalized model that can be used to predict diffusion across all three phases and over a diverse group of compounds. The importance of each input feature was ranked using a feature addition method revealing that the density of the compound has the largest impact on the ANN prediction of self-diffusion constants in pure compounds.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Characterization of Rare-Earth Elements in Lignite Coal of the Williston Basin: Past Efforts and Ongoing Work

Rare-earth elements (REEs) have been a subject area of high interest for their unique properties. REEs are crucial materials used in an incredible array of consumer goods, energy system components, and military defense applications. While the United States has one operating REE mine, the product is sent overseas for refining into usable metals making the United States 100% import-reliant on these critical materials. This has led the Federal Government to declare the REE market an issue of national security. The Energy and Environmental Research Center (EERC) under funding provided by the Department of Energy (DOE) has undertaken efforts to determine if the lignite coal found in the Williston Basin has the potential to be an ore body containing sufficient quantities of REEs and Critical Minerals (CM) for extraction and processing. In one such effort, the EERC collected over 400 samples from the Williston Basin lignite coal seams including outcrops as well as active mines. Those efforts have been followed up with ongoing work under the U.S. DOE’s Carbon Ore, Rare Earth and Critical Minerals Initiative (CORE-CM) currently ongoing in the Williston Basin as well as other basins in the United States. This ongoing work in the Williston Basin characterizing REEs in coal has focused on the collection of new sampling and analysis of REEs in coal and building upon previous characterization work. This information is being used to understand the spatial distribution of REEs through mapping and 3D modeling. The goals of these efforts are to better understand the mechanisms for distribution of the REEs in lignite coal as well as their concentration and determine knowledge gaps in characterization and to start to build the database required to understand the potential resource in the Williston Basin.

Feole, Ian K.↗

The Interactive Hydropower Generation Map

This interactive map shows the spatial distribution and temporal trends of net generation for hydropower facilities in the U.S., as included in the Existing Hydropower Assets (EHA) Net Generation Plant Database.

Schmidt, Erik [Oak Ridge National Laboratory (ORNL↗

Split phase inverter data

The increase in power electronic based generation sources require accurate modeling of inverters. Accurate modeling requires experimental data over wider operation range. We used 8.35 kW off-the-shelf grid following split phase PV inverter in the experiments. We used controllable AC supply and controllable DC supply to emulate AC and DC side characteristics. The experiments were performed at NREL's Energy Systems Integration Facility. Inverter is tested under 100%, 75%, 50%, 25% load conditions. In the first dataset, for each operating condition, controllable AC source voltage is varied from 0.9 to 1.1 per unit (p.u) with a step value of 0.025 p.u while keeping the frequency at 60 Hz. In the second dataset, under similar load conditions (100%, 75%, 50%, 25% ), the frequency of the controllable AC source voltage was varied from 59 Hz to 61 Hz with a step value of 0.2 Hz. Voltage and frequency range is chosen based on inverter protection. Voltages and currents on DC and AC side are included in the dataset.

24 POWER TRANSMISSION AND DISTRIBUTION↗