Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Geographic Access to Radiation Therapy Facilities in the United States

The current distribution of radiation therapy (RT) facilities in the United States is not well established. A comprehensive inventory of U.S. RT facilities was last assessed in 2005, based on data from state regulatory agencies and dosimetric quality assurance bodies. We updated this database to characterize population-level measures of geographic access to RT and analyze changes over the past 15 years.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Understanding the Spatial Organization of Simultaneous Heavy Precipitation Events Over the Conterminous United States

We introduce the idea of simultaneous heavy precipitation events (SHPEs) to understand whether extreme precipitation has a spatial organization manifested as specified tracks or contiguous fields with inherent scaling relationships. For this purpose, we created a database of SHPEs using ground-based precipitation observations recorded by the daily Global Historical Climatology Network across the conterminous United States during 1900–2014. SHPEs are examined for their seasonality, spatial manifestation, orientation, and areal extent. We quantified the spatial distribution of the centroids and principal axes of SHPEs and their quasi-elliptical manifestations, azimuthal orientations, and areal extents on the ground. Four seasons, December-January-February (DJF), March-April-May (MAM),June-July-August (JJA), and September-October-November (SON) are considered to examine the spatial patterns and associated large-scale atmospheric circulations. Results indicate that there are 54, 58, 103, and204 SHPEs in DJF, MAM, JJA, and SON seasons, and their longest stretches range on average between 650and 1,600, between 850 and 1,500, between 950 and 1,550, and between 750 and 1,450 km, respectively. SHPEs in the DJF, MAM, and JJA seasons occur mostly over the Pacific coast and central and midwestern United States, respectively. The atmospheric circulation patterns and mechanisms of precipitable water vapor and moisture transport in the atmosphere are also discussed in relation to these SHPEs. Power laws explain SHPEs' underlying area scaling behavior in all the four seasons, with stronger evidence in DJF and MAM. A seasonal spatial risk model is developed to predict the likelihood of SHPEs. Quantifying the characteristics of SHPEs and modeling their footprints can improve the projections of flood risk and understanding of damages to interconnected infrastructure systems.

54 ENVIRONMENTAL SCIENCES↗

Screening of Potential Biomass-Derived Streams as Fuel Blendstocks for Mixing Controlled Compression Ignition Combustion

Mixing controlled compression ignition or diesel engines are highly efficient and are likely to continue to be the primary means for movement of goods for many years to come. Low-carbon biofuels have the potential to significantly reduce the carbon footprint of diesel combustion and could potentially have advantageous properties for combustion such as high cetane number and reduced engine out particle emissions. In this study, we developed a list of potential biomass-derived diesel blendstocks. An online database of properties and characteristics of these bioblendstocks was populated with data. Fuel properties were determined by measurement, model prediction, or literature review. Screening criteria were developed to determine if a bioblendstock met the basic requirements for handling in the diesel distribution system and use as a blend with conventional diesel. Criteria included cetane number =40, flashpoint =52 degrees C, and boiling point or final boiling point <338 degrees C. Blendstocks needed to be soluble in diesel fuel, have a toxicity no worse than conventional diesel, and not be corrosive. Additionally, cloud point or freezing point below 0 degrees C was required. This screening produced a list of 27 potential bioblendstocks. Of these candidates, 13 were available commercially or could be synthesized by biofuels production researchers and included 11 nominally pure components and two mixtures. These 13 candidates were then subjected to further study based on how they impact fuel properties upon blending. Blend properties included cetane number, lubricity, conductivity, oxidation stability and viscosity. Results indicate that all thirteen candidates can meet the basic requirements for diesel fuel blending and are likely to reduce particle emissions from diesel combustion.

09 BIOMASS FUELS↗

Battery inverter experimental data

The increase in power electronic based generation sources require accurate modeling of inverters. Accurate modeling requires experimental data over wider operation range. We used 30 kW off-the-shelf grid following battery inverter in the experiments. We used controllable AC supply and controllable DC supply to emulate AC and DC side characteristics. The experiments were performed at NREL's Energy Systems Integration Facility. Inverter is tested under 100%, 75%, 50%, 25% load conditions. In the first dataset, for each operating condition, controllable AC source voltage is varied from 0.9 to 1.1 per unit (p.u) with a step value of 0.025 p.u while keeping the frequency at 60 Hz. In the second dataset, under similar load conditions (100%, 75%, 50%, 25% ), the frequency of the controllable AC source voltage was varied from 59 Hz to 61 Hz with a step value of 0.2 Hz. Voltage and frequency range is chosen based on inverter protection. Voltages and currents on DC and AC side are included in the dataset.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Evaluation of Graph Analytics Frameworks Using the GAP Benchmark Suite

The analysis of connected data is an increasingly important application in high-performance computing. Such analyses can reveal fraudulent patterns in financial transactions, optimize telecommunications networks, predict information flow in social networks, etc. However, the landscape of graph analytics is highly diverse. Graph algorithms stress processor architectures differently, and no one graph can represent all topologies. Consequently, no single approach or framework is expected to be optimal for all graph analytics problems. To help make sense of this diverse landscape, we evaluated four approaches to graph analytics: GraphBLAS, Galois, BGL17, GraphIt; and compare them against hand-tuned implementations that take advantage of hardware features on our test platform. Graph- BLAS formulates graph analytics as sparse linear algebra. Galois provides syntactic constructs for data parallelism over irregular data structures. BGL17 is a generic C++ template library for implementing graph algorithms. GraphIt provides a domain- specific language to describe and optimize graph algorithms. We use the GAP Benchmark Suite to establish baseline performance and guide the side-by-side evaluation of each framework. GAP consists of 30 tests: six graph analytics algorithms (breadth- first search, single-source shortest path, PageRank, betweenness centrality, connected components, and triangle counting) run on five graphs, each with different topological characteristics (e.g., high diameter, skewed degree distribution, high average degree). High-performance reference implementations are included for each benchmark algorithm. Because a graph can be loaded into memory a number of ways (e.g., flat file on disk, compressed sparse format, data frames, retrieved from SQL or NoSQL databases), our evaluation focused on computational performance rather than I/O. Our results show the relative strengths of each framework.

Graph algorithms, Benchmarking, shared-memory prog↗

Revealing the Statistics of Extreme Events Hidden in Short Weather Forecast Data

Extreme weather events have significant consequences, dominating the impact of climate on society. While high-resolution weather models can forecast many types of extreme events on synoptic timescales, long-term climatological risk assessment is an altogether different problem. A once-in-a-century event takes, on average, 100 years of simulation time to appear just once, far beyond the typical integration length of a weather forecast model. Therefore, this task is left to cheaper, but less accurate, low-resolution or statistical models. But there is untapped potential in weather model output: despite being short in duration, weather forecast ensembles are produced multiple times a week. Integrations are launched with independent perturbations, causing them to spread apart over time and broadly sample phase space. Collectively, these integrations add up to thousands of years of data. We establish methods to extract climatological information from these short weather simulations. Using ensemble hindcasts by the European Center for Medium-range Weather Forecasting archived in the subseasonal-to-seasonal (S2S) database, we characterize sudden stratospheric warming (SSW) events with multi-centennial return times. Consistent results are found between alternative methods, including basic counting strategies and Markov state modeling. By carefully combining trajectories together, we obtain estimates of SSW frequencies and their seasonal distributions that are consistent with reanalysis-derived estimates for moderately rare events, but with much tighter uncertainty bounds, and which can be extended to events of unprecedented severity that have not yet been observed historically. These methods hold potential for assessing extreme events throughout the climate system, beyond this example of stratospheric extremes.

58 GEOSCIENCES↗

Out-of-Pile Furnace Tests on Fast Reactor Metallic Fuels Conducted at the AGHCF

An extensive out-of-pile furnace test program was conducted at Argonne’s Alpha-Gamma Hot Cell Facility (AGHCF) from 1987-1994 to evaluate the fuel/clad compatibility and performance of fast reactor metallic fuels. This test program included over 150 tests on irradiated fuels conducted in two furnace apparatuses. The available records of these tests have been preserved with the support of the Advanced Reactor Technology program and organized in the OPTD (Out-of-Pile Transient Database). This report provides at-a-glance summary information for each of the out-of-pile furnace tests, including information about the tested fuel samples, test conditions, purpose of the tests, and key results. It is intended for open and unlimited distribution to allow all interested persons to view key information about the out-of-pile tests.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Predicting Rare Earth Element Potential in Produced and Geothermal Waters of the United States via Emergent Self-Organizing Maps

This work applies emergent self-organizing map (ESOM) techniques, a form of machine learning, in the multidimensional interpretation and prediction of rare earth element (REE) abundance in produced and geothermal waters in the United States. Visualization of the variables in the ESOM trained using the input data shows that each REE, with the exception of Eu, follows the same distribution patterns and that no single parameter appears to control their distribution. Cross-validation, using a random subsample of the starting data and only using major ions, shows that predictions are generally accurate to within an order of magnitude. Using the same approach, an abridged version of the U.S. Geological Survey Produced Waters Database, Version 2.3 (which includes both data from produced and geothermal waters) was mapped to the ESOM and predicted values were generated for samples that contained enough variables to be effectively mapped. Results show that in general, produced and geothermal waters are predicted to be enriched in REEs by an order of magnitude or more relative to seawater, with maximum predicted enrichments in excess of 1000-fold. Cartographic mapping of the resulting predictions indicates that maximum REE concentrations exceed values in seawater across the majority of geologic basins investigated and that REEs are typically spatially co-associated. The factors causing this co-association were not determined from ESOM analysis, but based on the information currently available, REE content in produced and geothermal waters is not directly controlled by lithology, reservoir temperature, or salinity.

Engle, Mark A. (ORCID:0000000152587374)↗

Enabling Data Exchange and Data Integration with the Common Information Model: An Introduction for Power Systems Engineers and Application Developers

The Common Information Model (CIM) is an open-source information model that is used to model an electrical network and the various equipment used on the network. CIM is widely used for data exchange of bulk transmission power systems and is finding increasing use for distribution systems. Use of a non-proprietary information model (such as CIM) that has been agreed upon and adopted by numerous utilities, vendors, and researchers allows significant reduction in the effort and cost of data integration. Likewise, adoption of open data platforms built around the CIM increases available functionalities for managing and optimizing the smart grid of the future. This report is intended as an introduction to CIM for utility engineers, power systems researchers, and application developers, providing a broad view of the CIM and how particular profiles can be adapted for various use cases. Unlike most other CIM introduction documents and the International Electrotechnical Commission (IEC) standards (which are mostly targeted to an audience of data scientists, enterprise database managers, and platform developers), this report is intended for users of traditional power systems analysis software and other readers without any prior experience with canonical information models, data profiles, or UML modeling.

24 POWER TRANSMISSION AND DISTRIBUTION↗

APACE: AlphaFold2 and advanced computing as a service for accelerated discovery in biophysics

The prediction of protein 3D structure from amino acid sequence is a computational grand challenge in biophysics and plays a key role in robust protein structure prediction algorithms, from drug discovery to genome interpretation. The advent of AI models, such as AlphaFold, is revolutionizing applications that depend on robust protein structure prediction algorithms. To maximize the impact, and ease the usability, of these AI tools we introduce APACE, AlphaFold2 and advanced computing as a service, a computational framework that effectively handles this AI model and its TB-size database to conduct accelerated protein structure prediction analyses in modern supercomputing environments. We deployed APACE in the Delta and Polaris supercomputers and quantified its performance for accurate protein structure predictions using four exemplar proteins: 6AWO, 6OAN, 7MEZ, and 6D6U. Using up to 300 ensembles, distributed across 200 NVIDIA A100 GPUs, we found that APACE is up to two orders of magnitude faster than off-the-self AlphaFold2 implementations, reducing time-to-solution from weeks to minutes. This computational approach may be readily linked with robotics laboratories to automate and accelerate scientific discovery.

97 MATHEMATICS AND COMPUTING↗

Genomic fingerprints of the world’s soil ecosystems

Despite the explosion of soil metagenomic data, we lack a synthesized understanding of patterns in the distribution and functions of soil microorganisms. These patterns are critical to predictions of soil microbiome responses to climate change and resulting feedbacks that regulate greenhouse gas release from soils. To address this gap, we assay 1,512 manually curated soil metagenomes using complementary annotation databases, read-based taxonomy, and machine learning to extract multidimensional genomic fingerprints of global soil microbiomes. Our objective is to uncover novel biogeographical patterns of soil microbiomes across environmental factors and ecological biomes with high molecular resolution. We reveal shifts in the potential for (i) microbial nutrient acquisition across pH gradients; (ii) stress-, transport-, and redox-based processes across changes in soil bulk density; and (iii) greenhouse gas emissions across biomes. We also use an unsupervised approach to reveal a collection of soils with distinct genomic signatures, characterized by coordinated changes in soil organic carbon, nitrogen, and cation exchange capacity and in bulk density and clay content that may ultimately reflect soil environments with high microbial activity. Genomic fingerprints for these soils highlight the importance of resource scavenging, plant-microbe interactions, fungi, and heterotrophic metabolisms. Across all analyses, we observed phylogenetic coherence in soil microbiomes—more closely related microorganisms tended to move congruently in response to soil factors. Collectively, the genomic fingerprints uncovered here present a basis for global patterns in the microbial mechanisms underlying soil biogeochemistry and help beget tractable microbial reaction networks for incorporation into process-based models of soil carbon and nutrient cycling.

59 BASIC BIOLOGICAL SCIENCES↗

Size distribution of polycyclic aromatic hydrocarbons in space: an old new light on the 11.2/3.3 μm intensity ratio

The intensity ratio of the 11.2/3.3 μm emission bands is considered to be a reliable tracer of the size distribution of polycyclic aromatic hydrocarbons (PAHs) in the interstellar medium (ISM). This paper describes the validation of the calculated intrinsic infrared (IR) spectra of PAHs that underlie the interpretation of the observed ratio. The comparison of harmonic calculations from the NASA Ames PAH IR spectroscopic database to gas-phase experimental absorption IR spectra reveals a consistent underestimation of the 11.2/3.3 μm intensity ratio by 34%. IR spectra based on higher level anharmonic calculations, on the other hand, are in very good agreement with the experiments. While there are indications that the 11.2/3.3 μm ratio increases systematically for PAHs in the relevant size range when using a larger basis set, it is unfortunately not yet possible to reliably calculate anharmonic spectra for large PAHs. Based on these considerations, we have adjusted the intrinsic ratio of these modes and incorporated this in an interstellar PAH emission model. This corrected model implies that typical PAH sizes in reflection nebulae such as NGC 7023 – previously inferred to be in the range of 50 to 70 carbon atoms per PAH are actually in the range of 40 to 55 carbon atoms. The higher limit of this range is close to the size of the C 60 fullerene (also detected in reflection nebulae), which would be in line with the hypothesis that, under appropriate conditions, large PAHs are converted into the more stable fullerenes in the ISM.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Viroid-like “obelisk” agents are widespread in the ocean and exceed the abundance of RNA viruses in the prokaryotic fraction

Abstract “Obelisks” are recently discovered ribonucleic acid (RNA) viroid-like elements present in diverse environments with no phylogenetic similarity to any known biological agent. obelisks were first identified in the human gut and in a commensal bacterium acting as a replicative host. They have a circular ∼1 kb RNA genome, rod-like secondary structures, and the encoding of a protein superfamily called “Oblins”. We performed a large-scale search of obelisks in the ocean using the Pebblescout program and the transcriptomic Sequence Archive Read databases, revealing the biogeography and abundance of these viroid-like RNA elements. We detected 55 obelisk genomes resulting in 35 marine clusters at the species level. These obelisks were detected in the prokaryotic fraction and to a lesser extent in the eukaryotic fraction, and distributed across all the oceans from surface to mesopelagic including the Arctic, and even in the coldest seawater of Earth beneath the Antarctic Ross Ice Shelf. The obelisk hallmark protein Oblin-1 confirmed by 3D models was found in various marine samples. Some of the detected marine obelisks harbor hammerhead self-cleaving ribozymes in both polarities. In the prokaryotic, but not the eukaryotic, fraction of the Tara Ocean dataset, relative abundance of obelisks calculated by transcriptomic fragment recruitment indicated that they are abundant in marine samples, reaching or even exceeding the relative abundance of the previously discovered uncultured RNA viruses. In conclusion, obelisks are abundant and widespread viroid-like elements that should be included in ocean biogeochemical models.

Environmental Sciences & Ecology↗

Scalable Knowledge Graph Analytics at 136 Petaflop/s

We are motivated by newly proposed methods for data mining large-scale corpora of scholarly publications, such as the full biomedical literature, which may consist of tens of millions of papers spanning decades of research. In this setting, analysts seek to discover how concepts relate to one another. They construct graph representations from annotated text databases and then formulate the relationship-mining problem as one of computing all-pairs shortest paths (APSP), which becomes a significant bottleneck. In this context, we present a new high-performance algorithm and implementation of the Floyd-Warshall algorithm for distributed-memory parallel computers accelerated by GPUs, which we call DSNAPSHOT (Distributed Accelerated Semiring All-Pairs Shortest Path). For our largest experiments, we ran DSNAPSHOT on a connected input graph with millions of vertices using 4, 096nodes (24,576GPUs) of the Oak Ridge National Laboratory's Summit supercomputer system. We find DSNAPSHOT achieves a sustained performance of 136×1015 floating-point operations per second (136petaflop/s) at a parallel efficiency of 90% under weak scaling and, in absolute speed, 70% of the best possible performance given our computation (in the single-precision tropical semiring or “min-plus” algebra). Looking forward, we believe this novel capability will enable the mining of scholarly knowledge corpora when embedded and integrated into artificial intelligence-driven natural language processing workflows at scale.

Kannan, Ramakrishnan {ramki}↗

A kinetic line-driven radiation operator and its application to Gyrokinetics

A velocity dependent, kinetic model for line radiation is developed for continuum kinetic codes. It has been implemented in the full-f gyrokinetic code Gkeyll. The total radiation for a charge state is modeled as an advection in velocity space with a form of $\nabla_v \cdot(v\nu(v)f(v))$, guaranteeing particle conservation. The velocity dependence (in the form of an effective frequency $\nu(v)$) is found through fitting the energy loss of the operator, i.e. the second velocity moment, to the radiation data in the OpenADAS database. Therefore, each individual transition does not need to be evaluated every time step, significantly reducing the computational cost of including line radiation in a kinetic model. The dependence on velocity instead of the usual, temperature, allows the radiation to be computed from non-Maxwellian electron distribution functions: We benchmark the model against a collisional radiative model using isotropic non-Maxwellian distribution functions. A velocity dependent model of radiation can more accurately describe the radiation in the more kinetic regimes expected in reactor-scale devices. The velocity dependence qualitatively captures the quantum mechanical need for a minimum velocity before any radiation occurs.

kinetic↗

A market-oriented database design for critical material research

Material databases are important tools to provide and store information from material research. Rising concerns about supply-chain risks to raw materials presents a need to incorporate raw-material market and end-use application data, beyond basic chemical and physical properties, into a material database. One key challenge for researchers working on critical materials is information scarcity and inconsistency. This paper introduces, as a result of a two-year project, a critical-material commodity database (CMCD) incorporated with a low-code web-based platform that allows easy access for users and simple updates for the authors. The main goal of this project was to educate material scientists on the applications having the most impact on the supply chain and current industrial specifications/markets for each application. The objective was to provide material researchers with harmonized information so that they could gain a better understanding of the market, focus their technologies on an application with a high potential for commercialization, and better contribute to supply-chain risk reduction. While the goal was met with high receptivity, several limitations stemmed from query design, distribution platform, and quality of data source. To overcome some of these limitations and expand on CMCD's potential, we are building a public webpage with an improved interface, better data organization, and higher extensibility.

36 MATERIALS SCIENCE↗

Site Characterization of the Highest-Priority Geologic Formations for CO2 Storage in Wyoming

The project Site Characterization of the Highest-Priority Geologic Formations for CO2 Storage in Wyoming is one of 9 site characterization projects that were implemented as part of ARRA (American Recovery and Reinvestment Act). Data from this project was used to improve resolution of data in NATCARB in the area of study. Data related to this study has already been incorporated in NATCARB Atlas. The Wyoming Carbon Underground Storage Project (WY-CUSP) consisted of CO2 storage site characterization and evaluation, focusing on Wyoming’s most promising CO2 storage reservoirs (the Pennsylvanian Weber/Tensleep Sandstone and Mississippian Madison Limestone) and premier CO2 storage site (Rock Springs Uplift). Results from the WY-CUSP project suggest the two reservoirs could store up to 17,000 million tons of CO2. The WY-CUSP team drilled a stratigraphic test well and acquired a 3-D seismic survey covering 25 square miles of the Rock Springs Uplift site. The team retrieved 916 feet of core from the 12,810-foot-deep well, along with a complete log suite, borehole images, fluid samples, and other data. Project partners (1) provided continuous visual documentation of the core, including grain size, mineralogy, facies distribution, and porosity; (2) performed continuous permeability and velocity scans of selected reservoir intervals; and (3) chemically analyzed the fluid samples. WY-CUSP scientists integrated seismic attributes with observations from log suites, a VSP survey, core, fluid samples, and laboratory analyses, including continuous permeability scans. From these integrations, researchers constructed 3-D spatial distribution volumes of reservoir and seal properties that represent geological heterogeneity at the targeted CO2 storage site. The WY-CUSP team used this data to perform new CO2 plume migration simulations. Baker Hughes, Inc., completed a series of small-scale, in-situ water injectivity measurements. A database was formed when observations, analyses, and experiments from the stratigraphic test well were integrated. Correlation of these data allowed petrophysical parameters to be extrapolated from the test well out into the storage domain (5x5 mile 3-D seismic survey volume). This resulted in an improved, realistic understanding of performance assessments for potential CO2 storage scenarios. The WY-CUSP team worked on (1) improving CO2 storage resource estimates, (2) establishing long-term integrity and permanence of confining layers, (3) designing a profitable strategy for pressure management, and (4) evaluating the utilization of stored CO2 at the Rock Spring Uplift. Finally, Baker Hughes developed a microseismic baseline for the test site using in-bore geophones to complete field operations.

3-D seismic↗

FAST.Farm load validation for single wake situations at alpha ventus

The main objective of the presented work is the validation of the simulation tool FAST.Farm for the calculation of power and structural loads in single wake situations; the basis for the validation is the measurement database of the operating offshore wind farm alpha ventus. The approach is described in detail and covers the calibration of the aeroelastic turbine model, transfer of environmental conditions to simulations, and comparison between simulations and adequately filtered measurements. It is shown that FAST.Farm accurately predicts power and structural load distributions over wind direction with discrepancies of less than 10 % for most of the cases compared to the measurements. Additionally, the frequency response of the structure is investigated, and it is calculated by FAST.Farm in good agreement with the measurements. In general, the calculation of fatigue loads is improved with a wake-added turbulence model added to FAST.Farm in the course of this study.

17 WIND ENERGY↗