Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “strong scaling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Examination of entrainment, combustibility, and heat transfer due to cell venting inside simplified battery energy storage enclosures

Battery energy storage systems (ESS) assembled of modules containing lithium-ion batteries can pose fire and explosion hazards during thermal runaway because vented gases may accumulate, combust, or escape enclosures. Here, we propose a prototypical configuration composed of the main components of an ESS rack, where modules are represented as rectangles within an enclosure and vent gas is injected steadily through the top of a module. This idealized geometry allows for exploration of how rack-level characteristics can influence smaller single-cell or module-level scales. Through this geometry, we conduct numerical simulations to understand how variations in the vent velocity, temperature, and vertical location affect entrainment rates, fraction of uncombusted fuel, rich-gas coverage, and heat flux to other modules. At lower vent velocities and vertical positions, buoyancy drives enough entrainment and mixing to react with most of the vent gases on the converging plenum. As jet momentum increases and/or buoyancy decreases via higher vent position in the rack, vent gases can propagate into and burn in the divergence plenum. Heat flux to the impinged module scales strongly with vent momentum and temperature, while heat transfer to upper modules peaks when flames in the diverging plenum reach the surface and declines when combustion is suppressed.

Computational fluid dynamics↗

Parallel quantum computing simulations via quantum accelerator platform virtualization

Quantum circuit execution is a central task in quantum computation. Due to inherent quantum-mechanical constraints, quantum computing workflows often involve a considerable number of independent measurements over a large set of slightly different quantum circuits. Here we discuss a simple model for parallelizing such quantum circuit executions that is based on introducing a large array of virtual quantum processing units (mapped to HPC nodes in our case) as a parallel quantum computing platform. Implemented within the XACC framework, the model can readily take advantage of its backend-agnostic features, enabling parallel quantum computing/simulation over any target backend supported by XACC. We illustrate the performance of this approach by demonstrating strong scaling in two pertinent domain science problems, namely in computing the gradients for the multi-contracted variational quantum eigensolver and in data-driven quantum circuit learning, where we vary the number of qubits and the number of circuit layers. Here, the latter simulation leverages the cuQuantum library to run efficiently on GPU-accelerated HPC platforms.

97 MATHEMATICS AND COMPUTING↗

Accelerating high-order continuum kinetic plasma simulations using multiple GPUs

Kinetic plasma simulations solve the Vlasov-Poisson or Vlasov-Maxwell equations to evolve scalar-variable distribution functions in position-velocity phase space and vector-variable electromagnetic fields in configuration space. The immense computational cost of evolving high-dimensional variables, and their large number of degrees of freedom, often limits the utility of continuum kinetic simulations and presents a challenge when it comes to accurately simulating real-world physical phenomena. To address this challenge, we present techniques that accelerate and minimize the computational work required for a scalable Vlasov-Poisson solver. We show theoretical hardware compute and communication bounds for solving a fourth-order finite-volume Vlasov-Poisson system. These bounds are then used to inform and evaluate the design of performance portable algorithms for a multiple graphics processing unit (GPU) accelerated version of the Vlasov-Poisson solver VCK-CPU [1]. We demonstrate that the multi-GPU Vlasov solver implementation, VCK-GPU, simultaneously minimizes required inter-process data transfer while also being bounded by the machine network performance limits. This results in an overall strong scaling speedup per timestep of up to 40x in three-dimensional phase space (one position, two velocity coordinates) and 54x in four dimensional phase space (two position, two velocity coordinates) and a 341x increase in simulation throughput of the GPU accelerated code over the existing CPU code. The GPU code is also able to weak scale up to 256 compute nodes and 1024 GPUs. In conclusion, we demonstrate that the improved compute performance enables exploring configurations which were previously computationally infeasible, including resolving fine-scale distribution function filamentation and multi-species dynamics with realistic electron-proton mass ratios.

Continuum kinetics↗

A High-Performance Discrete-Element Framework for Simulating Flow and Jamming of Moisture Bearing Biomass Feedstocks

We developed and verified a high-performance open-source discrete element method (DEM) solver with simultaneously-supported feedstock-specific interaction models, including bonded-sphere, liquid bridge, cohesion, and non-linear contact models. Our solver uses parallel data structures on hybrid central and graphics processing unit (CPU/GPU) architectures, with favorable strong scaling performance observed for large problem sizes comprised of (100 M particles), and 4X single-node GPU speedup. The particles for corn stover feedstock were conceptualized and calibrated based on experimental measurements and results. Sensitivity analyses demonstrate that the mass flow rate from a wedge hopper is governed primarily by moisture content, friction coefficient, and cohesion energy density. The model is used to reproduce experimentally observed hopper jamming results, highlighting that the experimental no-flow trends can only be achieved by using non-spherical particles, liquid bridge and cohesion models, highlighting the importance of using concurrent feedstock specialized models for the effective representation of biomass material handling problems.

bioenergy↗

Cross slip of extended dislocations in face-centered cubic metals through phase-field modeling

Cross slip is a dislocation mechanism that significantly impacts the mechanical behavior of engineering alloys. Here, in this work, we advance a 3D phase-field dislocation dynamics (PFDD) mesoscale technique to simulate cross slip across a broad range of face-centered cubic (FCC) metals. The formulation incorporates elastic anisotropy, an FCC numerical grid, and a high-fidelity representation of the entire γ -surface from density functional theory for eight FCC metals and no adjustable parameters or rules. The relaxed core structures under zero stress for all metals are predicted to extend in plane. The analytical model for stacking fault width agrees well with the PFDD result under the assumption of elastic isotropy but overestimates it under elastic anisotropy, when the degree of anisotropy is large. The dynamic simulations are designed to elucidate the material parameters that influence the propensity for cross slip. Whether cross slip occurs under a non-Schmid stress or to bypass a hard obstacle, the critical stress to cross slip scales strongly with the anisotropic energy coefficient for a screw dislocation.

36 MATERIALS SCIENCE↗

Physics-inspired spatiotemporal-graph AI ensemble for the detection of higher order wave mode signals of spinning binary black hole mergers

We present a new class of AI models for the detection of quasi-circular, spinning, non-precessing binary black hole mergers whose waveforms include the higher order gravitational wave modes ($\ell$, |m|) = {(2,2), (2,1), (3,3), (3,2), (4,4)}, and mode mixing effects in the $\ell$ = 3, |m| = 2 harmonics. These AI models combine hybrid dilated convolution neural networks to accurately model both short- and long-range temporal sequential information of gravitational waves; and graph neural networks to capture spatial correlations among gravitational wave observatories to consistently describe and identify the presence of a signal in a three detector network encompassing the Advanced LIGO and Virgo detectors. We first trained these spatiotemporal-graph AI models using synthetic noise, using 1.2 million modeled waveforms to densely sample this signal manifold, within 1.7 h using 256 NVIDIA A100 GPUs in the Polaris supercomputer at the Argonne Leadership Computing Facility. This distributed training approach exhibited optimal classification performance, and strong scaling up to 512 NVIDIA A100 GPUs. With these AI ensembles we processed data from a three detector network, and found that an ensemble of 4 AI models achieves state-of-the-art performance for signal detection, and reports two misclassifications for every decade of searched data. We distributed AI inference over 128 GPUs in the Polaris supercomputer and 128 nodes in the Theta supercomputer, and completed the processing of a decade of gravitational wave data from a three detector network within 3.5 h. Finally, we fine-tuned these AI ensembles to process the entire month of February 2020, which is part of the O3b LIGO/Virgo observation run, and found 6 gravitational waves, concurrently identified in Advanced LIGO and Advanced Virgo data, and zero false positives. This analysis was completed in one hour using one NVIDIA A100 GPU.

79 ASTRONOMY AND ASTROPHYSICS↗

Modeling the differential rate for signal interactions in coincidence with noise fluctuations or large rate backgrounds

The characteristic energy of a relic dark matter interaction with a detector scales strongly with the putative dark matter mass. Consequently, experimental search sensitivity at the lightest masses will always come from interactions whose size is similar to noise fluctuations and low energy backgrounds in the detector. In this paper, we correctly calculate the net change in measured differential event rate due to low rate signal interactions that overlap in time with noise and backgrounds, accounting for both periods of time when the signal is coincident with noise/backgrounds and for the decreased amount of time in which only noise/backgrounds occur and we show that the introduction of random simulated signal events into the continuous raw data stream (a form of “salting”) provides a correct and practical implementation to estimate this net linear signal sensitivity. Unfortunately this does not apply to the situation with non negligible pileups as we show through explicit examples. Consequently, this exclusion technique should be complemented by less sensitive techniques that can separately exclude large rate, high pileup signal parameter space. An analysis threshold above the Gaussian-like noise component of the measured differential rate spectrum should also likely produce conservative limits. Though previous light mass dark matter searches did not correctly account for decreased background only live time effects in their signal sensitivity, conservative analysis choices likely protected their published limits from being nonconservative.

Li, Xinran↗

FuseIM: Fusing Probabilistic Traversals for Influence Maximization on Exascale Systems

Probabilistic breadth-first traversals (BPTs) are used in many network science and graph machine learning applications. In this paper, we are motivated by the application of BPTs in stochastic diffusion-based graph problems such as influence maximization. These applications heavily rely on BPTs to implement a Monte-Carlo sampling step for their approximations. Given the large sampling complexity, stochasticity of the diffusion process, and the inherent irregularity in real-world graph topologies, efficiently parallelizing these BPTs remains significantly challenging. In this paper, we present a new algorithm to fuse massive number of concurrently executing BPTs with random starts on the input graph. Our algorithm is designed to fuse BPTs by combining separate traversals into a unified frontier on distributed multi-GPU systems. To show the general applicability of the fused BPT technique, we have incorporated it into two state-of-the-art influence maximization parallel implementations (gIM and Ripples). Our experiments on up to 4K nodes of the OLCF Frontier supercomputer (32,768 GPUs and 196K CPU cores) show strong scaling behavior, and that fused BPTs can improve the performance of these implementations up to 34x (for gIM) and ~360x (for Ripples).

Neff, Reece W.↗

Should We Conserve Entropy or Energy when Computing CAPE with Mixed-Phase Precipitation Physics?

Abstract The rapidly increasing resolution of global atmospheric reanalysis and climate model datasets necessitates finding methods for computing convective available potential energy (CAPE) both efficiently and accurately. To this end, this article compares two common methods for computing CAPE which conserve either energy or entropy. Inaccuracies in these computations arise from both physical and numerical errors. For instance, computing CAPE with entropy conserved results in physical errors from nonequilibrium phase transitions but minimizes numerical errors because solutions are analytic at each height. In contrast, computing CAPE with energy conserved avoids these physical errors, but accumulates numerical errors that are grid-resolution-dependent because the numerical integration of a differential equation is required. Analysis of CAPE computed with large databases of soundings from the tropical Amazon and midlatitude storm environments shows that physical errors from the entropy method are typically 1%–3% as large as CAPE, which is comparable to the numerical errors from conserving energy with grid spacing of 25 and 250 m using explicit first-order and second-order integration schemes, respectively. Errors in entropy-based CAPE calculations are also insensitive to vertical grid spacing, in contrast to energy-based calculations whose error strongly scales with the grid spacing. It is shown that entropy-based methods are advantageous when intercomparing datasets with differing vertical resolution because they produce accurate and reasonably fast results that are insensitive to grid resolution, whereas a second-order energy-based method is advantageous when analyzing data with a consistent vertical resolution because of its superior computational efficiency. Significance Statement Convective available potential energy (CAPE) is a measure of instability in the atmosphere that helps forecasters and researchers understand when and where thunderstorms will form. The purpose of this article is to identify the most efficient and accurate methods for computing CAPE. Two methods are considered here, one that relates to the entropy (a measure of thermodynamic disorder) of an air parcel and one that relates to the energy of an air parcel. Results indicate that the entropy method is most accurate and insensitive to the resolution of the data used for the calculation (which can vary considerably), whereas the energy method uses the least computation time.

Peters, John M.↗

Strong Scalability Analysis of the Albany Land Ice code on HPC Architectures

Scalability is a critical factor in High-Performance Computing (HPC), where optimizing resource usage has a direct impact on cost-effectiveness and time-efficiency. This report presents a strong scaling performance study of the Albany Land Ice (ALI) code across different HPC architectures, towards determining the best configuration to use when running large-scale simulation ensembles.

97 MATHEMATICS AND COMPUTING↗

Latitudinal variability of large-scale coronal temperature and its association with the density and the global magnetic field

In this paper we utilize the latitiude distribution of the coronal temperature during the period 1984-1992 that was derived in a paper by Guhathakurta et al, 1993, utilizing ground-based intensity observations of the green (5303 A Fe XIV) and red (6374 A Fe X) coronal forbidden lines from the National Solar Observatory at Sacramento Peak, and establish it association with the global magnetic field and the density distributions in the corona. A determination of plasma temperature, T, was estimated from the intensity ratio Fe X/Fe XIV (where T is inversely proportional to the ratio), since both emission lines come from ionized states of Fe, and the ratio is only weakly dependent on density. We observe that there is a large-scale organization of the inferred coronal temperature distribution that is associated with the large-scale, weak magnetic field structures and bright coronal features; this organization tends to persist through most of the magnetic activity cycle. These high-temperature structures exhibit time-space characteristics which are similar to those of the polar crown filaments. This distribution differs in spatial and temporal characterization from the traditional picture of sunspot and active region evolution over the range of the sunspot cycle, which are manifestations of the small-scale, strong magnetic field regions.

Guhathakurta, M.↗

The Sloan Lens ACS Survey. I. A Large Spectroscopically Selected Sample of Massive Early-Type Lens Galaxies

The Sloan Lens ACS (SLACS) Survey is an efficient Hubble Space Telescope (HST) Snapshot imaging survey for new galaxy-scale strong gravitational lenses. The targeted lens candidates are selected spectroscopically from the Sloan Digital Sky Survey (SDSS) database of galaxy spectra for having multiple nebular emission lines at a redshift significantly higher than that of the SDSS target galaxy. The SLACS survey is optimized to detect bright early-type lens galaxies with faint lensed sources in order to increase the sample of known gravitational lenses suitable for detailed lensing, photometric, and dynamical modeling. In this paper, the first in a series on the current results of our HST Cycle 13 imaging survey, we present a catalog of 19 newly discovered gravitational lenses, along with nine other observed candidate systems that are either possible lenses, nonlenses, or nondetections. The survey efficiency is thus >=68%. We also present Gemini 8 m and Magellan 6.5 m integral-field spectroscopic data for nine of the SLACS targets, which further support the lensing interpretation. A new method for the effective subtraction of foreground galaxy images to reveal faint background features is presented. We show that the SLACS lens galaxies have colors and ellipticities typical of the spectroscopic parent sample from which they are drawn (SDSS luminous red galaxies and quiescent MAIN sample galaxies), but are somewhat brighter and more centrally concentrated. Several explanations for the latter bias are suggested. The SLACS survey provides the first statistically significant and homogeneously selected sample of bright early-type lens galaxies, furnishing a powerful probe of the structure of early-type galaxies within the half-light radius. The high confirmation rate of lenses in the SLACS survey suggests consideration of spectroscopic lens discovery as an explicit science goal of future spectroscopic galaxy surveys.

galaxies↗

Does Aerosol Weaken or Strengthen the South Asian Monsoon?

Aerosols are known to have the ability to block off solar radiation reaching the earth surface, causing it to cool - the so-called solar dimming (SDM) effect. In the Asian monsoon region, the SDM effect by aerosol can produce differential cooling at the surface reducing the meridional thermal contrast between land and ocean, leading to a weakening of the monsoon (Ramanathan et al. 2005). On the other hand, absorbing aerosols such as black carbon and dust, when forced up against the steep slopes of the southern Tibetan Plateau can produce upper tropospheric heating, and induce convection-dynamic feedback leading to an advance of the rainy season over northern India and an enhancement of the South Asian monsoon through the "Elevated Heat Pump" (EHP) effect (Lau et al. 2006). In this paper, we present modeling results showing that in a coupled ocean-atmosphere-land system in which concentrations of greenhouse gases are kept constant, the response of the South Asian monsoon to dust and black carbon forcing is the net result of the two opposing effects of SDM and EHP. For the South Asian monsoon, if the increasing upper tropospheric thermal contrast between the Tibetan Plateau and region to the south spurred by the EHP overwhelms the reduction in surface temperature contrast due to SDM, the monsoon strengthens. Otherwise, the monsoon weakens. Preliminary observations are consistent with the above findings. We find that the two effects are strongly scale dependent. On interannual and shorter time scales, the EHP effect appears to dominate in the early summer season (May-June). On decadal or longer time scales, the SDM dominates for the mature monsoon (July-August). Better understanding the physical mechanisms underlying the SDM and the EHP effects, the local emission and transport of aerosols from surrounding deserts and arid-regions, and their interaction with monsoon water cycle dynamics are important in providing better prediction and assessment of climate change impacts on precipitation of the Asian monsoon land regions.

Lau, William K. M.↗

Does Aerosol Weaken or Strengthen the South Asian Monsoon?

Aerosols are known to have the ability to block off solar radiation reaching the earth surface, causing it to cool - the so-called solar dimming (SDM) effect. In the Asian monsoon region, the SDM effect by aerosol can produce differential cooling at the surface reducing the meridional thermal contrast between land and ocean, leading to a weakening of the monsoon. On the other hand, absorbing aerosols such as black carbon and dust, when forced up against the steep slopes of the southern Tibetan Plateau can produce upper tropospheric heating, and induce convection-dynamic feedback leading to an advance of the rainy season over northern India and an enhancement of the South Asian monsoon through the "Elevated Heat Pump" (EHP) effect. In this paper, we present modeling results showing that in a coupled ocean-atmosphere-land system in which concentrations of greenhouse gases are kept constant, the response of the South Asian monsoon to dust and black carbon forcing is the net result of the two opposing effects of SDM and EHP. For the South Asian monsoon, if the increasing upper tropospheric thermal contrast between the Tibetan Plateau and region to the south spurred by the EHP overwhelms the reduction in surface temperature contrast due to SDM, the monsoon strengthens. Otherwise, the monsoon weakens. Preliminary observations are consistent with the above findings. We find that the two effects are strongly scale dependent. On interannual and shorter time scales, the EHP effect appears to dominate in the early summer season (May-June). On decadal or longer time scales, the SDM dominates for the mature monsoon (July-August). Better understanding the physical mechanisms underlying the SDM and the EHP effects, the local emission and transport of aerosols from surrounding deserts and arid-regions, and their interaction with monsoon water cycle dynamics are important in providing better prediction and assessment of climate change impacts on precipitation of the Asian monsoon land regions.

Lau, William K.↗

Parallel Adaptive High-Order CFD Simulations Characterizing Cavity Acoustics for the Complete SOFIA Aircraft

This paper presents one-of-a-kind MPI-parallel computational fluid dynamics simulations for the Stratospheric Observatory for Infrared Astronomy (SOFIA). SOFIA is an airborne, 2.5-meter infrared telescope mounted in an open cavity in the aft of a Boeing 747SP. These simulations focus on how the unsteady flow field inside and over the cavity interferes with the optical path and mounting of the telescope. A temporally fourth-order Runge-Kutta, and spatially fifth-order WENO-5Z scheme was used to perform implicit large eddy simulations. An immersed boundary method provides automated gridding for complex geometries and natural coupling to a block-structured Cartesian adaptive mesh refinement framework. Strong scaling studies using NASA's Pleiades supercomputer with up to 32,000 cores and 4 billion cells shows excellent scaling. Dynamic load balancing based on execution time on individual AMR blocks addresses irregularities caused by the highly complex geometry. Limits to scaling beyond 32K cores are identified, and targeted code optimizations are discussed.

Acoustics↗

Early Lessons on Combining Lidar and Multi‑baseline SAR Measurements for Forest Structure Characterization

The estimation and monitoring of 3D forest structure at large scales strongly rely on the use of remote sensing techniques. Today, two of them are able to provide 3D forest structure estimates: lidar and synthetic aperture radar (SAR) configurations. The differences in wavelength, imaging geometry, and technical implementation make the measurements pro-vided by the two configurations different and, when it comes to the sensitivity to individual 3D forest structure components, complementary. Accordingly, the potential of combining lidar and SAR measurements toward an improved 3D forest structure estimation has been recognised from the very beginning. However, until today there is no established frame-work for this combination. This paper attempts to review differences, commonalities, and complementarities of lidar and SAR measurements. First, vertical lidar reflectance and SAR reflectivity profiles at different wavelengths are compared in different forest types. Then, current perspectives on their combination for the generation of enhanced structure products are discussed. Two promising frameworks for combining lidar and SAR measurements are reviewed. The first one is a model-based framework where lidar-derived parameters are used to initialize SAR scattering models, and relies on both the validity of the models and on the physical equivalence of the used lidar and SAR parameters. The second one is a structure-based framework based on the ability of lidar and SAR measurements to express physical forest structure by means of appropriate indices. These indices can then be used to establish a link between the two kind of measurements. The review is supported by experimental results achieved using space- and airborne data acquired in recent relevant mission and campaigns.

Matteo Pardini↗

Uncertainties, Limits, and Benefits of Climate Change Mitigation for Soil Moisture Drought in Southwestern North America

Over the last two decades, southwestern North America (SWNA) has been in the grip of one of the most severe droughts of the last 1,200 years, with one third to nearly one half of its severity attributable to climate change. We analyze how the risk of extreme soil moisture droughts in SWNA, analogous to the most severe 21-year (≥ in magnitude to 2000–2020) and single-year (≥ in magnitude to 2002) events of the last several decades, changes in projections from Phase 6 of the Coupled Model Intercomparison Project. By the end of the 21st century, SWNA experiences robust (R ≥ 0.80) soil moisture drying and substantial increases in extreme single-year drought risk that scale strongly with warming, spanning an 8%–26% probability of occurrence across +2–4 K. Notably, our results show that 21-year droughts analogous to 2000–2020 are up to 5 times more likely than extreme single-year droughts under all levels of warming (≈50%). These high levels of 21-year drought risk are largely invariant across scenarios because of large spring precipitation declines in half the models, shifting SWNA into a drier mean state. Despite projections of this sweeping and ostensibly inevitable increase in 21-year drought risk, climate mitigation reduces their severity by reducing the magnitude of extreme single-year droughts during these events. Our results emphasize both the importance of preparing SWNA for imminent increases in persistent drought events and constraining projected precipitation uncertainty to better resolve future long-term drought risk.

drought↗

Density fluctuations in strong Langmuir turbulence - Scalings, spectra, and statistics

A recently developed two-component model of strong Langmuir turbulence is applied to determine the scalings, spectra, and statistics of the associated density fluctuations. The predictions are found to be in excellent agreement with extensive results from numerical solution of the Zakharov equations in two and three dimensions.

Robinson, P. A.↗