Engineering PapersSearch

SEARCH · Engineering Papers

Results for “strong scaling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

COOL-LAMPS. VII. Quantifying Strong-lens Scaling Relations with 177 Cluster-scale Strong Gravitational Lenses in DECaLS

Abstract We estimate the Einstein-radius-enclosed total mass for 177 cluster-scale strong gravitational lenses identified by the ChicagO Optically selected Lenses Located At the Margins of Public Surveys (COOL-LAMPS) collaboration with lens redshifts ranging from 0.2 ⪅ z ⪅ 1.0 using the brightest-cluster-galaxy (BCG) redshift and an observable proxy for the Einstein radius. We constrain the Einstein-radius-enclosed luminosity and stellar mass by fitting parametric spectral energy distributions to aperture photometry from the Dark Energy Camera Legacy Survey (DECaLS) in the g -, r -, and z -band Dark Energy Camera filters. We find that the BCG redshift, enclosed total mass, and enclosed luminosity are strongly correlated and well described by a planar relationship in 3D space. We find that the enclosed total mass and stellar mass are correlated with a logarithmic slope of 0.50 0 − 0.031 + 0.029 , and the enclosed total mass and stellar-to-total mass fraction are correlated with a logarithmic slope of − 0.49 5 − 0.033 + 0.032 . In tandem with the small radii within which these slopes are constrained, this may suggest invariance in baryon conversion efficiency and feedback strength as a function of cluster-centric radii in galaxy clusters. Additionally, the correlations described here should have utility in ranking strong-lensing candidates in upcoming imaging surveys—such as Rubin/Legacy Survey of Space and Time—in which an algorithmic treatment of strong lenses will be needed due to the sheer volume of data these surveys will produce.

Mork, Simon D. (ORCID:0000000255739131)

Analyzing line-of-sight selection biases in galaxy-scale strong lensing with external convergence and shear

The upcoming Vera Rubin Observatory Legacy Survey of Space and Time (LSST) will dramatically increase the number of strong gravitational lensing systems, requiring precise modeling of line-of-sight (LOS) effects to mitigate biases in lensing observations and cosmological inferences. We develop a method to construct joint distributions of external convergence (κ ext ) and shear (γ ext ) for strong lensing LOS by aggregating large-scale structure simulations with high-resolution halo renderings and non-linear correction. Our approach captures both smooth background matter and perturbations from halos, enabling accurate modeling of LOS effects. Here, we apply non-linear LOS corrections to κ ext and γ ext that address the non-additive lensing effects caused by objects along the LOS in strong lensing. We find that, with a minimum image separation of 1.0'', non-linear LOS correction due to the presence of a dominant deflector slightly increases the ratio of quadruple to double lenses; this non-linear LOS correction also introduces systematic biases of ∼ 0.1% for galaxy-AGN lens in the inferred Hubble constant (H 0 ) if not accounted for. We also observe a 0.66% bias for galaxy-galaxy lenses on H 0 , and even larger biases 1.02% for galaxy-AGN systems if LOS effects are not accounted for. These results highlight the importance of LOS for precision cosmology. The publicly available code and datasets provide tools for incorporating LOS effects in future analyses.

Hubble constant

GPU-enabled extreme-scale turbulence simulations: Fourier pseudo-spectral algorithms at the exascale using OpenMP offloading

Fourier pseudo-spectral methods for nonlinear partial differential equations are of wide interest in many areas of advanced computational science, including direct numerical simulation of three-dimensional (3-D) turbulence governed by the Navier-Stokes equations in fluid dynamics. This paper presents a new capability for simulating turbulence at a new record resolution up to 35 trillion grid points, on the world's first exascale computer, Frontier, comprising AMD MI250x GPUs with HPE's Slingshot interconnect and operated by the US Department of Energy's Oak Ridge Leadership Computing Facility (OLCF). Key programming strategies designed to take maximum advantage of the machine architecture involve performing almost all computations on the GPU which has the same memory capacity as the CPU, performing all-to-all communication among sets of parallel processes directly on the GPU, and targeting GPUs efficiently using OpenMP offloading for intensive number-crunching including 1-D Fast Fourier Transforms (FFT) performed using AMD ROCm library calls. With 99% of computing power on Frontier being on the GPU, leaving the CPU idle leads to a net performance gain via avoiding the overhead of data movement between host and device except when needed for some I/O purposes. Memory footprint including the size of communication buffers for MPI_ALLTOALL is managed carefully to maximize the largest problem size possible for a given node count. Detailed performance data including separate contributions from different categories of operations to the elapsed wall time per step are reported for five grid resolutions, from 2048 3 on a single node to 32768 3 on 4096 or 8192 nodes out of 9408 on the system. Both 1D and 2D domain decompositions which divide a 3D periodic domain into slabs and pencils respectively are implemented. The present code suite (labeled by the acronym GESTS, GPUs for Extreme Scale Turbulence Simulations) achieves a figure of merit (in grid points per second) exceeding goals set in the Center for Accelerated Application Readiness (CAAR) program for Frontier. The performance attained is highly favorable in both weak scaling and strong scaling, with notable departures only for 2048 3 where communication is entirely intra-node, and for 32768 3 , where a challenge due to small message sizes does arise. Communication performance is addressed further using a lightweight test code that performs all-to-all communication in a manner matching the full turbulence simulation code. Performance at large problem sizes is affected by both small message size due to high node counts as well as dragonfly network topology features on the machine, but is consistent with official expectations of sustained performance on Frontier. Overall, although not perfect, the scalability achieved at the extreme problem size of 32768 3 (and up to 8192 nodes — which corresponds to hardware rated at just under 1 exaflop/sec of theoretical peak computational performance) is arguably better than the scalability observed using prior state-of-the-art algorithms on Frontier's predecessor machine (Summit) at OLCF. New science results for the study of intermittency in turbulence enabled by this code and its extensions are to be reported separately in the near future.

3D fast Fourier transform

OpenACC offloading of the MFC compressible multiphase flow solver on AMD and NVIDIA GPUs

GPUs are the heart of the latest generations of supercomputers. We efficiently accelerate a compressible multiphase flow solver via OpenACC on NVIDIA and AMD Instinct GPUs. Optimization is accomplished by specifying the directive clauses gang vector and collapse. Further speedups of six and ten times are achieved by packing user-defined types into coalesced multidimensional arrays and manual inlining via metaprogramming. Additional optimizations yield seven-times speedup of array packing and thirty-times speedup of select kernels on Frontier. Weak scaling efficiencies of 97% and 95% are observed when scaling to 50% of Summit and 87% of Frontier. Strong scaling efficiencies of 84% and 81% are observed when increasing the device count by a factor of 8 and 16 on V100 and MI250X hardware. The strong scaling efficiency of AMD’s MI250X increases to 92% when increasing the device count by a factor of 16 when GPU-aware MPI is used for communication.

Wilfong, Benjamin

Comprehensive Neural Posterior Estimation for Galaxy-Galaxy Strong Lensing

We present a deep learning model based on neural posterior estimation (NPE) for comprehensive extraction of astrophysical parameters from galaxy-scale strong gravitational lenses. The unprecedentedly large amount of galaxy-scale strong lenses expected in future cosmological surveys (${\cal O}(10^5)$) promises to enable valuable statistical constraints in various studies ranging from galaxy formation to the nature of dark matter, but it also poses a significant challenge for traditional modelling pipelines. To this end, our automated model includes several new, state-of-the-art features and approaches leveraging the framework of simulation-based inference (SBI). We infer a total of 20 parameters describing the mass and light profiles of both lens and source galaxies, using simulated raw multi-band data modelled under noise and observing conditions expected by the Legacy Survey of Space and Time (LSST), with its summary statistics generated by a residual network. We examine the efficacy of multi-band data in extracting nearly 20 model parameters simultaneous from strong lensing images including lens light. Finally, We perform a comprehensive set of diagnostics for SBI models, evaluating the model's prediction accuracy, stability, and uncertainty quantification.

Zhao, Roy J. [Chicago U., KICP]

Refactoring the elastic–viscous–plastic solver from the sea ice model CICE v6.5.1 for improved performance

This study focuses on the performance of the elastic–viscous–plastic (EVP) dynamical solver within the sea ice model, CICE v6.5.1. The study has been conducted in two steps. First, the standard EVP solver was extracted from CICE for experiments with refactored versions, which are used for performance testing. Second, one refactored version was integrated and tested in the full CICE model to demonstrate that the new algorithms do not significantly impact the physical results. The study reveals two dominant bottlenecks, namely (1) the number of Message Parsing Interface (MPI) and Open Multi-Processing (OpenMP) synchronization points required for halo exchanges during each time step combined with the irregular domain of active sea ice points and (2) the lack of single-instruction, multiple-data (SIMD) code generation. The standard EVP solver has been refactored based on two generic patterns. The first pattern exposes how general finite differences on masked multi-dimensional arrays can be expressed in order to produce significantly better code generation by changing the memory access pattern from random access to direct access. The second pattern takes an alternative approach to handle static grid properties. The measured single-core performance improvement is more than a factor of 5 compared to the standard implementation. The refactored implementation of strong scales on the Intel® Xeon® Scalable Processors series node until the available bandwidth of the node is used. For the Intel® Xeon® CPU Max series, there is sufficient bandwidth to allow the strong scaling to continue for all the cores on the node, resulting in a single-node improvement factor of 35 over the standard implementation. This study also demonstrates improved performance on GPU processors.

58 GEOSCIENCES

Runtime performance of a GAMESS quantum chemistry application offloaded to GPUs

Summary Computational chemistry is at the forefront of solving urgent societal problems, such as polymer upcycling and carbon capture. The complexity of modeling these processes at appropriate length and time scales is mainly manifested in the number and types of chemical species involved in the reactions and may require models of several thousand atoms and large basis sets to accurately capture the chemical complexity and heterogeneity in the physical and chemical processes. The quantum chemistry package General Atomic and Molecular Electronic Structure System (GAMESS) has a wide array of methods that can efficiently and accurately treat complex chemical systems. In this work, we have used the GAMESS Effective Fragment Molecule Orbital (EFMO) method for electronic structure calculation of a challenging mesoporous silica nanoparticle (MSN) model surrounded by about 4700 water molecules to investigate the strong scaling and GPU offloading on hybrid CPU‐GPU nodes. Experiments were performed on the Perlmutter platform at the National Energy Research Scientific Computing Center. Good strong scaling and load balancing have been observed on up to 88 hybrid nodes for different settings of the execution parameters for the calculation considered here. When GPUs are oversubscribed by offloading work from multiple CPU processes, using the NVIDIA multi‐process service (MPS) has consistently reduced time to solution and energy consumed. Additionally, for some configuration parameter settings, oversubscription with MPS improved performance by up to 5.8% over the case without oversubscription.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Multi-GPU porting of a phase-change cascaded lattice Boltzmann method for three-dimensional pool boiling simulations

The Lattice Boltzmann method (LBM) has proven effective in simulating phase-change phenomena, such as melting, solidification, evaporation, and boiling. In this work, we develop a highly parallelized multi-GPU implementation of LBM for three-dimensional pool boiling simulations. The code is based on the OpenACC programming model, which enables the code to be deployed efficiently on multi-core CPUs, GPUs, and potentially other accelerators, without the need for architecture-specific rewrites. To support large-scale simulations, the domain is decomposed and distributed across multiple compute nodes using MPI. We demonstrate that the code exhibits excellent scaling properties, with ideal strong-scaling running with up to 256 GPUs on the MareNostrum5 cluster.

97 MATHEMATICS AND COMPUTING

Static Subspace Approximation for Random Phase Approximation Correlation Energies: Implementation and Performance

Developing theoretical understanding of complex reactions and processes at interfaces requires using methods that go beyond semilocal density functional theory to accurately describe the interactions between solvent, reactants and substrates. Methods based on many-body perturbation theory, such as the random phase approximation (RPA), have previously been limited due to their computational complexity. However, this is now a surmountable barrier due to the advances in computational power available, in particular through modern GPU-based supercomputers. In this work, we describe the implementation of RPA calculations within BerkeleyGW and show its favorable computational performance on large complex systems relevant for catalysis and electrochemistry applications. Our implementation builds off of the static subspace approximation which, by employing a compressed representation of the frequency dependent polarizability, enables the evaluation of the RPA correlation energy with significant acceleration and systematically controllable accuracy. We find that the computational cost of calculating the RPA correlation energy scales only linearly with system size for systems containing up to 50 thousand bands, and is expected to scale quadratically thereafter. We also show excellent strong scaling results across several supercomputers, demonstrating the performance and portability of this implementation.

algorithmic development

Nek5000/RS performance on advanced GPU architectures

The authors explore performance scalability of the open-source thermal-fluids code, NekRS, on the U.S. Department of Energy's leadership computers, Crusher, Frontier, Summit, Perlmutter, and Polaris. Particular attention is given to analyzing performance and time-to-solution at the strong-scale limit for a target efficiency of 80%, which is typical for production runs on the DOE's high-performance computing systems. Several examples of anomalous behavior are also discussed and analyzed.

97 MATHEMATICS AND COMPUTING

TorchBraid: High-Performance Layer-Parallel Training of Deep Neural Networks with MPI and GPU Acceleration

TorchBraid is a high-performance implementation of layer-parallel training for deep neural networks (DNNs) supporting MPI-based parallelism and GPU acceleration. Layer-parallel training has been developed to overcome the serialization inherent in forward and backward propagation of DNNs that limits utilization of computational resources in the strong scaling limit. To achieve this, TorchBraid integrates the PyTorch neural network framework with the state-of-the-art XBraid time-parallel library. Furthermore, this article presents the use and performance of TorchBraid, in addition to solutions for overcoming the algorithmic challenges inherent in combining automatic differentiation with layer-parallel. Results are presented with and without GPU acceleration for the Tiny ImageNet and MNIST image classification data sets, as well as recurrent neural networks. Overall, TorchBraid enables fast training of DNNs, both in a strong and weak scaling context. In addition to the TorchBraid software, several new advances in applying layer-parallel algorithms are detailed. Integration of layer-parallel with data-parallel algorithms is presented for the first time, showing the computational advantages of the combination. Standard deep learning techniques, like batch-normalization, are developed for layer-parallel training. Finally, a new approach combining layer-parallel with spatial coarsening in order to accelerate training for 3D image classification shows roughly a 10× speedup over serial execution.

Layer-parallel

Sensitivity of tropical orographic precipitation to wind speed with implications for future projections

Abstract. Some of the rainiest regions on Earth lie upstream of tropical mountains, where the interaction of prevailing winds with orography produces frequent precipitating convection. Yet the response of tropical orographic precipitation to the large-scale wind and temperature variations induced by anthropogenic climate change remains largely unconstrained. Here, we quantify the sensitivity of tropical orographic precipitation to background cross-slope wind using theory, idealized simulations, and observations. We build on a recently developed theoretical framework that characterizes the orographic enhancement of seasonal mean precipitation, relative to upstream regions, as a response of convection to cooling and moistening of the lower free troposphere by stationary orographic gravity waves. Using this framework and convection-permitting simulations, we show that higher cross-slope wind speeds deepen the penetration of the cool and moist gravity wave perturbation upstream of orography, resulting in a mean rainfall increase of 20 % (m s−1)−1 to 30 % (m s−1)−1 increase in cross-slope wind speed. Additionally, we show that orographic precipitation in five tropical regions exhibits a similar dependence on changes in cross-slope wind at both seasonal and daily timescales. Given next-century changes in large-scale winds around tropical orography projected by global climate models, this strong scaling rate implies wind-induced changes in some of Earth's rainiest regions that are comparable with any produced directly by increases in global mean temperature and humidity.

Nicolas, Quentin (ORCID:000000024116973X)

An experimental investigation of the trailing edge noise mechanism

An experimental investigation has been conducted to understand the physical mechanism of noise generation from a turbulent wall jet discharging over a flat plate and interacting with its sharp trailing edge. An aspect ratio 10 rectangular nozzle was used to provide the wall jet. Measurements made consist of farfield noise, surface pressure fluctuations, turbulent velocity fluctuations, and two-point space-time cross-correlations among these quantities. Results are presented which suggest strongly that the generation mechanism is the interaction of the convecting large scale quasi-orderly disturbance in the upper free shear layer of the wall jet with the trailing edge. The interaction also excites large scale strong vortical motion in the trailing edge wake. The dominant part of the sound field is highly coherent and in phase opposition across the trailing edge.

Yu, J. C.

Breaking the mold: Overcoming the time constraints of molecular dynamics on general-purpose hardware

The evolution of molecular dynamics (MD) simulations has been intimately linked to that of computing hardware. For decades following the creation of MD, simulations have improved with computing power along the three principal dimensions of accuracy, atom count (spatial scale), and duration (temporal scale). Since the mid-2000s, computer platforms have, however, failed to provide strong scaling for MD, as scale-out central processing unit (CPU) and graphics processing unit (GPU) platforms that provide substantial increases to spatial scale do not lead to proportional increases in temporal scale. Important scientific problems therefore remained inaccessible to direct simulation, prompting the development of increasingly sophisticated algorithms that present significant complexity, accuracy, and efficiency challenges. While bespoke MD-only hardware solutions have provided a path to longer timescales for specific physical systems, their impact on the broader community has been mitigated by their limited adaptability to new methods and potentials. In this work, we show that a novel computing architecture, the Cerebras wafer scale engine, completely alters the scaling path by delivering unprecedentedly high simulation rates up to 1.144 M steps/s for 200 000 atoms whose interactions are described by an embedded atom method potential. This enables direct simulations of the evolution of materials using general-purpose programmable hardware over millisecond timescales, dramatically increasing the space of direct MD simulations that can be carried out. In this paper, we provide an overview of advances in MD over the last 60 years and present our recent result in the context of historical MD performance trends.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

GR-Athena++: General-relativistic Magnetohydrodynamics Simulations of Neutron Star Spacetimes

We present the extension of GR-Athena++ to general-relativistic magnetohydrodynamics (GRMHD) for applications to neutron star spacetimes. The new solver couples the constrained transport implementation of Athena++ to the Z4c formulation of the Einstein equations to simulate dynamical spacetimes with GRMHD using oct-tree adaptive mesh refinement. We consider benchmark problems for isolated and binary neutron star spacetimes demonstrating stable and convergent results at relatively low resolutions and without grid symmetries imposed. The code correctly captures magnetic field instabilities in nonrotating stars with total relative violation of the divergence-free constraint of 10 –16 . It handles evolutions with a microphysical equation of state and black hole formation in the gravitational collapse of a rapidly rotating star. For binaries, we demonstrate correctness of the evolution under the gravitational radiation reaction and show convergence of gravitational waveforms. We showcase the use of adaptive mesh refinement to resolve the Kelvin–Helmholtz instability at the collisional interface in a merger of magnetised binary neutron stars. GR-Athena++ shows strong scaling efficiencies above 80% in excess of 10 5 CPU cores and excellent weak scaling is shown up to ~5 × 10 5 CPU cores in a realistic production setup. GR-Athena++ allows for the robust simulation of GRMHD flows in strong and dynamical gravity with exa-scale computers.

79 ASTRONOMY AND ASTROPHYSICS

Simulating many-engine spacecraft: Exceeding 1 quadrillion degrees of freedom via information geometric regularization

We present an optimized implementation of the recently proposed information geometric regularization (IGR) for unprecedented scale simulation of compressible fluid flows applied to multi-engine spacecraft boosters. We improve upon state-of-the-art computational fluid dynamics (CFD) techniques in terms of computational cost, memory footprint, and energy-to-solution metrics. Unified memory on coupled CPU–GPU or APU platforms increases problem size with negligible overhead. Mixed half/single-precision storage and computation are used on well-conditioned numerics. We simulate flow at 200 trillion grid points and 1 quadrillion degrees of freedom, exceeding the current record by a factor of 20. A factor of 4 wall-time speedup is achieved over optimized baselines. Ideal weak scaling is observed on OLCF Frontier, LLNL El Capitan, and CSCS Alps using the full systems. Strong scaling is near ideal at extreme conditions, including 80% efficiency on CSCS Alps with an 8 node baseline and stretching to the full system.

Wilfong, Benjamin [Georgia Institute of Technology

Mars Global Surveyor Thermal Emission Spectrometer (TES) Observations: Atmospheric Temperatures During Aerobraking and Science Phasing

Between September 1997, when the Mars Global Surveyor spacecraft arrived at Mars, and September 1998 when the final aerobraking phase of the mission began, the Thermal Emission Spectrometer (TES) has acquired an extensive data set spanning approximately half of a Martian year. Nadir-viewing spectral measurements from this data set within the 15-micrometers CO2 absorption band are inverted to obtain atmospheric temperature profiles from the surface up to about the 0.1 mbar level. The computational procedure used to retrieve the temperatures is presented. Mean meridional cross sections of thermal structure are calculated for periods of time near northern hemisphere fall equinox, winter solstice, and spring equinox, as well as for a time interval immediately following the onset of the Noachis Terra dust storm. Gradient thermal wind cross sections are calculated from the thermal structure. Regions of possible wave activity are identified using cross sections of rms temperature deviations from the mean. Results from both near-equinox periods show some hemispheric asymmetry with peak eastward thermal winds in the north about twice the magnitude of those in the south. The results near solstice show an intense circumpolar vortex at high northern latitudes and waves associated with the vortex jet core. Warming of the atmosphere aloft at mid-northern latitudes suggests the presence of a strong cross-equatorial Hadley circulation. Although the Noachis dust storm did not become global in scale, strong perturbations to the atmospheric structure are found, including an enhanced temperature maximum aloft at high northern latitudes resulting from intensification of the Hadley circulation. TES results for the various seasonal conditions are compared with published results from Mars general circulation models, and generally good qualitative agreement is found.

Conrath, Barney J.

Open quantum system approach to inclusive jet production in heavy-ion collisions

We derive a factorization formula for inclusive jet production in heavy-ion collisions using the tools of Effective Field Theory (EFT). We show how physics at widely separated scales in this process can be systematically separated by matching to EFTs at successively lower virtualities. Owing to a strong scale separation, we recover a vacuum-like DGLAP evolution above the jet scale, while the additional low-energy scales induced by the medium effectively probe the internal structure of the jet. As a result, the cross section can be written as a series with an increasing number of subjets characterized by perturbative matching coefficients each of which is convolved with a distinct function. These functions encode broadening, medium-induced radiations as well as quantum interference such as the Landau-Pomeranchuk-Migdal effect and color coherence dynamics to all orders in perturbation theory. As a first application of this EFT framework, we investigate the case of an unresolved jet and show how the cross section can be factorized and fully separate the jet dynamics from the universal physics of the medium. To compare to the existing literature, we explicitly compute the medium jet function at next-to-leading order in the coupling and leading order in medium opacity.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS