Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “GPU accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Development of $\mathrm{AMOEBA}$ Polarizable Force Field for Rare-Earth La 3+ Interaction with Bioinspired Ligands

Rare-earth metals (REMs) are crucial for many important industries, such as power generation and storage, in addition to cancer treatment and medical imaging. One promising new REM refinement approach involves mimicking the highly selective and efficient binding of REMs observed in relatively recently discovered proteins. However, realizing any such bioinspired approach requires an understanding of the biological recognition mechanisms. In this report we developed a new classical polarizable force field based on the AMOEBA framework for modeling a lanthanum ion (La 3+ ) interacting with water, acetate, and acetamide, which have been found to coordinate the ion in proteins. The parameters were derived by comparing to high-level ab initio quantum mechanical (QM) calculations that include relativistic effects. The AMOEBA model, with advanced atomic multipoles and electronic polarization, is successful in capturing both the QM distance-dependent La 3+ –ligand interaction energies and experimental hydration free energy. A new scheme for pairwise polarization damping (POLPAIR) was developed to describe the polarization energy in La 3+ interactions with both charged and neutral ligands. Simulations of La3+ in water showed water coordination numbers and ion–water distances consistent with previous experimental and theoretical findings. Water residence time analysis revealed both fast and slow kinetics in water exchange around the ion. This new model will allow investigation of fully solvated lanthanum ion–protein systems using GPU-accelerated dynamics simulations to gain insights on binding selectivity, which may be applied to the design of synthetic analogues.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Molecular Dynamics Simulation of Complex Reactivity with the Rapid Approach for Proton Transport and Other Reactions (RAPTOR) Software Package

Simulating chemically reactive phenomena such as proton transport on nanosecond to microsecond and beyond time scales is a challenging task. Ab initio methods are unable to currently access these time scales routinely, and traditional molecular dynamics methods feature fixed bonding arrangements that cannot account for changes in the system’s bonding topology. The Multiscale Reactive Molecular Dynamics (MS-RMD) method, as implemented in the Rapid Approach for Proton Transport and Other Reactions (RAPTOR) software package for the LAMMPS molecular dynamics code, offers a method to routinely sample longer time scale reactive simulation data with statistical precision. RAPTOR may also be interfaced with enhanced sampling methods to drive simulations toward the analysis of reactive rare events, and a number of collective variables (CVs) have been developed to facilitate this. Key advances to this methodology, including GPU acceleration efforts and novel CVs to model water wire formation are reviewed, along with recent applications of the method which demonstrate its versatility and robustness.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Adsorption Hysteresis Under Control: Tuning Host–Guest Interactions via a Genetic Algorithm

Mesoporous adsorbent materials offer a large volumetric capacity; however, cyclic adsorption/desorption processes in these systems often suffer from hysteresis and may require a significant pressure swing to access this capacity. To mitigate hysteresis, a proposed strategy is to include nucleation sites on the walls of the mesoporous material to facilitate droplet and bubble formation, lowering the free energy barriers to the respective phase transitions. It is unclear, however, what combination of adsorbate− adsorbent interactions and spatial patterning would be beneficial for a given application, considering that improvements to some sorption properties may come at the expense of other attributes. To understand these interconnected observables, we examine two model systems, planar-slit and cylindrical pores with tunable interaction sites, using GPU-accelerated transition matrix Monte Carlo simulations. The simulations provide a free energy map of the pressure−adsorption space in a matter of minutes, which we use to track adsorption isotherm characteristics as a function of adsorbent properties. We then leverage the rapid acquisition of simulation data to construct a genetic algorithm to iteratively modify interaction sites of the slit-pore wall to minimize the hysteresis of this system without sacrificing uptake. We find that the adsorption branch of the isotherm is easily modulated via the average host−guest interaction strength, but desorption is only adjustable if there is a suitable bubble nucleation site. Within the context of a slit-pore system, we identify relative interaction strengths and patch sizes required to gain control over both branches of the hysteresis loop.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Scalable freeform optimization of wide-aperture 3D metalenses by zoned discrete axisymmetry

We introduce a novel framework for design and optimization of 3D freeform metalenses that attains nearly linear scaling of computational cost with diameter, by breaking the lens into a sequence of radial “zones” with 𝑛-fold discrete axisymmetry, where 𝑛 increases with radius. This allows vastly more design freedom than imposing continuous axisymmetry, while avoiding the compromises of the locally periodic approximation (LPA) or scalar diffraction theory. Using a GPU-accelerated finite-difference time-domain (FDTD) solver in cylindrical coordinates, we perform full-wave simulation and topology optimization within each supra-wavelength zone. We validate our approach by designing millimeter and centimeter-scale, poly-achromatic, 3D freeform metalenses which outperform the state of the art. By demonstrating the scalability and resulting optical performance enabled by our “zoned discrete axisymmetry” (ZDA) and supra-wavelength domain decomposition, we highlight the potential of our framework to advance large-scale meta-optics and next-generation photonic technologies.

Sun, Mengdi [Wesleyan University]↗

The Role of Topography in Controlling Evapotranspiration Age

Abstract Evapotranspiration (ET) age is a key metric of water sustainability but a major unknown partly due to the extreme difficulty in modeling it. Groundwater is found to be important in ET age variations in small‐scale studies, yet our understanding is insufficient because groundwater systems are nested across scales. Here, we conducted GPU‐accelerated particle tracking with integrated hydrologic modeling to quantify the variations in ET age at a regional scale of ∼0.4 M km 2 . Simulation results reveal topography‐driven flow paths shaping the spatial and temporal patterns of ET age variations. On ridges, where root zone decoupling with deep subsurface storage, ET age is generally young, with seasonal variations dominated by meteorological conditions. In the valley bottom, ET age is generally old, with significant subseasonal variations caused by the convergence of subsurface flow paths. On hillslopes with water table depths ranging from 1 to 10 m, ET age shows strong seasonal variations caused by the connections with lateral groundwater regulated by ET demand. Our modeling approach provides insights into the basic linkages between ET age and topography at large scale. Our work highlights the perspective of multiscale studies of ET age, suggesting new field experiments to test these process connections and to determine if such linkages warrant inclusion in Earth System Models.

54 ENVIRONMENTAL SCIENCES↗

BMX: Biological modelling and interface exchange

Abstract High performance computing has a great potential to provide a range of significant benefits for investigating biological systems. These systems often present large modelling problems with many coupled subsystems, such as when studying colonies of bacteria cells. The aim to understand cell colonies has generated substantial interest as they can have strong economic and societal impacts through their roles in in industrial bioreactors and complex community structures, called biofilms, found in clinical settings. Investigating these communities through realistic models can rapidly exceed the capabilities of current serial software. Here, we introduce BMX, a software system developed for the high performance modelling of large cell communities by utilising GPU acceleration. BMX builds upon the AMRex adaptive mesh refinement package to efficiently model cell colony formation under realistic laboratory conditions. Using simple test scenarios with varying nutrient availability, we show that BMX is capable of correctly reproducing observed behavior of bacterial colonies on realistic time scales demonstrating a potential application of high performance computing to colony modelling. The open source software is available from the zenodo repository https://doi.org/10.5281/zenodo.8084270 under the BSD-2-Clause licence.

97 MATHEMATICS AND COMPUTING↗

GPAW: An open Python package for electronic structure calculations

We review the GPAW open-source Python package for electronic structure calculations. GPAW is based on the projector-augmented wave method and can solve the self-consistent density functional theory (DFT) equations using three different wave-function representations, namely real-space grids, plane waves, and numerical atomic orbitals. The three representations are complementary and mutually independent and can be connected by transformations via the real-space grid. This multi-basis feature renders GPAW highly versatile and unique among similar codes. By virtue of its modular structure, the GPAW code constitutes an ideal platform for the implementation of new features and methodologies. Moreover, it is well integrated with the Atomic Simulation Environment (ASE), providing a flexible and dynamic user interface. In addition to ground-state DFT calculations, GPAW supports many-body GW band structures, optical excitations from the Bethe–Salpeter Equation, variational calculations of excited states in molecules and solids via direct optimization, and real-time propagation of the Kohn–Sham equations within time-dependent DFT. A range of more advanced methods to describe magnetic excitations and non-collinear magnetism in solids are also now available. In addition, GPAW can calculate non-linear optical tensors of solids, charged crystal point defects, and much more. Recently, support for graphics processing unit (GPU) acceleration has been achieved with minor modifications to the GPAW code thanks to the CuPy library. We end the review with an outlook, describing some future plans for GPAW.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Toward digital design at the exascale: An overview of project ICECap

High performance computing has entered the Exascale Age. Capable of performing over 1018 floating point operations per second, exascale computers, such as El Capitan, the National Nuclear Security Administration's first, have the potential to revolutionize the detailed in-depth study of highly complex science and engineering systems. However, in addition to these kind of whole machine “hero” simulations, exascale systems could also enable new paradigms in digital design by making petascale hero runs routine. Currently, untenable problems in complex system design, optimization, model exploration, and scientific discovery could all become possible. Motivated by the challenge of uncovering the next generation of robust high-yield inertial confinement fusion (ICF) designs, project ICECap (Inertial Confinement on El Capitan) attempts to integrate multiple advances in machine learning (ML), scientific workflows, high performance computing, GPU-acceleration, and numerical optimization to prototype such a future. Built on a general framework, ICECap is exploring how these technologies could broadly accelerate scientific discovery on El Capitan. In addition to our requirements, system-level design, and challenges, we describe some of the key technologies in ICECap, including ML replacements for multiphysics packages, tools for human-machine teaming, and algorithms for multifidelity design optimization under uncertainty. As a test of our prototype pre-El Capitan system, we advance the state-of-the art for ICF hohlraum design by demonstrating the optimization of a 17-parameter National Ignition Facility experiment and show that our ML-assisted workflow makes design choices that are consistent with physics intuition, but in an automated, efficient, and mathematically rigorous fashion.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A fast, matrix-based method to perform omnidirectional pressure integration

Abstract Experimentally-measured pressure fields play an important role in understanding many fluid dynamics problems. Unfortunately, pressure fields are difficult to measure directly with non-invasive, spatially resolved diagnostics, and calculations of pressure from velocity have proven sensitive to error in the data. Omnidirectional line integration methods are usually more accurate and robust to these effects as compared to implicit Poisson equations, but have seen slower uptake due to the higher computational and memory costs, particularly in 3D domains. This paper demonstrates how omnidirectional line integration approaches can be converted to a matrix inversion problem. This novel formulation uses an iterative approach so that the boundary conditions are updated each step, preserving the convergence behavior of omnidirectional schemes while also keeping the computational efficiency of Poisson solvers. This method is implemented in Matlab and also as a GPU-accelerated code in CUDA-C++. The behavior of the new method is demonstrated on 2D and 3D synthetic and experimental data. Three-dimensional grid sizes of up to 125 million grid points are tractable with this method, opening exciting opportunities to perform volumetric pressure field estimation from 3D PIV measurements.

42 ENGINEERING↗

Omnigenous stellarators with improved ideal and kinetic ballooning stability

Omnigenity is a property of a magnetic field which ensures confinement of trapped particles. It is a necessary requirement for any high-performance stellarator. After creating an omnigenous equilibrium, one must also ensure reduced transport resulting from kinetic and magnetohydrodynamic (MHD) instabilities. To this end, we leverage the GPU-accelerated DESC optimization suite, which is used to design stable, finite-β omnigenous equilibria with poloidal, toroidal, and helical symmetry, achieving Mercier, ideal ballooning, and as a consequence, improved kinetic ballooning stability. We discover stellarators with second stability, a regime of large pressure gradient where an equilibrium becomes ideal ballooning stable, and demonstrate and explore both using theory and gyrokinetic simulations the connection between ideal and kinetic ballooning stability.

optimization↗

Enhancing predictive capabilities in fusion burning plasmas through surrogate-based optimization in core transport solvers

Abstract This work presents the PORTALS framework (Rodriguez-Fernandez et al 2022 Nucl. Fusion 62 076036), which leverages surrogate modeling and optimization techniques to enable the prediction of core plasma profiles and performance with nonlinear gyrokinetic simulations at significantly reduced cost, with no loss of accuracy. The efficiency of PORTALS is benchmarked against standard methods, and its full potential is demonstrated on a unique, simultaneous 5-channel (electron temperature, ion temperature, electron density, impurity density and angular rotation) prediction of steady-state profiles in a DIII-D ITER Similar Shape plasma with GPU-accelerated, nonlinear CGYRO (Candy et al 2016 J. Comput. Phys. 324 73–93). This paper also provides general guidelines for accurate performance predictions in burning plasmas and the impact of transport modeling in fusion pilot plants studies.

Physics↗

Performance of modern color decompositions for standard candle LHC tree amplitudes

In the last decade, developments of matrix element and phase space generators have focused on providing good efficiency and maximal flexibility and automation for a wide range of physical processes. However, as recent studies have shown, they are a major bottleneck in the established Monte Carlo event generator toolchains. With the advent of the HL-LHC and ever rising precision requirements, future developments will need to focus on computational performance, especially at intermediate to large jet multiplicities. We present the novel BlockGen family of fast matrix element algorithms that are amenable for GPU acceleration, making use of modern, minimal color decompositions. Moreover, we discuss the performance achieved for standard candle processes such as V +jets and tt̄+jets production.

Bothmann, E. [Gottingen U.]↗

Phoebe: a high-performance framework for solving phonon and electron Boltzmann transport equations

Understanding the electrical and thermal transport properties of materials is critical to the design of electronics, sensors, and energy conversion devices. Computational modeling can accurately predict material properties but, in order to be reliable, requires accurate descriptions of electron and phonon states and their interactions. While first-principles methods are capable of describing the energy spectrum of each carrier, using them to compute transport properties is still a formidable task, both computationally demanding and memory intensive, requiring integration of fine microscopic scattering details for estimation of macroscopic transport properties. To address this challenge, we present Phoebe—a newly developed software package that includes the effects of electron–phonon, phonon–phonon, boundary, and isotope scattering in computations of electrical and thermal transport properties of materials with a variety of available methods and approximations. This open source C++ code combines MPI-OpenMP hybrid parallelization with GPU acceleration and distributed memory structures to manage computational cost, allowing Phoebe to effectively take advantage of contemporary computing infrastructures. We demonstrate that Phoebe accurately and efficiently predicts a wide range of transport properties, opening avenues for accelerated computational analysis of complex crystals.

36 MATERIALS SCIENCE↗

Crystal generation using the fully differentiable pipeline and latent space optimization

We present a materials generation framework that couples a symmetry-conditioned variational autoencoder with a differentiable SO(3) power spectrum objective to steer candidates toward a specified local environment under the crystallographic constraints. In particular, we implement a fully differentiable pipeline that performs batch-wise optimization on both direct and latent crystallographic representations. Using the GPU acceleration, the implementation achieves about fivefold speed compared to our previous CPU workflow, while yielding comparable outcomes. In addition, we introduce the optimization strategy that alternatively performs optimization on the direct and latent crystal representations. This dual-level relaxation approach can effectively escape local minima defined by different objective gradients, thus increasing the success rate of generating complex structures satisfying the target local environments. This framework can be extended to systems consisting of multi-components and multi-environments, providing a scalable route to generate material structures with the target local environment.

conditional VAE↗

Global adjoint tomography—model GLAD-M25

SUMMARY Building on global adjoint tomography model GLAD-M15, we present transversely isotropic global model GLAD-M25, which is the result of 10 quasi-Newton tomographic iterations with an earthquake database consisting of 1480 events in the magnitude range 5.5 ≤ Mw ≤ 7.2, an almost sixfold increase over the first-generation model. We calculated fully 3-D synthetic seismograms with a shortest period of 17 s based on a GPU-accelerated spectral-element wave propagation solver which accommodates effects due to 3-D anelastic crust and mantle structure, topography and bathymetry, the ocean load, ellipticity, rotation and self-gravitation. We used an adjoint-state method to calculate Fréchet derivatives in 3-D anelastic Earth models facilitated by a parsimonious storage algorithm. The simulations were performed on the Cray XK7 ‘Titan’ and the IBM Power 9 ‘Summit’ at the Oak Ridge Leadership Computing Facility. We quantitatively evaluated GLAD-M25 by assessing misfit reductions and traveltime anomaly histograms in 12 measurement categories. We performed similar assessments for a held-out data set consisting of 360 earthquakes, with results comparable to the actual inversion. We highlight the new model for a variety of plumes and subduction zones.

58 GEOSCIENCES↗

General relativistic MHD simulations of non-thermal flaring in Sagittarius A*

Sgr A* exhibits regular variability in its multiwavelength emission, including daily X-ray flares and roughly continuous near-infrared (NIR) flickering. The origin of this variability is still ambiguous since both inverse Compton and synchrotron emission are possible radiative mechanisms. The underlying particle distributions are also not well constrained, particularly the non-thermal contribution. In this work, we employ the GPU-accelerated general relativistic magnetohydrodynamics code H-AMR to perform a study of flare flux distributions, including the effect of particle acceleration for the first time in high-resolution 3D simulations of Sgr A*. For the particle acceleration, we use the general relativistic ray-tracing code bhoss to perform the radiative transfer, assuming a hybrid thermal+non-thermal electron energy distribution. We extract ~60 h light curves in the sub-millimetre, NIR and X-ray wavebands, and compare the power spectra and the cumulative flux distributions of the light curves to statistical descriptions for Sgr A* flares. Our results indicate that non-thermal populations of electrons arising from turbulence-driven reconnection in weakly magnetized accretion flows lead to moderate NIR and X-ray flares and reasonably describe the X-ray flux distribution while fulfilling multiwavelength flux constraints. These models exhibit high rms per cent amplitudes, $\gtrsim 150{{\ \rm per\ cent}}$ both in the NIR and the X-rays, with changes in the accretion rate driving the 230 GHz flux variability, in agreement with Sgr A* observations.

79 ASTRONOMY AND ASTROPHYSICS↗

encore : an O ( N g2) estimator for galaxy N -point correlation functions

ABSTRACT We present a new algorithm for efficiently computing the N-point correlation functions (NPCFs) of a 3D density field for arbitrary N. This can be applied both to a discrete spectroscopic galaxy survey and a continuous field. By expanding the statistics in a separable basis of isotropic functions built from spherical harmonics, the NPCFs can be estimated by counting pairs of particles in space, leading to an algorithm with complexity $\mathcal {O}(N_\mathrm{g}^2)$ for Ng particles, or $\mathcal {O}(N_\mathrm{FFT}\log N_\mathrm{FFT})$ when using a Fast Fourier Transform with NFFT grid-points. In practice, the rate-limiting step for N > 3 will often be the summation of the histogrammed spherical harmonic coefficients, particularly if the number of radial and angular bins is large. In this case, the algorithm scales linearly with Ng. The approach is implemented in the encore code, which can compute the 3PCF, 4PCF, 5PCF, and 6PCF of a BOSS-like galaxy survey in ${\sim}100$ CPU-hours, including the corrections necessary for non-uniform survey geometries. We discuss the implementation in depth, along with its GPU acceleration, and provide practical demonstration on realistic galaxy catalogues. Our approach can be straightforwardly applied to current and future data sets to unlock the potential of constraining cosmology from the higher point functions.

79 ASTRONOMY AND ASTROPHYSICS↗

Tidal disruption discs formed and fed by stream–stream and stream–disc interactions in global GRHD simulations

When a star passes close to a supermassive black hole (BH), the BH’s tidal forces rip it apart into a thin stream, leading to a tidal disruption event (TDE). In this work, we study the post-disruption phase of TDEs in general relativistic hydrodynamics (GRHD) using our GPU-accelerated code h-amr. We carry out the first grid-based simulation of a deep-penetration TDE (β = 7) with realistic system parameters: a black hole-to-star mass ratio of 10 6 , a parabolic stellar trajectory, and a non-zero BH spin. We also carry out a simulation of a tilted TDE whose stellar orbit is inclined relative to the BH midplane. We show that for our aligned TDE, an accretion disc forms due to the dissipation of orbital energy with ~20 percent of the infalling material reaching the BH. The dissipation is initially dominated by violent self-intersections and later by stream–disc interactions near the pericentre. The self-intersections completely disrupt the incoming stream, resulting in five distinct self-intersection events separated by approximately 12 h and a flaring in the accretion rate. We also find that the disc is eccentric with mean eccentricity e ≈ 0.88. For our tilted TDE, we find only partial self-intersections due to nodal precession near pericentre. Although these partial intersections eject gas out of the orbital plane, an accretion disc still forms with a similar accreted fraction of the material to the aligned case. These results have important implications for disc formation in realistic tidal disruptions. For instance, the periodicity in accretion rate induced by the complete stream disruption may explain the flaring events from Swift J1644+57.

79 ASTRONOMY AND ASTROPHYSICS↗