Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graphic processing units”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

TAMM: Tensor algebra for many-body methods

Tensor algebra operations such as contractions in computational chemistry consume a significant fraction of the computing time on large-scale computing platforms. The widespread use of tensor contractions between large multi-dimensional tensors in describing electronic structure theory has motivated the development of multiple tensor algebra frameworks targeting heterogeneous computing platforms. In this paper, we present Tensor Algebra for Many-body Methods (TAMM), a framework for productive and performance-portable development of scalable computational chemistry methods. TAMM decouples the specification of the computation from the execution of these operations on available high-performance computing systems. With this design choice, the scientific application developers (domain scientists) can focus on the algorithmic requirements using the tensor algebra interface provided by TAMM, whereas high-performance computing developers can direct their attention to various optimizations on the underlying constructs, such as efficient data distribution, optimized scheduling algorithms, and efficient use of intra-node resources (e.g., graphics processing units). The modular structure of TAMM allows it to support different hardware architectures and incorporate new algorithmic advances. We describe the TAMM framework and our approach to the sustainable development of scalable ground- and excited-state electronic structure methods. We present case studies highlighting the ease of use, including the performance and productivity gains compared to other frameworks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Massively parallel GPU enabled third-order cluster perturbation excitation energies for cost-effective large scale excitation energy calculations

We present here a massively parallel implementation of the recently developed CPS(D-3) excitation energy model that is based on cluster perturbation theory. The new algorithm extends the one developed in Baudin et al. [J. Chem. Phys., 150, 134110 (2019)] to leverage multiple nodes and utilize graphical processing units for the acceleration of heavy tensor contractions. Furthermore, we show that the extended algorithm scales efficiently with increasing amounts of computational resources and that the developed code enables CPS(D-3) excitation energy calculations on large molecular systems with a low time-to-solution. More specifically, calculations on systems with over 100 atoms and 1000 basis functions are possible in a few hours of wall clock time. This establishes CPS(D-3) excitation energies as a computationally efficient alternative to those obtained from the coupled-cluster singles and doubles model.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Porting fragmentation methods to GPUs using an OpenMP API: Offloading the resolution-of-the-identity second-order Møller–Plesset perturbation method

Here, using an OpenMP Application Programming Interface, the resolution-of-the-identity second-order Møller–Plesset perturbation (RI-MP2) method has been off-loaded onto graphical processing units (GPUs), both as a standalone method in the GAMESS electronic structure program and as an electron correlation energy component in the effective fragment molecular orbital (EFMO) framework. First, a new scheme has been proposed to maximize data digestion on GPUs that subsequently linearizes data transfer from central processing units (CPUs) to GPUs. Second, the GAMESS Fortran code has been interfaced with GPU numerical libraries (e.g., NVIDIA cuBLAS and cuSOLVER) for efficient matrix operations (e.g., matrix multiplication, matrix decomposition, and matrix inversion). The standalone GPU RI-MP2 code shows an increasing speedup of up to 7.5× using one NVIDIA V100 GPU with one IBM 42-core P9 CPU for calculations on fullerenes of increasing size from 40 to 260 carbon atoms using the 6-31G(d)/cc-pVDZ-RI basis sets. A single Summit node with six V100s can compute the RI-MP2 correlation energy of a cluster of 175 water molecules using the correlation consistent basis sets cc-pVDZ/cc-pVDZ-RI containing 4375 atomic orbitals and 14 700 auxiliary basis functions in ~0.85 h. In the EFMO framework, the GPU RI-MP2 component shows near linear scaling for a large number of V100s when computing the energy of an 1800-atom mesoporous silica nanoparticle in a bath of 4000 water molecules. The parallel efficiencies of the GPU RI-MP2 component with 2304 and 4608 V100s are 98.0% and 96.1%, respectively.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

GPU acceleration of local and semilocal density functional calculations in the SPARC electronic structure code

We present a Graphics Processing Unit (GPU)-accelerated version of the real-space SPARC electronic structure code for performing Kohn–Sham density functional theory calculations within the local density and generalized gradient approximations. In particular, we develop a modular math-kernel based implementation for NVIDIA architectures wherein the computationally expensive operations are carried out on the GPUs, with the remainder of the workload retained on the central processing units (CPUs). Here, using representative bulk and slab examples, we show that relative to CPU-only execution, GPUs enable speedups of up to 6× and 60× in node and core hours, respectively, bringing time to solution down to less than 30 s for a metallic system with over 14 000 electrons and enabling significant reductions in computational resources required for a given wall time.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Distributed memory, GPU accelerated Fock construction for hybrid, Gaussian basis density functional theory

With the growing reliance of modern supercomputers on accelerator-based architecture such a graphics processing units (GPUs), the development and optimization of electronic structure methods to exploit these massively parallel resources has become a recent priority. While significant strides have been made in the development GPU accelerated, distributed memory algorithms for many modern electronic structure methods, the primary focus of GPU development for Gaussian basis atomic orbital methods has been for shared memory systems with only a handful of examples pursing massive parallelism. Here in this work, we present a set of distributed memory algorithms for the evaluation of the Coulomb and exact exchange matrices for hybrid Kohn–Sham DFT with Gaussian basis sets via direct density-fitted (DF-J-Engine) and seminumerical (sn-K) methods, respectively. The absolute performance and strong scalability of the developed methods are demonstrated on systems ranging from a few hundred to over one thousand atoms using up to 128 NVIDIA A100 GPUs on the Perlmutter supercomputer.

97 MATHEMATICS AND COMPUTING↗

GPAW: An open Python package for electronic structure calculations

We review the GPAW open-source Python package for electronic structure calculations. GPAW is based on the projector-augmented wave method and can solve the self-consistent density functional theory (DFT) equations using three different wave-function representations, namely real-space grids, plane waves, and numerical atomic orbitals. The three representations are complementary and mutually independent and can be connected by transformations via the real-space grid. This multi-basis feature renders GPAW highly versatile and unique among similar codes. By virtue of its modular structure, the GPAW code constitutes an ideal platform for the implementation of new features and methodologies. Moreover, it is well integrated with the Atomic Simulation Environment (ASE), providing a flexible and dynamic user interface. In addition to ground-state DFT calculations, GPAW supports many-body GW band structures, optical excitations from the Bethe–Salpeter Equation, variational calculations of excited states in molecules and solids via direct optimization, and real-time propagation of the Kohn–Sham equations within time-dependent DFT. A range of more advanced methods to describe magnetic excitations and non-collinear magnetism in solids are also now available. In addition, GPAW can calculate non-linear optical tensors of solids, charged crystal point defects, and much more. Recently, support for graphics processing unit (GPU) acceleration has been achieved with minor modifications to the GPAW code thanks to the CuPy library. We end the review with an outlook, describing some future plans for GPAW.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cluster perturbation theory X: A parallel implementation of Lagrangian perturbation series for the coupled cluster singles and doubles ground-state energy through fifth order

We describe an efficient implementation of cluster perturbation and Møller-Plesset Lagrangian energy series through fifth order that target the coupled cluster singles and doubles energy utilizing the resolution of the identity approximation. We illustrate the computational performance of the implementation by performing ground state energy calculations on systems with up to 1200 basis functions using a single node and by comparison to conventional CCSD calculations. We further show that our hybrid MPI/OMP parallel implementation that also utilizes graphical processing units can be used to obtain fifth order energies on systems with almost 1200 basis functions with a 90 minute "time to solution" running on Frontier at Oak Ridge National Laboratory.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Force Field X: A computational microscope to study genetic variation and organic crystals using theory and experiment

Force Field X (FFX) is an open-source software package for atomic resolution modeling of genetic variants and organic crystals that leverages advanced potential energy functions and experimental data. FFX currently consists of nine modular packages with novel algorithms that include global optimization via a many-body expansion, acid–base chemistry using polarizable constant-pH molecular dynamics, estimation of free energy differences, generalized Kirkwood implicit solvent models, and many more. Applications of FFX focus on the use and development of a crystal structure prediction pipeline, biomolecular structure refinement against experimental datasets, and estimation of the thermodynamic effects of genetic variants on both proteins and nucleic acids. The use of Parallel Java and OpenMM combines to offer shared memory, message passing, and graphics processing unit parallelization for high performance simulations. Overall, the FFX platform serves as a computational microscope to study systems ranging from organic crystals to solvated biomolecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

LibERI—A portable and performant multi-GPU accelerated library for electron repulsion integrals via OpenMP offloading and standard language parallelism

A portable and performant graphics processing unit (GPU)-accelerated library for electron repulsion integral (ERI) evaluation, named LibERI, has been developed and implemented via directive-based (e.g., OpenMP and OpenACC) and standard language parallelism (e.g., Fortran DO CONCURRENT). Offloaded ERIs consist of integrals over low and high contraction s, p, and d functions using the rotated-axis and Rys quadrature methods. GPU codes are factorized based on previous developments with two layers of integral screening and quartet presorting. In this work, the density screening is moved to the GPU to enhance the computational efficacy for large molecular systems. Here, the L-shells in the Pople basis set are also separated into pure S and P shells to increase the ERI homogeneity and reduce atomic operations and the memory footprint. LibERI is compatible with any quantum chemistry drivers supporting the MolSSI Driver Interface. Benchmark calculations of LibERI interfaced with the GAMESS software package were carried out on various GPU architectures and molecular systems. The results show that the LibERI performance is comparable to other state-of-the-art GPU-accelerated codes (e.g., TeraChem and GMSHPC) and, in some cases, outperforms conventionally developed ERI CUDA kernels (e.g., QUICK) while fully maintaining portability.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

3-center and 4-center 2-particle Gaussian AO integrals on modern accelerated processors

We report an implementation of the McMurchie–Davidson (MD) algorithm for 3-center and 4-center 2-particle integrals over Gaussian atomic orbitals (AOs) with low and high angular momenta l and varying degrees of contraction for graphical processing units (GPUs). This work builds upon our recent implementation of a matrix form of the MD algorithm that is efficient for GPU evaluation of 4-center 2-particle integrals over Gaussian AOs of high angular momenta (l ≥ 4) [A. Asadchev and E. F. Valeev, J. Phys. Chem. A 127, 10889–10895 (2023)]. The use of unconventional data layouts and three variants of the MD algorithm allow for the evaluation of integrals with double precision and sustained performance between 25% and 70% of the theoretical hardware peak. Performance assessment includes integrals over AOs with l ≤ 6 (a higher l is supported). Preliminary implementation of the Hartree–Fock exchange operator is presented and assessed for computations with up to a quadruple-zeta basis and more than 20 000 AOs. The corresponding C++ code is part of the experimental open-source LibintX library available at https://github.com/ValeevGroup/libintx.

Chemistry↗

Extending GPU-accelerated Gaussian integrals in the TeraChem software package to f type orbitals: Implementation and applications

Here, the increasing availability of graphics processing units (GPUs) for scientific computing has prompted interest in accelerating quantum chemical calculations through their use. However, the complexity of integral kernels for high angular momentum basis functions often limits the utility of GPU implementations with large basis sets or for metal containing systems. In this work, we report the implementation of f function support in the GPU-accelerated TeraChem software package through the development of efficient kernels for the evaluation of Hamiltonian integrals. The high efficiency of the resulting code is demonstrated through density functional theory (DFT) calculations on increasingly large organic molecules and transition metal complexes, as well as coupled cluster singles and doubles calculations on water clusters. Preliminary investigations into Ni(I) catalysis with DFT and the photochemistry of MnH(CH 3 ) with complete active space self-consistent field are also carried out. Overall, our GPU-accelerated software appears to be well-suited for fast simulation of large transition metal containing systems, as well as organic molecules.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Polariton spectra under the collective coupling regime. II. 2D non-linear spectra

In our previous work [Mondal et al., J. Chem. Phys. 162, 014114 (2025)], we developed several efficient computational approaches to simulate exciton–polariton dynamics described by the Holstein–Tavis–Cummings (HTC) Hamiltonian under the collective coupling regime. Here, we incorporated these strategies into the previously developed Lindblad-partially linearized density matrix (⁠$\mathscr{L}$-PLDM) approach for simulating 2D electronic spectroscopy (2DES) of exciton–polariton under the collective coupling regime. In particular, we apply the efficient quantum dynamics propagation scheme developed in Paper I to both the forward and the backward propagations in the PLDM and develop an efficient importance sampling scheme and graphics processing unit vectorization scheme that allow us to reduce the computational costs from $\mathscr{O}$($\mathscr{K}$ 2 )$\mathscr{O}$(T 3 ) to $\mathscr{O}$($\mathscr{K}$)$\mathscr{O}$(T 0 ) for the 2DES simulation, where $\mathscr{K}$ is the number of states and T is the number of time steps of propagation. As a result, we further simulated the 2DES for an HTC Hamiltonian under the collective coupling regime and analyzed the signal from both rephasing and non-rephasing contributions of the ground state bleaching, excited state emission, and stimulated emission pathways.

2D non-linear spectra↗

Initial position optimization in molecular dynamics simulations for a Coulomb system

A new algorithm for molecular dynamics (MD) simulations is developed to optimize plasma particle distributions at given initial temperatures. By combining velocity scaling and reassignment, the method effectively eliminates the initial rise and oscillation in temperatures observed with randomly distributed positions. These rises and oscillations are undesired numerical artifacts observed in conventional plasma MD simulations, arising from unoptimized particle positions. The algorithm demonstrates temperature relaxation without initial rises or oscillations, as well as precise flow velocity relaxation, enabling accurate measurement of relaxation times. The code is accelerated using graphics processing units for parallel processing, enhancing the study of plasma dynamics. The proposed method for distributing physically valid particles in MD simulations enables accurate studies of intrinsic collision processes in plasmas, including the dynamics of strongly coupled plasmas, plasma–wave interactions, and transport phenomena in magnetized plasmas. The paper concludes with a discussion of potential applications and future enhancements to the algorithm.

Jo, Jawon (ORCID:0009000924193285)↗

pyRMG: A framework for high-throughput, large-cell DFT calculations on supercomputers

Exascale computing delivers the raw power to simulate ever larger and more chemically realistic systems, but realizing this potential requires codes that can efficiently use thousands of processors. Our real-space multigrid (RMG) density functional theory (DFT) code’s grid-decomposition approach scales nearly linearly with the number of graphics processing units (GPUs), even for simulations exceeding thousands of atoms. This scalability makes RMG a compelling tool for high-throughput DFT studies of materials that would otherwise be bottlenecked in other codes (for example, by global fast Fourier transforms in plane-wave DFT). However, the limited workflow infrastructure for RMG has thus far constrained its adoption to a small user community. In this work, we present pyRMG, a Python package designed to streamline the setup and execution of RMG DFT calculations. Built on the pymatgen and ASE (Atomic Simulation Environment) computational materials science Python packages, pyRMG automates input generation and convergence checking, and it integrates with modern job schedulers (e.g., Flux) on leadership-class platforms such as Frontier and Perlmutter. Here, we demonstrate pyRMG for a high-throughput study of strain effects in 2D 2L-Bi 2 Se 3 /2L-NbSe 2 heterostructures, which offers chemical insights into this system and shows that RMG-based workflows can converge with limited user intervention.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

SHarmonic: A fast and accurate implementation of spherical harmonics for electronic-structure calculations

The authors present SHarmonic, a new implementation of the spherical harmonics targeted for electronic-structure calculations. Their approach is to use explicit formulas for the harmonics written in terms of normalized Cartesian coordinates. This approach results in a code that is as precise as other implementations while being at least one order of magnitude more computationally efficient. The library can run on graphics processing units as well, achieving an additional order of magnitude in execution speed. This new implementation is simple to use and is provided under an open-source license; it can be readily used by other codes to avoid the error-prone and cumbersome implementation of the spherical harmonics.

Mathematics and Computing↗

Ground and excited state gradients with end-to-end differentiable semiempirical quantum chemistry

Accurate and efficient gradients of molecular energy with respect to nuclear degrees of freedom are essential for geometry optimization and molecular dynamics, including simulations that go beyond the Born–Oppenheimer regime. A common approach involves deriving analytical formulas for new electronic structure methods, which is often conceptually difficult and requires tedious coding. Here, we implement analytical, semi-numerical, and automatic differentiation (AD)-based gradient pathways for semiempirical Hamiltonian models in the PYSEQM software package, leveraging both graphics processing unit (GPU) and central processing unit (CPU) architectures. We further extend these capabilities to excited states calculated using the configuration interaction singles and time-dependent Hartree–Fock ansätze. We benchmark wall time, peak memory usage, and accuracy across three molecular families of varying chemical complexity, including systems of up to a thousand atoms. For ground-state simulations, analytical and AD gradients achieve near-identical GPU runtimes, while semi-numerical gradients are slower on GPU but remain competitive on CPU. For excited states, both analytical and custom AD approaches using implicit differentiation show similar performance and low memory requirements, whereas gradients with full AD are memory-limited. AD gradients match analytical ones in accuracy across all tested systems, aided by a quaternion-based diatomic frame rotation for two-center quantities that ensures smooth energy surfaces. Overall, automatic differentiation emerges as a practical alternative to analytical gradients in semiempirical quantum chemistry, offering high accuracy while allowing seamless integration in AI-driven workflows and popular packages, such as PyTorch and JAX. Our results provide actionable guidance for selecting optimal gradient strategies in large-scale ground- and excited-state molecular dynamics simulations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cardinal: A Lower-Length-Scale Multiphysics Simulator for Pebble-Bed Reactors

This paper demonstrates a multiphysics solver for pebble-bed reactors, in particular, for Berkeley’s pebble-bed -fluoride-salt-cooled high-temperature reactor (PB-FHR) (Mark I design). The FHR is a class of advanced nuclear reactors that combines the robust coated particle fuel form from high-temperature gas-cooled reactors, the direct reactor auxiliary cooling system passive decay removal of liquid-metal fast reactors, and the transparent, high-volumetric heat capacitance liquid-fluoride salt working fluids (e.g., FLiBe) from molten salt reactors. This fuel and coolant combination enables FHRs to operate in a high-temperature, low-pressure design space that has beneficial safety and economic implications. The PB-FHR relies on a pebble-bed approach, and pebble-bed reactors are, in a sense, the poster child for multiscale analysis. Relying heavily on the MultiApp capability of the Multiphysics Object-Oriented Simulation Environment (MOOSE), we have developed Cardinal, a new platform for lower-length-scale simulation of pebble-bed cores. The lower-length-scale simulator comprises three physics: neutronics (OpenMC), thermal fluids (Nek5000/NekRS), and fuel performance (BISON). Cardinal tightly couples all three physics and leverages advances in MOOSE, such as the MultiApp system and the concept of MOOSE-wrapped applications. Moreover, Cardinal can utilize graphics processing units for accelerating solutions. In this paper, we discuss the development of Cardinal and the verification and validation and demonstration simulations.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Data-Driven RANS Turbulence Closures for Forced Convection Flow in Reactor Downcomer Geometry

Recent progress in data-driven turbulence modeling has shown its potential to enhance or replace traditional equation-based Reynolds-averaged Navier-Stokes (RANS) turbulence models. Here, this work utilizes invariant neural network (NN) architectures to model Reynolds stresses and turbulent heat fluxes in forced convection flows (when the models can be decoupled). As the considered flow is statistically one dimensional, the invariant NN architecture for the Reynolds stress model reduces to the linear eddy viscosity model. To develop the data-driven models, direct numerical and RANS simulations in vertical planar channel geometry mimicking a part of the reactor downcomer are performed. Different conditions and fluids relevant to advanced reactors (sodium, lead, unitary-Prandtl-number fluid, and molten salt) constitute the training database. The models enabled accurate predictions of velocity and temperature, and compared to the baseline k–τ turbulence model with the simple gradient diffusion hypothesis, do not require tuning of the turbulent Prandtl number. The data-driven framework is implemented in the open-source graphics processing unit–accelerated spectral element solver nekRS and has shown the potential for future developments and consideration of more complex mixed convection flows.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗