Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graphics processing units”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

GPAW: An open Python package for electronic structure calculations

We review the GPAW open-source Python package for electronic structure calculations. GPAW is based on the projector-augmented wave method and can solve the self-consistent density functional theory (DFT) equations using three different wave-function representations, namely real-space grids, plane waves, and numerical atomic orbitals. The three representations are complementary and mutually independent and can be connected by transformations via the real-space grid. This multi-basis feature renders GPAW highly versatile and unique among similar codes. By virtue of its modular structure, the GPAW code constitutes an ideal platform for the implementation of new features and methodologies. Moreover, it is well integrated with the Atomic Simulation Environment (ASE), providing a flexible and dynamic user interface. In addition to ground-state DFT calculations, GPAW supports many-body GW band structures, optical excitations from the Bethe–Salpeter Equation, variational calculations of excited states in molecules and solids via direct optimization, and real-time propagation of the Kohn–Sham equations within time-dependent DFT. A range of more advanced methods to describe magnetic excitations and non-collinear magnetism in solids are also now available. In addition, GPAW can calculate non-linear optical tensors of solids, charged crystal point defects, and much more. Recently, support for graphics processing unit (GPU) acceleration has been achieved with minor modifications to the GPAW code thanks to the CuPy library. We end the review with an outlook, describing some future plans for GPAW.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cluster perturbation theory X: A parallel implementation of Lagrangian perturbation series for the coupled cluster singles and doubles ground-state energy through fifth order

We describe an efficient implementation of cluster perturbation and Møller-Plesset Lagrangian energy series through fifth order that target the coupled cluster singles and doubles energy utilizing the resolution of the identity approximation. We illustrate the computational performance of the implementation by performing ground state energy calculations on systems with up to 1200 basis functions using a single node and by comparison to conventional CCSD calculations. We further show that our hybrid MPI/OMP parallel implementation that also utilizes graphical processing units can be used to obtain fifth order energies on systems with almost 1200 basis functions with a 90 minute "time to solution" running on Frontier at Oak Ridge National Laboratory.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Force Field X: A computational microscope to study genetic variation and organic crystals using theory and experiment

Force Field X (FFX) is an open-source software package for atomic resolution modeling of genetic variants and organic crystals that leverages advanced potential energy functions and experimental data. FFX currently consists of nine modular packages with novel algorithms that include global optimization via a many-body expansion, acid–base chemistry using polarizable constant-pH molecular dynamics, estimation of free energy differences, generalized Kirkwood implicit solvent models, and many more. Applications of FFX focus on the use and development of a crystal structure prediction pipeline, biomolecular structure refinement against experimental datasets, and estimation of the thermodynamic effects of genetic variants on both proteins and nucleic acids. The use of Parallel Java and OpenMM combines to offer shared memory, message passing, and graphics processing unit parallelization for high performance simulations. Overall, the FFX platform serves as a computational microscope to study systems ranging from organic crystals to solvated biomolecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

LibERI—A portable and performant multi-GPU accelerated library for electron repulsion integrals via OpenMP offloading and standard language parallelism

A portable and performant graphics processing unit (GPU)-accelerated library for electron repulsion integral (ERI) evaluation, named LibERI, has been developed and implemented via directive-based (e.g., OpenMP and OpenACC) and standard language parallelism (e.g., Fortran DO CONCURRENT). Offloaded ERIs consist of integrals over low and high contraction s, p, and d functions using the rotated-axis and Rys quadrature methods. GPU codes are factorized based on previous developments with two layers of integral screening and quartet presorting. In this work, the density screening is moved to the GPU to enhance the computational efficacy for large molecular systems. Here, the L-shells in the Pople basis set are also separated into pure S and P shells to increase the ERI homogeneity and reduce atomic operations and the memory footprint. LibERI is compatible with any quantum chemistry drivers supporting the MolSSI Driver Interface. Benchmark calculations of LibERI interfaced with the GAMESS software package were carried out on various GPU architectures and molecular systems. The results show that the LibERI performance is comparable to other state-of-the-art GPU-accelerated codes (e.g., TeraChem and GMSHPC) and, in some cases, outperforms conventionally developed ERI CUDA kernels (e.g., QUICK) while fully maintaining portability.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

3-center and 4-center 2-particle Gaussian AO integrals on modern accelerated processors

We report an implementation of the McMurchie–Davidson (MD) algorithm for 3-center and 4-center 2-particle integrals over Gaussian atomic orbitals (AOs) with low and high angular momenta l and varying degrees of contraction for graphical processing units (GPUs). This work builds upon our recent implementation of a matrix form of the MD algorithm that is efficient for GPU evaluation of 4-center 2-particle integrals over Gaussian AOs of high angular momenta (l ≥ 4) [A. Asadchev and E. F. Valeev, J. Phys. Chem. A 127, 10889–10895 (2023)]. The use of unconventional data layouts and three variants of the MD algorithm allow for the evaluation of integrals with double precision and sustained performance between 25% and 70% of the theoretical hardware peak. Performance assessment includes integrals over AOs with l ≤ 6 (a higher l is supported). Preliminary implementation of the Hartree–Fock exchange operator is presented and assessed for computations with up to a quadruple-zeta basis and more than 20 000 AOs. The corresponding C++ code is part of the experimental open-source LibintX library available at https://github.com/ValeevGroup/libintx.

Chemistry↗

Extending GPU-accelerated Gaussian integrals in the TeraChem software package to f type orbitals: Implementation and applications

Here, the increasing availability of graphics processing units (GPUs) for scientific computing has prompted interest in accelerating quantum chemical calculations through their use. However, the complexity of integral kernels for high angular momentum basis functions often limits the utility of GPU implementations with large basis sets or for metal containing systems. In this work, we report the implementation of f function support in the GPU-accelerated TeraChem software package through the development of efficient kernels for the evaluation of Hamiltonian integrals. The high efficiency of the resulting code is demonstrated through density functional theory (DFT) calculations on increasingly large organic molecules and transition metal complexes, as well as coupled cluster singles and doubles calculations on water clusters. Preliminary investigations into Ni(I) catalysis with DFT and the photochemistry of MnH(CH 3 ) with complete active space self-consistent field are also carried out. Overall, our GPU-accelerated software appears to be well-suited for fast simulation of large transition metal containing systems, as well as organic molecules.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Polariton spectra under the collective coupling regime. II. 2D non-linear spectra

In our previous work [Mondal et al., J. Chem. Phys. 162, 014114 (2025)], we developed several efficient computational approaches to simulate exciton–polariton dynamics described by the Holstein–Tavis–Cummings (HTC) Hamiltonian under the collective coupling regime. Here, we incorporated these strategies into the previously developed Lindblad-partially linearized density matrix (⁠$\mathscr{L}$-PLDM) approach for simulating 2D electronic spectroscopy (2DES) of exciton–polariton under the collective coupling regime. In particular, we apply the efficient quantum dynamics propagation scheme developed in Paper I to both the forward and the backward propagations in the PLDM and develop an efficient importance sampling scheme and graphics processing unit vectorization scheme that allow us to reduce the computational costs from $\mathscr{O}$($\mathscr{K}$ 2 )$\mathscr{O}$(T 3 ) to $\mathscr{O}$($\mathscr{K}$)$\mathscr{O}$(T 0 ) for the 2DES simulation, where $\mathscr{K}$ is the number of states and T is the number of time steps of propagation. As a result, we further simulated the 2DES for an HTC Hamiltonian under the collective coupling regime and analyzed the signal from both rephasing and non-rephasing contributions of the ground state bleaching, excited state emission, and stimulated emission pathways.

2D non-linear spectra↗

Initial position optimization in molecular dynamics simulations for a Coulomb system

A new algorithm for molecular dynamics (MD) simulations is developed to optimize plasma particle distributions at given initial temperatures. By combining velocity scaling and reassignment, the method effectively eliminates the initial rise and oscillation in temperatures observed with randomly distributed positions. These rises and oscillations are undesired numerical artifacts observed in conventional plasma MD simulations, arising from unoptimized particle positions. The algorithm demonstrates temperature relaxation without initial rises or oscillations, as well as precise flow velocity relaxation, enabling accurate measurement of relaxation times. The code is accelerated using graphics processing units for parallel processing, enhancing the study of plasma dynamics. The proposed method for distributing physically valid particles in MD simulations enables accurate studies of intrinsic collision processes in plasmas, including the dynamics of strongly coupled plasmas, plasma–wave interactions, and transport phenomena in magnetized plasmas. The paper concludes with a discussion of potential applications and future enhancements to the algorithm.

Jo, Jawon (ORCID:0009000924193285)↗

pyRMG: A framework for high-throughput, large-cell DFT calculations on supercomputers

Exascale computing delivers the raw power to simulate ever larger and more chemically realistic systems, but realizing this potential requires codes that can efficiently use thousands of processors. Our real-space multigrid (RMG) density functional theory (DFT) code’s grid-decomposition approach scales nearly linearly with the number of graphics processing units (GPUs), even for simulations exceeding thousands of atoms. This scalability makes RMG a compelling tool for high-throughput DFT studies of materials that would otherwise be bottlenecked in other codes (for example, by global fast Fourier transforms in plane-wave DFT). However, the limited workflow infrastructure for RMG has thus far constrained its adoption to a small user community. In this work, we present pyRMG, a Python package designed to streamline the setup and execution of RMG DFT calculations. Built on the pymatgen and ASE (Atomic Simulation Environment) computational materials science Python packages, pyRMG automates input generation and convergence checking, and it integrates with modern job schedulers (e.g., Flux) on leadership-class platforms such as Frontier and Perlmutter. Here, we demonstrate pyRMG for a high-throughput study of strain effects in 2D 2L-Bi 2 Se 3 /2L-NbSe 2 heterostructures, which offers chemical insights into this system and shows that RMG-based workflows can converge with limited user intervention.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

SHarmonic: A fast and accurate implementation of spherical harmonics for electronic-structure calculations

The authors present SHarmonic, a new implementation of the spherical harmonics targeted for electronic-structure calculations. Their approach is to use explicit formulas for the harmonics written in terms of normalized Cartesian coordinates. This approach results in a code that is as precise as other implementations while being at least one order of magnitude more computationally efficient. The library can run on graphics processing units as well, achieving an additional order of magnitude in execution speed. This new implementation is simple to use and is provided under an open-source license; it can be readily used by other codes to avoid the error-prone and cumbersome implementation of the spherical harmonics.

Mathematics and Computing↗

Ground and excited state gradients with end-to-end differentiable semiempirical quantum chemistry

Accurate and efficient gradients of molecular energy with respect to nuclear degrees of freedom are essential for geometry optimization and molecular dynamics, including simulations that go beyond the Born–Oppenheimer regime. A common approach involves deriving analytical formulas for new electronic structure methods, which is often conceptually difficult and requires tedious coding. Here, we implement analytical, semi-numerical, and automatic differentiation (AD)-based gradient pathways for semiempirical Hamiltonian models in the PYSEQM software package, leveraging both graphics processing unit (GPU) and central processing unit (CPU) architectures. We further extend these capabilities to excited states calculated using the configuration interaction singles and time-dependent Hartree–Fock ansätze. We benchmark wall time, peak memory usage, and accuracy across three molecular families of varying chemical complexity, including systems of up to a thousand atoms. For ground-state simulations, analytical and AD gradients achieve near-identical GPU runtimes, while semi-numerical gradients are slower on GPU but remain competitive on CPU. For excited states, both analytical and custom AD approaches using implicit differentiation show similar performance and low memory requirements, whereas gradients with full AD are memory-limited. AD gradients match analytical ones in accuracy across all tested systems, aided by a quaternion-based diatomic frame rotation for two-center quantities that ensures smooth energy surfaces. Overall, automatic differentiation emerges as a practical alternative to analytical gradients in semiempirical quantum chemistry, offering high accuracy while allowing seamless integration in AI-driven workflows and popular packages, such as PyTorch and JAX. Our results provide actionable guidance for selecting optimal gradient strategies in large-scale ground- and excited-state molecular dynamics simulations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cardinal: A Lower-Length-Scale Multiphysics Simulator for Pebble-Bed Reactors

This paper demonstrates a multiphysics solver for pebble-bed reactors, in particular, for Berkeley’s pebble-bed -fluoride-salt-cooled high-temperature reactor (PB-FHR) (Mark I design). The FHR is a class of advanced nuclear reactors that combines the robust coated particle fuel form from high-temperature gas-cooled reactors, the direct reactor auxiliary cooling system passive decay removal of liquid-metal fast reactors, and the transparent, high-volumetric heat capacitance liquid-fluoride salt working fluids (e.g., FLiBe) from molten salt reactors. This fuel and coolant combination enables FHRs to operate in a high-temperature, low-pressure design space that has beneficial safety and economic implications. The PB-FHR relies on a pebble-bed approach, and pebble-bed reactors are, in a sense, the poster child for multiscale analysis. Relying heavily on the MultiApp capability of the Multiphysics Object-Oriented Simulation Environment (MOOSE), we have developed Cardinal, a new platform for lower-length-scale simulation of pebble-bed cores. The lower-length-scale simulator comprises three physics: neutronics (OpenMC), thermal fluids (Nek5000/NekRS), and fuel performance (BISON). Cardinal tightly couples all three physics and leverages advances in MOOSE, such as the MultiApp system and the concept of MOOSE-wrapped applications. Moreover, Cardinal can utilize graphics processing units for accelerating solutions. In this paper, we discuss the development of Cardinal and the verification and validation and demonstration simulations.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Data-Driven RANS Turbulence Closures for Forced Convection Flow in Reactor Downcomer Geometry

Recent progress in data-driven turbulence modeling has shown its potential to enhance or replace traditional equation-based Reynolds-averaged Navier-Stokes (RANS) turbulence models. Here, this work utilizes invariant neural network (NN) architectures to model Reynolds stresses and turbulent heat fluxes in forced convection flows (when the models can be decoupled). As the considered flow is statistically one dimensional, the invariant NN architecture for the Reynolds stress model reduces to the linear eddy viscosity model. To develop the data-driven models, direct numerical and RANS simulations in vertical planar channel geometry mimicking a part of the reactor downcomer are performed. Different conditions and fluids relevant to advanced reactors (sodium, lead, unitary-Prandtl-number fluid, and molten salt) constitute the training database. The models enabled accurate predictions of velocity and temperature, and compared to the baseline k–τ turbulence model with the simple gradient diffusion hypothesis, do not require tuning of the turbulent Prandtl number. The data-driven framework is implemented in the open-source graphics processing unit–accelerated spectral element solver nekRS and has shown the potential for future developments and consideration of more complex mixed convection flows.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Verification and Validation of Spectral Element Code for Supercritical CO2 Flow in Vertical Heated Tubes

The investigation of heat transfer in supercritical CO2 (sCO2) has garnered considerable attention in recent decades, given sCO2's potential as a promising working fluid for advanced power conversion cycles. Despite previous research efforts, there are still gaps in our understanding of sCO2 heat transfer, particularly in conditions associated with heat transfer deterioration. To delve into sCO2 heat transfer more comprehensively, we propose employing the high-fidelity computational fluid dynamics code NekRS to simulate sCO2 flow using the large eddy simulation technique. Through graphics processing unit acceleration, NekRS achieves a higher computational speed than traditional CPU-based systems. However, before using NekRS in practical applications involving sCO2, it is imperative to perform verification and validation. Here, this paper presents our efforts to verify and validate the NekRS code's capability for simulating sCO2 using heated vertical tubes, where heat transfer deterioration usually happens. To accommodate the unique properties of sCO2, we have modified the NekRS code by integrating third-party property modules, such as REFPROP and PROPATH. Our simulations are compared with experimental and numerical data from the literature, instilling confidence in leveraging NekRS for future engineering applications. Our simulations also reveal that the accuracy of the property module significantly impacts the results, with REFPROP outperforming PROPATH for sCO2 properties. Additionally, we observed that, depending on the flow direction, buoyancy can either enhance or suppress turbulence in sCO2 flow. In upward flow, under certain conditions, the suppressed turbulence leads to heat transfer deterioration, resulting in elevated wall temperatures.

NekRS↗

High-Fidelity CFD Simulation of Mixed Convection and Forced Convection in a Pebble Bed Test Reactor Core

The Hermes low-power [35-MW(thermal)] reactor will be built and operated by Kairos Power LLC (KP) to demonstrate its fluoride salt-cooled high-temperature reactor (FHR) technology. In the KP FHR, the reactor core is composed of randomly packed pebbles with TRISO fuel particles inside with FLiBe flow upward through the core acting as a coolant. Previous numerical and experimental studies have been limited to either a small-size bed or to a lack of detailed measurements for heat transfer. Here, to address the lack of high-fidelity heat transfer data in a real-size FHR core, in this study, we simulated a pebble bed core with 34 374 pebbles randomly packed, similar to the Hermes reactor's size. The core radius was 14 times that of the pebble diameter, while the core height was 45 times. In this work, we were particularly interested in a mixed convection regime, where buoyancy is important. Therefore, we performed several large-eddy simulations at different Reynolds numbers (160 to 1000) with gravitational force included. The spectral element computational fluid dynamics code NekRS with graphics processing unit acceleration was used for this study. The low-Mach number approximation was applied to address property changes in the FLiBe and to account for buoyancy. A pure hexahedral mesh with 60 million elements was generated by the Voronoi cell method. At the polynomial order of 5, the total degrees of freedom was 7.5 billion. The developed case in this work is the first of its kind in terms of size and complexity. The local numerical data across the domain were obtained and compared with empirical correlations. After examining the data, we found the following conclusions. For pressure drop, the Reger correlation predicted less than a 5% error. On the other hand, for heat transfer, the Wakao correlation outperformed the others. Based on our findings, we recommend the use of the Wakao correlation for the Nusselt number calculation, and for pressure drop, the KTA (Kerntechnischer Ausschuss) correclation, among the available experimental correlations. In conclusion, the Reger direct numerical simulation-driven correlation for pressure drops should also be considered, given its best agreement with our calculations.

Mixed Convection↗

HyKKT: a hybrid direct-iterative method for solving KKT linear systems

Here, we propose a solution strategy for the large indefinite linear systems arising in interior methods for nonlinear optimization. The method is suitable for implementation on hardware accelerators such as graphical processing units (GPUs). The current gold standard for sparse indefinite systems is the LBLT factorization where L is a lower triangular matrix and B is 1×1 or 2×2 block diagonal. However, this requires pivoting, which substantially increases communication cost and degrades performance on GPUs. Our approach solves a large indefinite system by solving multiple smaller positive definite systems, using an iterative solver on the Schur complement and an inner direct solve (via Cholesky factorization) within each iteration. Cholesky is stable without pivoting, thereby reducing communication and allowing reuse of the symbolic factorization. We demonstrate the practicality of our approach on large optimal power flow problems and show that it can efficiently utilize GPUs and outperform LBL T factorization of the full system.

97 MATHEMATICS AND COMPUTING↗

A Discrete Hankel Transform Approach to Nuclear Data Processing for Fusion Applications

This study introduces advancements to the numerical solutions employed in the processing of nuclear data for fusion applications. It leverages the convolution theorem and Fourier transform techniques to enhance computational efficiency and broaden applicability. Building upon a previously reported discrete Hankel transform approach for Doppler broadening, this work refines the solution of convolution integrals central to these applications. The methodology provides a general and unified framework for evaluating any convolution operation, regardless of whether the underlying problem involves temperature effects in nuclear reactions. The applicability to the nuclear data processing for fusion is demonstrated by deriving the convolution integrals for some of the fusion-related quantities. As before, the convolution operation utilizes a Gaussian-based kernel; however, the discrete Hankel transform of order $𝛼$ = $\frac{1}{2}$ is now applied to the forward Fourier transform of the nonkernel argument, rather than the inverse Fourier transform. This modification eliminates the need for the integration of the nonkernel, cross section–based function, which is a step that posed challenges for certain pointwise cross-section representations. It also removes the requirement for cross-section linearization. Optimized for graphics processing unit architectures, the approach significantly improves computational performance. These advancements are currently under evaluation as the foundation for the next-generation thermonuclear data file processing codes being developed at Lawrence Livermore National Laboratory.

Nuclear science and engineering↗

Near real-time streaming analysis of big fusion data

Experiments on fusion plasmas produce high-dimensional data time series with ever-increasing magnitude and velocity, but turn-around times for analysis of this data have not kept up. For example, many data analysis tasks are often performed in a manual, ad-hoc manner some time after an experiment. In this article, we introduce the Delta framework that facilitates near real-time streaming analysis of big and fast fusion data. By streaming measurement data from fusion experiments to a high-performance compute center, Delta allows computationally expensive data analysis tasks to be performed in between plasma pulses. This article describes the modular and expandable software architecture of Delta and presents performance benchmarks of individual components as well as of an example workflow. Focusing on a streaming analysis workflow where electron cyclotron emission imaging (ECEi) data is measured at KSTAR on the National Energy Research Scientific Computing Center's (NERSC's) supercomputer we routinely observe data transfer rates of about 4 Gigabit per second. In NERSC, a demanding turbulence analysis workflow effectively utilizes multiple nodes and graphical processing units and executes them in under 5 min. We further discuss how Delta uses modern database systems and container orchestration services to provide web-based real-time data visualization. For the case of ECEi data we demonstrate how data visualizations can be augmented with outputs from machine learning models. Here, by providing session leaders and physics operators, results of higher-order data analysis using live visualizations may make more informed decisions on how to configure the machine for the next shot.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗