Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “GPU accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

A Multi-Architecture Approach for Implicit Computational Fluid Dynamics on Unstructured Grids

High-performance computing (HPC) architectures are trending toward manycore paradigms such as graphics processing units (GPUs). Approximately half of the top 100 publicly disclosed supercomputers in the world utilize GPU accelerators for performance. This is in contrast to a decade ago, where there were only a few such machines in the top 100. It is not currently possible to compile and run legacy central processing unit (CPU) software efficiently on GPUs without significant refactoring. Though a number of frameworks offering performance portability exist, none offer a standardized specification that is supported by all major hardware vendors. Additionally, experiences show that obtaining a high percentage of peak performance often requires architecture-specific code. This work details a pragmatic multi-architecture computational fluid dynamics library focused on aerospace problems across the speed range from low subsonic to hypersonic flows involving thermochemical nonequilibrium. A thin abstraction layer above NVIDIA CUDA C++ is utilized, which enables primarily single-source software currently capable of running efficiently on multicore CPUs, NVIDIA GPUs, AMD GPUs, and Intel GPUs. Results on various problems of interest across the speed range are presented and performance is compared between various architectures.

GPU↗

A Multi-Architecture Approach for Implicit Computational Fluid Dynamics on Unstructured Grids

High-performance computing (HPC) architectures are trending toward manycore paradigms such as graphics processing units (GPUs). Approximately half of the top 100 publicly disclosed supercomputers in the world utilize GPU accelerators for performance. This is in contrast to a decade ago, where there were only a few such machines in the top 100. It is not currently possible to compile and run legacy central processing unit (CPU) software efficiently on GPUs without significant refactoring. Though a number of frameworks offering performance portability exist, none offer a standardized specification that is supported by all major hardware vendors. Additionally, experiences show that obtaining a high percentage of peak performance often requires architecture-specific code. This work details a pragmatic multi-architecture computational fluid dynamics library focused on aerospace problems across the speed range from low subsonic to hypersonic flows involving thermochemical nonequilibrium. A thin abstraction layer above NVIDIA CUDA C++ is utilized, which enables primarily single-source software currently capable of running efficiently on multicore CPUs, NVIDIA GPUs, AMD GPUs, and Intel GPUs. Results on various problems of interest across the speed range are presented and performance is compared between various architectures.

GPU↗

Development of a CPU/GPU portable software library for Lagrangian–Eulerian simulations of liquid sprays

The Lagrangian–Eulerian method is widely used for simulations of fuel sprays in turbulent combustion because of the advantage of treating the spray droplets as discrete points. One challenge of the Lagrangian–Eulerian method is the intense computational requirement when tracking the large number of Lagrangian particles needed for high fidelity. We have developed a performance-portable library, Grit, to track the Lagrangian particles in parallel on central processing unit (CPU) and graphics processing unit (GPU) accelerated high performance computing (HPC) architectures. Grit is a C++ library which employs Message Passing Interface (MPI) for distributed memory parallelism and Kokkos programming model for on-node shared memory parallelism with performance portability across different architectures of GPUs and multi-core/manycore CPUs. The parallel algorithms, key parallel kernels, and their performances on the pre-exascale supercomputer, Summit, are presented. Grit is coupled with a direct numerical simulation (DNS) solver, S3D, for multiphase simulations. A conservative formulation has been developed and implemented in Grit for phase coupling with thermodynamic consistency. The formulation separates the conservation of mass, momentum and energy from the physical models to prevent accidental violation of conservation laws due to inconsistent models. The formulation also enforces consistent definitions of enthalpies of the fuel for both phases and the latent heat of evaporation. Finally, simulations of turbulent particle-laden flow and the evaporation of dilute turbulent spray jet are performed to verify the software implementation, and to demonstrate the scalability of Grit for large-scale multiphase simulations.

97 MATHEMATICS AND COMPUTING↗

GIGA-Lens: Fast Bayesian Inference for Strong Gravitational Lens Modeling

We present GIGA-Lens: a gradient-informed, GPU-accelerated Bayesian framework for modeling strong gravitational lensing systems, implemented in TensorFlow and JAX. The three components, optimization using multistart gradient descent, posterior covariance estimation with variational inference, and sampling via Hamiltonian Monte Carlo, all take advantage of gradient information through automatic differentiation and massive parallelization on graphics processing units (GPUs). We test our pipeline on a large set of simulated systems and demonstrate in detail its high level of performance. The average time to model a single system on four Nvidia A100 GPUs is 105 s. The robustness, speed, and scalability offered by this framework make it possible to model the large number of strong lenses found in current surveys and present a very promising prospect for the modeling of ${ \mathcal O }({10}^{5})$ lensing systems expected to be discovered in the era of the Vera C. Rubin Observatory, Euclid, and the Nancy Grace Roman Space Telescope.

79 ASTRONOMY AND ASTROPHYSICS↗

Multi-task Parallelism for Robust Pre-training of Graph Foundation Models on Multi-source, Multi-fidelity Atomistic Modeling Data

Graph foundation models using graph neural networks promise sustainable, efficient atomistic modeling. To tackle challenges of processing multi-source, multi-fidelity data during pre-training, recent studies employ multi-task learning, in which shared message passing layers initially process input atomistic structures regardless of source, then route them to multiple decoding heads that predict data-specific outputs. This approach stabilizes pre-training and enhances a model’s transferability to unexplored chemical regions. Preliminary results on approximately four million structures are encouraging, yet questions remain about generalizability to larger, more diverse datasets and scalability on supercomputers. We propose a multi-task parallelism method that distributes each head across computing resources with GPU acceleration. Implemented in the open-source HydraGNN architecture, our method was trained on over 24 million structures from five datasets and tested on the Perlmutter, Aurora, and Frontier supercomputers, demonstrating efficient scaling on all three highly heterogeneous super-computing architectures.

Lupo Pasini, Massimiliano [ORNL] (ORCID:0000000249↗

Formulation, Implementation and Validation of a 1D Boundary Layer Inflow Scheme for the QUIC Modeling System

Recent studies have highlighted the importance of accurate meteorological conditions for urban transport and dispersion calculations. In this work, we present a novel scheme to compute the meteorological input in the Quick Urban & Industrial Complex () diagnostic urban wind solver to improve the characterization of upstream wind veer and shear in the Atmospheric Boundary Layer (ABL). The new formulation is based on a coupled set of Ordinary Differential Equations (ODEs) derived from the Reynolds Averaged Navier–Stokes (RANS) equations, and is fast to compute. Building upon recent progress in modeling the idealized ABL, we include effects from surface roughness, turbulent stress, Coriolis force, buoyancy and baroclinicity. We verify the performance of the new scheme with canonical Large Eddy Simulation (LES) tests with the GPU-accelerated FastEddy"Equation missing" solver in neutral, stable, unstable and baroclinic conditions with different surface roughness. Furthermore, we evaluate QUIC calculations with and without the new inflow scheme with real data from the Urban Threat Dispersion (UTD) field experiment, which includes Lidar-based wind measurements as well as concentration observations from multiple outdoor releases of a non-reactive tracer in downtown New York City. Compared to previous inflow capabilities that were limited to a constant wind direction with height, we show that the new scheme can model wind veer in the ABL and enhance the prediction of the surface cross-isobaric angle, improving evaluation statistics of simulated concentrations paired in time and space with UTD measurements.

54 ENVIRONMENTAL SCIENCES↗

FlowPM: Distributed TensorFlow implementation of the FastPM cosmological N-body solver

Here, we present FlowPM, a Particle-Mesh (PM) cosmological N-body code implemented in Mesh-TensorFlow for GPU-accelerated, distributed, and differentiable simulations. We implement and validate the accuracy of a novel multi-grid scheme based on multiresolution pyramids to compute large-scale forces efficiently on distributed platforms. We explore the scaling of the simulation on large-scale supercomputers and compare it with corresponding Python based PM code, finding on an average 10x speed-up in terms of wallclock time. We also demonstrate how this novel tool can be used for efficiently solving large scale cosmological inference problems, in particular reconstruction of cosmological fields in a forward model Bayesian framework with hybrid PM and neural network forward model. We provide skeleton code for these examples and the entire code is publicly available at https://github.com/modichirag/flowpm[Formula presented].

79 ASTRONOMY AND ASTROPHYSICS↗

From atomistic models to machine learning: Predictive design of nanocarbons under extreme conditions

The formation of technologically valuable nanocarbon structures under extreme conditions, such as those produced during high-explosive detonations, remains poorly understood but holds significant potential for the development of controlled synthesis pathways. While detonation shockwaves provide the high-pressure, high-temperature environment required for nanodiamond formation, subsequent cooling and decompression dictate whether the diamond phase is preserved or transformed into other nanocarbon structures. Here, in this study, we employ GPU-accelerated reactive molecular dynamics (ReaxFF) simulations to investigate the graphitization and structural remodeling of detonation nanodiamond under nonlinear quench and pressure-release trajectories. We further investigate how the initial nanodiamond morphology; cuboctahedral, octahedral, or hexagonal prism influences the resulting transformation products. Evolution of nanostructure, allotrope (via simulated x-ray diffraction), carbon hybridization, and ring statistics are tracked during a two-stage quench from 5000 K to 60 GPa. Rapid cooling combined with slow decompression optimizes cubic diamond retention, whereas slow cooling with rapid pressure release promotes surface-to-core graphitization, producing concentric sp 2 -hybridized layers and hollowed inner shells. Octahedral nanodiamonds evolve into carbon nano-onions, initially forming bucky diamonds that progressively transform into fully sp 2 -hybridized structures, while hexagonal prisms preferentially form parallel-stacked graphite layers resembling carbon dots. Transient hexagonal diamond (lonsdaleite) emerges as an interfacial phase, suggesting potential reversibility in the shock-induced graphite-to-diamond transformation pathway transformation route. To extend predictive capabilities, we trained machine learning (ML) regressors on over 10 5 node-hours of molecular dynamics (MD) trajectories. A multilayer perceptron (MLP) model reliably predicts the number of graphitized layers from temperature–pressure trajectories with a coefficient of determination (R 2 ) exceeding 0.90. This high predictive fidelity enables efficient, high-throughput mapping of the synthesis parameter space for optimized graphitization outcomes. Collectively, morphological control combined with optimized quench–decompression conditions promote the selective synthesis of nanocarbon allotropes. This work establishes a data-driven framework for the rational, a priori design of carbon nanomaterials for applications in energy storage, sensing, and biomedicine.

Detonation nanodiamond remodeling↗

A fast matrix-free approach to the high-order control volume finite element method with application to low-Mach flow

Here, a fast matrix-free formulation of the control volume finite element method is presented, requiring much less memory and computational work than previous efforts. The method is implemented and evaluated as a solver for low-Mach flow, including the evaluation of a preconditioning strategy for the pressure Poisson equation. The efficiency and scaling with polynomial order is evaluated on simple turbulent flows of interest, with appropriate solution quality metrics, and compared with a reference node-centered finite volume discretization. For a turbulent channel flow test, we show improvement in computational work for a given accuracy with the high-order scheme. The performance on a GPU accelerated platform is also investigated, with benefit shown for the matrix-free discretization.

42 ENGINEERING↗

A Feynman-Kac based numerical method for the exit time probability of a class of transport problems

The exit time probability, which gives the likelihood that an initial condition leaves a prescribed region of the phase space of a dynamical system at, or before, a given time, is arguably one of the most natural and important transport problems. In this work, we present an accurate and efficient numerical method for computing this probability for systems described by non-autonomous (time-dependent) stochastic differential equations (SDEs) or their equivalent Fokker-Planck partial differential equations. The method is based on the direct approximation of the Feynman-Kac formula that establishes a link between the adjoint Fokker-Planck equation and the forward SDE. The Feynman-Kac formula is approximated using the Gauss-Hermite quadrature rules and piecewise cubic Hermite interpolating polynomials, and a GPU accelerated matrix representation is used to compute the entire time evolution of the exit time probability using a single pass of the algorithm. The method is unconditionally stable, exhibits second order convergence in space, first order convergence in time, and it is straightforward to parallelize. Applications are presented to the advection diffusion of a passive tracer in a fluid flow exhibiting chaotic advection, and to the runaway acceleration of electrons in a plasma in the presence of an electric field, collisions, and radiation damping. Benchmarks against analytical solutions as well as comparisons with explicit and implicit finite difference standard methods for the adjoint Fokker-Planck equation are presented.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Machine learning for the redox potential prediction of molecules in organic redox flow battery

Here, organic redox flow batteries (ORFB) are recognized as an innovative technology for the large-scale storage of renewable energy. The redox potential of organic redox-active molecules plays a vital role in their performance. Advanced screening techniques like high-throughput experiment and machine learning (ML) have significantly enhanced organic material performance and transformed the field of ORFB. However, the scarcity of experimental data poses a considerable challenge for ML model development in this domain. In our study, we developed lightweight graph-based Gaussian process regression (GPR) models with GPU-accelerated marginalized graph kernel and hybrid kernel to predict the redox potentials of organic redox-active molecules for ORFBs, specifically focusing on small datasets. To evaluate model accuracy, we created a new experimental database of organic redox-active molecules by the data from hundreds of published papers and assembled previous computational datasets. We also considered some key parameters, such as pH conditions and solvent type, to assess their impact on redox potential prediction. Our GPR model predicted redox potentials with high accuracy across all datasets using minimal training data. The study provides powerful tools for molecule screening and design and delivers valuable guidance on designing training datasets for costly experiments.

25 ENERGY STORAGE↗

SPH modeling of biomass granular flow: Theoretical implementation and experimental validation

The commercialization of biomass-derived energy is impeded by flowability challenges arising from the feeding and handling of granular biomass materials in full-scale biorefineries. To overcome these obstacles, a robust and accurate model to simulate the flow of granular biomass is indispensable. However, conventional mesh-based numerical codes are limited by inherent mesh distortion in simulating large deformation that commonly occurs in granular biomass handling. Here, in this study, we propose a graphics processing unit (GPU)-accelerated meshless Smoothed Particle Hydrodynamics (SPH) code to model the flow of granular biomass materials. A modified void ratio-based mass conversation, a hybrid particle-to-particle/surface frictional boundary treatment, and a hypoplastic constitutive model are implemented. Four numerical examples, an elastic block sliding on inclined planes, sand column collapse, Angle of Repose, and axial compression tests for pine chips, were simulated using the developed SPH code. The results demonstrate good agreement between numerical predictions and analytical and experimental data for all four examples, validating the SPH code and increasing confidence that it can be applied to simulate more complex granular biomass handling processes, such as hopper feeding or auger conveyance.

09 BIOMASS FUELS↗

Uncertainty quantification for Multiphase-CFD simulations of bubbly flows: a machine learning-based Bayesian approach supported by high-resolution experiments

In this paper, we developed a machine learning-based Bayesian approach to inversely quantify and reduce the uncertainties of multiphase computational fluid dynamics (MCFD) simulations for bubbly flows. The proposed approach is supported by high-resolution two-phase flow measurements, including those by double-sensor conductivity probes, high-speed imaging, and particle image velocimetry. Local distributions of key physical quantities of interest (QoIs), including the void fraction and phasic velocities, are obtained to support the Bayesian inference. In the process, the epistemic uncertainties of the closure relations are inversely quantified while the aleatory uncertainties from stochastic fluctuations of the system are evaluated based on experimental uncertainty analysis. The combined uncertainties are then propagated through the MCFD solver to obtain uncertainties of the QoIs, based on which probability-boxes are constructed for validation. The proposed approach relies on three machine learning methods: feedforward neural networks and principal component analysis for surrogate modeling, and Gaussian processes for model form uncertainty modeling. The whole process is implemented within the framework of an open-source deep learning library PyTorch with graphics processing unit (GPU) acceleration, thus ensuring the efficiency of the computation. The results demonstrate that with the support of high-resolution data, the uncertainties of MCFD simulations can be significantly reduced. The proposed approach has the potential for other applications that involve numerical models with empirical parameters.

42 ENGINEERING↗

GX: a GPU-native gyrokinetic turbulence code for tokamak and stellarator design

GX is a code designed to solve the nonlinear gyrokinetic system for low-frequency turbulence in magnetized plasmas, particularly tokamaks and stellarators. In GX, our primary motivation and target is a fast gyrokinetic solver that can be used for fusion reactor design and optimization along with wide-ranging physics exploration. Here, this has led to several code and algorithm design decisions, specifically chosen to prioritize time to solution. First, we have used a discretization algorithm that is pseudospectral in the entire phase space, including a Laguerre–Hermite pseudospectral formulation of velocity space, which allows for smooth interpolation between coarse gyrofluid-like resolutions and finer conventional gyrokinetic resolutions and efficient evaluation of a model collision operator. Additionally, we have built GX to natively target graphics processors (GPUs), which are among the fastest computational platforms available today. Finally, we have taken advantage of the reactor-relevant limit of small $\rho _*$ by using the radially local flux-tube approach. In this paper we present details about the gyrokinetic system and the numerical algorithms used in GX to solve the system. We then present several numerical benchmarks against established gyrokinetic codes in both tokamak and stellarator magnetic geometries to verify that GX correctly simulates gyrokinetic turbulence in the small $\rho _*$. Moreover, we show that the convergence properties of the Laguerre–Hermite spectral velocity formulation are quite favourable for nonlinear problems of interest. Coupled with GPU acceleration, which we also investigate with scaling studies, this enables GX to be able to produce useful turbulence simulations in minutes on one (or a few) GPUs and higher fidelity results in a few hours using several GPUs. GX is open-source software that is ready for fusion reactor design studies.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

NEXTorch: A Design and Bayesian Optimization Toolkit for Chemical Sciences and Engineering

Automation and optimization of chemical systems require well-informed decisions on what experiments to run to reduce time, materials, and/or computations. Data-driven active learning algorithms have emerged as valuable tools to solve such tasks. Bayesian optimization, a sequential global optimization approach, is a popular active-learning framework. Past studies have demonstrated its efficiency in solving chemistry and engineering problems. Here we introduce NEXTorch, a library in Python/PyTorch, to facilitate laboratory or computational design using Bayesian optimization. NEXTorch offers fast predictive modeling, flexible optimization loops, visualization capabilities, easy interfacing with legacy software, and multiple types of parameters and data type conversions. It provides GPU acceleration, parallelization, and state-of-the-art Bayesian optimization algorithms and supports both automated an d human-in-the-loop optimization. The comprehensive online documentation introduces Bayesian optimization theory and several examples from catalyst synthesis, reaction condition optimization, parameter estimation, and reactor geometry optimization. NEXTorch is open-source and available on GitHub

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Using a Coarse-Grained Modeling Framework to Identify Oligomeric Motifs with Tunable Secondary Structure

Coarse-grained modeling can be used to explore general theories that are independent of specific chemical detail. In this paper, we present cg_openmm, a Python-based simulation framework for modeling coarse-grained hetero-oligomers and screening them for structural and thermodynamic characteristics of cooperative secondary structures. cg_openmm facilitates the building of coarse-grained topology and random starting configurations, setup of GPU-accelerated replica exchange molecular dynamics simulations with the OpenMM software package, and features a suite of postprocessing thermodynamic and structural analysis tools. In particular, native contact analysis, heat capacity calculations, and free energy of folding calculations are used to identify and characterize cooperative folding transitions and stable secondary structures. In this work, we demonstrate the capabilities of cg_openmm on a simple 1–1 Lennard-Jones coarse-grained model, in which each residue contains 1 backbone and 1 side-chain bead. By scanning both nonbonded and bonded force-field parameter spaces at the coarse-grained level, we identify and characterize sets of parameters which result in the formation of stable helices through cooperative folding transitions. Furthermore, we show that the geometries and stabilities of these helices can be tuned by manipulating the force-field parameters.

36 MATERIALS SCIENCE↗

Coarse-Grained Density Functional Theory Predictions via Deep Kernel Learning

Scalable electronic predictions are critical for soft materials design. Recently, the Electronic Coarse-Graining (ECG) method was introduced to renormalize all-atom quantum chemical (QC) predictions to coarse-grained (CG) resolutions using deep neural networks (DNNs). While DNNs can learn complex representations that prove challenging for kernel-based methods, they are susceptible to overfitting and the overconfidence of uncertainty estimations. Here, we develop ECG within a GPU-accelerated Deep Kernel Learning (DKL) framework to enable CG QC predictions using range-separated hybrid density functional theory (DFT), obtaining a 107 speedup relative to naive all-atom QC. By treating the predicted electronic properties as random Gaussian Processes, DKL incorporates CG mapping degeneracy by learning the distribution of electronic energies as a function of CG configuration. DKL-ECG accurately reproduces molecular orbital energies from range-separated DFT while facilitating efficient training via active learning using the uncertainties provided by DKL. Further, we show that while active learning algorithms enable efficient sampling of a more diverse configurational space relative to random sampling, all explored query methods exhibit comparable performance for the examined system. We attribute this result to the significant overlap of the feature space and output property distributions across multiple temperatures.

97 MATHEMATICS AND COMPUTING↗

I ntera C hem : Exploring Excited States in Virtual Reality with Ab Initio Interactive Molecular Dynamics

InteraChem is an ab initio interactive molecular dynamics (AI-IMD) visualizer that leverages recent advances in virtual reality hardware and software, as well as the graphical processing unit (GPU)-accelerated TeraChem electronic structure package, in order to render quantum chemistry in real time. We introduce the exploration of electronically excited states via AI-IMD using the floating occupation molecular orbital-complete active space configuration interaction method. The optimization tools in InteraChem enable identification of excited state minima as well as minimum energy conical intersections for further characterization of excited state chemistry in small- to medium-sized systems. We demonstrate that finite-temperature Hartree–Fock theory is an efficient method to perform ground state AI-IMD. InteraChem allows users to track electronic properties such as molecular orbitals and bond order in real time, resulting in an interactive visualization tool that aids in the interpretation of excited state chemistry data and makes quantum chemistry more accessible for both research and educational purposes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗