Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gpu”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Divergence Reduction in Monte Carlo Neutron Transport with On-GPU Asynchronous Scheduling

While Monte Carlo Neutron Transport (MCNT) is near-embarrasingly parallel, the effectively unpredictable lifetime of neutrons can lead to divergence when MCNT is evaluated on GPUs. Divergence is the phenomenon of adjacent threads in a warp executing different control flow paths; on GPUS, it reduces performance because each work group may only execute one path at a time. The process of Thread Data Remapping (TDR) resolves these discrepancies by moving data across hardware such that data in the same warp will be processed through similar paths. A common issue among prior implementations of TDR is the synchronous nature of its remapping and processing cycles, which exhaustively sort data produced by prior processing passes and exhaustively evaluate the sorted data. In another work, we defined a method of remapping data through an asynchronous scheduler which allows for work to be stored in shared memory and deferred arbitrarily until that work is a viable option for low-divergence evaluation. This article surveys a wider set of cases, with the goal of characterizing performance trends across a more comprehensive set of parameters. These parameters include cross sections of scattering/capturing/fission, use of implicit capture, source neutron counts, simulation time spans, and tuned memory allocations. Across these cases, we have recorded minimum and average execution times, as well as a heuristically tuned near-optimal memory allocation size for both synchronous and asynchronous scheduling. Across the collected data, it is shown that the asynchronous method is faster and more memory efficient in the majority of cases, and that it requires less tuning to achieve competitive performance.

Computer Science↗

Demonstration and performance testing of extreme-resolution simulations with static meshes on Summit (CPU & GPU) for a parked-turbine configuration and an actuator-line (mid-fidelity model) wind farm configuration (ECP-Q4 FY2020 Milestone Report)

The goal of the ExaWind project is to enable predictive simulations of wind farms comprised of many megawatt-scale turbines situated in complex terrain. Predictive simulations will require computational fluid dynamics (CFD) simulations for which the mesh resolves the geometry of the turbines and captures the rotation and large deflections of blades. Whereas such simulations for a single turbine are arguably petascale class, multi-turbine wind farm simulations will require exascale-class resources. The primary physics codes in the ExaWind simulation environment are Nalu-Wind, an unstructured-grid solver for the acoustically incompressible Navier-Stokes equations, AMR-Wind, a block-structured-grid solver with adaptive mesh refinement capabilities, and OpenFAST, a wind-turbine structural dynamics solver. The Nalu-Wind model consists of the mass-continuity Poisson-type equation for pressure and Helmholtz-type equations for transport of momentum and other scalars. For such modeling approaches, simulation times are dominated by linear-system setup and solution for the continuity and momentum systems. For the ExaWind challenge problem, the moving meshes greatly affect overall solver costs as reinitialization of matrices and recomputation of preconditioners is required at every time step. The choice of overset-mesh methodology to model the moving and non-moving parts of the computational domain introduces constraint equations in the elliptic pressure-Poisson solver. The presence of constraints greatly affects the performance of algebraic multigrid preconditioners.

17 WIND ENERGY↗

Enhanced Beam Diagnostics with Existing BPPMs via GPU-powered Multi-Particle Simulation

This research aims to utilize the multi-particle code, High-Performance Simulator (HPSim), to realistically model the Side-Coupled-Cavity Linac (CCL) lattice of the LANSCE accelerator. This new model would allow us to predict the beam’s bunch length (the longitudinal spread), which is unavailable for individual accelerating modules or only accessible at the end of the linac. However, a correct bunch length is critical for the high-energy beam transport after the CCL. Its impact would be most significant in the Proton Storage Ring (PSR), where we should be able to reduce losses for the circulating beam. The PSR is scheduled to have a 25% current increase for the neutron spallation target upgrade at the Lujan Center. A highly bunched beam would be necessary to reduce the particle losses and lower the radiation levels produced from the ring. A realistic HPSim model with >1M macro-particles can help tackle the beam losses at the sub-percent level. This new work would also create a realistic surrogate model for future machine learning projects.

43 PARTICLE ACCELERATORS↗

Spiner-EOSPAC Comparison: performance and accuracy on Power9 CPU and GPU

The goal of this report is to concisely compare EOSPAC to Spiner EOS for a selection of commonly-used materials over as much of their domain as is possible. Importantly, this comparison does not address the way that the Euler equations may amplify interpolation errors and we take as given that the rational function interpolator gives the best possible representation of the EOS represented by the Sesame data. Neither does the report address all, or even majority of, the capabilities of each code nor does it explore performance or accuracy over multiple compilers and platforms.

97 MATHEMATICS AND COMPUTING↗

Celeritas: GPU-accelerated particle transport for detector simulation in High Energy Physics experiments

Within the next decade, experimental High Energy Physics (HEP) will enter a new era of scientific discovery through a set of targeted programs recommended by the Particle Physics Project Prioritization Panel (P5), including the upcoming High Luminosity Large Hadron Collider (LHC) HL-LHC upgrade and the Deep Underground Neutrino Experiment (DUNE). These efforts in the Energy and Intensity Frontiers will require an unprecedented amount of computational capacity on many fronts including Monte Carlo (MC) detector simulation. In order to alleviate this impending computational bottleneck, the Celeritas MC particle transport code is designed to leverage the new generation of heterogeneous computer architectures, including the exascale computing power of U.S. Department of Energy (DOE) Leadership Computing Facilities (LCFs), to model targeted HEP detector problems at the full fidelity of Geant4. This paper presents the planned roadmap for Celeritas, including its proposed code architecture, physics capabilities, and strategies for integrating it with existing and future experimental HEP computing workflows.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗