Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

Performance Enhancements for the Lattice-Boltzmann Solver in the LAVA Framework

Performance enhancements in NASA's recently developed Lattice Boltzmann solver within the Launch Ascent and Vehicle Aerodynamics (LAVA) framework are presented. Two key algorithmic developments are highlighted. A coarse-fine interface treatment that discretely conserves mass and momentum has been implemented and successfully verified and validated. Code optimizations targeting improved serial and parallel performance were presented. For a simple turbulent Taylor-Green Vortex problem, we were able to demonstrate a 2.3 times speedup over the baseline code for a single Skylake-SP node containing 40 physical cores, and a 2.14 times speedup for 64 nodes containing 2560 physical cores. In addition, we were able to show that the optimizations enabled us to scale the code almost perfectly to 20480 physical cores where, including ghost cells, the problem size was 10 billion cells.

Barad, Michael↗

Performance Enhancements for the Lattice-Boltzmann Solver in the LAVA Framework

Performance enhancements in NASA's recently developed Lattice Boltzmann solver within the Launch Ascent and Vehicle Aerodynamics (LAVA) framework are presented. Two key algorithmic developments are highlighted. A coarse-fine interface treatment that discretely conserves mass and momentum has been implemented and successfully verified and validated. Code optimizations targeting improved serial and parallel performance were presented. For a simple turbulent Taylor-Green Vortex problem, we were able to demonstrate a 2.3 times speedup over the baseline code for a single Skylake-SP node containing 40 physical cores, and a 2.14 times speedup for 64 nodes containing 2560 physical cores. In addition, we were able to show that the optimizations enabled us to scale the code almost perfectly to 20480 physical cores where, including ghost cells, the problem size was 10 billion cells.

LAVA↗

Spinor $GW$ Bethe-Salpeter calculations in BerkeleyGW: Implementation, symmetries, benchmarking, and performance

Computing the GW quasiparticle band structure and Bethe-Salpeter equation (BSE) absorption spectra for materials with spin-orbit coupling have commonly been done by treating GW corrections and spin-orbit coupling (SOC) as separate perturbations to density-functional theory. However, accurate treatment of materials with strong spin-orbit coupling (such as many topological materials of recent interest, and thermoelectrics) often requires a nonperturbative approach using spinor wave functions in the Kohn-Sham equation and GW/BSE. Such calculations have only recently become available, in particular for the BSE. Here, we have implemented this approach in the plane-wave pseudopotential GW/BSE code BerkeleyGW, which is highly parallelized and widely used in the electronic-structure community. We present reference results for quasiparticle band structures and optical absorption spectra of solids with different strengths of spin-orbit coupling, including Si, Ge, GaAs, GaSb, CdSe, Au, and Bi 2 Se 3 . The calculated quasiparticle band gaps of these systems are found to agree with experiment to within a few tens of meV. SOC splittings are found to be generally in better agreement with experiment, including quasiparticle corrections to band energies. The absorption spectrum of GaAs is not significantly impacted by the inclusion of spin-orbit coupling due to its relatively small value (0.2 eV) in the Λ direction, while the absorption spectrum of GaSb calculated with the spinor GW/BSE captures the large spin-orbit splitting of peaks in the spectrum. For the prototypical topological insulator Bi 2 Se 3 , we find a drastic change in the low-energy band structure compared to that of DFT, with the spinorial treatment of the GW approximation correctly capturing the parabolic nature of the valence and conduction bands after including off-diagonal self-energy matrix elements. We present the detailed methodology, approach to spatial symmetries for spinors, comparison against other codes, and performance compared to spinless GW/BSE calculations and perturbative approaches to SOC. This work aims to spur further development of spinor GW/BSE methodology in excited-state research software and enables a more accurate and detailed exploration of electronic and optical properties of materials containing elements with large atomic numbers.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Improvements to the Unstructured Mesh Generator MESH3D

The AIRPLANE process starts with an aircraft geometry stored in a CAD system. The surface is modeled with a mesh of triangles and then the flow solver produces pressures at surface points which may be integrated to find forces and moments. The biggest advantage is that the grid generation bottleneck of the CFD process is eliminated when an unstructured tetrahedral mesh is used. MESH3D is the key to turning around the first analysis of a CAD geometry in days instead of weeks. The flow solver part of AIRPLANE has proven to be robust and accurate over a decade of use at NASA. It has been extensively validated with experimental data and compares well with other Euler flow solvers. AIRPLANE has been applied to all the HSR geometries treated at Ames over the course of the HSR program in order to verify the accuracy of other flow solvers. The unstructured approach makes handling complete and complex geometries very simple because only the surface of the aircraft needs to be discretized, i.e. covered with triangles. The volume mesh is created automatically by MESH3D. AIRPLANE runs well on multiple platforms. Vectorization on the Cray Y-MP is reasonable for a code that uses indirect addressing. Massively parallel computers such as the IBM SP2, SGI Origin 2000, and the Cray T3E have been used with an MPI version of the flow solver and the code scales very well on these systems. AIRPLANE can run on a desktop computer as well. AIRPLANE has a future. The unstructured technologies developed as part of the HSR program are now targeting high Reynolds number viscous flow simulation. The pacing item in this effort is Navier-Stokes mesh generation.

Thomas, Scott D.↗

Forward Monte Carlo Computations of Polarized Microwave Radiation

Microwave radiative transfer computations continue to acquire greater importance as the emphasis in remote sensing shifts towards the understanding of microphysical properties of clouds and with these to better understand the non linear relation between rainfall rates and satellite-observed radiance. A first step toward realistic radiative simulations has been the introduction of techniques capable of treating 3-dimensional geometry being generated by ever more sophisticated cloud resolving models. To date, a series of numerical codes have been developed to treat spherical and randomly oriented axisymmetric particles. Backward and backward-forward Monte Carlo methods are, indeed, efficient in this field. These methods, however, cannot deal properly with oriented particles, which seem to play an important role in polarization signatures over stratiform precipitation. Moreover, beyond the polarization channel, the next generation of fully polarimetric radiometers challenges us to better understand the behavior of the last two Stokes parameters as well. In order to solve the vector radiative transfer equation, one-dimensional numerical models have been developed, These codes, unfortunately, consider the atmosphere as horizontally homogeneous with horizontally infinite plane parallel layers. The next development step for microwave radiative transfer codes must be fully polarized 3-D methods. Recently a 3-D polarized radiative transfer model based on the discrete ordinate method was presented. A forward MC code was developed that treats oriented nonspherical hydrometeors, but only for plane-parallel situations.

Battaglia, A.↗

PyCDFT: A Python package for constrained density functional theory

In this paper, we present PyCDFT, a Python package to compute diabatic states using constrained density functional theory (CDFT). PyCDFT provides an object-oriented, customizable implementation of CDFT, and allows for both single-point self-consistent-field calculations and geometry optimizations. PyCDFT is designed to interface with existing density functional theory (DFT) codes to perform CDFT calculations where constraint potentials are added to the Kohn–Sham Hamiltonian. Here, we demonstrate the use of PyCDFT by performing calculations with a massively parallel first-principles molecular dynamics code, Qbox, and we benchmark its accuracy by computing the electronic coupling between diabatic states for a set of organic molecules. We show that PyCDFT yields results in agreement with existing implementations and is a robust and flexible package for performing CDFT calculations. The program is available at https://dx.doi.org/10.5281/zenodo.3821097.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multigrid reduction in time with Richardson extrapolation

The advent of exascale computing will leave many users with access to more computational resources than they can simultaneously use, e.g., billion-way parallelism. In particular, this is true for time-dependent simulations that limit parallelism to the spatial domain. One method to add parallelism in time to existing simulation codes and thus take advantage of ever larger compute resources is Multigrid Reduction in Time (MGRIT). The goal is to achieve a smaller time-to-solution through parallelism in time. In this paper, MGRIT is enhanced with Richardson extrapolation in a cost-efficient way to produce a parallel-in-time method with improved accuracy. Overall, this leads to a large improvement in the accuracy per computational cost of MGRIT.

97 MATHEMATICS AND COMPUTING↗

Support for Debugging Automatically Parallelized Programs

This viewgraph presentation provides information on the technical aspects of debugging computer code that has been automatically converted for use in a parallel computing system. Shared memory parallelization and distributed memory parallelization entail separate and distinct challenges for a debugging program. A prototype system has been developed which integrates various tools for the debugging of automatically parallelized programs including the CAPTools Database which provides variable definition information across subroutines as well as array distribution information.

Hood, Robert↗

Charon User Manual (V.2.1) (Rev.01)

This manual gives usage information for the Charon semiconductor device simulator. Charon was developed to meet the modeling needs of Sandia National Laboratories and to improve on the capabilities of the commercial TCAD simulators; in particular, the additional capabilities are running very large simulations on parallel computers and modeling displacement damage and other radiation effects in significant detail. The parallel capabilities are based around the MPI interface which allows the code to be ported to a large number of parallel systems, including linux clusters and proprietary "big iron" systems found at the national laboratories and in large industrial settings.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Charon User Manual: v. 2.2 (revision1)

This manual gives usage information for the Charon semiconductor device simulator. Charon was developed to meet the modeling needs of Sandia National Laboratories and to improve on the capabilities of the commercial TCAD simulators; in particular, the additional capabilities are running very large simulations on parallel computers and modeling displacement damage and other radiation effects in significant detail. The parallel capabilities are based around the MPI interface which allows the code to be ported to a large number of parallel systems, including linux clusters and proprietary “big iron” systems found at the national laboratories and in large industrial settings.

42 ENGINEERING↗

Efficient Parallel Kernel Solvers for Computational Fluid Dynamics Applications

Distributed-memory parallel computers dominate today's parallel computing arena. These machines, such as Intel Paragon, IBM SP2, and Cray Origin2OO, have successfully delivered high performance computing power for solving some of the so-called "grand-challenge" problems. Despite initial success, parallel machines have not been widely accepted in production engineering environments due to the complexity of parallel programming. On a parallel computing system, a task has to be partitioned and distributed appropriately among processors to reduce communication cost and to attain load balance. More importantly, even with careful partitioning and mapping, the performance of an algorithm may still be unsatisfactory, since conventional sequential algorithms may be serial in nature and may not be implemented efficiently on parallel machines. In many cases, new algorithms have to be introduced to increase parallel performance. In order to achieve optimal performance, in addition to partitioning and mapping, a careful performance study should be conducted for a given application to find a good algorithm-machine combination. This process, however, is usually painful and elusive. The goal of this project is to design and develop efficient parallel algorithms for highly accurate Computational Fluid Dynamics (CFD) simulations and other engineering applications. The work plan is 1) developing highly accurate parallel numerical algorithms, 2) conduct preliminary testing to verify the effectiveness and potential of these algorithms, 3) incorporate newly developed algorithms into actual simulation packages. The work plan has well achieved. Two highly accurate, efficient Poisson solvers have been developed and tested based on two different approaches: (1) Adopting a mathematical geometry which has a better capacity to describe the fluid, (2) Using compact scheme to gain high order accuracy in numerical discretization. The previously developed Parallel Diagonal Dominant (PDD) algorithm and Reduced Parallel Diagonal Dominant (RPDD) algorithm have been carefully studied on different parallel platforms for different applications, and a NASA simulation code developed by Man M. Rai and his colleagues has been parallelized and implemented based on data dependency analysis. These achievements are addressed in detail in the paper.

Sun, Xian-He↗

Toward exascale whole-device modeling of fusion devices: Porting the GENE gyrokinetic microturbulence code to GPU

GENE solves the five-dimensional gyrokinetic equations to simulate the development and evolution of plasma microturbulence in magnetic fusion devices. The plasma model used is close to first principles and computationally very expensive to solve in the relevant physical regimes. In order to use the emerging computational capabilities to gain new physics insights, several new numerical and computational developments are required. Here, we focus on the fact that it is crucial to efficiently utilize GPUs (graphics processing units) that provide the vast majority of the computational power on such systems. In this paper, we describe the various porting approaches considered and given the constraints of the GENE code and its development model, justify the decisions made, and describe the path taken in porting GENE to GPUs. We introduce a novel library called gtensor that was developed along the way to support the process. Performance results are presented for the ported code, which in a single node of the Summit supercomputer achieves a speed-up of almost 15× compared to running on central processing unit (CPU) only. Typical GPU kernels are memory-bound, achieving about 90% of peak. Our analysis shows that there is still room for improvement if we can refactor/fuse kernels to achieve higher arithmetic intensity. We also performed a weak parallel scalability study, which shows that the code runs well on a massively parallel system, but communication costs start becoming a significant bottleneck.

Germaschewski, K. (ORCID:0000000284956354)↗

Three-dimensional Skyrme Hartree-Fock-Bogoliubov solver in coordinate-space representation

The coordinate-space representation of the Hartree-Fock-Bogoliubov theory is the method of choice to study weakly bound nuclei whose properties are affected by the quasiparticle continuum space. To describe such systems, we developed a three-dimensional Skyrme-Hartree-Fock-Bogoliubov solver HFBFFT based on the existing, highly optimized and parallelized Skyrme-Hartree-Fock code Sky3D. The code does not impose any self-consistent spatial symmetries such as mirror inversions or parity. The underlying equations are solved in HFBFFT directly in the canonical basis using the fast Fourier transform. To remedy the problems with pairing collapse, we implemented the soft energy cutoff and pairing annealing. The convergence of HFB solutions was improved by a sub-iteration method. The Hermiticity violation of differential operators brought by Fourier-transform-based differentiation has also been solved. Furthermore, the accuracy and performance of HFBFFT were tested by benchmarking it against other HFB codes, both spherical and deformed, for a set of nuclei, both well-bound and weakly-bound.

3D coordinate-space representation↗

DART-PFLOTRAN: An ensemble-based data assimilation system for estimating subsurface flow and transport model parameters

Ensemble-based Data Assimilation (EDA), based on the Monte Carlo approach, has been effectively applied to estimate model parameters through inverse modeling in subsurface flow and transport problems. However, implementation of EDA approach involves a complicated workflow that include setting up and executing ensemble forward model simulations, processing observations and model simulation results for parameter updates, and repeat for sequential or iterative EDA. To facilitate the management of such workflow and lower the barriers for adopting EDA-based parameter estimation in subsurface science, we develop a generic software frame-work linking the Data Assimilation Research Testbed (DART) with a massively parallel subsurface FLOw and TRANsport code PFLOTRAN. The new DART-PFLOTRAN leverages both the core data assimilation engines in DART and the computational power afforded by PFLOTRAN. In addition to the standard smoother and filtering options, DART-PFLOTRAN enables an iterative EDA workflow based on the Ensemble Smoother for Multiple Data Assimilation method (ES-MDA) to improve estimation accuracy for nonlinear forward problems. Here, we verify the implementation of ES-MDA in DART-PFLOTRAN using two synthetic cases designed to estimate static permeability and dynamic exchange fluxes across the riverbed, respectively, from continuous temperature measurements made across a depth profile. One-dimensional hydro-thermal simulations are performed in both cases to relate temperature responses with the parameters of interest. In the case of estimating dynamic parameters, we demonstrate the flexibility of DART-PFLOTRAN in automating sequential ES-MDA workflow, which will significantly reduce the time researchers spend on managing complex workflows in similar applications. Both studies yield accurate estimations of the parameters compared to their synthetic truth, while ES-MDA leads to more accurate estimation when a high level of nonlinearity exist between observed responses and unknown parameters. With a code base in Python and Fortran, DART-PFLOTRAN paves the way for applications in large-scale subsurface inverse modeling by automating the complex workflow of sequential ES-MDA that can be executed on various computing platforms.

97 MATHEMATICS AND COMPUTING↗

High-fidelity wind farm simulation methodology with experimental validation

The complexity and associated uncertainties involved with atmospheric-turbine-wake interactions produce challenges for accurate wind farm predictions of generator power and other important quantities of interest (QoIs), even with state-of-the-art high-fidelity atmospheric and turbine models. A comprehensive computational study was undertaken with consideration of simulation methodology, parameter selection, and mesh refinement on atmospheric, turbine, and wake QoIs to identify capability gaps in the validation process. For neutral atmospheric boundary layer conditions, the massively parallel large eddy simulation (LES) code Nalu-Wind was used to produce high-fidelity computations for experimental validation using high-quality meteorological, turbine, and wake measurement data collected at the Department of Energy/Sandia National Laboratories Scaled Wind Farm Technology (SWiFT) facility located at Texas Tech University’s National Wind Institute. The wake analysis showed the simulated lidar model implemented in Nalu-Wind was successful at capturing wake profile trends observed in the experimental lidar data.

17 WIND ENERGY↗

Encoder–decoder neural network for solving the nonlinear Fokker–Planck–Landau collision operator in XGC

An encoder–decoder neural network has been used to examine the possibility for acceleration of a partial integro-differential equation, the Fokker–Planck–Landau collision operator. This is part of the governing equation in the massively parallel particle-in-cell code XGC, which is used to study turbulence in fusion energy devices. The neural network emphasizes physics-inspired learning, where it is taught to respect physical conservation constraints of the collision operator by including them in the training loss, along with the ℓ 2 loss. In particular, network architectures used for the computer vision task of semantic segmentation have been used for training. A penalization method is used to enforce the ‘soft’ constraints of the system and integrate error in the conservation properties into the loss function. During training, quantities representing the particle density, momentum and energy for all species of the system are calculated at each configuration vertex, mirroring the procedure in XGC. This simple training has produced a median relative loss, across configuration space, of the order of 10 –4 , which is low enough if the error is of random nature, but not if it is of drift nature in time steps. The run time for the current Picard iterative solver of the operator is O(n 2 ), where n is the number of plasma species. As the XGC1 code begins to attack problems including a larger number of species, the collision operator will become expensive computationally, making the neural network solver even more important, especially since its training only scales as O(n). Here, a wide enough range of collisionality has been considered in the training data to ensure the full domain of collision physics is captured. An advanced technique to decrease the losses further will be subject of a subsequent report. Eventual work will include expansion of the network to include multiple plasma species.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Ab-Initio Investigation of Finite Size Effects in Rutile Titania Nanoparticles with Semilocal and Nonlocal Density Functionals

In this work, we employ hybrid and generalized gradient approximation (GGA) level density functional theory (DFT) calculations to investigate the convergence of surface properties and electronic gap of rutile titania nanoparticles with particle size. The surface energies and electronic gaps are calculated for cuboidal particles with minimum dimension ranging from 3.7 Angstrom (24 atoms) to 10.3 Angstrom (384 atoms) using a highly-parallel real-space DFT code to enable hybrid level DFT calculations of larger nanoparticles than are typically practical. We deconvolute the geometric and electronic finite size effects in surface energy, and evaluate the influence of defects on electronic gap and density of states (DOS). The electronic finite size effects in surface energy vanish when the minimum length scale of the nanoparticles becomes greater than 10 Angstrom. We show that this length scale is consistent with a computationally efficient numerical analysis of the characteristic length scale of electronic interactions. The surface energy of nanoparticles having minimum dimension beyond this characteristic length can be approximated using slab calculations that account for the geometric defects. In contrast, the finite size effects on the electronic gap and DOS is highly dependent on the shape and size of these particles. Furthermore, the DOS for cuboidal particles and more realistic particles constructed using the Wulff algorithm reveal that defect states within the electronic gap play a key role in determining the eigen value distribution of nanoparticles and the electronic gap does not converge to the bulk limit for the particle sizes investigated.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

High-fidelity parallel entangling gates on a neutral-atom quantum computer

The ability to perform entangling quantum operations with low error rates in a scalable fashion is a central element of useful quantum information processing. Neutral-atom arrays have recently emerged as a promising quantum computing platform, featuring coherent control over hundreds of qubits and any-to-any gate connectivity in a flexible, dynamically reconfigurable architecture. The main outstanding challenge has been to reduce errors in entangling operations mediated through Rydberg interactions. Here we report the realization of two-qubit entangling gates with 99.5% fidelity on up to 60 atoms in parallel, surpassing the surface-code threshold for error correction. Our method uses fast, single-pulse gates based on optimal control, atomic dark states to reduce scattering and improvements to Rydberg excitation and atom cooling. We benchmark fidelity using several methods based on repeated gate applications, characterize the physical error sources and outline future improvements. Finally, we generalize our method to design entangling gates involving a higher number of qubits, which we demonstrate by realizing low-error three-qubit gates. By enabling high-fidelity operation in a scalable, highly connected system, these advances lay the groundwork for large-scale implementation of quantum algorithms, error-corrected circuits and digital simulations.

97 MATHEMATICS AND COMPUTING↗