Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multigrid optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

51 records · Page 3

Exascale Algorithms and Software for Lattice Field Theory in High Energy Physics: Searching Beyond the Standard Model

The Boston University component has focused on the algorithmic development of new Multigrid solver for the critical kernel for the Dirac propagators that dominated the both simulation require for lattice ensemble and the analysis of physical correlation functions. Progress on this has meet the above objects, even exceeding them a bit. The result is the beginning if multiscale lattice QCD applicable to future Exascale hardware and the development of the QUDA software for NVIDIA GPUs to give near optimal performance. As we approach exascale hardware and computation at that scale this provides the infrastructure for further advances.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Preparing an Incompressible-Flow Fluid Dynamics Code for Exascale-Class Wind Energy Simulations

The U.S. Department of Energy has identified exascale-class wind farm simulation as critical to wind energy scientific discovery. A primary objective of the ExaWind project is to build high-performance, predictive computational fluid dynamics (CFD) tools that satisfy these modeling needs. GPU accelerators will serve as the computational thoroughbreds of next-generation, exascale-class supercomputers. Here, we report on our efforts in preparing the ExaWind unstructured mesh solver, Nalu-Wind, for exascale-class machines. For computing at this scale, a simple port of the incompressible-flow algorithms to GPUs is insufficient. To achieve high performance, one needs novel algorithms that are application aware, memory efficient, and optimized for the latest-generation GPU devices. The result of our efforts are unstructured-mesh simulations of wind turbines that can effectively leverage thousands of GPUs. In particular, we demonstrate a first-of-its-kind, incompressible-flow simulation using Algebraic Multigrid solvers that strong scales to more than 4000 GPUs on the Summit supercomputer.

algebraic multigrid↗

Preparing an Incompressible-Flow Fluid Dynamics Code for Exascale-Class Wind Energy Simulations: Preprint

The US Department of Energy has identified Exascale-Class wind farm simulation tools as critical to wind energy scientific discovery. A primary objective of the Exawind project is to build high-performance, predictive Computational Fluid Dynamics tools that satisfy these modeling needs. GPU accelerators will serve as the computational thoroughbreds of next generation, Exascale-Class, platforms. Here, we report on our efforts for preparing the Exawind unstructured mesh solver, Nalu-Wind, for Exascale-Class machines. For computing at this scale, a simple port of the incompressible-flow algorithms to GPUs is not sufficient. One needs novel algorithms that are application aware, memory efficient, and optimized for latest generation GPU devices to get high-performance. The result of our efforts are unstructured mesh simulations of wind turbines that use 1/6 the compute resources of Summit supercomputer at Oak Ridge National Lab. In particular, we demonstrate a first-of-its-kind, simulation using Algebraic Multigrid solvers on over 4000 GPUs.

algebraic multigrid↗

Matrix-free preconditioning for high-order H (curl) discretizations

The greater arithmetic intensity of high-order finite element discretizations makes them attractive for implementation on next-generation hardware, but assembly of high-order finite element operators as matrices is prohibitively expensive. As a result, the development of general algebraic solvers for such operators has been an open research challenge. Fast matrix-free application of high-order operators has received significant attention in the literature in the context of Poisson-type problems, but preconditioners and solvers for inverting more general operators are not very well-developed. In this paper, we consider the problem of preconditioning a definite Maxwell operator at high polynomial order without assembling a matrix. We show that given efficient preconditioners for high-order H 1 finite element problems on the same mesh, efficient H(curl) preconditioners can be constructed in an auxiliary space framework. We demonstrate the resulting preconditioners in a practical setting with tensor-product basis functions on an unstructured mesh of quadrilaterals. Overall, our approach uses a sparsified H 1 solver constructed on a low-order mesh of the nodal points of the underlying high-order space, and we show that the resulting H(curl) preconditioner is effective at very high polynomial orders for two-dimensional model problems with complicated geometry, varying piecewise constant coefficients, and curved elements. The resulting preconditioner scales with nearly optimal O(p d+1 ) floating point operation count and optimal O(p d ) memory transfer requirements, outperforming existing Maxwell preconditioners in the high-order regime.

97 MATHEMATICS AND COMPUTING↗

Implementation of a Mesh refinement algorithm into the quasi-static PIC code QuickPIC

Plasma-based acceleration (PBA) has emerged as a promising candidate for the accelerator technology used to build a future linear collider and/or an advanced light source. In PBA, a trailing or witness particle beam is accelerated in the plasma wave wakefield (WF) created by a laser or particle beam driver. The WF is often nonlinear and involves the crossing of plasma particle trajectories in real space and thus particle-in-cell methods are used. The distance over which the drive beam evolves is several orders of magnitude larger than the wake wavelength. This large disparity in length scales is amenable to the quasi-static approach. Three-dimensional (3D), quasi-static (QS), particle-in-cell (PIC) codes, e.g., QuickPIC, have been shown to provide high fidelity simulation capability with 2-4 orders of magnitude speedup over 3D fully explicit PIC codes. In PBA, the witness beam needs to be matched to the focusing forces of the WF to reduce the emittance growth. In some linear collider designs, the matched spot size of the witness beam can be 2 to 3 orders of magnitude smaller than the spot size (and wavelength) of the wakefield. Such an additional disparity in length scales is ideal for mesh refinement where the WF within the witness beam is described on a finer mesh than the rest of the WF. A mesh refinement scheme is described that has been implemented into the 3D QS PIC code, QuickPIC. Very fine (high) resolution is used in a small spatial region that includes the witness beam and progressively coarser resolutions in the rest of the simulation domain. A fast multigrid Poisson solver has been implemented for the field solve on the refined meshes and a Fast Fourier Transform (FFT) based Poisson solver is used for the coarse mesh. The code has been parallelized with both MPI and OpenMP, and the parallel scalability has also been improved by using pipelining. A preliminary adaptive mesh refinement technique is described to optimize the computational time for simulations with an evolving witness beam size. Several test problems are used to verify that the mesh refinement algorithm provides accurate results. Additionally, the results are benchmarked against highly resolved simulations exhibiting near-azimuthal symmetry, performed using QPAD—a novel hybrid QS PIC code that uses a PIC description in the coordinates (r, ct – z) and a gridless description in the azimuthal angle, Φ.

Linear collider↗

Extending PETSc’s Composable, Hierarchical, Nested Solvers (Final Report)

For this project, I have focused mainly on developing discretization tech nology in PETSc in order to allow us to support optimal solvers for com plex, multiphysics problems, and also outer-loop problems, such as PDE constrained optimization. There have been improvements to the unstruc tured mesh support in DMPlex and particle discretizations in DMSwarm. In addition, we have produced a number of physical examples, tutorials, and tools for understanding performance.

97 MATHEMATICS AND COMPUTING↗

RAPIDS: Reconciling Availability, Accuracy, and Performance in Managing Geo-Distributed Scientific Data

In modern science, big data plays an increasingly important role. Many scientific applications, such as running simulations on supercomputers or conducting experiments on advanced instruments, produce huge amount of data at unprecedented speed. Analyzing and understanding such big data is the key for scientists to make scientific breakthroughs. However, data might become unavailable for scientists to access when outages or maintenance of the storage system occur, which severely hinders scientific discovery. To improve the data availability, data duplication and erasure coding (EC) are often used. But as the scientific data gets larger, using these two methods can cause considerable storage and network overhead.In this paper, we propose RAPIDS, a hybrid approach that combines the multigrid-based error-bounded lossy compression with erasure coding, to significantly reduce the storage and network overhead required for maintaining high data availability. Our experiments show that RAPIDS reduces the storage overhead by up to 7.5x and network overhead by up to 3x to achieve the same level of availability compared to the regular EC method. We improve RAPIDS by building two models to optimize the fault tolerance configurations and data gathering strategy. We demonstrate that RAPIDS significantly improves performance when running on many CPU cores in parallel or on GPUs.

Wan, Lipeng↗

Non-Intrusive Parallel-in-Time Solvers for Partial Differential Equations (Final Report)

Many time-dependent problems and simulations are often modeled using Partial Differential Equations. Traditional modeling approaches that use sequential time-stepping are reaching a bottleneck in optimizing efficiency. The Center of Applied Science and Computing at Lawrence Livermore National Laboratory extensively works on parallelizing these algorithms to leverage the increasing computational power from the growing number of processors in computer hardware. In particular, they aim to design non-intrusive algorithms that can generalize to a variety of problems and sizes without requiring additional information from or modifications on the original problems. Multigrid Reduction in Time (MGRIT) is a parallel-in-time algorithm that is designed to be non-intrusive. This project focuses on increasing the efficiency of MGRIT by approximating the coarse-grid operator using machine learning approaches as a means to find the most non-intrusive, or general, solution.

97 MATHEMATICS AND COMPUTING↗

Parallel-in-Time Simulation of Lindblad's Equation

Constructing fast quantum logic gates is critical to building a scalable quantum computer. We consider a qudit, a quantum version of a bit that can take an arbitrary number of states, coupled with a cavity. In this project, we wish to force the qudit to reach the 0-state, for any possible initial state. The coupled system changes in time according to Lindblad’s equation, an ordinary differential equation on the density matrix of the quantum system. Lindblad’s equation contains some parameters that we can control, so-called control functions. We seek control functions which force the qudit to the 0-state within 2 microseconds, which is much faster than what is currently done in practice. The search method is gradient descent, a numerical optimization method that uses gradient information to iteratively improve the control parameters. My contribution to this project is an attempt to speed up the computation of the gradient. It currently takes about 40 seconds to compute the gradient which involves solving a set of ODEs sequentially. Current supercomputers have thousands of cores, but sequential computations can only make use of 1 core at a time. We wish to divide up the work better, so that we can use many more cores at once. To this end, we have implemented the Multigrid Reduction in Time (MGRIT) algorithm. We perform a systematic parameter search on how to best apply this algorithm. Results indicate a 25 percent speed up for solving Lindblad’s equation and determining how close the final state is the 0-state.

97 MATHEMATICS AND COMPUTING↗

Performance portable ice-sheet modeling with MALI

High-resolution simulations of polar ice sheets play a crucial role in the ongoing effort to develop more accurate and reliable Earth system models for probabilistic sea-level projections. These simulations often require a massive amount of memory and computation from large supercomputing clusters to provide sufficient accuracy and resolution; therefore, it has become essential to ensure performance on these platforms. Many of today’s supercomputers contain a diverse set of computing architectures and require specific programming interfaces in order to obtain optimal efficiency. In an effort to avoid architecture-specific programming and maintain productivity across platforms, the ice-sheet modeling code known as MPAS-Albany Land Ice (MALI) uses high-level abstractions to integrate Trilinos libraries and the Kokkos programming model for performance portable code across a variety of different architectures. In this article, we analyze the performance portable features of MALI via a performance analysis on current CPU-based and GPU-based supercomputers. The analysis highlights not only the performance portable improvements made in finite element assembly and multigrid preconditioning within MALI with speedups between 1.26 and 1.82x across CPU and GPU architectures but also identifies the need to further improve performance in software coupling and preconditioning on GPUs. We perform a weak scalability study and show that simulations on GPU-based machines perform 1.24–1.92x faster when utilizing the GPUs. The best performance is found in finite element assembly, which achieved a speedup of up to 8.65x and a weak scaling efficiency of 82.6% with GPUs. We additionally describe an automated performance testing framework developed for this code base using a changepoint detection method. The framework is used to make actionable decisions about performance within MALI. We provide several concrete examples of scenarios in which the framework has identified performance regressions, improvements, and algorithm differences over the course of 2 years of development.

54 ENVIRONMENTAL SCIENCES↗

Extending Petsc's Composable Hierarchically Nested Linear Solvers

The Contributions from the RELACS group at both Rice University and the University at Buffalo in this phase of the PETSc Composable Solvers effort have centered around four main areas: scalable mesh processing, mesh adaptivity, solvers for subsurface flow, and performance modeling. The prominence of mesh processing demonstrates the tight relationship between meshing and discretization on the one hand, and optimal solvers on the other. All optimal solvers that we consider depend on some notion of hierarchy, and we express this using the DMPlex abstraction in PETSc. This relationship demands tight integration between the DM and SNES/TS components in PETSc that is the foundation of much of this work. In addition, interpretation of performance results for scalable solvers necessitates that information from the discretization and solver enter the performance model. Without this, comparing different solvers can be a fruitless exercise. Some major accomplishment of the past three years in these areas include: scalable mesh loading in PETSc on more than 10K cores, integrated mesh adaptivity using both p4est and Pragmatic, scalable multigrid for DG discretizations of subsurface flow, and predictive performance modeling incorporating error estimates.

79 ASTRONOMY AND ASTROPHYSICS↗

Optimizing shift selection in multilevel Monte Carlo for disconnected diagrams in lattice QCD

The calculation of disconnected diagram contributions to physical signals is a computationally expensive task in Lattice QCD. To extract the physical signal, the trace of the inverse Lattice Dirac operator, a large sparse matrix, must be stochastically estimated. Because the variance of the stochastic estimator is typically large, variance reduction techniques must be employed. Multilevel Monte Carlo (MLMC) methods reduce the variance of the trace estimator by utilizing a telescoping sequence of estimators. Frequency Splitting is one such method that uses a sequence of inverses of shifted operators to estimate the trace of the inverse lattice Dirac operator, however there is no a priori way to select the shifts that minimize the cost of the multilevel trace estimation. Here we present a sampling and interpolation scheme that is able to predict the variances associated with Frequency Splitting under displacements of the underlying space time lattice. The interpolation scheme is able to predict the variances to high accuracy and therefore chooses shifts that correspond to an approximate minimum of the cost for the trace estimation. We show that Frequency Splitting with the chosen shifts displays significant speedups over multigrid deflation, and that these shifts can be used for multiple configurations within the same ensemble with no penalty to performance.

97 MATHEMATICS AND COMPUTING↗

NCCS High Performance GMRES Mixed Precision

HPG-MxP is a software package that performs a fixed number of multigrid preconditioned (using a Gauss-Seidel smoother) Generalized minimal residual (PGMRES) iterations in order to solve a possibly nonsymmetric large sparse linear system of equations. It is designed to be a benchmark to measure a computer's performance for sparse linear algebra workloads typical in scientific computing while allowing the use of mixed precision methods. The solution is required to have convergence characteristics and accuracy similar to double precision GMRES. It is based on the High Performance Conjugate Gradient Benchmark (HPCG) which restricts all implementations to use only the IEEE double precision format (FP64). The original implementation (https://github.com/hpg-mxp/hpg-mxp) was written by Ichitaro Yamazaki, Jennifer Loe, Christian Glusa, Sivasankaran Rajamanickam, Piotr Luszczek, and Jack Dongarra. Please refer to that repository for documentation on the original implementation. This version is maintained by the National Center for Computational Sciences at Oak Ridge National Laboratory. It is highly scalable and optimized for Oak Ridge Leadership Computing Facility (OLCF) systems, particularly Frontier.

Kashi, Aditya [Oak Ridge National Laboratory (ORNL↗

Extensive analysis of reconstruction algorithms for DESI 2024 baryon acoustic oscillations

Reconstruction of the baryon acoustic oscillation (BAO) signal has been a standard procedure in BAO analyses over the past decade and has helped to improve the BAO parameter precision by a factor of ∼2 on average. The Dark Energy Spectroscopic Instrument (DESI) BAO analysis for the first year (DR1) data uses the “standard” reconstruction framework, in which the displacement field is estimated from the observed density field by solving the linearized continuity equation in redshift space, and galaxy and random positions are shifted in order to partially remove non-linearities. There are several approaches to solving for the displacement field in real survey data, including the multigrid (MG), iterative Fast Fourier Transform (iFFT), and iterative Fast Fourier Transform particle (iFFTP) algorithms. In this work, we analyze these algorithms and compare them with various metrics including two-point statistics and the displacement itself using realistic DESI mocks. We focus on three representative DESI samples, the emission line galaxies (ELG), quasars (QSO), and the bright galaxy sample (BGS), which cover the extreme redshifts and number densities, and potential wide-angle effects. We conclude that the MG and iFFT algorithms agree within 0.4% in post-reconstruction power spectrum on BAO scales with the RecSym convention, which does not remove large-scale redshift space distortions (RSDs), in all three tracers. The RecSym convention appears to be less sensitive to displacement errors than the RecIso convention, which attempts to remove large-scale RSDs. However, iFFTP deviates from the first two; thus, we recommend against using iFFTP without further development. In addition, we provide the optimal settings for reconstruction for five years of DESI observation. The analyses presented in this work pave the way for DESI DR1 analysis as well as future BAO analyses.

79 ASTRONOMY AND ASTROPHYSICS↗

A Block-Structured Adaptive Mesh Framework to Solve Radiation Transfer Equation in Irregular Embedded Geometries

Radiation transport arises in various scientific, industrial, and medical fields, and understanding its effect in applications is needed to make accurate predictions, safety assessments and performance optimizations. Solving the Radiation Transport Equation (RTE) is challenging due to its integro-differential nature, which involves both differential and integral terms. The differential term describes the change in radiation intensity due to absorption and emission, while the integral term accounts for scattering. The accurate modeling of radiation is further complicated in many applications due to the complex, irregular geometries. Various methods exist for solving the RTE, including the zonal, Monte Carlo, spherical harmonics, discrete ordinates, and finite volume methods. Traditional mesh-based approaches, which rely on structured or unstructured meshes, struggle with irregular geometries due to: a) the difficulty of conforming structured grids to irregular domains, b) challenges in enforcing boundary conditions correctly, and c) the additional computational cost of unstructured mesh methods. This work presents a second-order accurate method for solving the RTE in irregular geometries. The radiation intensity is discretized using the finite-volume method in both spatial and angular directions on regular Cartesian grid blocks. Leveraging the block-structured adaptive mesh refinement (AMR) framework provided by AMReX, our method refines the grid locally to reduce spatial discretization error, ensuring a converged numerical solution while minimizing computational costs elsewhere. A two-stage deferred correction approach is employed: First, a first-order discretization on grid blocks is solved using an algebraic multigrid method in HYPRE. Second, a correction term is applied explicitly to achieve second-order accuracy. The correction term is calculated by approximating the radiation flux on cell faces using a Total Variation Diminishing (TVD) scheme. This approach ensures quick convergence of the multigrid method while preserving higher-order accuracy of the numerical solution. Irregular geometries are resolved as embedded boundaries (EB), resulting in both cut cells and regular cells. In cut cells, we modify the fluxes using face fractions and incorporate additional contributions from EB boundary conditions. To ensure higher-order convergence near the EB interface, the correction term is modified by interpolating the radiation intensity to fictitious ghost points. The implementation takes advantage of modern supercomputers by leveraging AMReX’sMPI/X parallelization strategy where X can be MPI or a GPU accelerator including CUDA, HIP and DPC++. We validate our solver using classical test cases, both with and without EB, demonstrating accuracy and efficiency. Additionally, we analyze the impact of adaptive mesh refinement on solution accuracy and computational cost, highlighting the advantages of our approach for high-resolution radiation transport simulations.

computational fluid dynamics (CFD)↗