Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “solvers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Evaluating asynchronous Schwarz solvers on GPUs

With the commencement of the exascale computing era, we realize that the majority of the leadership supercomputers are heterogeneous and massively parallel. Even a single node can contain multiple co-processors such as GPUs and multiple CPU cores. For example, ORNL’s Summit accumulates six NVIDIA Tesla V100 GPUs and 42 IBM Power9 cores on each node. Synchronizing across compute resources of multiple nodes can be prohibitively expensive. Hence, it is necessary to develop and study asynchronous algorithms that circumvent this issue of bulk-synchronous computing. In this study, we examine the asynchronous version of the abstract Restricted Additive Schwarz method as a solver. We do not explicitly synchronize, but allow the communication between the sub-domains to be completely asynchronous, thereby removing the bulk synchronous nature of the algorithm. We accomplish this by using the one-sided Remote Memory Access (RMA) functions of the MPI standard. We study the benefits of using such an asynchronous solver over its synchronous counterpart. We also study the communication patterns governed by the partitioning and the overlap between the sub-domains on the global solver. Finally, we show that this concept can render attractive performance benefits over the synchronous counterparts even for a well-balanced problem.

Nayak, Pratik↗

Batched sparse direct solver design and evaluation in SuperLU_DIST

Over the course of interactions with various application teams, the need for batched sparse linear algebra functions has emerged in order to make more efficient use of the GPUs for many small and sparse linear algebra problems. In this paper, we present our recent work on a batched sparse direct solver for GPUs. The sparse LU factorization is computed by the levels of the elimination tree, leveraging the batched dense operations at each level and a new batched Scatter GPU kernel. The sparse triangular solve is computed by the level sets of the directed acyclic graph (DAG) of the triangular matrix. Batched operations overcome the large overhead associated with launching many small kernels. For medium sized matrix batches with not-so-small bandwidth, using an NVIDIA A100 GPU, our new batched sparse direct solver is orders of magnitude faster than a batched banded solver and uses less than one-tenth of the memory.

Boukaram, Wajih↗

A GPU-based compressible combustion solver for applications exhibiting disparate space and time scales

High-speed chemically active flows pose significant computational challenges due to their disparate space and time scales, with stiff chemistry often dominating simulation time. While modern scientific computing programs achieve exascale performance by leveraging graphics processing units (GPUs), existing GPU-based compressible combustion solvers face critical limitations in memory management, load balancing, and handling the highly localized nature of chemical reactions. To this end, we present a high-performance compressible reacting flow solver built on the AMReX framework and optimized for multi-GPU settings. Here, our approach addresses three GPU performance bottlenecks: memory access patterns through column-major storage optimization, computational workload variability via a bulk-sparse integration strategy for chemical kinetics, and multi-GPU load distribution for adaptive mesh refinement applications. The solver adapts existing matrix-based chemical kinetics formulations to multi-grid contexts. Using representative combustion applications, including 2D and 3D detonations and a 3D jet-in-crossflow configuration, we demonstrate 1.4–5× performance improvements over initial implementations on an in-house cluster of NVIDIA H100 GPUs, and near-ideal weak scaling on the Frontier supercomputer (Oak Ridge Leadership Computing Facility) with up to 1024 AMD Instinct MI250X GPUs. Roofline analysis reveals substantial improvements in arithmetic intensity for both convection (∼ 10 ×) and chemistry (∼ 4 ×) routines, confirming efficient utilization of GPU memory bandwidth and computational resources.

42 ENGINEERING↗

Asynchronous Iterative Solvers for Extreme-Scale Computing (Final Report)

This is the final report for the project: Asynchronous Iterative Solvers for Extreme-Scale Computing. This was a collaborative project. This report only covers the activities specific to Georgia Institute of Technology. The project investigated and developed iterative solvers that operate asynchronously, thereby avoiding the high cost of synchronization that is apparent when using standard, synchronous iterative solvers at extreme levels of parallelism.

97 MATHEMATICS AND COMPUTING↗

Performance Evaluation of hypre Solvers

This report compares GPU and CPU performance of structured and unstructured hypre solvers on Lassen and Summit. We applied the solvers to various problems of different sizes varying the number of nodes and processes. We provide the individual results for setup phase and solve phase, since preconditioned solvers are often used in different settings. While for a single system solve the total time is determined by the sum of setup and solve time, this can change for different scenarios. For example, in time dependent problems the preconditioner might have to be setup only occasionally or possibly even only once. In that situation solve times will dominate. Thus, since often better GPU speedups can be attained in the solve phase, this will lead to generally better speedups.

97 MATHEMATICS AND COMPUTING↗

MOSCATO Solver Development and Integration Plan

During FY21, we conducted ongoing development work for the MOSCATO (Molten Salt Chemistry and Transport) solver. The code development work primarily consisted of transitioning capabilities from the original version of the solver, which was written in OpenFOAM, into Nek5000. In doing so, a fast, highly parallelizable solver was created that is capable of complex chemistry and corrosion simulations for engineering-scale molten salt systems. The Nek5000 version of MOSCATO is now fully featured and capable of higher-fidelity simulations than were previously possible. Demonstration cases including a thermal convection loop have been simulated to test these new capabilities. Although capable of large-scale simulations, MOSCATO is not well-suited to parametric studies of complete reactor geometries. These types of simulations are instead better handled by reduced-order modeling codes such as ORNL’s Mole code. Reduced-order simulation tools like Mole, however, are dependent on high fidelity correlations to account for complex, coupled three-dimensional phenomena that they do not directly simulate. Tools such as MOSCATO must therefore be used to create these correlations, as suitable empirical relationships are not available for most molten salt systems. Toward that end, we used the Nek-derived version of MOSCATO to create new mass transfer correlations for three relevant cases including tubular, tube bank, and subchannel geometries. These new correlations are more accurate than any existing ones and can be readily integrated into any reduced-order modeling tools that are targeting full-scale MSR simulations.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Refining Processing Engines from SAPHIRE: Initialization of Fault Tree/Event Tree Solver

SAPHIRE has been extensively employed for over 35 years to model risk-important systems and quantify risk models. As a well-established and thoroughly documented Probabilistic Risk Assessment (PRA) tool, SAPHIRE has continuously tracked computational trends and received regular updates. Despite its ongoing evolution, there remains a need for further enhancements, particularly in dealing with the quantification of exceptionally large models. These improvements could take the form of algorithmic advancements, harnessing the power of parallel computing, and exploring the potential benefits of cloud computing solutions. Considering these aspirations, the notion of a remote solve option was introduced and subsequently evolved into a dedicated project within the SAPHIRE development team. A significant outcome of this initiative is SAPHSOLVE, an engine extracted from the SageRisk API designed specifically for remote solving capabilities. The ongoing project is nearing its culmination, marked by a series of discoveries that have brought undocumented aspects to light. Among these revelations is the intricacy of the input and output format for the SAPHSOLVE engine. This document serves the crucial purpose of meticulously delineating the precise formats for both input and output, as they form an indispensable foundation. The importance of documenting these formats cannot be overstated, as it is a pivotal step in facilitating rigorous testing and comparison. Whether it involves scrutinizing SAPHSOLVE results against those of the internal integrated solver or other external solvers, the ability to construct models or transform existing ones into a compatible SAPHSOLVE format is imperative. Chapter 1 offers a succinct introduction to both SAPHIRE and SAPHSOLVE, followed by Chapter 2 which outlines the roadmap for enhancing SAPHSOLVE. The core of this report is Chapter 3, which intricately elucidates the intricacies of the input and output file formats. To provide a tangible illustration of these formats, a rudimentary example has been compiled and is available in Appendix. SAPHSOLVE represents a novel external solving mechanism developed by the SAPHIRE team, although it has not yet reached the full spectrum of capabilities possessed by SAPHIRE's internal solver. However, the SAPHIRE team has set a comprehensive course for incorporating the functionalities of SAPHSOLVE. A comprehensive outlook on the future of SAPHSOLVE is expounded upon in Chapter 4.

97 MATHEMATICS AND COMPUTING↗

Variational quantum solver employing the PDS energy functional

In our previous work (J. Chem. Phys. 2020, 153, 201102),we reported a new class of quantum algorithms that are based on the quantum computation of the connected moment expansion to find the ground and excited state energies. In particular, the Peeters-Devreese-Soldatov (PDS) formulation is found variational and bearing the potential for further combining with the existing variational quantum infrastructure. Following this direction, here we propose a variational quantum solver employing the PDS energy gradient. In comparison with the usual variational quantum eigensolver (VQE) and the original static PDS approach, the proposed variational quantum solver offers an effective approach to achieve high ac-curacy at finding the ground state and its energy through the rotation of the trial wave function of modest quality guided by the low order PDS energy gradients, thus improves the ac-curacy and efficiency of the quantum simulation. We demonstrate the performance of the proposed variational quantum solver for toy models, H2molecule, and strongly correlated planar H4system in some challenging situations. In all the case studies, the proposed variational quantum approach out-performs the usual VQE and static PDS calculations even at the lowest order.

Peng, Bo↗

Current Status of the Finite-Element Fluid Solver (COFFE) within HPCMP CREATE™-AV Kestrel

COFFE is the finite-element flow solver within HPCMP CREATE-AV™ Kestrel. Kestrel supports a range of flow solver fidelity options to support the DoD acquisition community, and COFFE targets the need for high-fidelity flow solutions. The COFFE solver, like all components of Kestrel, is under continual development as part of the CREATE-AV program. This paper documents the usage of COFFE on several workshop cases, which demonstrate new Kestrel capabilities and exercise features unique to COFFE. It concludes with a brief discussion of upcoming features in development for COFFE.

Holst, Kevin R.↗

Developing a Vorticity-Velocity-Based Off-Body Solver to Perform Multifidelity Simulations of Wind Farms

Wind power has become a key player in satisfying the global energy needs. With increased market penetration, unanticipated unsteady loading induced failures, installation related reductions in power generation, and significant maintenance costs have underscored the need to predict the unsteady fluid-structure interactions related to turbine layout and off-design wind conditions. Contemporary turbine design tools are incapable of accounting for such loadings. As a result, researchers have started utilizing high-Performance-Computing (HPC) based Computational Fluid Dynamics (CFD) solvers, such as the U.S. Department of Energy sponsored ExaWind software package, to investigate these phenomena. Unfortunately, such HPC tools are computationally expensive for routine industrial use, often because of the sheer number of cells required to resolve the wake flowfield. This paper describes a preliminary effort to address this issue by developing a vorticity-velocity based CFD off-body solver, VorTran-M2-AMReX, that integrates directly with DOE's ExaWind wind turbine analysis system to perform accurate and reliable simulations of wind turbine/farm at a lower computational cost than ExaWind alone. This article summarizes work undertaken to date concerning the assembly of the proposed analysis tool, and provides preliminary validation and verification of the VorTran-M2-AMReX off-body solver.

adaptive mesh refinement↗

A Scaling Study for Incompressible Multispecies Solver in Vertex-CFD

Multispecies incompressible flows occur widely in engineering and environmental applications, such as chemical reactors, fuel cells, ocean mixing, and biomedical systems. However, accurately resolving the complex transport and mixing phenomena associated with multiple interacting species remains computationally challenging, especially for large-scale problems. In this study, we present a robust, high-performance computing--enabled multispecies incompressible Navier–Stokes solver integrated within the Vertex-CFD framework. Our solver employs a fully coupled, implicit, finite element--based formulation that accurately captures the advection, diffusion, and interaction of multiple species in incompressible flows by leveraging the Kokkos library for parallel computing to achieve high computational efficiency. For pressure coupling, the entropically damped artificial compressibility method is utilized. We validated the solver against canonical test cases, including multispecies advection, diffusion, and Bateman systems; the results demonstrate second- and third-order spatial accuracy and consistent convergence. Additionally, we demonstrated the strong and weak scaling study results obtained on the leadership-class high-performance computing system, Frontier at Oak Ridge National Laboratory.

Oz, Furkan [ORNL] (ORCID:0000000265831724)↗

Refactoring the elastic–viscous–plastic solver from the sea ice model CICE v6.5.1 for improved performance

This study focuses on the performance of the elastic–viscous–plastic (EVP) dynamical solver within the sea ice model, CICE v6.5.1. The study has been conducted in two steps. First, the standard EVP solver was extracted from CICE for experiments with refactored versions, which are used for performance testing. Second, one refactored version was integrated and tested in the full CICE model to demonstrate that the new algorithms do not significantly impact the physical results. The study reveals two dominant bottlenecks, namely (1) the number of Message Parsing Interface (MPI) and Open Multi-Processing (OpenMP) synchronization points required for halo exchanges during each time step combined with the irregular domain of active sea ice points and (2) the lack of single-instruction, multiple-data (SIMD) code generation. The standard EVP solver has been refactored based on two generic patterns. The first pattern exposes how general finite differences on masked multi-dimensional arrays can be expressed in order to produce significantly better code generation by changing the memory access pattern from random access to direct access. The second pattern takes an alternative approach to handle static grid properties. The measured single-core performance improvement is more than a factor of 5 compared to the standard implementation. The refactored implementation of strong scales on the Intel® Xeon® Scalable Processors series node until the available bandwidth of the node is used. For the Intel® Xeon® CPU Max series, there is sufficient bandwidth to allow the strong scaling to continue for all the cores on the node, resulting in a single-node improvement factor of 35 over the standard implementation. This study also demonstrates improved performance on GPU processors.

58 GEOSCIENCES↗

Approximate Inverse Chain Preconditioner: Iteration Count Case Study for Spectral Support Solvers

As the growing availability of computational power slows, there has been an increasing reliance on algorithmic advances. However, faster algorithms alone will not necessarily bridge the gap in allowing computational scientists to study problems at the edge of scientific discovery in the next several decades. Often, it is necessary to simplify or precondition solvers to accelerate the study of large systems of linear equations commonly seen in a number of scientific fields. Preconditioning a problem to increase efficiency is often seen as the best approach; yet, preconditioners which are fast, smart, and efficient do not always exist. Following the progress of [1], we present a new preconditioner for symmetric diagonally dominant (SDD) systems of linear equations. These systems are common in certain PDEs, network science, and supervised learning among others. Based on spectral support graph theory, this new preconditioner builds off of the work of [2], computing and applying a V-cycle chain of approximate inverse matrices. This preconditioner approach is both algebraic in nature as well as hierarchically-constrained depending on the condition number of the system to be solved. Due to its generation of an Approximate Inverse Chain of matrices, we refer to this as the AIC preconditioner. We further accelerate the AIC preconditioner by utilizing precomputations to simplify setup and multiplications in the con-text of an iterative Krylov-subspace solver. While these iterative solvers can greatly reduce solution time, the number of iterations can grow large quickly in the absence of good preconditioners. Initial results for the AIC preconditioner have shown a very large reduction in iteration counts for SDD systems as compared to standard preconditioners such as Incomplete Cholesky (ICC) and Multigrid (MG). We further show significant reduction in iteration counts against the more advanced Combinatorial Multigrid (CMG) preconditioner. We have further developed no-fill sparsification techniques to ensure that the computational cost of applying the AIC preconditioner does not grow prohibitively large as the depth of the V-cycle grows for systems with larger condition numbers. Our numerical results have shown that these sparsifiers maintain the sparsity structure of our system while also displaying significant reductions in iteration counts.1 2

97 MATHEMATICS AND COMPUTING↗

Approximate Inverse Chain Preconditioner: Iteration Count Case Study for Spectral Support Solvers

As the growing availability of computational power slows, there has been an increasing reliance on algorithmic advances. However, faster algorithms alone will not necessarily bridge the gap in allowing computational scientists to study problems at the edge of scientific discovery in the next several decades. Often, it is necessary to simplify or precondition solvers to accelerate the study of large systems of linear equations commonly seen in a number of scientific fields. Preconditioning a problem to increase efficiency is often seen as the best approach; yet, preconditioners which are fast, smart, and efficient do not always exist. Following the progress of [1], we present a new preconditioner for symmetric diagonally dominant (SDD) systems of linear equations. These systems are common in certain PDEs, network science, and supervised learning among others. Based on spectral support graph theory, this new preconditioner builds off of the work of [2], computing and applying a V-cycle chain of approximate inverse matrices. This preconditioner approach is both algebraic in nature as well as hierarchically-constrained depending on the condition number of the system to be solved. Due to its generation of an Approximate Inverse Chain of matrices, we refer to this as the AIC preconditioner. We further accelerate the AIC preconditioner by utilizing precomputations to simplify setup and multiplications in the con-text of an iterative Krylov-subspace solver. While these iterative solvers can greatly reduce solution time, the number of iterations can grow large quickly in the absence of good preconditioners. Initial results for the AIC preconditioner have shown a very large reduction in iteration counts for SDD systems as compared to standard preconditioners such as Incomplete Cholesky (ICC) and Multigrid (MG). We further show significant reduction in iteration counts against the more advanced Combinatorial Multigrid (CMG) preconditioner. We have further developed no-fill sparsification techniques to ensure that the computational cost of applying the AIC preconditioner does not grow prohibitively large as the depth of the V-cycle grows for systems with larger condition numbers. Our numerical results have shown that these sparsifiers maintain the sparsity structure of our system while also displaying significant reductions in iteration counts.1 2

97 MATHEMATICS AND COMPUTING↗

An Orthogonal Recursive Bisection (ORB) Based Time Advancement Algorithm for CFD-DEM Solvers

The time integration of the granular phase in coupled computational fluid dynamics (CFD) – discrete element method (DEM) simulations presents a unique computational challenge brought about by the large variations in particle collisional time scales. Particles in the dilute regions of the computational domain can be advanced with large time steps while dense regions require much smaller time increments. However, the time step size in most solvers is globally set as the limit for accuracy and stability imposed by the collisions and is typically orders of magnitude less than that required away from collisions. This work addresses this precise issue and provides a strategy to avoid the use of a global conservative small time step size for the entire set of particles.A novel time stepping algorithm for CFD-DEM solvers using a partitioning approach using orthogonal recursive bisection (ORB) that allows for variable time steps among particles is described and its computational performance is compared against baseline explicit methods, typically used in several CFD-DEM solvers. ORB has advantages of being relatively quick and easy to update incrementally and has the required heuristic behavior (i.e., it will split the region in half with a cluster on each side) when groups of particles are well separated (clustered). The algorithm presented in this work uses a local time stepping approach to resolve collisional time scales for subsets of particles that are present at the leaves of the ORB, thereby resulting in substantial reduction of computational cost. The parallel implementation of this method where a ``knapsack” algorithm is used in tandem with ORB for effective load-balancing is also presented, where a best possible partitioning is obtained based on number of particles and local time-stepping costs. The algorithm is tested against benchmark problems with varying particle distributions that include fluidized bed and riser flow scenarios. Preliminary results indicate that the approach is 2-3X faster than traditional explicit methods for problems that involve both dense and dilute regions, while maintaining the same level of accuracy.

adaptive timestepping↗

Mesoflow: An Open-Source Reacting Flow Solver for Catalysis at Mesoscale

We present the capabilities and software performance metrics of our open-source continuum solver for catalysis, Mesoflow, developed specifically for modeling transport and chemistry at the mesoscale. Our solver utilizes Cartesian block-structured adaptive mesh refinement to resolve complex catalyst surface morphologies directly obtained from X-ray tomography data. An immersed boundary based formulation enables rapid representation of complex geometries prevalent in most mesoporous catalyst interfaces. The solver is developed on top of open-source performance portable library, AMReX, providing parallel execution capabilities on current and upcoming high-performance-computing (HPC) architectures. Our flexible software framework enables integration of complex chemical mechanisms at heterogenous interfaces and time-split algorithms for circumventing highly disparate reaction and flow time-scales. Our current studies indicate a ten-fold performance gain by using graphics-processing-units (GPUs) compared to a single processor for representative problem sizes (2 million cell mesh). We will also present a brief introduction on how to build and use this software for application problems pertaining to catalytic upgrading and gas transport within porous catalyst particles.

adaptive meshing↗

Implementation of hybrid finite element method based transport solver in GRIFFIN

A new transport solver option based on the hybrid FEM (HFEM) was implemented in GRIFFIN, the MOOSE-based reactor analysis code, as an effort to support routine core design calculations for advanced reactor applications. The HFEM formulation with P{sub N} (spherical harmonics expansion), akin to the variational nodal method, is effective for solving a spatially homogenized problem with strong transport effect. The residual and Jacobian evaluations of the HFEM weak form were derived and successfully implemented in GRIFFIN, having the diffusion and the PN options available in the new HFEM based transport solver. The performance was tested with the simplified ABTR benchmark problems. The results indicate that the HFEM-based transport solver is a feasible option for solving problems with spatially homogenized and strong streaming by providing superior accuracy with a proper p-refinement. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Implementation of Surface Tension on a Reacting Flow Solver, PeleLM: Preprint

In liquid rocket engines, the fuel is supplied to the combustion chamber in the liquid state though injectors. Such fuel undergoes atomization, vaporization, and combustion processes. To design reliable and efficient injectors, it is required to understand the full processes. This research is part of an effort to develop a full atomization-vaporization-combustion solver from first principles. As an initial step to tackle the atomization process, a multiphase flow solver is under development. For the development, a library of the volume of fluid scheme for multiphase, IRL is coupled with a reacting Navier-Stokes equation solver, PeleLM. Furthermore, as the surface tension has considerable effects on spray breakup. surface tension is implemented in the momentum equation using the continuum surface force model and the improved height function technique.

height function↗