Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Solvers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Refining Processing Engines from SAPHIRE: Initialization of Fault Tree/Event Tree Solver

SAPHIRE has been extensively employed for over 35 years to model risk-important systems and quantify risk models. As a well-established and thoroughly documented Probabilistic Risk Assessment (PRA) tool, SAPHIRE has continuously tracked computational trends and received regular updates. Despite its ongoing evolution, there remains a need for further enhancements, particularly in dealing with the quantification of exceptionally large models. These improvements could take the form of algorithmic advancements, harnessing the power of parallel computing, and exploring the potential benefits of cloud computing solutions. Considering these aspirations, the notion of a remote solve option was introduced and subsequently evolved into a dedicated project within the SAPHIRE development team. A significant outcome of this initiative is SAPHSOLVE, an engine extracted from the SageRisk API designed specifically for remote solving capabilities. The ongoing project is nearing its culmination, marked by a series of discoveries that have brought undocumented aspects to light. Among these revelations is the intricacy of the input and output format for the SAPHSOLVE engine. This document serves the crucial purpose of meticulously delineating the precise formats for both input and output, as they form an indispensable foundation. The importance of documenting these formats cannot be overstated, as it is a pivotal step in facilitating rigorous testing and comparison. Whether it involves scrutinizing SAPHSOLVE results against those of the internal integrated solver or other external solvers, the ability to construct models or transform existing ones into a compatible SAPHSOLVE format is imperative. Chapter 1 offers a succinct introduction to both SAPHIRE and SAPHSOLVE, followed by Chapter 2 which outlines the roadmap for enhancing SAPHSOLVE. The core of this report is Chapter 3, which intricately elucidates the intricacies of the input and output file formats. To provide a tangible illustration of these formats, a rudimentary example has been compiled and is available in Appendix. SAPHSOLVE represents a novel external solving mechanism developed by the SAPHIRE team, although it has not yet reached the full spectrum of capabilities possessed by SAPHIRE's internal solver. However, the SAPHIRE team has set a comprehensive course for incorporating the functionalities of SAPHSOLVE. A comprehensive outlook on the future of SAPHSOLVE is expounded upon in Chapter 4.

97 MATHEMATICS AND COMPUTING↗

Variational quantum solver employing the PDS energy functional

In our previous work (J. Chem. Phys. 2020, 153, 201102),we reported a new class of quantum algorithms that are based on the quantum computation of the connected moment expansion to find the ground and excited state energies. In particular, the Peeters-Devreese-Soldatov (PDS) formulation is found variational and bearing the potential for further combining with the existing variational quantum infrastructure. Following this direction, here we propose a variational quantum solver employing the PDS energy gradient. In comparison with the usual variational quantum eigensolver (VQE) and the original static PDS approach, the proposed variational quantum solver offers an effective approach to achieve high ac-curacy at finding the ground state and its energy through the rotation of the trial wave function of modest quality guided by the low order PDS energy gradients, thus improves the ac-curacy and efficiency of the quantum simulation. We demonstrate the performance of the proposed variational quantum solver for toy models, H2molecule, and strongly correlated planar H4system in some challenging situations. In all the case studies, the proposed variational quantum approach out-performs the usual VQE and static PDS calculations even at the lowest order.

Peng, Bo↗

Current Status of the Finite-Element Fluid Solver (COFFE) within HPCMP CREATE™-AV Kestrel

COFFE is the finite-element flow solver within HPCMP CREATE-AV™ Kestrel. Kestrel supports a range of flow solver fidelity options to support the DoD acquisition community, and COFFE targets the need for high-fidelity flow solutions. The COFFE solver, like all components of Kestrel, is under continual development as part of the CREATE-AV program. This paper documents the usage of COFFE on several workshop cases, which demonstrate new Kestrel capabilities and exercise features unique to COFFE. It concludes with a brief discussion of upcoming features in development for COFFE.

Holst, Kevin R.↗

Developing a Vorticity-Velocity-Based Off-Body Solver to Perform Multifidelity Simulations of Wind Farms

Wind power has become a key player in satisfying the global energy needs. With increased market penetration, unanticipated unsteady loading induced failures, installation related reductions in power generation, and significant maintenance costs have underscored the need to predict the unsteady fluid-structure interactions related to turbine layout and off-design wind conditions. Contemporary turbine design tools are incapable of accounting for such loadings. As a result, researchers have started utilizing high-Performance-Computing (HPC) based Computational Fluid Dynamics (CFD) solvers, such as the U.S. Department of Energy sponsored ExaWind software package, to investigate these phenomena. Unfortunately, such HPC tools are computationally expensive for routine industrial use, often because of the sheer number of cells required to resolve the wake flowfield. This paper describes a preliminary effort to address this issue by developing a vorticity-velocity based CFD off-body solver, VorTran-M2-AMReX, that integrates directly with DOE's ExaWind wind turbine analysis system to perform accurate and reliable simulations of wind turbine/farm at a lower computational cost than ExaWind alone. This article summarizes work undertaken to date concerning the assembly of the proposed analysis tool, and provides preliminary validation and verification of the VorTran-M2-AMReX off-body solver.

adaptive mesh refinement↗

A Scaling Study for Incompressible Multispecies Solver in Vertex-CFD

Multispecies incompressible flows occur widely in engineering and environmental applications, such as chemical reactors, fuel cells, ocean mixing, and biomedical systems. However, accurately resolving the complex transport and mixing phenomena associated with multiple interacting species remains computationally challenging, especially for large-scale problems. In this study, we present a robust, high-performance computing--enabled multispecies incompressible Navier–Stokes solver integrated within the Vertex-CFD framework. Our solver employs a fully coupled, implicit, finite element--based formulation that accurately captures the advection, diffusion, and interaction of multiple species in incompressible flows by leveraging the Kokkos library for parallel computing to achieve high computational efficiency. For pressure coupling, the entropically damped artificial compressibility method is utilized. We validated the solver against canonical test cases, including multispecies advection, diffusion, and Bateman systems; the results demonstrate second- and third-order spatial accuracy and consistent convergence. Additionally, we demonstrated the strong and weak scaling study results obtained on the leadership-class high-performance computing system, Frontier at Oak Ridge National Laboratory.

Oz, Furkan [ORNL] (ORCID:0000000265831724)↗

Refactoring the elastic–viscous–plastic solver from the sea ice model CICE v6.5.1 for improved performance

This study focuses on the performance of the elastic–viscous–plastic (EVP) dynamical solver within the sea ice model, CICE v6.5.1. The study has been conducted in two steps. First, the standard EVP solver was extracted from CICE for experiments with refactored versions, which are used for performance testing. Second, one refactored version was integrated and tested in the full CICE model to demonstrate that the new algorithms do not significantly impact the physical results. The study reveals two dominant bottlenecks, namely (1) the number of Message Parsing Interface (MPI) and Open Multi-Processing (OpenMP) synchronization points required for halo exchanges during each time step combined with the irregular domain of active sea ice points and (2) the lack of single-instruction, multiple-data (SIMD) code generation. The standard EVP solver has been refactored based on two generic patterns. The first pattern exposes how general finite differences on masked multi-dimensional arrays can be expressed in order to produce significantly better code generation by changing the memory access pattern from random access to direct access. The second pattern takes an alternative approach to handle static grid properties. The measured single-core performance improvement is more than a factor of 5 compared to the standard implementation. The refactored implementation of strong scales on the Intel® Xeon® Scalable Processors series node until the available bandwidth of the node is used. For the Intel® Xeon® CPU Max series, there is sufficient bandwidth to allow the strong scaling to continue for all the cores on the node, resulting in a single-node improvement factor of 35 over the standard implementation. This study also demonstrates improved performance on GPU processors.

58 GEOSCIENCES↗

Approximate Inverse Chain Preconditioner: Iteration Count Case Study for Spectral Support Solvers

As the growing availability of computational power slows, there has been an increasing reliance on algorithmic advances. However, faster algorithms alone will not necessarily bridge the gap in allowing computational scientists to study problems at the edge of scientific discovery in the next several decades. Often, it is necessary to simplify or precondition solvers to accelerate the study of large systems of linear equations commonly seen in a number of scientific fields. Preconditioning a problem to increase efficiency is often seen as the best approach; yet, preconditioners which are fast, smart, and efficient do not always exist. Following the progress of [1], we present a new preconditioner for symmetric diagonally dominant (SDD) systems of linear equations. These systems are common in certain PDEs, network science, and supervised learning among others. Based on spectral support graph theory, this new preconditioner builds off of the work of [2], computing and applying a V-cycle chain of approximate inverse matrices. This preconditioner approach is both algebraic in nature as well as hierarchically-constrained depending on the condition number of the system to be solved. Due to its generation of an Approximate Inverse Chain of matrices, we refer to this as the AIC preconditioner. We further accelerate the AIC preconditioner by utilizing precomputations to simplify setup and multiplications in the con-text of an iterative Krylov-subspace solver. While these iterative solvers can greatly reduce solution time, the number of iterations can grow large quickly in the absence of good preconditioners. Initial results for the AIC preconditioner have shown a very large reduction in iteration counts for SDD systems as compared to standard preconditioners such as Incomplete Cholesky (ICC) and Multigrid (MG). We further show significant reduction in iteration counts against the more advanced Combinatorial Multigrid (CMG) preconditioner. We have further developed no-fill sparsification techniques to ensure that the computational cost of applying the AIC preconditioner does not grow prohibitively large as the depth of the V-cycle grows for systems with larger condition numbers. Our numerical results have shown that these sparsifiers maintain the sparsity structure of our system while also displaying significant reductions in iteration counts.1 2

97 MATHEMATICS AND COMPUTING↗

Approximate Inverse Chain Preconditioner: Iteration Count Case Study for Spectral Support Solvers

As the growing availability of computational power slows, there has been an increasing reliance on algorithmic advances. However, faster algorithms alone will not necessarily bridge the gap in allowing computational scientists to study problems at the edge of scientific discovery in the next several decades. Often, it is necessary to simplify or precondition solvers to accelerate the study of large systems of linear equations commonly seen in a number of scientific fields. Preconditioning a problem to increase efficiency is often seen as the best approach; yet, preconditioners which are fast, smart, and efficient do not always exist. Following the progress of [1], we present a new preconditioner for symmetric diagonally dominant (SDD) systems of linear equations. These systems are common in certain PDEs, network science, and supervised learning among others. Based on spectral support graph theory, this new preconditioner builds off of the work of [2], computing and applying a V-cycle chain of approximate inverse matrices. This preconditioner approach is both algebraic in nature as well as hierarchically-constrained depending on the condition number of the system to be solved. Due to its generation of an Approximate Inverse Chain of matrices, we refer to this as the AIC preconditioner. We further accelerate the AIC preconditioner by utilizing precomputations to simplify setup and multiplications in the con-text of an iterative Krylov-subspace solver. While these iterative solvers can greatly reduce solution time, the number of iterations can grow large quickly in the absence of good preconditioners. Initial results for the AIC preconditioner have shown a very large reduction in iteration counts for SDD systems as compared to standard preconditioners such as Incomplete Cholesky (ICC) and Multigrid (MG). We further show significant reduction in iteration counts against the more advanced Combinatorial Multigrid (CMG) preconditioner. We have further developed no-fill sparsification techniques to ensure that the computational cost of applying the AIC preconditioner does not grow prohibitively large as the depth of the V-cycle grows for systems with larger condition numbers. Our numerical results have shown that these sparsifiers maintain the sparsity structure of our system while also displaying significant reductions in iteration counts.1 2

97 MATHEMATICS AND COMPUTING↗

An Orthogonal Recursive Bisection (ORB) Based Time Advancement Algorithm for CFD-DEM Solvers

The time integration of the granular phase in coupled computational fluid dynamics (CFD) – discrete element method (DEM) simulations presents a unique computational challenge brought about by the large variations in particle collisional time scales. Particles in the dilute regions of the computational domain can be advanced with large time steps while dense regions require much smaller time increments. However, the time step size in most solvers is globally set as the limit for accuracy and stability imposed by the collisions and is typically orders of magnitude less than that required away from collisions. This work addresses this precise issue and provides a strategy to avoid the use of a global conservative small time step size for the entire set of particles.A novel time stepping algorithm for CFD-DEM solvers using a partitioning approach using orthogonal recursive bisection (ORB) that allows for variable time steps among particles is described and its computational performance is compared against baseline explicit methods, typically used in several CFD-DEM solvers. ORB has advantages of being relatively quick and easy to update incrementally and has the required heuristic behavior (i.e., it will split the region in half with a cluster on each side) when groups of particles are well separated (clustered). The algorithm presented in this work uses a local time stepping approach to resolve collisional time scales for subsets of particles that are present at the leaves of the ORB, thereby resulting in substantial reduction of computational cost. The parallel implementation of this method where a ``knapsack” algorithm is used in tandem with ORB for effective load-balancing is also presented, where a best possible partitioning is obtained based on number of particles and local time-stepping costs. The algorithm is tested against benchmark problems with varying particle distributions that include fluidized bed and riser flow scenarios. Preliminary results indicate that the approach is 2-3X faster than traditional explicit methods for problems that involve both dense and dilute regions, while maintaining the same level of accuracy.

adaptive timestepping↗

Mesoflow: An Open-Source Reacting Flow Solver for Catalysis at Mesoscale

We present the capabilities and software performance metrics of our open-source continuum solver for catalysis, Mesoflow, developed specifically for modeling transport and chemistry at the mesoscale. Our solver utilizes Cartesian block-structured adaptive mesh refinement to resolve complex catalyst surface morphologies directly obtained from X-ray tomography data. An immersed boundary based formulation enables rapid representation of complex geometries prevalent in most mesoporous catalyst interfaces. The solver is developed on top of open-source performance portable library, AMReX, providing parallel execution capabilities on current and upcoming high-performance-computing (HPC) architectures. Our flexible software framework enables integration of complex chemical mechanisms at heterogenous interfaces and time-split algorithms for circumventing highly disparate reaction and flow time-scales. Our current studies indicate a ten-fold performance gain by using graphics-processing-units (GPUs) compared to a single processor for representative problem sizes (2 million cell mesh). We will also present a brief introduction on how to build and use this software for application problems pertaining to catalytic upgrading and gas transport within porous catalyst particles.

adaptive meshing↗

Implementation of hybrid finite element method based transport solver in GRIFFIN

A new transport solver option based on the hybrid FEM (HFEM) was implemented in GRIFFIN, the MOOSE-based reactor analysis code, as an effort to support routine core design calculations for advanced reactor applications. The HFEM formulation with P{sub N} (spherical harmonics expansion), akin to the variational nodal method, is effective for solving a spatially homogenized problem with strong transport effect. The residual and Jacobian evaluations of the HFEM weak form were derived and successfully implemented in GRIFFIN, having the diffusion and the PN options available in the new HFEM based transport solver. The performance was tested with the simplified ABTR benchmark problems. The results indicate that the HFEM-based transport solver is a feasible option for solving problems with spatially homogenized and strong streaming by providing superior accuracy with a proper p-refinement. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Implementation of Surface Tension on a Reacting Flow Solver, PeleLM: Preprint

In liquid rocket engines, the fuel is supplied to the combustion chamber in the liquid state though injectors. Such fuel undergoes atomization, vaporization, and combustion processes. To design reliable and efficient injectors, it is required to understand the full processes. This research is part of an effort to develop a full atomization-vaporization-combustion solver from first principles. As an initial step to tackle the atomization process, a multiphase flow solver is under development. For the development, a library of the volume of fluid scheme for multiphase, IRL is coupled with a reacting Navier-Stokes equation solver, PeleLM. Furthermore, as the surface tension has considerable effects on spray breakup. surface tension is implemented in the momentum equation using the continuum surface force model and the improved height function technique.

height function↗

A fully implicit, scalable, conservative nonlinear relativistic Fokker–Planck 0D-2P solver for runaway electrons

Upon application of a sufficiently strong electric field, electrons break away from thermal equilibrium and approach relativistic speeds. These highly energetic ‘runaway’ electrons (~ MeV) play a significant role in tokamak disruption physics, and therefore their accurate understanding is essential to develop reliable mitigation strategies. As such, we have developed a fully implicit solver for the 0D-2P (i.e., including two momenta coordinates) relativistic nonlinear Fokker–Planck equation (rFP). As in earlier implicit rFP studies (NORSE, CQL3D), electron–ion interactions are modeled using the Lorentz operator, and synchrotron damping using the Abraham–Lorentz–Dirac reaction term. However, our implementation improves on these earlier studies by (1) ensuring exact conservation properties for electron collisions, (2) strictly preserving positivity, and (3) being scalable algorithmically and in parallel. Key to our proposed approach is an efficient multigrid preconditioner for the linearized rFP equation, a multigrid elliptic solver for the Braams–Karney potentials, and a novel adaptive technique to determine the associated boundary values. We verify the accuracy and efficiency of the proposed scheme with numerical results ranging from small electric-field electrical conductivity measurements to the accurate reproduction of runaway tail dynamics when strong electric fields are applied.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Matrix-Free High-Performance Saddle-Point Solvers for High-Order Problems in \(\boldsymbol{H}(\operatorname{\textbf{div}})\)

Here, this work describes the development of matrix-free GPU-accelerated solvers for high-order finite element problems in H(div). The solvers are applicable to grad-div and Darcy problems in saddle-point formulation, and have applications in radiation diffusion and porous media flow problems, among others. Using the interpolation–histopolation basis, efficient matrix-free preconditioners can be constructed for the (1, 1)-block and Schur complement of the block system. With these approximations, block-preconditioned MINRES converges in a number of iterations that is independent of the mesh size and polynomial degree. The approximate Schur complement takes the form of an M-matrix graph Laplacian and therefore can be well-preconditioned by highly scalable algebraic multigrid methods. High-performance GPU-accelerated algorithms for all components of the solution algorithm are developed, discussed, and benchmarked. Numerical results are presented on a number of challenging test cases, including the “crooked pipe” grad-div problem, the SPE10 reservoir modeling benchmark problem, and a nonlinear radiation diffusion test case.

97 MATHEMATICS AND COMPUTING↗

An efficient, conservative, time-implicit solver for the fully kinetic arbitrary-species 1D-2V Vlasov-Ampère system

In this paper, we consider the solution of the fully kinetic (including electrons) Vlasov-Ampère system in a one-dimensional physical space and two-dimensional velocity space (1D-2V) for an arbitrary number of species with a time-implicit Eulerian algorithm. The problem of velocity-space meshing for disparate thermal and bulk velocities is dealt with by an adaptive coordinate transformation of the Vlasov equation for each species, which is then discretized, including the resulting inertial terms. Mass, momentum, and energy are conserved, and Gauss's law is enforced to within the nonlinear convergence tolerance of the iterative solver through a set of nonlinear constraint functions while permitting significant flexibility in choosing discretizations in time, configuration, and velocity space. We mitigate the temporal stiffness introduced by, e.g., the plasma frequency through the use of high-order/low-order (HOLO) acceleration of the iterative implicit solver. We present several numerical results for canonical problems of varying degrees of complexity, including the multiscale ion-acoustic shock wave problem, which demonstrate the efficacy, accuracy, and efficiency of the scheme.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Fast nonlinear iterative solver for an implicit, energy-conserving, asymptotic-preserving charged-particle orbit integrator

Here recently, an asymptotic-preserving (AP) particle orbit integrator has been proposed with remarkable properties including exact energy conservation, the ability to capture of all first-order drifts (including the ∇B-drift), the ability to capture trapped-passing boundaries with parallel velocity extremely close to the critical velocity, and the ability to transition from strongly to weakly magnetized spatial regions. The new AP orbit integrator is implicit, employing a Crank-Nicolson (CN) temporal discretization to ensure exact energy conservation. This, in turn, requires a local nonlinear iteration involving particle velocities and positions, and the local electromagnetic fields, to obtain the new-time solution. Ref. [1] did not attempt to provide an efficient solver for this system, and employed a brute-force GMRES-driven Jacobian-free Newton-Krylov (JFNK) solver to invert the particle orbit equations at every timestep for expediency. While JFNK is robust and reliable, it is also expensive and very intrusive for practical implementations of the method (it requires having the JFNK machinery available and solving a 6 x 6 Jacobian system iteratively once per iteration per particle).

97 MATHEMATICS AND COMPUTING↗

High performance sparse multifrontal solvers on modern GPUs

Here, we have ported the numerical factorization and triangular solve phases of the sparse direct solver STRUMPACK to GPU. STRUMPACK implements sparse LU factorization using the multifrontal algorithm, which performs most of its operations in dense linear algebra operations on so-called frontal matrices of various sizes. Our GPU implementation off-loads these dense linear algebra operations, as well as the sparse scatter–gather operations between frontal matrices. For the larger frontal matrices, our GPU implementation relies on vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs and rocBLAS and rocSOLVER for AMD GPUs. For the smaller frontal matrices we developed custom CUDA and HIP kernels to reduce kernel launch overhead. Overall, high performance is achieved by identifying submatrix factorizations corresponding to sub-trees of the multifrontal assembly tree which fit entirely in GPU memory. The multi-GPU setting uses SLATE (Software for Linear Algebra Targeting Exascale) as a modern GPU-aware replacement for ScaLAPACK. On 4 nodes of SUMMIT the code runs ~10X faster when using all 24 V100 GPUs compared to when it only uses the 168 POWER9 cores. On 8 SUMMIT nodes, using 48 V100 GPUs, the sparse solver reaches over 50TFlop/s. Compared to SuperLU, on a single V100, for a set of 17 matrices our implementation is faster for all but one matrix, and is on average 5X (median 4X) faster

97 MATHEMATICS AND COMPUTING↗

Low-synch Gram–Schmidt with delayed reorthogonalization for Krylov solvers

The parallel strong-scaling of iterative methods is often determined by the number of global reductions at each iteration. Low-synch Gram-Schmidt algorithms are applied here to the Arnoldi algorithm to reduce the number of global reductions and therefore to improve the parallel strong-scaling of iterative solvers for nonsymmetric matrices such as the GMRES and the Krylov-Schur iterative methods. In the Arnoldi context, the factorization is "left-looking" and processes one column at a time. Among the methods for generating an orthogonal basis for the Arnoldi algorithm, the classical Gram-Schmidt algorithm, with reorthogonalization (CGS2) requires three global reductions per iteration. A new variant of CGS2 that requires only one reduction per iteration is presented and applied to the Arnoldi algorithm. Delayed CGS2 (DCGS2) employs the minimum number of global reductions per iteration (one) for a one-column at-a-time algorithm. The main idea behind the new algorithm is to group global reductions by rearranging the order of operations. DCGS2 must be carefully integrated into an Arnoldi expansion or a GMRES solver. Numerical stability experiments assess robustness for Krylov-Schur eigenvalue computations. Performance experiments on the ORNL Summit supercomputer then establish the superiority of DCGS2 over CGS2.

97 MATHEMATICS AND COMPUTING↗