Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “solver”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Efficient proximal subproblem solvers for a nonsmooth trust-region method

In [R. J. Baraldi and D. P. Kouri, Mathematical Programming, (2022), pp. 1-40], we introduced an inexact trust-region algorithm for minimizing the sum of a smooth nonconvex and nonsmooth convex function. The principle expense of this method is in computing a trial iterate that satisfies the so-called fraction of Cauchy decrease condition—a bound that ensures the trial iterate produces sufficient decrease of the subproblem model. In this paper, we expound on various proximal trust-region subproblem solvers that generalize traditional trust-region methods for smooth unconstrained and convex-constrained problems. We introduce a simplified spectral proximal gradient solver, a truncated nonlinear conjugate gradient solver, and a dogleg method. Finally, we compare algorithm performance on examples from data science and PDE-constrained optimization.

97 MATHEMATICS AND COMPUTING↗

A Finite Difference informed Random Walk solver for simulating radiation defect evolution in polycrystalline structures with strongly inhomogeneous diffusivity

Diffusivity of species and defects on grain boundaries is usually several orders of magnitude larger than that inside grains. Such strongly inhomogeneous diffusivity requires prohibitively high computational demands for modeling microstructural evolution. Here, this paper presents a highly-efficient numerical solver, combining the Finite Difference method and Random Walk model, designed for accurately modeling strongly inhomogeneous diffusion within polycrystalline structures. The proposed solver, termed Finite Difference informed Random Walk (FDiRW), integrates a customized Finite Difference (cFD) scheme tailored for fast diffusion along thin grain boundaries represented by a single-layer of nodes. Numerical experiments demonstrate that the FDiRW solver achieves an impressive efficiency gain of 1560x compared to traditional Finite Difference methods while maintaining accuracy, making it feasible for personal computer machines to handle diffusional systems with strongly inhomogeneous diffusivity across static polycrystalline microstructures. The model has been successfully applied to simulate radiation defect evolution, showcasing its scalability to engineering scales in both length and time dimensions.

36 MATERIALS SCIENCE↗

Implicit and coupled fluid plasma solver with adaptive Cartesian mesh and its applications to non-equilibrium gas discharges

In this work, we present a new fluid plasma solver with adaptive Cartesian mesh (ACM) based on a full-Newton (nonlinear, implicit) scheme for non-equilibrium gas discharge plasma. The electrons and ions are described using drift-diffusion approximation coupled to Poisson equation for the electric field. The electron-energy transport equation is solved to account for electron thermal conductivity, Joule heating, and energy loss of electrons in collisions with neutral species. The rate of electron-induced ionization is a function of electron temperature and could also depend on electron density (important for plasma stratification). The ion and gas temperature are kept constant. The transport equations are discretized using a non-isothermal Scharfetter-Gummel scheme to resolve possible large temperature gradients in the sheaths. We demonstrate the new solver for simulations of direct current (DC) and radiofrequency (RF) discharges. The implicit treatment of the coupled equations allows using large time steps. The full-Newton method (FNM) enables fast nonlinear convergence at each time step, offering significantly improved simulation efficiency. We discuss the selection of time steps for solving different plasma problems. The new solver enables solving several problems we could not solve before with existing software: two- and three-dimensional structures of the entire DC discharges including cathode and anode regions, electric field reversals and double-layer formation, the normal cathode spot and an anode ring, moving striations in diffuse and constricted DC discharges, and standing striations in RF discharges. The developed FNM-ACM technique offers many benefits for tackling the disparity of gas discharge plasma systems' time scales and nonlinearity.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Quantum solver for single-impurity Anderson models with particle-hole symmetry

Quantum embedding methods, such as dynamical mean-field theory (DMFT), provide a powerful framework for investigating strongly correlated materials. A central computational bottleneck in DMFT is in solving the Anderson impurity model (AIM), whose exact solution is classically intractable for large bath sizes. In this work, we benchmark a quantum-classical hybrid solver tailored for particle-hole symmetric AIMs, using the variational quantum eigensolver to prepare the ground state of the model with shallow quantum circuits. The solver uses shallow quantum ansätze and one set of variational parameters to prepare the ground state and its particle and hole excitations, enabling the construction of the impurity Green’s function through a continued-fraction expansion. We evaluate the performance of this approach across a few bath sizes and interaction strengths under noisy, shot-limited conditions. We compare three optimization routines (COBYLA, Adam, and L-BFGS-B) in terms of convergence and fidelity, assess the benefits of estimating a quantum-computed moment correction to the variational energies, and benchmark the approach by comparing the density of states computed from the impurity Green’s function against that obtained using a classical pipeline. Our results demonstrate the feasibility of Green’s function construction on near-term devices and establish practical benchmarks for quantum impurity solvers embedded within self-consistent DMFT loops.

Karabin, Mariia [ORNL]↗

Batched sparse direct solver design and evaluation in SuperLU_DIST

Over the course of interactions with various application teams, the need for batched sparse linear algebra functions has emerged in order to make more efficient use of the GPUs for many small and sparse linear algebra problems. In this paper, we present our recent work on a batched sparse direct solver for GPUs. The sparse LU factorization is computed by the levels of the elimination tree, leveraging the batched dense operations at each level and a new batched Scatter GPU kernel. The sparse triangular solve is computed by the level sets of the directed acyclic graph (DAG) of the triangular matrix. Batched operations overcome the large overhead associated with launching many small kernels. For medium sized matrix batches with not-so-small bandwidth, using an NVIDIA A100 GPU, our new batched sparse direct solver is orders of magnitude faster than a batched banded solver and uses less than one-tenth of the memory.

Boukaram, Wajih↗

A GPU-based compressible combustion solver for applications exhibiting disparate space and time scales

High-speed chemically active flows pose significant computational challenges due to their disparate space and time scales, with stiff chemistry often dominating simulation time. While modern scientific computing programs achieve exascale performance by leveraging graphics processing units (GPUs), existing GPU-based compressible combustion solvers face critical limitations in memory management, load balancing, and handling the highly localized nature of chemical reactions. To this end, we present a high-performance compressible reacting flow solver built on the AMReX framework and optimized for multi-GPU settings. Here, our approach addresses three GPU performance bottlenecks: memory access patterns through column-major storage optimization, computational workload variability via a bulk-sparse integration strategy for chemical kinetics, and multi-GPU load distribution for adaptive mesh refinement applications. The solver adapts existing matrix-based chemical kinetics formulations to multi-grid contexts. Using representative combustion applications, including 2D and 3D detonations and a 3D jet-in-crossflow configuration, we demonstrate 1.4–5× performance improvements over initial implementations on an in-house cluster of NVIDIA H100 GPUs, and near-ideal weak scaling on the Frontier supercomputer (Oak Ridge Leadership Computing Facility) with up to 1024 AMD Instinct MI250X GPUs. Roofline analysis reveals substantial improvements in arithmetic intensity for both convection (∼ 10 ×) and chemistry (∼ 4 ×) routines, confirming efficient utilization of GPU memory bandwidth and computational resources.

42 ENGINEERING↗

Asynchronous Iterative Solvers for Extreme-Scale Computing (Final Report)

This is the final report for the project: Asynchronous Iterative Solvers for Extreme-Scale Computing. This was a collaborative project. This report only covers the activities specific to Georgia Institute of Technology. The project investigated and developed iterative solvers that operate asynchronously, thereby avoiding the high cost of synchronization that is apparent when using standard, synchronous iterative solvers at extreme levels of parallelism.

97 MATHEMATICS AND COMPUTING↗

Performance Evaluation of hypre Solvers

This report compares GPU and CPU performance of structured and unstructured hypre solvers on Lassen and Summit. We applied the solvers to various problems of different sizes varying the number of nodes and processes. We provide the individual results for setup phase and solve phase, since preconditioned solvers are often used in different settings. While for a single system solve the total time is determined by the sum of setup and solve time, this can change for different scenarios. For example, in time dependent problems the preconditioner might have to be setup only occasionally or possibly even only once. In that situation solve times will dominate. Thus, since often better GPU speedups can be attained in the solve phase, this will lead to generally better speedups.

97 MATHEMATICS AND COMPUTING↗

MOSCATO Solver Development and Integration Plan

During FY21, we conducted ongoing development work for the MOSCATO (Molten Salt Chemistry and Transport) solver. The code development work primarily consisted of transitioning capabilities from the original version of the solver, which was written in OpenFOAM, into Nek5000. In doing so, a fast, highly parallelizable solver was created that is capable of complex chemistry and corrosion simulations for engineering-scale molten salt systems. The Nek5000 version of MOSCATO is now fully featured and capable of higher-fidelity simulations than were previously possible. Demonstration cases including a thermal convection loop have been simulated to test these new capabilities. Although capable of large-scale simulations, MOSCATO is not well-suited to parametric studies of complete reactor geometries. These types of simulations are instead better handled by reduced-order modeling codes such as ORNL’s Mole code. Reduced-order simulation tools like Mole, however, are dependent on high fidelity correlations to account for complex, coupled three-dimensional phenomena that they do not directly simulate. Tools such as MOSCATO must therefore be used to create these correlations, as suitable empirical relationships are not available for most molten salt systems. Toward that end, we used the Nek-derived version of MOSCATO to create new mass transfer correlations for three relevant cases including tubular, tube bank, and subchannel geometries. These new correlations are more accurate than any existing ones and can be readily integrated into any reduced-order modeling tools that are targeting full-scale MSR simulations.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Refining Processing Engines from SAPHIRE: Initialization of Fault Tree/Event Tree Solver

SAPHIRE has been extensively employed for over 35 years to model risk-important systems and quantify risk models. As a well-established and thoroughly documented Probabilistic Risk Assessment (PRA) tool, SAPHIRE has continuously tracked computational trends and received regular updates. Despite its ongoing evolution, there remains a need for further enhancements, particularly in dealing with the quantification of exceptionally large models. These improvements could take the form of algorithmic advancements, harnessing the power of parallel computing, and exploring the potential benefits of cloud computing solutions. Considering these aspirations, the notion of a remote solve option was introduced and subsequently evolved into a dedicated project within the SAPHIRE development team. A significant outcome of this initiative is SAPHSOLVE, an engine extracted from the SageRisk API designed specifically for remote solving capabilities. The ongoing project is nearing its culmination, marked by a series of discoveries that have brought undocumented aspects to light. Among these revelations is the intricacy of the input and output format for the SAPHSOLVE engine. This document serves the crucial purpose of meticulously delineating the precise formats for both input and output, as they form an indispensable foundation. The importance of documenting these formats cannot be overstated, as it is a pivotal step in facilitating rigorous testing and comparison. Whether it involves scrutinizing SAPHSOLVE results against those of the internal integrated solver or other external solvers, the ability to construct models or transform existing ones into a compatible SAPHSOLVE format is imperative. Chapter 1 offers a succinct introduction to both SAPHIRE and SAPHSOLVE, followed by Chapter 2 which outlines the roadmap for enhancing SAPHSOLVE. The core of this report is Chapter 3, which intricately elucidates the intricacies of the input and output file formats. To provide a tangible illustration of these formats, a rudimentary example has been compiled and is available in Appendix. SAPHSOLVE represents a novel external solving mechanism developed by the SAPHIRE team, although it has not yet reached the full spectrum of capabilities possessed by SAPHIRE's internal solver. However, the SAPHIRE team has set a comprehensive course for incorporating the functionalities of SAPHSOLVE. A comprehensive outlook on the future of SAPHSOLVE is expounded upon in Chapter 4.

97 MATHEMATICS AND COMPUTING↗

Variational quantum solver employing the PDS energy functional

In our previous work (J. Chem. Phys. 2020, 153, 201102),we reported a new class of quantum algorithms that are based on the quantum computation of the connected moment expansion to find the ground and excited state energies. In particular, the Peeters-Devreese-Soldatov (PDS) formulation is found variational and bearing the potential for further combining with the existing variational quantum infrastructure. Following this direction, here we propose a variational quantum solver employing the PDS energy gradient. In comparison with the usual variational quantum eigensolver (VQE) and the original static PDS approach, the proposed variational quantum solver offers an effective approach to achieve high ac-curacy at finding the ground state and its energy through the rotation of the trial wave function of modest quality guided by the low order PDS energy gradients, thus improves the ac-curacy and efficiency of the quantum simulation. We demonstrate the performance of the proposed variational quantum solver for toy models, H2molecule, and strongly correlated planar H4system in some challenging situations. In all the case studies, the proposed variational quantum approach out-performs the usual VQE and static PDS calculations even at the lowest order.

Peng, Bo↗

Current Status of the Finite-Element Fluid Solver (COFFE) within HPCMP CREATE™-AV Kestrel

COFFE is the finite-element flow solver within HPCMP CREATE-AV™ Kestrel. Kestrel supports a range of flow solver fidelity options to support the DoD acquisition community, and COFFE targets the need for high-fidelity flow solutions. The COFFE solver, like all components of Kestrel, is under continual development as part of the CREATE-AV program. This paper documents the usage of COFFE on several workshop cases, which demonstrate new Kestrel capabilities and exercise features unique to COFFE. It concludes with a brief discussion of upcoming features in development for COFFE.

Holst, Kevin R.↗

Developing a Vorticity-Velocity-Based Off-Body Solver to Perform Multifidelity Simulations of Wind Farms

Wind power has become a key player in satisfying the global energy needs. With increased market penetration, unanticipated unsteady loading induced failures, installation related reductions in power generation, and significant maintenance costs have underscored the need to predict the unsteady fluid-structure interactions related to turbine layout and off-design wind conditions. Contemporary turbine design tools are incapable of accounting for such loadings. As a result, researchers have started utilizing high-Performance-Computing (HPC) based Computational Fluid Dynamics (CFD) solvers, such as the U.S. Department of Energy sponsored ExaWind software package, to investigate these phenomena. Unfortunately, such HPC tools are computationally expensive for routine industrial use, often because of the sheer number of cells required to resolve the wake flowfield. This paper describes a preliminary effort to address this issue by developing a vorticity-velocity based CFD off-body solver, VorTran-M2-AMReX, that integrates directly with DOE's ExaWind wind turbine analysis system to perform accurate and reliable simulations of wind turbine/farm at a lower computational cost than ExaWind alone. This article summarizes work undertaken to date concerning the assembly of the proposed analysis tool, and provides preliminary validation and verification of the VorTran-M2-AMReX off-body solver.

adaptive mesh refinement↗

A Scaling Study for Incompressible Multispecies Solver in Vertex-CFD

Multispecies incompressible flows occur widely in engineering and environmental applications, such as chemical reactors, fuel cells, ocean mixing, and biomedical systems. However, accurately resolving the complex transport and mixing phenomena associated with multiple interacting species remains computationally challenging, especially for large-scale problems. In this study, we present a robust, high-performance computing--enabled multispecies incompressible Navier–Stokes solver integrated within the Vertex-CFD framework. Our solver employs a fully coupled, implicit, finite element--based formulation that accurately captures the advection, diffusion, and interaction of multiple species in incompressible flows by leveraging the Kokkos library for parallel computing to achieve high computational efficiency. For pressure coupling, the entropically damped artificial compressibility method is utilized. We validated the solver against canonical test cases, including multispecies advection, diffusion, and Bateman systems; the results demonstrate second- and third-order spatial accuracy and consistent convergence. Additionally, we demonstrated the strong and weak scaling study results obtained on the leadership-class high-performance computing system, Frontier at Oak Ridge National Laboratory.

Oz, Furkan [ORNL] (ORCID:0000000265831724)↗

Refactoring the elastic–viscous–plastic solver from the sea ice model CICE v6.5.1 for improved performance

This study focuses on the performance of the elastic–viscous–plastic (EVP) dynamical solver within the sea ice model, CICE v6.5.1. The study has been conducted in two steps. First, the standard EVP solver was extracted from CICE for experiments with refactored versions, which are used for performance testing. Second, one refactored version was integrated and tested in the full CICE model to demonstrate that the new algorithms do not significantly impact the physical results. The study reveals two dominant bottlenecks, namely (1) the number of Message Parsing Interface (MPI) and Open Multi-Processing (OpenMP) synchronization points required for halo exchanges during each time step combined with the irregular domain of active sea ice points and (2) the lack of single-instruction, multiple-data (SIMD) code generation. The standard EVP solver has been refactored based on two generic patterns. The first pattern exposes how general finite differences on masked multi-dimensional arrays can be expressed in order to produce significantly better code generation by changing the memory access pattern from random access to direct access. The second pattern takes an alternative approach to handle static grid properties. The measured single-core performance improvement is more than a factor of 5 compared to the standard implementation. The refactored implementation of strong scales on the Intel® Xeon® Scalable Processors series node until the available bandwidth of the node is used. For the Intel® Xeon® CPU Max series, there is sufficient bandwidth to allow the strong scaling to continue for all the cores on the node, resulting in a single-node improvement factor of 35 over the standard implementation. This study also demonstrates improved performance on GPU processors.

58 GEOSCIENCES↗

A multigrid solver for semi-implicit global shallow-water models

A multigrid solver is developed for the discretized two-dimensional elliptic equation on the sphere that arises from a semiimplicit time discretization of the global shallow-water equations. Different formulations of the semiimplicit scheme result in variable-coefficient Helmholtz-type equations for which no fast direct solvers are available. The efficiency of the multigrid solver is optimal, in the sense that the total operation count is proportional to the number of unknowns. Numerical experiments using initial data derived from actual 300-mb height and wind velocity fields indicate that the present model has very good accuracy and stability properties.

Barros, Saulo R. M.↗

An approximate Riemann solver for hypervelocity flows

We describe an approximate Riemann solver for the computation of hypervelocity flows in which there are strong shocks and viscous interactions. The scheme has three stages, the first of which computes the intermediate states assuming isentropic waves. A second stage, based on the strong shock relations, may then be invoked if the pressure jump across either wave is large. The third stage interpolates the interface state from the two initial states and the intermediate states. The solver is used as part of a finite-volume code and is demonstrated on two test cases. The first is a high Mach number flow over a sphere while the second is a flow over a slender cone with an adiabatic boundary layer. In both cases the solver performs well.

Jacobs, Peter A.↗

An installed nacelle design code using a multiblock Euler solver. Volume 1: Theory document

An efficient multiblock Euler design code was developed for designing a nacelle installed on geometrically complex airplane configurations. This approach employed a design driver based on a direct iterative surface curvature method developed at LaRC. A general multiblock Euler flow solver was used for computing flow around complex geometries. The flow solver used a finite-volume formulation with explicit time-stepping to solve the Euler Equations. It used a multiblock version of the multigrid method to accelerate the convergence of the calculations. The design driver successively updated the surface geometry to reduce the difference between the computed and target pressure distributions. In the flow solver, the change in surface geometry was simulated by applying surface transpiration boundary conditions to avoid repeated grid generation during design iterations. Smoothness of the designed surface was ensured by alternate application of streamwise and circumferential smoothings. The capability and efficiency of the code was demonstrated through the design of both an isolated nacelle and an installed nacelle at various flow conditions. Information on the execution of the computer program is provided in volume 2.

Chen, H. C.↗