Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “solvers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Speedup of UEDGE Parameter Scans Using Machine-Learning Optimized OpenMP Parallelization and a Continuation Solver

This article presents the OpenMP parallelization of the preconditioning Jacobian assembly and right‐hand side residual evaluation in UEDGE. A continuation algorithm, utilizing the internal NKSOL implicit Jacobian‐Free Newton‐Krylov solver to efficiently scan physical parameters, is also presented. The implemented parallelization reduces the computational time for a benchmark scan run on 32 threads by compared to the serial version when using trained random forest regression models to identify the optimal decomposition of the system of equations. Random forest regression models applied to the UEDGE time‐dependent and continuation solver algorithms did not yield meaningful improvement in computational performance. A benchmark DIII‐D gas injection rate scan in the 0.35–0.75 kA interval, performed on a test cluster using the parallelized code and continuation solver, produced 1066 steady‐state solutions with a 22 s average wall‐clock computational time per steady‐state solution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A new field solver for modeling of relativistic particle-laser interactions using the particle-in-cell algorithm

A customized finite-difference field solver for the particle-in-cell (PIC) algorithm that provides higher fidelity for wave-particle interactions in intense electromagnetic waves is presented. In many problems of interest, particles with relativistic energies interact with intense electromagnetic fields that have phase velocities near the speed of light. Numerical errors can arise due to (1) dispersion errors in the phase velocity of the wave, (2) the staggering in time between the electric and magnetic fields and between particle velocity and position and (3) errors in the time derivative in the momentum advance. Errors of the first two kinds are analyzed in detail. It is shown that by using field solvers with different -space operators in Faraday’s and Ampere’s law, the dispersion errors and magnetic field time-staggering errors in the particle pusher can be simultaneously removed for electromagnetic waves moving primarily in a specific direction. Here, the new algorithm was implemented into Osiris by using customized higher-order finite-difference operators. Schemes using the proposed solver in combination with different particle pushers are compared through PIC simulation. It is shown that the use of the new algorithm, together with an analytic particle pusher (assuming constant fields over a time step), can lead to accurate modeling of the motion of a single electron in an intense laser field with normalized vector potentials, eA / mc 2 , exceeding for typical cell sizes and time steps.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

GPU-resident sparse direct linear solvers for alternating current optimal power flow analysis

Integrating renewable resources within the transmission grid at a wide scale poses significant challenges for economic dispatch as it requires analysis with more optimization parameters, constraints, and sources of uncertainty. This motivates the investigation of more efficient computational methods, especially those for solving the underlying linear systems, which typically take more than half of the overall computation time. In this paper, we present our work on sparse linear solvers that take advantage of hardware accelerators, such as graphical processing units (GPUs), and improve the overall performance when used within economic dispatch computations. We treat the problems as sparse, which allows for faster execution but also makes the implementation of numerical methods more challenging. We present the first GPU-native sparse direct solver that can execute on both AMD and NVIDIA GPUs. We demonstrate significant performance improvements when using high-performance linear solvers within alternating current optimal power flow (ACOPF) analysis. Furthermore, we demonstrate the feasibility of getting significant performance improvements by executing the entire computation on GPU-based hardware. Finally, we identify outstanding research issues and opportunities for even better utilization of heterogeneous systems, including those equipped with GPUs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

jaxhps: An elliptic PDE solver built with machine learning in mind

Elliptic partial differential equations (PDEs) can model many physical phenomena, such as electrostatics, acoustics, wave propagation, and diffusion. In scientific machine learning settings, a high-throughput PDE solver may be required to generate a training dataset, run in the inner loop of an iterative algorithm, or interface directly with a deep neural network. To provide value to machine learning users, such a PDE solver must be compatible with standard automatic differentiation frameworks, scale efficiently when run on graphics processing units (GPUs), and maintain high accuracy for a large range of input parameters. We have designed the jaxhps package with these use-cases in mind by implementing a highly efficient and accurate solver for elliptic problems with native hardware acceleration and automatic differentiation support.

97 MATHEMATICS AND COMPUTING↗

Asynchronous Iterative Solvers for Extreme-Scale Computing

The Asynchronous Iterative Solvers for Extreme-Scale Computing (AsyncIS) project aims to explore more efficient numerical algorithms by decreasing their overhead. AsyncIS does this by replacing the outer Krylov subspace solver with an asynchronous optimized Schwarz method, thereby removing the global synchronization and bulk synchronous operations typically used in numerical codes. AsyncIS—a U.S. Department of Energy (DOE)-funded collaboration between Georgia Tech, the University of Tennessee, Knoxville, Temple University, and Sandia National Laboratories—also focuses on the development and optimization of asynchronous preconditioners (i.e., preconditioners that are generated and/or applied in an asynchronous fashion). The novel preconditioning algorithms that provide fine-grained parallelism enable preconditioned Krylov solvers to run efficiently on large-scale distributed systems and manycore accelerators like GPUs.

97 MATHEMATICS AND COMPUTING↗

Neural Network Solver for Small Quantum Clusters

Machine learning approaches have recently been applied to the study of various problems in physics. Most of these studies are focused on interpreting the data generated by conventional numerical methods or the data on an existing experimental database. An interesting question is whether it is possible to use a machine learning approach, in particular a neural network, for solving the many-body problem. In this paper, we present a neural network solver for the single impurity Anderson model, the paradigm of an interacting quantum problem in small clusters. We demonstrate that the neural-network-based solver provides quantitative accurate results for the spectral function as compared to the exact diagonalization method. This opens the possibility of utilizing the neural network approach as an impurity solver for other many-body numerical approaches, such as the dynamical mean field theory.

36 MATERIALS SCIENCE↗

The curled wake model: a three-dimensional and extremely fast steady-state wake solver for wind plant flows

Abstract. Wind turbine wake models typically require approximations, such as wake superposition and deflection models, to accurately describe wake physics. However, capturing the phenomena of interest, such as the curled wake and interaction of multiple wakes, in wind power plant flows comes with an increased computational cost. To address this, we propose a new hybrid method that uses analytical solutions with an approximate form of the Reynolds-averaged Navier–Stokes equations to solve the time-averaged flow over a wind plant. We compare results from the solver to supervisory control and data acquisition data from the Lillgrund wind plant obtaining wake model predictions which are generally within 1 standard deviation of the mean power data. We perform simulations of flow over the Columbia River Gorge to demonstrate the capabilities of the model in complex terrain. We also apply the solver to a case with wake steering, which agreed well with large-eddy simulations. This new solver reduces the time – and therefore the related cost – it takes to simulate a steady-state wind plant flow (on the order of seconds using one core). Because the model is computationally efficient, it can also be used for different applications including wake steering for wind power plants and layout optimization.

17 WIND ENERGY↗

A Parallel Incompressible Navier-Stokes Solver with Multigrid Iterations

We developed a parallel, numerically accurate and stable, and computationally efficient finate-difference incompressible Navier-Stokes (N-S) fluid flow solver. The solver runs on both sequential and massively parallel computers. The numerical method used here is a second-order projection method on a staggered grid. The code is highly modular and it can be used either as a stand-alone flow solver and or a template code which can be adapted or expanded to a specific application. Numerical results and parallel performances of our code on Intel Delta and Paragon are reported.

Navier-Stokes solver↗

Implementation of a First-Order Quadratic Program Solver in C

This paper details a translation of a first order quadratic program (QP) solver from MATLAB to C. NASA could use this QP solver to generate online flight path trajectories for powered descent vehicles during landing. Over 12 weeks, the team designed, implemented, and tested two iterations of the QP solver for accuracy and runtime on 104 benchmark QP tests. The final iteration was 541.07% faster than the first, handling most tests in under one second. Additionally, it solved four more QP tests for N≥1383, and all outputs for cost and D_x matched the MATLAB reference values.

Optimization↗

Comparison of Some RANS Solvers

We will take a look at solving the Reynolds-averaged Navier-Stokes (RANS) equations that are encountered in the context of wind farm performance simulations and optimizations. We will compare some of the more popular ways to solve these equations with a focus on using iterative solvers for the linear solve. We will compare their performance and reliability to a direct solve as we scale the problem both by adding more and parallel resources and by increasing the size of the domain, both in two and three dimensions. There are many strategies that can be applied to solving the RANS equations, some are very efficient, while others are very insensitive to, for example, the Reynolds number. The first contender we will consider is the Pressure-Convection-Diffusion (PCD) preconditioner. Early results suggest that PCD is indeed a very efficient solver, in particular in two dimensions, as long as the Reynold's number remains small. Next we will try to reorder our degrees of freedom such that we can use GMRES with ILU for our linear solve. Another popular choice we will consider for solving the RANS equations is SIMPLE (and its derivatives). For all of our implementations we make use of either FEniCS or Firedrake, basing our work on both existing implementations of some of these solvers while also writing new extensions for others.

CFD↗

Efficient proximal subproblem solvers for a nonsmooth trust-region method

In [R. J. Baraldi and D. P. Kouri, Mathematical Programming, (2022), pp. 1-40], we introduced an inexact trust-region algorithm for minimizing the sum of a smooth nonconvex and nonsmooth convex function. The principle expense of this method is in computing a trial iterate that satisfies the so-called fraction of Cauchy decrease condition—a bound that ensures the trial iterate produces sufficient decrease of the subproblem model. In this paper, we expound on various proximal trust-region subproblem solvers that generalize traditional trust-region methods for smooth unconstrained and convex-constrained problems. We introduce a simplified spectral proximal gradient solver, a truncated nonlinear conjugate gradient solver, and a dogleg method. Finally, we compare algorithm performance on examples from data science and PDE-constrained optimization.

97 MATHEMATICS AND COMPUTING↗

A Finite Difference informed Random Walk solver for simulating radiation defect evolution in polycrystalline structures with strongly inhomogeneous diffusivity

Diffusivity of species and defects on grain boundaries is usually several orders of magnitude larger than that inside grains. Such strongly inhomogeneous diffusivity requires prohibitively high computational demands for modeling microstructural evolution. Here, this paper presents a highly-efficient numerical solver, combining the Finite Difference method and Random Walk model, designed for accurately modeling strongly inhomogeneous diffusion within polycrystalline structures. The proposed solver, termed Finite Difference informed Random Walk (FDiRW), integrates a customized Finite Difference (cFD) scheme tailored for fast diffusion along thin grain boundaries represented by a single-layer of nodes. Numerical experiments demonstrate that the FDiRW solver achieves an impressive efficiency gain of 1560x compared to traditional Finite Difference methods while maintaining accuracy, making it feasible for personal computer machines to handle diffusional systems with strongly inhomogeneous diffusivity across static polycrystalline microstructures. The model has been successfully applied to simulate radiation defect evolution, showcasing its scalability to engineering scales in both length and time dimensions.

36 MATERIALS SCIENCE↗

HTR solver: An open-source exascale-oriented task-based multi-GPU high-order code for hypersonic aerothermodynamics

In this study, the open-source Hypersonics Task-based Research (HTR) solver for hypersonic aerothermodynamics is described. The physical formulation of the code includes thermochemical effects induced by high temperatures (vibrational excitation and chemical dissociation). The HTR solver uses high-order TENO-based spatial discretization on structured grids and efficient time integrators for stiff systems, is highly scalable in GPU-based supercomputers as a result of its implementation in the Regent/Legion stack, and is designed for direct numerical simulations of canonical hypersonic flows at high Reynolds numbers. Additionally, the performance of the HTR solver is tested with benchmark cases including inviscid vortex advection, low- and high-speed laminar boundary layers, inviscid one-dimensional compressible flows in shock tubes, supersonic turbulent channel flows, and hypersonic transitional boundary layers of both calorically perfect gases and dissociating air.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Implicit and coupled fluid plasma solver with adaptive Cartesian mesh and its applications to non-equilibrium gas discharges

In this work, we present a new fluid plasma solver with adaptive Cartesian mesh (ACM) based on a full-Newton (nonlinear, implicit) scheme for non-equilibrium gas discharge plasma. The electrons and ions are described using drift-diffusion approximation coupled to Poisson equation for the electric field. The electron-energy transport equation is solved to account for electron thermal conductivity, Joule heating, and energy loss of electrons in collisions with neutral species. The rate of electron-induced ionization is a function of electron temperature and could also depend on electron density (important for plasma stratification). The ion and gas temperature are kept constant. The transport equations are discretized using a non-isothermal Scharfetter-Gummel scheme to resolve possible large temperature gradients in the sheaths. We demonstrate the new solver for simulations of direct current (DC) and radiofrequency (RF) discharges. The implicit treatment of the coupled equations allows using large time steps. The full-Newton method (FNM) enables fast nonlinear convergence at each time step, offering significantly improved simulation efficiency. We discuss the selection of time steps for solving different plasma problems. The new solver enables solving several problems we could not solve before with existing software: two- and three-dimensional structures of the entire DC discharges including cathode and anode regions, electric field reversals and double-layer formation, the normal cathode spot and an anode ring, moving striations in diffuse and constricted DC discharges, and standing striations in RF discharges. The developed FNM-ACM technique offers many benefits for tackling the disparity of gas discharge plasma systems' time scales and nonlinearity.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Quantum solver for single-impurity Anderson models with particle-hole symmetry

Quantum embedding methods, such as dynamical mean-field theory (DMFT), provide a powerful framework for investigating strongly correlated materials. A central computational bottleneck in DMFT is in solving the Anderson impurity model (AIM), whose exact solution is classically intractable for large bath sizes. In this work, we benchmark a quantum-classical hybrid solver tailored for particle-hole symmetric AIMs, using the variational quantum eigensolver to prepare the ground state of the model with shallow quantum circuits. The solver uses shallow quantum ansätze and one set of variational parameters to prepare the ground state and its particle and hole excitations, enabling the construction of the impurity Green’s function through a continued-fraction expansion. We evaluate the performance of this approach across a few bath sizes and interaction strengths under noisy, shot-limited conditions. We compare three optimization routines (COBYLA, Adam, and L-BFGS-B) in terms of convergence and fidelity, assess the benefits of estimating a quantum-computed moment correction to the variational energies, and benchmark the approach by comparing the density of states computed from the impurity Green’s function against that obtained using a classical pipeline. Our results demonstrate the feasibility of Green’s function construction on near-term devices and establish practical benchmarks for quantum impurity solvers embedded within self-consistent DMFT loops.

Karabin, Mariia [ORNL]↗

Investigating the use of field solvers for simulating classical systems

We explore the use of field solvers as approximations of classical Vlasov-Poisson systems. This correspondence is investigated in both electrostatic and gravitational contexts. We demonstrate the ability of field solvers to be excellent approximations of problems with cold initial condition into the nonlinear regime. We also investigate extensions of the Schrödinger-Poisson system that employ multiple stacked cold streams, and the von Neumann–Poisson equation as methods that can successfully reproduce the classical evolution of warm initial conditions. We then discuss how appropriate simulation parameters need to be chosen to avoid interference terms, aliasing, and wave behavior in the field solver solutions. Finally, we present a series of criteria clarifying how parameters need to be chosen in order to effectively approximate classical solutions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Multidimensional Tests of a Finite-Volume Solver for MHD With a Real-Gas Equation of State

This work considers two algorithms of a finite-volume solver for the MHD equations with a real-gas equation of state (EOS). Both algorithms use a multistate form of Harten-Lax-Van Leer approximate Riemann solver as formulated for MHD discontinuities. This solver is modified to use the generalized sound speed from the real-gas EOS. Two methods are tested: EOS evaluation at cell centers and at flux interfaces where the former is more computationally efficient. A battery of 1D and 2D tests are employed: convergence of 1D and 2D linearized waves, shock tube Riemann problems, a 2D nonlinear circularly polarized Alfvén wave, and a 2D magneto-Rayleigh-Taylor instability test. The cell-centered EOS evaluation algorithm produces unresolvable thermodynamic inconsistencies in the intermediate states leading to spurious solutions while the flux-interface EOS evaluation algorithm robustly produces the correct solution. The linearized wave tests show this inconsistency is associated with the magnetosonic waves and the magneto-Rayleigh-Taylor instability test demonstrates simulation findings where the spurious solution leads to an unphysical simulation.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Evaluating asynchronous Schwarz solvers on GPUs

With the commencement of the exascale computing era, we realize that the majority of the leadership supercomputers are heterogeneous and massively parallel. Even a single node can contain multiple co-processors such as GPUs and multiple CPU cores. For example, ORNL’s Summit accumulates six NVIDIA Tesla V100 GPUs and 42 IBM Power9 cores on each node. Synchronizing across compute resources of multiple nodes can be prohibitively expensive. Hence, it is necessary to develop and study asynchronous algorithms that circumvent this issue of bulk-synchronous computing. In this study, we examine the asynchronous version of the abstract Restricted Additive Schwarz method as a solver. We do not explicitly synchronize, but allow the communication between the sub-domains to be completely asynchronous, thereby removing the bulk synchronous nature of the algorithm. We accomplish this by using the one-sided Remote Memory Access (RMA) functions of the MPI standard. We study the benefits of using such an asynchronous solver over its synchronous counterpart. We also study the communication patterns governed by the partitioning and the overlap between the sub-domains on the global solver. Finally, we show that this concept can render attractive performance benefits over the synchronous counterparts even for a well-balanced problem.

Nayak, Pratik↗