Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “solver”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Reconstruction of 2D line-integrated electron density using angular filter refractometry and a fast marching Eikonal solver

Refraction of an optical probe beam by a plasma can be measured with angular filter refractometry (AFR), which produces an image of the beam’s 2D spatial profile that contains intensity contours corresponding to curves of constant refraction angle. Further analysis is required to reconstruct the underlying line-integrated electron density. Most prior efforts to calculate density from AFR data have been limited to 1D analysis or forward-fitting techniques. Here, in this paper, we detail the use of a fast-marching Eikonal solver to directly invert AFR data and obtain the full 2D line-integrated electron density. The analysis method is first verified with synthetic data and then applied to experimental measurements of single and colliding plasma plumes collected at the OMEGA EP Laser Facility. The calculated densities agree with 1D results and are shown to be consistent with the original AFR measurements via forward modeling. We also discuss ways to improve the precision of this technique.

McCluskey, B. [Princeton Univ., NJ (United States)↗

Parallel interior-point solver for block-structured nonlinear programs on SIMD/GPU architectures

Here, we investigate how to port the standard interior-point method to new exascale architectures for block-structured nonlinear programs with state equations. Computationally, we decompose the interior-point algorithm into two successive operations: the evaluation of the derivatives and the solution of the associated Karush-Kuhn-Tucker (KKT) linear system. Our method accelerates both operations using two levels of parallelism. First, we distribute the computations on multiple processes using coarse parallelism. Second, each process uses SIMD/GPU accelerators locally to accelerate the operations using fine-grained parallelism. The KKT system is reduced by eliminating the inequalities and the state variables from the corresponding equations. We demonstrate our method's capability on the supercomputer Polaris, a testbed for the future exascale Aurora system. Each node is equipped with four GPUs, a setup amenable to our two-level approach. Our experiments on the stochastic optimal power flow problem show that the reduction method is 50x faster than the sparse linear solver HSL MA57 running in serial on the CPU, and 6x faster than Pardiso running in parallel on CPU on the same number of processes.

97 MATHEMATICS AND COMPUTING↗

An Unsteady Actuator Line Solver to Enable Adjoint Sensitivity Studies for Wake Steering

This study demonstrates the sensitivity of wind turbine wake steering performance to blade design. An actuator line model was implemented within an unsteady adjoint solver that enables efficient execution of gradient-based optimization and sensitivity studies. After first confirming the feasibility of wake steering by controlling actuator line chord profiles and formulating a suitable objective function for wake position, a sensitivity study was conducted to determine the relative importance of chord length as a function of spanwise position on the resulting turbine wake deflection. The results presented here support the idea that blade design choices play a role in wake control. In a larger context, this study demonstrates a computational framework in which turbine and blade designs can be studied at the individual and farm-wide level to enhance wind plant controllability and manage power output.

17 WIND ENERGY↗

SAGIPS: a physics-inspired scalable asynchronous generative inverse-problem solver

Abstract Solving large-scale inverse problems using deep-learning algorithms have become an essential part of modern research and industrial applications. The complexity of the underlying inverse problem may require the utilization of high performance computing systems which poses a challenge on the algorithmic design of the inverse problem solver. Most deep learning algorithms require, due to their design, custom parallelization techniques in order to be resource efficient while showing a reasonable convergence. In this paper we introduce a S calable A synchronous G enerative I nverse P roblem S olver (SAGIPS) on high-performance computing systems. We present a workflow that utilizes an asynchronous ring-allreduce algorithm to transfer the gradients of the generator network across multiple GPUs. Experiments with a scientific proxy application demonstrate that SAGIPS shows near linear weak scaling, together with a convergence quality that is comparable to traditional methods. The approach presented here allows leveraging Generative Adverserial Network across multiple GPUs, promising advancements in solving complex inverse problems at scale.

97 MATHEMATICS AND COMPUTING↗

Randomized Adiabatic Quantum Linear Solver Algorithm with Optimal Complexity Scaling and Detailed Running Costs

Solving linear systems of equations is a fundamental problem with a wide variety of applications across many fields of science, and there is increasing effort to develop quantum linear solver algorithms. Subaşı et al. [Phys. Rev. Lett. 122, 060504 (2019)] proposed a randomized algorithm inspired by adiabatic quantum computing, based on a sequence of random Hamiltonian simulation steps, with suboptimal scaling in the condition number 𝜅 of the linear system and the target error 𝜖. Here we go beyond these results in several ways. Firstly, using filtering [Lin and Tong, Quantum 4, 361 (2020)] and Poissonization techniques [Cunningham and Roland, ArXiv:2406.03972 (2024)], the algorithm complexity is improved to the optimal scaling 𝑂⁡(𝜅⁢log (1/𝜖))—an exponential improvement in 𝜖, and a shaving of a log 𝜅 scaling factor in 𝜅. Secondly, the algorithm is further modified to achieve constant factor improvements, which are vital as we progress towards hardware implementations on fault-tolerant devices. We introduce a cheaper randomized walk operator method replacing Hamiltonian simulation—which also removes the need for potentially challenging classical precomputations; randomized routines are sampled over optimized random variables; circuit constructions are improved. We obtain a closed formula rigorously upper bounding the expected number of times one needs to apply a block-encoding of the linear system matrix to output a quantum state encoding the solution to the linear system. The upper bound is 837⁢𝜅 at 𝜖 = 10 −10 for Hermitian matrices.

97 MATHEMATICS AND COMPUTING↗

Interaction-expansion inchworm Monte Carlo solver for lattice and impurity models

Multiorbital quantum impurity models with general interaction and hybridization terms appear in a wide range of applications, including embedding, quantum transport, and nanoscience. However, most quantum impurity solvers are restricted to a few impurity orbitals, discretized baths, diagonal hybridizations, or density-density interactions. We generalize the inchworm quantum Monte Carlo method to the interaction expansion, and we explore its application to typical single- and multiorbital problems encountered in investigations of impurity and lattice models. Our implementation generically outperforms bare and bold-line quantum Monte Carlo algorithms in the interaction expansion. For the systems studied here, our implementation remains inferior to the more specialized hybridization expansion and auxiliary field algorithms. So, the problem of convergence to unphysical fixed points, which hampers so-called bold-line methods, is not encountered in inchworm Monte Carlo.

36 MATERIALS SCIENCE↗

Advanced Quantum Poisson Solver in the NISQ era

The Poisson equation has many applications across the broad areas of science and engineering. Most quantum algorithms for the Poisson solver presented so far, either suffer from lack of accuracy and/or are limited to very small sizes of the problem, and thus have no practical usage. Here we present an advanced quantum algorithm for solving the Poisson equation with high accuracy and dynamically tunable problem size. After converting the Poisson equation to the linear systems through the finite difference method, we adopt the Harrow-Hassidim-Lloyd (HHL) algorithm as the basic framework. Particularly, in this work we present an advanced circuit that ensures the accuracy of the solution by implementing non-truncated eigenvalues through eigenvalue amplification as well as by increasing the accuracy of the controlled rotation angular coefficients, which are the critical factors in the HHL algorithm. We show that our algorithm not only increases the accuracy of the solutions, but also composes more practical and scalable circuits by dynamically controlling problem size in the NISQ devices. We present both simulated and experimental results, and discuss the sources of errors. Finally, we conclude that overall results on the quantum hardware are dominated by the error in the CNOT gates.

Robson, Walter↗

OpenACC offloading of the MFC compressible multiphase flow solver on AMD and NVIDIA GPUs

GPUs are the heart of the latest generations of supercomputers. We efficiently accelerate a compressible multiphase flow solver via OpenACC on NVIDIA and AMD Instinct GPUs. Optimization is accomplished by specifying the directive clauses gang vector and collapse. Further speedups of six and ten times are achieved by packing user-defined types into coalesced multidimensional arrays and manual inlining via metaprogramming. Additional optimizations yield seven-times speedup of array packing and thirty-times speedup of select kernels on Frontier. Weak scaling efficiencies of 97% and 95% are observed when scaling to 50% of Summit and 87% of Frontier. Strong scaling efficiencies of 84% and 81% are observed when increasing the device count by a factor of 8 and 16 on V100 and MI250X hardware. The strong scaling efficiency of AMD’s MI250X increases to 92% when increasing the device count by a factor of 16 when GPU-aware MPI is used for communication.

Wilfong, Benjamin↗

Limitations of Fault-Tolerant Quantum Linear System Solvers for Quantum Power Flow

Quantum computers hold promise for solving problems intractable for classical computers, especially those with high time or space complexity. Practical quantum advantage can be said to exist for such problems when the end-to-end time for solving such a problem using a classical algorithm exceeds that required by a quantum algorithm. Reducing the power flow (PF) problem into a linear system of equations allows for the formulation of quantum PF (QPF) algorithms, which are based on solving methods for quantum linear systems such as the Harrow-Hassidim-Lloyd (HHL) algorithm. Speedup from using QPF algorithms is often claimed to be exponential when compared to classical PF solved by state-of-the-art algorithms. Here, we investigate the potential for practical quantum advantage in solving QPF compared to classical methods on gate-based quantum computers. Notably, this paper does not present a new QPF solving algorithm but scrutinizes the end-to-end complexity of the QPF approach, providing a nuanced evaluation of the purported quantum speedup in this problem. Our analysis establishes a best-case bound for the HHL-based quantum power flow complexity, conclusively demonstrating that the HHL-based method has higher runtime complexity compared to the classical algorithm for solving the direct current power flow (DCPF) and fast decoupled load flow (FDLF) problem. Notably, our analysis and conclusions can be extended to any quantum linear system solver with rigorous performance guarantees, based on the known complexity lower bounds for this problem. Additionally, we establish that for potential practical quantum advantage (PQA) to exist it is necessary to consider DCPF-type problems with a very narrow range of condition number values and readout requirements.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

PeleMP: The Multiphysics Solver for the Combustion Pele Adaptive Mesh Refinement Code Suite

Combustion encompasses multiscale, multiphase reacting flow physics spanning a wide range of scales from the molecular scales, where chemical reactions occur, to the device scales, where the turbulent flow is affected by the geometry of the combustor. This scale disparity and the limited measurement capabilities from experiments make modeling combustion a significant challenge. Recent advancements in high-performance computing (HPC), particularly with the Department of Energy's Exascale Computing Project (ECP), have enabled high-fidelity simulations of practical applications to be performed. The major physics submodels, including chemical reactions, turbulence, sprays, soot, and thermal radiation, exhibit distinctive computational characteristics that need to be examined separately to ensure efficient utilization of computational resources. This paper presents the multiphysics solver for the Pele code suite, called PeleMP, which consists of models for spray, soot, and thermal radiation. Here, the mathematical and algorithmic aspects of the model implementations are described in detail as well as the verification process. The computational performance of these models is benchmarked on multiple supercomputers, including Frontier, an exascale machine. Results are presented from production simulations of a turbulent sooting ethylene flame and a bluff-body swirl stabilized spray flame with sustainable aviation fuels to demonstrate the capability of the Pele codes for modeling practical combustion problems with multiphysics. This work is an important step toward the exascale computing era for high-fidelity combustion simulations providing physical insights and data for predictive modeling of real-world devices.

42 ENGINEERING↗

Single Grid Error Estimation for Neutron Transport Solvers

The method of nearby problems (MNP) is a solution verification technique that does not require the use of multiple spatial grids. To estimate spatial discretization error without requiring a high-fidelity spatial grid, an analytical curve fit is interpolated from the numerical solution. The residual between the curve fit solution and numerical solution is calculated and added as an additional source term to the governing equation. The nearby solution is estimated using the updated source term and boundary conditions to remain consistent with the curve fit interpolation. The nearby solution can be compared to the curve fit solution as a discretization error estimation while using a single spatial grid. Without the use of higher fidelity spatial grids, the MNP is able to approximate the spatial discretization error, a facet of solution verification. The application of the method of nearby problems is presented for one- and two-dimensional neutron transport problems for both fixed source and criticality problems on the spatial variable. The fixed source results demonstrate the effectiveness of nearby problems for spatial error identification using the discrete ordinates method. Criticality results are shown to identify area of high spatial error for the C5G7 problem as well as for the discrete ordinates solver. A novel approach of combining the capabilities of Monte Carlo with the discrete ordinates nearby problems is presented for one- and two-dimensional fixed source problems. In conclusion, the MNP demonstrates its effectiveness at identifying spatial error on a single structured grid with a wide variety of neutron transport problems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Fast and Scalable Sparse Triangular Solver for Multi-GPU Based HPC Architectures

Designing efficient and scalable sparse linear algebra kernels on modern multi-GPU based HPC systems is a daunting task due to significant irregular memory references and workload imbalance across the GPUs. This is particularly the case for \textit{Sparse Triangular Solver (SpTRSV)} which introduces additional two-dimensional computation dependencies among subsequent computation steps. Dependency information is exchanged and shared among GPUs, thus warrant for efficient memory allocation, data partitioning, and workload distribution as well as fine-grained communication and synchronization support. In this work, we demonstrate that directly adopting unified memory can adversely affect the performance of SpTRSV on multi-GPU architectures, despite linking via fast interconnect like NVLinks and NVSwitches. Alternatively, we employ the latest NVSHMEM technology based on Partitioned Global Address Space programming model to enable efficient fine-grained communication and drastic synchronization overhead reduction. Furthermore, to handle workload imbalance, we propose a malleable task-pool execution model which can further enhance the utilization of GPUs. By applying these techniques, our experiments on the NVIDIA multi-GPU supernode V100-DGX-1 and DGX-2 systems demonstrate that our design can achieve on average 3.53x (up to 9.86x) speedup on a DGX-1 system and 3.66x (up to 9.64x) speedup on a DGX-2 system with 4-GPUs over the Unified-Memory design. The comprehensive sensitivity and scalability studies also show that the proposed zero-copy SpTRSV is able to fully utilize the computing and communication resources of the multi-GPU system.

Xie, Chenhao↗

ARKODE: A Flexible IVP Solver Infrastructure for One-step Methods

We describe the ARKODE library of one-step time integration methods for ordinary differential equation (ODE) initial-value problems (IVPs). In addition to providing standard explicit and diagonally implicit Runge–Kutta methods, ARKODE supports one-step methods designed to treat additive splittings of the IVP, including implicit-explicit (ImEx) additive Runge–Kutta methods and multirate infinitesimal (MRI) methods. We present the role of ARKODE within the SUNDIALS suite of time integration and nonlinear solver libraries, the core ARKODE infrastructure for utilities common to large classes of one-step methods, as well as its use of “time stepper” modules enabling easy incorporation of novel algorithms into the library. Numerical results show example problems of increasing complexity, highlighting the algorithmic flexibility afforded through this infrastructure, and include a larger multiphysics application leveraging multiple algorithmic features from ARKODE and SUNDIALS.

97 MATHEMATICS AND COMPUTING↗

Kohn-Sham Solver (KSSOLV) v2.0

KSSOLV is a MATLAB toolbox for solving Kohn-Sham density functional theory based electronic structure eigenvalue problems. It uses an object oriented features of MATLAB to represent atom, molecules, wavefunctions and Hamiltonians and their operations. It is designed to make it easier for users to prototype and test new algorithms for solving the Kohn-Sham problem. KSSOLV2.0 contains significant improvement over the original KSSOLV described in a paper published in ACM Transaction on Mathematical Software (attached). In addition to performing ground state calculation for small molecules, it can also perform geometry optimization for both molecules and solids. It uses standard pseudopotentials and implements local density approximation, generalized gradient approximation and hybrid functionals. Future releases will also include time-dependent DFT and post DFT calculations such as the GW quasi-particle energy calculation and Bethe-Salpeter equation solver for optical absorption.

Yang, Chao↗

MARBLES (Multi-scale Adaptively Refined Boltzmann LatticE Solver) [SWR-23-37]

MARBLES (Multi-scale Adaptively Refined Boltzmann LatticE Solver) is an open-source computational fluid dynamics package powered by the lattice Boltzmann equations and built on AMReX. In the lattice Boltzmann method, local collisions between meso-scale fictitious particles drive the governing equations which enables MARBLES to easily simulate flow around complex and/or moving geometry without the generation of a body-conforming mesh. Using AMReX data structures and operations ensures a high level of computational performance and parallel scaling on heterogenous architectures while also naturally supporting locally enhanced grid resolution and fidelity through automatic mesh refinement. New domains and problem definitions are easily specified through an input file with examples and guidance on all options and variables provided in the MARBLES documentation.

Henry de Frahan, Marc↗

LANL contribution to ryujin, an open source finite element solver

Ryujin (https://github.com/conservation-laws/ryujin) is a high-performance finite-element software for solving mathematical partial differential equations (PDEs) with dominant hyperbolic structures. The author of this request, Eric Tovar, is using Ryujin as a high-performance tool for his Mark Kac postdoctoral fellowship research at LANL. Eric would like to contribute openly to the ryujin software without changing its core functionality. This includes: (i) bug fixes; (ii) re-organization of code for performance and syntactic updates including documentation; (iii) implementation of new PDE numerical methods that align with the core solver; (iv) implementation of new initial state configurations for target applications.

Tovar, Eric↗

A massively parallel time-domain coupled electrodynamics–micromagnetics solver

We present a high-performance coupled electrodynamics–micromagnetics solver for full physical modeling of signals in microelectronic circuitry. The overall strategy couples a finite-difference time-domain approach for Maxwell’s equations to a magnetization model described by the Landau–Lifshitz–Gilbert equation. The algorithm is implemented in the Exascale Computing Project software framework, AMReX, which provides effective scalability on manycore and GPU-based supercomputing architectures. Furthermore, the code leverages ongoing developments of the Exascale Application Code, WarpX, which is primarily being developed for plasma wakefield accelerator modeling. Our temporal coupling scheme provides second-order accuracy in space and time by combining the integration steps for the magnetic field and magnetization into an iterative sub-step that includes a trapezoidal temporal discretization for the magnetization. The performance of the algorithm is demonstrated by the excellent scaling results on NERSC multicore and GPU systems, with a significant (59×) speedup on the GPU using a node-by-node comparison. We demonstrate the utility of our code by performing simulations of an electromagnetic waveguide and a magnetically tunable filter.

97 MATHEMATICS AND COMPUTING↗