Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preconditioner”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

An adaptive Hessian approximated stochastic gradient MCMC method

Bayesian approaches have been successfully integrated into training deep neural networks. One popular family is stochastic gradient Markov chain Monte Carlo methods (SG-MCMC), which have gained increasing interest due to their ability to handle large datasets and the potential to avoid overfitting. Although standard SG-MCMC methods have shown great performance in a variety of problems, they may be inefficient when the random variables in the target posterior densities have scale differences or are highly correlated. Here, we present an adaptive Hessian approximated stochastic gradient MCMC method to incorporate local geometric information while sampling from the posterior. The idea is to apply stochastic approximation (SA) to sequentially update a preconditioning matrix at each iteration. The preconditioner possesses second-order information and can guide the random walk of a sampler efficiently. Instead of computing and saving the full Hessian of the log posterior, we use limited memory of the samples and their stochastic gradients to approximate the inverse Hessian-vector multiplication in the updating formula. Moreover, by smoothly optimizing the preconditioning matrix via SA, our proposed algorithm can asymptotically converge to the target distribution with a controllable bias under mild conditions. To reduce the training and testing computational burden, we adopt a magnitude-based weight pruning method to enforce the sparsity of the network. Our method is user-friendly and demonstrates better learning results compared to standard SG-MCMC updating rules. The approximation of inverse Hessian alleviates storage and computational complexities for large dimensional models. Numerical experiments are performed on several problems, including sampling from 2D correlated distribution, synthetic regression problems, and learning the numerical solutions of heterogeneous elliptic PDE. The numerical results demonstrate great improvement in both the convergence rate and accuracy.

97 MATHEMATICS AND COMPUTING↗

A family of independent Variable Eddington Factor methods with efficient preconditioned iterative solvers

We present a family of discretizations for the Variable Eddington Factor (VEF) equations that have high-order accuracy on curved meshes and efficient preconditioned iterative solvers. The VEF discretizations are combined with the Discontinuous Galerkin transport discretization from to form effective high-order, linear transport methods. The VEF discretizations are derived by extending the unified analysis of Discontinuous Galerkin methods for elliptic problems presented by Arnold et al. to the VEF equations. This framework is used to define analogs of the interior penalty, second method of Bassi and Rebay, minimal dissipation local Discontinuous Galerkin, and continuous finite element methods. The analysis of subspace correction preconditioners, which use a continuous operator to iteratively precondition the discontinuous discretization, is extended to the case of the non-symmetric VEF system. Numerical results demonstrate that the VEF discretizations have arbitrary-order accuracy on curved meshes, preserve the thick diffusion limit, and are effective on a proxy problem from thermal radiative transfer in both outer transport iterations and inner preconditioned linear solver iterations. We demonstrate that the VEF solution converges to the S N transport solution as the mesh is refined on both problems with smooth and non-smooth behavior in angle. Parallel performance studies show that the interior penalty VEF discretization's linear solve weak scales out to 1024 processors and strong scales well on a single node. Particular attention is paid to the parallel performance of the VEF algorithm when used in combination with a parallel block Jacobi transport sweep.

97 MATHEMATICS AND COMPUTING↗

A scalable DG solver for the electroneutral Nernst-Planck equations

The robust, scalable simulation of flowing electrochemical systems is increasingly important due to the synergy between intermittent renewable energy and electrochemical technologies such as energy storage and chemical manufacturing. The high Péclet regime of many such applications prevents the use of off-the-shelf discretization methods. In this work, we present a high-order Discontinuous Galerkin scheme for the electroneutral Nernst-Planck equations. The chosen charge conservation formulation allows for the specific treatment of the different physics: upwinding for advection and migration, and interior penalty for diffusion of ionic species as well the electric potential. Similarly, the formulation enables different treatments in the preconditioner: AMG for the potential blocks and ILU-based methods for the advection-dominated concentration blocks. Here we evaluate the convergence rate of the discretization scheme through numerical tests. Strong scaling results for two preconditioning approaches are shown for a large 3D flow-plate reactor example.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A fully implicit, asymptotic-preserving, semi-Lagrangian algorithm for the time dependent anisotropic heat transport equation

In this paper, we extend the operator-split asymptotic-preserving, semi-Lagrangian algorithm for time dependent anisotropic heat transport equation proposed in Chacón et al. (2014) [18] to use a fully implicit time integration with backward differentiation formulas. The proposed implicit method can deal with arbitrary heat-transport anisotropy ratios $\mathcal{X}$∥ /$ \mathcal{X}$⟂ $\ggg$ 1 (with $\mathcal{X}$∥, $ \mathcal{X}$⟂ the parallel and perpendicular heat diffusivities, respectively) in complicated magnetic field topologies in an accurate and efficient manner. Further, the implicit algorithm is second-order accurate temporally and demonstrates an accurate treatment at boundary layers (e.g., island separatrices), which was not ensured by the operator-split implementation. The condition number of the resulting algebraic system is independent of the anisotropy ratio, and is inverted with preconditioned GMRES. We propose a simple preconditioner that renders the finite-dimensional linear operator compact, resulting in mesh-independent convergence rates for topologically simple magnetic fields, and convergence rates scaling as ~ (NΔt) 1/4 (with N the total mesh size and Δt the timestep) in topologically complex magnetic-field configurations. We demonstrate the accuracy and performance of the approach with test problems of varying complexity, including an analytically tractable boundary-layer problem in a straight magnetic field, and a topologically complex magnetic field featuring magnetic islands with extreme anisotropy ratios $\mathcal{X}$∥ /$ \mathcal{X}$⟂ = 10 10 ) .

97 MATHEMATICS AND COMPUTING↗

A scalable multidimensional fully implicit solver for Hall magnetohydrodynamics

We propose an optimally performant fully implicit algorithm for the Hall magnetohydrodynamics (HMHD) equations based on multigrid-preconditioned Jacobian-free Newton-Krylov methods. HMHD is a challenging system to solve numerically because it supports stiff fast dispersive waves. The preconditioner is formulated using an operator-split approximate block factorization (Schur complement), informed by physics insight. We use a vector-potential formulation (instead of a magnetic field one) to allow a clean segregation of the problematic $\nabla$ x $\nabla$ x operator in the electron Ohm's law subsystem. This segregation allows the formulation of an effective damped block-Jacobi smoother for multigrid. We demonstrate by analysis that our proposed block-Jacobi iteration is convergent and has the smoothing property. The resulting HMHD solver is verified linearly with wave propagation examples, and nonlinearly with the GEM challenge reconnection problem by comparison against another HMHD code. We demonstrate the excellent algorithmic and parallel performance of the algorithm up to 16384 MPI tasks in two dimensions.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Leveraging explainable AI to characterize floating-point exceptions in linear solvers

Linear solver packages are central to many scientific, engineering, and machine learning applications. When floating-point exceptions occur in these solvers, e.g., division by zero or overflow, numerical results are compromised and become unreliable. Existing static and dynamic analysis tools can detect such exceptions, but they do not explain why the exceptions occur in terms of the solver inputs. Here, we present a study to characterize the inputs that cause numerical exceptions in linear solver packages. Our approach uses explainable AI (XAI) to find the most relevant characteristics of input matrices that explain the occurrence of exceptions in the solvers. Since training data in this domain is scarce, we perform extensive data gathering and data augmentation to obtain exception-inducing inputs. Our approach uses a repair strategy on the features blamed by XAI to validate that such features indeed explain the exceptions. We compare the LIME and SHAP XAI techniques using a dozen matrix features with three classifiers. We evaluate the approach on three widely used linear solver packages and find that some input characteristics can explain the occurrence of exceptions 100% of the time, in specific solvers and preconditioners.

Explainable AI↗

Recent Advances for Improving the Accuracy, Transferability, and Efficiency of Reactive Force Fields

Reactive force fields provide an affordable model for simulating chemical reactions at a fraction of the cost of quantum mechanical approaches. However, classically accounting for chemical reactivity often comes at the expense of accuracy and transferability, while computational cost is still large relative to nonreactive force fields. Here, we summarize recent efforts for improving the performance of reactive force fields in these three areas with a focus on the ReaxFF theoretical model. To improve accuracy, we describe recent reformulations of charge equilibration schemes to overcome unphysical long-range charge transfer, new ReaxFF models that account for explicit electrons, and corrections for energy conservation issues of the ReaxFF model. To enhance transferability we also highlight new advances to include explicit treatment of electrons in the ReaxFF and hybrid nonreactive/reactive simulations that make it possible to model charge transfer, redox chemistry, and large systems such as reverse micelles within the framework of a reactive force field. To address the computational cost, we review recent work in extended Lagrangian schemes and matrix preconditioners for accelerating the charge equilibration method component of ReaxFF and improvements in its software performance in LAMMPS.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Extended Lagrangian Born–Oppenheimer molecular dynamics for orbital-free density-functional theory and polarizable charge equilibration models

We report extended Lagrangian Born–Oppenheimer molecular dynamics (XL-BOMD) is formulated for orbital-free Hohenberg–Kohn density-functional theory and for charge equilibration and polarizable force-field models that can be derived from the same orbital-free framework. The purpose is to introduce the most recent features of orbital-based XL-BOMD to molecular dynamics simulations based on charge equilibration and polarizable force-field models. These features include a metric tensor generalization of the extended harmonic potential, preconditioners, and the ability to use only a single Coulomb summation to determine the fully equilibrated charges and the interatomic forces in each time step for the shadow Born–Oppenheimer potential energy surface. The orbital-free formulation has a charge-dependent, short-range energy term that is separate from long-range Coulomb interactions. This enables local parameterizations of the short-range energy term, while the long-range electrostatic interactions can be treated separately. The theory is illustrated for molecular dynamics simulations of an atomistic system described by a charge equilibration model with periodic boundary conditions. The system of linear equations that determines the equilibrated charges and the forces is diagonal, and only a single Ewald summation is needed in each time step. The simulations exhibit the same features in accuracy, convergence, and stability as are expected from orbital-based XL-BOMD.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Implicit full-F simulations of neoclassical ion transport

The development of implicit time integration capabilities for axisymmetric full-F continuum simulations of ion neoclassical transport is reported. The approach involves the implicit treatment of the gyrokinetic Vlasov equation coupled to the nonlinear Fokker–Planck collision model in the long-wavelength limit approximation. To facilitate implicit simulations, advanced preconditioning of individual physics operators is developed, and a global multi-physics preconditioner is constructed by adopting an operator splitting methodology. The algorithm is implemented in the finite-volume code COGENT and is applied to study neoclassical transport properties for both the main ion species and the lithium impurity species in the closed-field-line region of the LTX- β tokamak. The implicit COGENT simulations elucidate the role of non-local transport effects, while demonstrating substantial speedup over the corresponding explicit approach.

Dorf, Mikhail [Lawrence Livermore National Laborat↗

Linearized frequency domain Landau-Lifshitz-Gilbert equation formulation

We present a general finite element linearized Landau-Lifshitz-Gilbert equation (LLGE) solver for magnetic systems under weak time-harmonic excitation field. The linearized LLGE is obtained by assuming a small deviation around the equilibrium state of the magnetic system. Inserting such expansion into LLGE and keeping only first order terms gives the linearized LLGE, which gives a frequency domain solution for the complex magnetization amplitudes under an external time-harmonic applied field of a given frequency. We solve the linear system with an iterative solver using generalized minimal residual method. We construct a preconditioner matrix to effectively solve the linear system. The validity, effectiveness, speed, and scalability of the linear solver are demonstrated via numerical examples.

36 MATERIALS SCIENCE↗

Efficient 3-D velocity model building using joint inline and crossline plane-wave wave-equation migration velocity analyses

SUMMARY Wave-equation migration velocity analysis (WEMVA) is an image-domain inversion method for velocity model building. Automatic plane-wave WEMVA (PWEMVA) calculates the moveouts of plane-wave common-image gathers (CIGs) by searching a best-fitting parabola with semblance analysis and backprojects residual CIG moveouts into wavefield wave paths with a reflection tomographic kernel. However, 3-D PWEMVA is very computationally expensive because 3-D reflection tomographic inversion requires at least five 3-D reverse-time migrations per iteration and stores two types of source wavefields at model boundaries. We develop a joint inline and crossline PWEMVA method for efficient 3-D velocity model building. We alternatively implement the inline and crossline PWEMVAs with a constraint for each other, in which we iteratively construct the 3-D velocity model update through 1-D spline interpolation of 2-D gradients. The inline and crossline joint inversion is practical since PWEMVA only inverts for low-wavenumber velocity perturbations along wave paths, and the method can take less than 1 per cent of the computational cost of full 3-D PWEMVA. To construct unaliased plane waves for our joint inline and crossline PWEMVA, we develop a 3-D data interpolation method in the frequency–wavenumber (FK) domain to recover regularly and randomly missing traces. The method minimizes the misfit on sufficiently localized data subsets with iterative optimal step lengths and a gradient preconditioner that iteratively selects dominant dips along different azimuths. In numerical experiments, we use a 3-D synthetic seismic data set and a land 3-D field seismic data set acquired at the Farnsworth CO2-EOR (enhanced oil recovery) field to demonstrate the efficacy of our velocity model building and data interpolation methods.

Liu, Xuejian↗

Quantum state preparation by adiabatic evolution with custom gates

Quantum state preparation by adiabatic evolution is currently rendered ineffective by the long implementation times of the underlying quantum circuits, comparable to the decoherence time of present and near-term quantum devices. These implementation times can be significantly reduced by realizing these circuits with custom gates. Using classical computing, we model the output of a realistic two-qubit processor implementing the adiabatic evolution of a two-spin system by means of custom gates. This modeled output is then compared with the results of quantum simulations solving the same problem on IBM Quantum (IBMQ) systems. Additionally, when used to emulate the behavior of the IBMQ quantum circuit, our realistic model yields state fidelities ranging from 65% to 85%, similar to the actual performance of a diverse set of IBMQ devices. When we reduced the implementation time by using custom gates, however, the loss of fidelity was reduced by at least a factor of 4, allowing us to accurately extract the energy of the target state. This improvement is enough to render adiabatic evolution useful for quantum state preparation for small systems or as a preconditioner for other state preparation methods.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Fast and scalable quantum Monte Carlo simulations of electron-phonon models

We introduce methodologies for highly scalable quantum Monte Carlo simulations of electron-phonon models, and report benchmark results for the Holstein model on the square lattice. The determinant quantum Monte Carlo (DQMC) method is a widely used tool for simulating simple electron-phonon models at finite temperatures, but incurs a computational cost that scales cubically with system size. Alternatively, near-linear scaling with system size can be achieved with the hybrid Monte Carlo (HMC) method and an integral representation of the Fermion determinant. Here, we introduce a collection of methodologies that make such simulations even faster. To combat "stiffness" arising from the bosonic action, we review how Fourier acceleration can be combined with time-step splitting. To overcome phonon sampling barriers associated with strongly-bound bipolaron formation, we design global Monte Carlo updates that approximately respect particle-hole symmetry. To accelerate the iterative linear solver, we introduce a preconditioner that becomes exact in the adiabatic limit of infinite atomic mass. Finally, we demonstrate how stochastic measurements can be accelerated using fast Fourier transforms. Here, these methods are all complementary and, combined, may produce multiple orders of magnitude speedup, depending on model details.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Exploiting voxel-sparsity for bone imaging with sparse-view cone-beam computed tomography

An optimization-based image reconstruction frame work is developed specifically for bone imaging. This framework exploits voxel-sparsity by use of ℓ 1 -norm image regularization and it enables image reconstruction from sparse-view cone-beam computed tomography (CBCT) acquisition. The effectiveness of the voxel-sparsity regularization is enhanced by using a blurred image representation. Ramp-filtering is included in the data discrepancy term and it has the effect of acting as a preconditioner, reducing the necessary number of iterations. The bone image reconstruction framework is demonstrated on CBCT data taken from an equine metacarpal condyle specimen.

Bone imaging↗

A Scalable Multigrid Reduction Framework for Multiphase Poromechanics of Heterogeneous Media

Simulation of multiphase poromechanics involves solving a multiphysics problem in which multiphase flow and transport are tightly coupled with the porous medium deformation. To capture this dynamic interplay, fully implicit methods, also known as monolithic approaches, are usually preferred. The main bottleneck of a monolithic approach is that it requires solution of large linear systems that result from the discretization and linearization of the governing balance equations. Because such systems are nonsymmetric, indefinite, and highly ill-conditioned, preconditioning is critical for fast convergence. Recently, most efforts in designing efficient preconditioners for multiphase poromechanics have been dominated by physics-based strategies. Current state-of-the-art “black-box” solvers such as algebraic multigrid (AMG) are ineffective because they cannot effectively capture the strong coupling between the mechanics and the flow subproblems, as well as the coupling inherent in the multiphase flow and transport process. In this work, we develop an algebraic framework based on multigrid reduction (MGR) that is suited for tightly coupled systems of PDEs. Using this framework, the decoupling between the equations is done algebraically through defining appropriate interpolation and restriction operators. One can then employ existing solvers for each of the decoupled blocks or design a new solver based on knowledge of the physics. We demonstrate the applicability of our framework when used as a “black-box” solver for multiphase poromechanics. Here, we show that the framework is flexible to accommodate a wide range of scenarios, as well as efficient and scalable for large problems.

97 MATHEMATICS AND COMPUTING↗

Polynomial Preconditioned Arnoldi with Stability Control

Polynomial preconditioning can improve the convergence of the Arnoldi method for computing eigenvalues. Such preconditioning significantly reduces the cost of orthogonalization; for difficult problems, it can also reduce the number of matrix-vector products. Parallel computations can particularly benefit from the reduction of communication-intensive operations. Additoinally, the GMRES algorithm provides a simple and effective way of generating the preconditioning polynomial. For some problems high degree polynomials are especially effective, but they can lead to stability problems that must be mitigated. A two-level “double polynomial preconditioning” strategy provides an effective way to generate high-degree preconditioners.

97 MATHEMATICS AND COMPUTING↗

Monolithic Multigrid Methods for Magnetohydrodynamics

The magnetohydrodynamics equations model a wide range of plasma physics applications and are characterized by a nonlinear system of partial differential equations that strongly couples a charged fluid with the evolution of electromagnetic fields. After discretization and linearization, the resulting system of equations is generally difficult to solve due to the coupling between variables and the heterogeneous coefficients induced by the linearization process. In this paper, we investigate multigrid preconditioners for this system based on specialized relaxation schemes that properly address the system structure and coupling. Here, three extensions of Vanka relaxation are proposed and applied to problems with up to 170 million degrees of freedom and fluid and magnetic Reynolds numbers up to 400 for stationary problems and up to 20,000 for time-dependent problems.

97 MATHEMATICS AND COMPUTING↗

Parallel Element-based Algebraic Multigrid for H (c url ) and H (div) Problems Using the ParELAG Library

This paper presents the use of element-based algebraic multigrid (AMGe) hierarchies, implemented in the ParELAG (Parallel Element Agglomeration Algebraic Multigrid Upscaling and Solvers) library, to produce multilevel preconditioners and solvers for H (c url ) and H (div) formulations. ParELAG constructs hierarchies of compatible nested spaces, forming an exact de Rham sequence on each level. This allows the application of hybrid smoothers on all levels and AMS (Auxiliary-space Maxwell Solver) or ADS (Auxiliary-space Divergence Solver) on the coarsest levels, obtaining complete multigrid cycles. Numerical results are presented, showing the parallel performance of the proposed methods. As a part of the exposition, this paper demonstrates some of the capabilities of ParELAG and outlines some of the components and procedures within the library.

97 MATHEMATICS AND COMPUTING↗