Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “iterative solvers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Parallel Preconditioning for CFD Problems on the CM-5

Up to today, preconditioning methods on massively parallel systems have faced a major difficulty. The most successful preconditioning methods in terms of accelerating the convergence of the iterative solver such as incomplete LU factorizations are notoriously difficult to implement on parallel machines for two reasons: (1) the actual computation of the preconditioner is not very floating-point intensive, but requires a large amount of unstructured communication, and (2) the application of the preconditioning matrix in the iteration phase (i.e. triangular solves) are difficult to parallelize because of the recursive nature of the computation. Here we present a new approach to preconditioning for very large, sparse, unsymmetric, linear systems, which avoids both difficulties. We explicitly compute an approximate inverse to our original matrix. This new preconditioning matrix can be applied most efficiently for iterative methods on massively parallel machines, since the preconditioning phase involves only a matrix-vector multiplication, with possibly a dense matrix. Furthermore the actual computation of the preconditioning matrix has natural parallelism. For a problem of size n, the preconditioning matrix can be computed by solving n independent small least squares problems. The algorithm and its implementation on the Connection Machine CM-5 are discussed in detail and supported by extensive timings obtained from real problem data.

Simon, Horst D.↗

Iterative methods for large scale static analysis of structures on a scalable multiprocessor supercomputer

A parallel Preconditioned Conjugate Gradient (PCG) iterative solver has been developed and implemented on the iPSC-860 scalable hypercube. This new implementation makes use of the Parallel Automated Runtime Toolkit at ICASE (PARTI) primitives to efficiently program irregular communications patterns that exist in general sparse matrices and in particular in the finite element sparse stiffness matrices. The iterative PCG has been used to solve the finite element equations that result from discretizing large scale aerospace structures. In particular, the static response of the High Speed Civil Transport (HSCT) finite element model is solved on the iPSC-860.

Sobh, Nahil Atef↗

ANTS

The ANTS code (Alternate Non-Linear Two-phase Solver) is based on a novel non-linear solution algorithm for the solution of the two-phase, subchannel fluid equations. It achieves its performance through decoupling of the two-phase momentum equations (axial and transverse) from the axial phasic mass and energy equations which allows for a nested non-linear iteration scheme. This enables a plane-by-plane solution which the inner iteration focuses on a reduced non-linear equation set for the primitives in phasic mass flow rate, enthalpy and void for each node edge. Single node edges are coupled as part of the outer iteration via surface mass fluxes which appear as source terms in the inner iteration scheme. The outer iteration readily accommodates two-phase flow phenomena closure relationships for subchannel mixing and void drift. A primary feature is the use of a non-staggered mesh computational mesh and steady-state iterative solver in contrast to all existing subchannel codes.

Kropaczek, David J↗

Technique for Solving Electrically Small to Large Structures for Broadband Applications

Fast iterative algorithms are often used for solving Method of Moments (MoM) systems, having a large number of unknowns, to determine current distribution and other parameters. The most commonly used fast methods include the fast multipole method (FMM), the precorrected fast Fourier transform (PFFT), and low-rank QR compression methods. These methods reduce the O(N) memory and time requirements to O(N log N) by compressing the dense MoM system so as to exploit the physics of Green s Function interactions. FFT-based techniques for solving such problems are efficient for spacefilling and uniform structures, but their performance substantially degrades for non-uniformly distributed structures due to the inherent need to employ a uniform global grid. FMM or QR techniques are better suited than FFT techniques; however, neither the FMM nor the QR technique can be used at all frequencies. This method has been developed to efficiently solve for a desired parameter of a system or device that can include both electrically large FMM elements, and electrically small QR elements. The system or device is set up as an oct-tree structure that can include regions of both the FMM type and the QR type. The system is enclosed with a cube at a 0- th level, splitting the cube at the 0-th level into eight child cubes. This forms cubes at a 1st level, recursively repeating the splitting process for cubes at successive levels until a desired number of levels is created. For each cube that is thus formed, neighbor lists and interaction lists are maintained. An iterative solver is then used to determine a first matrix vector product for any electrically large elements as well as a second matrix vector product for any electrically small elements that are included in the structure. These matrix vector products for the electrically large and small elements are combined, and a net delta for a combination of the matrix vector products is determined. The iteration continues until a net delta is obtained that is within the predefined limits. The matrix vector products that were last obtained are used to solve for the desired parameter. The solution for the desired parameter is then presented to a user in a tangible form; for example, on a display.

Jandhyala, Vikram↗

Combined fast multipole-QR compression technique for solving electrically small to large structures for broadband applications

An approach that efficiently solves for a desired parameter of a system or device that can include both electrically large fast multipole method (FMM) elements, and electrically small QR elements. The system or device is setup as an oct-tree structure that can include regions of both the FMM type and the QR type. An iterative solver is then used to determine a first matrix vector product for any electrically large elements, and a second matrix vector product for any electrically small elements that are included in the structure. These matrix vector products for the electrically large elements and the electrically small elements are combined, and a net delta for a combination of the matrix vector products is determined. The iteration continues until a net delta is obtained that is within predefined limits. The matrix vector products that were last obtained are used to solve for the desired parameter.

Jandhyala, Vikram↗

Linearized frequency domain Landau-Lifshitz-Gilbert equation formulation

We present a general finite element linearized Landau-Lifshitz-Gilbert equation (LLGE) solver for magnetic systems under weak time-harmonic excitation field. The linearized LLGE is obtained by assuming a small deviation around the equilibrium state of the magnetic system. Inserting such expansion into LLGE and keeping only first order terms gives the linearized LLGE, which gives a frequency domain solution for the complex magnetization amplitudes under an external time-harmonic applied field of a given frequency. We solve the linear system with an iterative solver using generalized minimal residual method. We construct a preconditioner matrix to effectively solve the linear system. The validity, effectiveness, speed, and scalability of the linear solver are demonstrated via numerical examples.

36 MATERIALS SCIENCE↗

On High-Order/Low-Order and Micro-Macro Methods for Implicit Time-Stepping of the BGK Model

In this paper, a high-order/low-order (HOLO) method is combined with a micro-macro (MM) decomposition to accelerate iterative solvers in fully implicit time-stepping of the Bhatnagar–Gross–Krook (BGK) equation for gas dynamics. The MM formulation represents a kinetic distribution as the sum of a local Maxwellian and a perturbation. In highly collisional regimes, the perturbation away from initial and boundary layers is small and can be compressed to reduce the overall storage cost of the distribution. The convergence behavior of the MM methods, the usual HOLO method, and the standard source iteration method is analyzed on a linear BGK model. Both the HOLO and MM methods are implemented using a discontinuous Galerkin (DG) discretization in phase space, which naturally preserves the consistency between high- and low-order models required by the HOLO approach. Furthermore, the accuracy and performance of these methods are compared on the Sod shock tube problem and a sudden wall heating boundary layer problem. Overall, the results demonstrate the robustness of the MM and HOLO approaches and illustrate the compression benefits enabled by the MM formulation when the kinetic distribution is near equilibrium.

BGK model↗

Scalable Multiphysics Block Preconditioning for Low Mach Number Compressible Resistive MHD with Application to Magnetic Confinement Fusion

This study investigates multiphysics block preconditioners that are critical in devising scalable Newton–Krylov iterative solvers for longer time-scale fully implicit fluid plasma models. The specific model of interest is the visco-resistive, low Mach number, compressible magnetohydrodynamics (MHD) model. This model describes the dynamics of conducting fluids in the presence of electromagnetic fields and can be used to study aspects of astrophysical phenomena, important science and technology applications, and basic plasma physics. The specific application of interest that motivates this study is the macroscopic simulation of longer time-scale stability and disruptions of magnetic confinement fusion devices, specifically the ITER Tokamak. The computational solution of the governing balance equations for mass, momentum, heat transfer, and magnetic induction for resistive MHD systems can be extremely challenging. These difficulties arise from both the strong nonlinear, nonsymmetric coupling of fluid and electromagnetic phenomena as well as the significant range of time and length scales that the interactions of these physical mechanisms produce. To handle the range of time and spatial scales of interest, a fully implicit unstructured variational multiscale finite element formulation is employed. For the scalable solution of the Newton linearized systems, fully coupled block preconditioners are designed to leverage algebraic multigrid subsolves. In conclusion, results are presented for the strong and weak scaling of the method as well as the robustness of these techniques for a large range of Lundquist numbers.

97 MATHEMATICS AND COMPUTING↗

Parallel-vector computation for linear structural analysis and non-linear unconstrained optimization problems

Several parallel-vector computational improvements to the unconstrained optimization procedure are described which speed up the structural analysis-synthesis process. A fast parallel-vector Choleski-based equation solver, pvsolve, is incorporated into the well-known SAP-4 general-purpose finite-element code. The new code, denoted PV-SAP, is tested for static structural analysis. Initial results on a four processor CRAY 2 show that using pvsolve reduces the equation solution time by a factor of 14-16 over the original SAP-4 code. In addition, parallel-vector procedures for the Golden Block Search technique and the BFGS method are developed and tested for nonlinear unconstrained optimization. A parallel version of an iterative solver and the pvsolve direct solver are incorporated into the BFGS method. Preliminary results on nonlinear unconstrained optimization test problems, using pvsolve in the analysis, show excellent parallel-vector performance indicating that these parallel-vector algorithms can be used in a new generation of finite-element based structural design/analysis-synthesis codes.

Nguyen, D. T.↗

MAPPRAISER: A massively parallel map-making framework for multi-kilo pixel CMB experiments

Forthcoming cosmic microwave background (CMB) polarized anisotropy experiments have the potential to revolutionize our understanding of the Universe and fundamental physics. The sought-after, tale-telling signatures will be however distributed over voluminous data sets which these experiments will collect. These data sets will need to be efficiently processed and unwanted contributions due to astrophysical, environmental, and instrumental effects characterized and efficiently mitigated in order to uncover the signatures. This poses a significant challenge to data analysis methods, techniques, and software tools which will not only have to be able to cope with huge volumes of data but to do so with unprecedented precision driven by the demanding science goals posed for the new experiments. A keystone of efficient CMB data analysis is solvers of very large linear systems of equations. Such systems appear in very diverse contexts throughout CMB data analysis pipelines, however they typically display similar algebraic structures and can therefore be solved using similar numerical techniques. Linear systems arising in the so-called map-making problem are one of the most prominent and common ones. In this work we present a massively parallel, flexible and extensible framework, comprised of a numerical library, MIDAPACK, and a high level code, MAPPRAISER, which provide tools for solving efficiently such systems. Here, the framework implements iterative solvers based on conjugate gradient techniques: enlarged and preconditioned using different preconditioners. We demonstrate the framework on simulated examples reflecting basic characteristics of the forthcoming data sets issued by ground-based and satellite-borne instruments, executing it on as many as 16,384 compute cores. The software is developed as an open source project freely available to the community at: https://github.com/B3Dcmb/midapack.

79 ASTRONOMY AND ASTROPHYSICS↗

A Segregated Approach for Modeling the Electrochemistry in the 3-D Microstructure of Li-Ion Batteries and Its Acceleration Using Block Preconditioners

Abstract Battery performance is strongly correlated with electrode microstructure. Electrode materials for lithium-ion batteries have complex microstructure geometries that require millions of degrees of freedom to solve the electrochemical system at the microstructure scale. A fast-iterative solver with an appropriate preconditioner is then required to simulate large representative volume in a reasonable time. In this work, a finite element electrochemical model is developed to resolve the concentration and potential within the electrode active materials and the electrolyte domains at the microstructure scale, with an emphasis on numerical stability and scaling performances. The block Gauss-Seidel (BGS) numerical method is implemented because the system of equations within the electrodes is coupled only through the nonlinear Butler–Volmer equation, which governs the electrochemical reaction at the interface between the domains. The best solution strategy found in this work consists of splitting the system into two blocks—one for the concentration and one for the potential field—and then performing block generalized minimal residual preconditioned with algebraic multigrid, using the FEniCS and the Portable, Extensible Toolkit for Scientific Computation libraries. Significant improvements in terms of time to solution (six times faster) and memory usage (halving) are achieved compared with the MUltifrontal Massively Parallel sparse direct Solver. Additionally, BGS experiences decent strong parallel scaling within the electrode domains. Last, the system of equations is modified to specifically address numerical instability induced by electrolyte depletion, which is particularly valuable for simulating fast-charge scenarios relevant for automotive application.

25 ENERGY STORAGE↗

Multilevel Graph Partitioning for Three-Dimensional Discrete Fracture Network Flow Simulations

We present a topology-based method for mesh-partitioning in three-dimensional discrete fracture network (DFN) simulations that takes advantage of the intrinsic multi-level nature of a DFN. DFN models are used to simulate flow and transport through low-permeability fractured media in the subsurface by explicitly representing fractures as discrete entities. The governing equations for flow and transport are numerically integrated on computational meshes generated on the interconnected fracture networks. Modern high-fidelity DFN simulations require high-performance computing on multiple processors where performance and scalability depends partially on obtaining a high-quality partition of the mesh to balance work-loads and minimize communication across all processors. The discrete structure of a DFN naturally lends itself to various graph representations, which can be thought of as coarse-scale representations of the computational mesh. Using this concept, we develop two applications of the multilevel graph partitioning algorithm to partition the mesh of a DFN. In the first, we project a partition of the graph based on the DFN topology onto the mesh of the DFN and in the second, this DFN-based projection is used as the initial condition for further partitioning refinement of the mesh. We compare the performance of these methods with standard multi-level graph partitioning using graph-based metrics (cut, imbalance, partitioning time), computational-based metrics (FLOPS, iterations, solver time), and total run time. The DFN-based and the mesh-based partitioning methods are comparable in terms of the graph-based metrics, but the time required to obtain the partition is several orders of magnitude faster using the DFN-based partitions. The computation-based metrics show comparable performance between both methods so, in combination, the DFN-based partitions are several orders of magnitude faster than the mesh-based partition. Furthermore, the method which uses the DFN-partition solution as the initial condition of the mesh partition provided cut and imbalance values that were close to the mesh-based partition but in a fraction of the time. In turn, this hybrid method outperformed both of the other methods in terms of the total run time.

58 GEOSCIENCES↗

Effect of non-uniform void distributions on the yielding of metals

High-throughput (several thousand) calculations have been carried out to investigate the yield behavior of porous materials with randomly distributed pores, porosity levels over four orders of magnitude and up to a hundred pores per simulation box. To this end, a Galerkin based fast Fourier transform (FFT) formulation was enhanced to deal with high phase contrast materials. In addition, GPU parallelization was employed in solving the governing equation for strain fluctuations using a Krylov iterative solver. Emphasis is laid on the conditions under which percolation of plastically non-deforming zones through the porous network emerge, a regime termed unhomogeneous yielding. By way of contrast, the regime where the plastic strain fluctuations (associated with the heterogeneous void-matrix aggregate) fall below the percolation threshold is defined as homogeneous yielding. Here, we find that nonuniform pore distributions only affect unhomogeneous yielding and have a universal softening effect. The extent of this distribution softening is analyzed as a function of porosity, cell size and number of realizations. Whether the uncovered universal distribution softening has direct implications on failure resistance of porous materials is discussed.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

A thermodynamic model of integrated liquid-to-liquid thermoelectric heat pump systems

Thermoelectric (TE) heat pumps (TEHPs) are advantageous for heating and cooling in various applications because of their modularity and simple design. A TEHP system includes the TE modules with p- and n-type materials bonded to substrates, plus heat exchangers, thermal interfaces to the heat exchangers, and heat transfer fluids. Although modeling an individual TE module has been extensively studied, limited studies have reported performance at the larger system-level. Furthermore, no prior study has addressed the impact of temperature-dependent TE material properties (e.g., electric resistivity, thermal conductivity, and Seebeck coefficient) on overall heat-pump-system-level performance. Here, this work presents a mathematical model for TEHP system performance based on Goldsmid's approach for TE material performance, “effective” TE material properties, Gnielinski's correlation for convective heat transfer, and thermal balance theory for a heat exchange network. This combined approach provides an accurate model of the liquid-to-liquid TEHP system. Three different approaches—one empirical, one based on the manufacturer's specifications, and one drawn from the literature—were then used to determine values for TE material properties. The first two methods treated properties as constants, while the last approach treated properties as surface-temperature-based functions. Finally, experimental TEHP data was used to validate the models, all with relative absolute deviations of approximately 10% when predicting heating capacity and 10%–25% when forecasting cooling capacity up to a 30 K surface temperature lift. The results demonstrated that, at the TEHP system level, the TE material properties could be treated as constants, avoiding solver iterations and reducing the performance uncertainty by up to 95%.

42 ENGINEERING↗

DG-IMEX method for a two-moment model for radiation transport in the $\mathscr{O}$($v$/$c$) limit

Here, we consider neutral particle systems described by moments of a phase-space density and propose a realizability-preserving numerical method to evolve a spectral two-moment model for particles interacting with a background fluid moving with nonrelativistic velocities. The system of nonlinear moment equations, with special relativistic corrections to $\mathscr{O}$($v$/$c$), expresses a balance between phase-space advection and collisions and includes velocity-dependent terms that account for spatial advection, Doppler shift, and angular aberration. The model is conservative for the correct $\mathscr{O}$($v$/$c$) Eulerian-frame number density and is consistent, to $\mathscr{O}$($v$/$c$), with Eulerian-frame energy and momentum conservation. This model is closely related to the one promoted by Lowrie et al. and similar to models currently used to study transport phenomena in large-scale simulations of astrophysical environments. The proposed numerical method is designed to preserve moment realizability, which guarantees that the moments correspond to a nonnegative phase-space density. The realizability-preserving scheme consists of the following key components: (i) a strong stability-preserving implicit-explicit (IMEX) time-integration method; (ii) a discontinuous Galerkin (DG) phase-space discretization with carefully constructed numerical uxes; (iii) a realizability-preserving implicit collision update; and(iv) a realizability-enforcing limiter. In time integration, nonlinearity of the moment model necessitates solution of nonlinear equations, which we formulate as fixed-point problems and solve with tailored iterative solvers that preserve moment realizability with guaranteed global convergence. We also analyze the simultaneous Eulerian-frame number and energy conservation properties of the semi-discrete DG scheme and propose a "spectral redistribution" scheme that promotes Eulerian-frame energy conservation. Through numerical experiments, we demonstrate the accuracy and robustness of this DG-IMEX method and investigate its Eulerian-frame energy conservation properties.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Simulating self-powered neutron detector responses to infer burnup-induced power distribution perturbations in next-generation light water reactors

Understanding how 3D power distribution will be monitored throughout reactor core volumetric space in next-generation nuclear power reactors is crucial to the design, deployment, and licensing of these reactors. Although numerous techniques exist for 3D power distribution monitoring based on the response of both in situ and ex situ sensors currently implemented or proposed for use in the US reactor fleet, crucial details about these techniques are often unclear. The publicly available documentation does not include information such as how well these techniques are characterized and optimized in their implementations and the levels of uncertainty in the inferred 3D power distribution. The work described herein investigated a recently developed 3D power distribution inferencing method as applied to two next-generation reactor simulations: (1) the NuScale small modular reactor design and (2) the Westinghouse AP1000 design, both of which contain in-core strings of vanadium self-powered neutron detectors (SPNDs). This investigation considered a range of SPND string sensor densities, as well as a range of 3D power distribution axial segment sizes. In this work, SPND response simulation is informed by neutron flux calculations in representative homogenized cores. For the different sensor densities and power distribution axial segment sizes in these simulations, the average solution error, solver iterations, and run time were tracked to parameterize the sensor-core configuration.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Robust scalable initialization for Bayesian variational inference with multi-modal Laplace approximations

Predictive modeling typically relies on Bayesian model calibration to provide uncertainty quantification. Variational inference utilizing fully independent (“mean-field”) Gaussian distributions are often used as approximate probability density functions. This simplification is attractive since the number of variational parameters grows only linearly with the number of unknown model parameters. However, the resulting diagonal covariance structure and unimodal behavior can be too restrictive to provide useful approximations of intractable Bayesian posteriors that exhibit highly non-Gaussian behavior, including multimodality. High-fidelity surrogate posteriors for these problems can be obtained by considering the family of Gaussian mixtures. Gaussian mixtures are capable of capturing multiple modes and approximating any distribution to an arbitrary degree of accuracy, while maintaining some analytical tractability. Unfortunately, variational inference using Gaussian mixtures with full-covariance structures suffers from a quadratic growth in variational parameters with the number of model parameters. The existence of multiple local minima due to strong nonconvex trends in the loss functions often associated with variational inference present additional complications, These challenges motivate the need for robust initialization procedures to improve the performance and computational scalability of variational inference with mixture models. In this work, we propose a method for constructing an initial Gaussian mixture model approximation that can be used to warm-start the iterative solvers for variational inference. The procedure begins with a global optimization stage in model parameter space. In this step, local gradient-based optimization, globalized through multistart, is used to determine a set of local maxima, which we take to approximate the mixture component centers. Around each mode, a local Gaussian approximation is constructed via the Laplace approximation. Finally, the mixture weights are determined through constrained least squares regression. The robustness and scalability of the proposed methodology is demonstrated through application to an ensemble of synthetic tests using high-dimensional, multimodal probability density functions. Here, the practical aspects of the approach are demonstrated with inversion problems in structural dynamics.

97 MATHEMATICS AND COMPUTING↗

Encoder–decoder neural network for solving the nonlinear Fokker–Planck–Landau collision operator in XGC

An encoder–decoder neural network has been used to examine the possibility for acceleration of a partial integro-differential equation, the Fokker–Planck–Landau collision operator. This is part of the governing equation in the massively parallel particle-in-cell code XGC, which is used to study turbulence in fusion energy devices. The neural network emphasizes physics-inspired learning, where it is taught to respect physical conservation constraints of the collision operator by including them in the training loss, along with the ℓ 2 loss. In particular, network architectures used for the computer vision task of semantic segmentation have been used for training. A penalization method is used to enforce the ‘soft’ constraints of the system and integrate error in the conservation properties into the loss function. During training, quantities representing the particle density, momentum and energy for all species of the system are calculated at each configuration vertex, mirroring the procedure in XGC. This simple training has produced a median relative loss, across configuration space, of the order of 10 –4 , which is low enough if the error is of random nature, but not if it is of drift nature in time steps. The run time for the current Picard iterative solver of the operator is O(n 2 ), where n is the number of plasma species. As the XGC1 code begins to attack problems including a larger number of species, the collision operator will become expensive computationally, making the neural network solver even more important, especially since its training only scales as O(n). Here, a wide enough range of collisionality has been considered in the training data to ensure the full domain of collision physics is captured. An advanced technique to decrease the losses further will be subject of a subsequent report. Eventual work will include expansion of the network to include multiple plasma species.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗