Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “linear equation systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Toward efficient polynomial preconditioning for GMRES

Here, we present a polynomial preconditioner for solving large systems of linear equations. The polynomial is derived from the minimum residual polynomial (the GMRES polynomial) and is more straightforward to compute and implement than many previous polynomial preconditioners. Our current implementation of this polynomial using its roots is naturally more stable than previous methods of computing the same polynomial. We implement further stability control using added roots, and this allows for high degree polynomials. We discuss the effectiveness and challenges of root-adding and give an additional check for stability. In this article, we study the polynomial preconditioner applied to GMRES; however it could be used with any Krylov solver. This polynomial preconditioning algorithm can dramatically improve convergence for some problems, especially for difficult problems, and can reduce dot products by an even greater margin.

97 MATHEMATICS AND COMPUTING↗

A smeared crack modeling framework accommodating multi-directional fracture at finite strains

A generic smeared crack modeling framework predicated on the deformation gradient decomposition (DGD) approach is proposed for use in dynamic fracture problems at finite strains, accommodating failure along multiple mutually orthogonal fracture planes embedded within an independently defined bulk material model. Within this constitutive framework, the traction equilibrium conditions imposed at each failure surface are used to determine the associated crack displacements stored as internal state variables. In general, the enforcement of interfacial equilibrium entails the implicit solution of a non-linear system of equations within the constitutive update procedure. However, if inertial effects arising due to the relative motion of the fractured material are incorporated within the model, the traction equilibrium conditions are shown to give rise to corresponding dynamic equations of motion governing the time-evolution of the crack opening displacements. For dynamic problems, an explicit time-integration procedure is devised to efficiently update the material state, subject to a set of internal frictionless contact constraints to prevent material inversion. Finally, the efficacy of the proposed modeling framework is investigated through several benchmark dynamic fracture problems run within the explicit finite element code DYNA3D.

42 ENGINEERING↗

Asynchronous Richardson iterations: theory and practice

We consider asynchronous versions of the first- and second-order Richardson methods for solving linear systems of equations. These methods depend on parameters whose values are chosen a priori. We explore the parameter values that can be proven to give convergence of the asynchronous methods. This is the first such analysis for asynchronous second-order methods. We find that for the first-order method, the optimal parameter value for the synchronous case also gives an asynchronously convergent method. For the second-order method, the parameter ranges for which we can prove asynchronous convergence do not contain the optimal parameter values for the synchronous iteration. In practice, however, the asynchronous second-order iterations may still converge using the optimal parameter values, or parameter values close to the optimal ones, despite this result. We explore this behavior with a multithreaded parallel implementation of the asynchronous methods.

97 MATHEMATICS AND COMPUTING↗

Leveraging operator learning to accelerate convergence of the preconditioned conjugate gradient method

We propose a new deflation strategy to accelerate the convergence of the preconditioned conjugate gradient (PCG) method for solving parametric large-scale linear systems of equations. Unlike traditional deflation techniques that rely on eigenvector approximations or recycled Krylov subspaces, we generate the deflation subspaces using operator learning, specifically the Deep Operator Network (DeepONet). To this aim, we introduce two complementary approaches for assembling the deflation operators. The first approach approximates near-null space vectors of the discrete PDE operator using the basis functions learned by the DeepONet. The second approach directly leverages solutions predicted by the DeepONet. To further enhance convergence, we also propose several strategies for prescribing the sparsity pattern of the deflation operator. Here, a comprehensive set of numerical experiments encompassing steady-state, time-dependent, scalar, and vector-valued problems posed on both structured and unstructured geometries is presented and demonstrates the effectiveness of the proposed DeepONet-based deflated PCG method, as well as its generalization across a wide range of model parameters and problem resolutions.

Deflation↗

Simulation toolkit for digital material characterization of large image-based microstructures

In this paper, an efficient image-based simulation toolkit for material characterization is presented, which is scalable to work from personal computers to workstations. The effective thermal conductivity, elasticity, and permeability are evaluated employing a computational homogenization framework based on the Finite Element Method (FEM). Two complementary open-source packages are presented: one developed in Python, which can convert digital images into voxel meshes (pyTomoviewer); the other developed in Julia, that can run numerical simulations to compute effective material properties (chpack). Also, a CUDA C version of chpack is provided (chfem_gpu). They were designed to deal with large multi-phase models, so strategies were devised to minimize their memory footprint, while avoiding a high toll on execution time. The voxel-based approach significantly simplifies the FEM meshes and allows efficient matrix-free implementations. In that sense, to handle large linear systems of equations, the element-by-element (EBE) technique is adopted, in conjunction with a low-memory implementation of the Preconditioned Conjugate Gradient (PCG) method. Finally, the code was thoroughly tested on an artificial geometry made of a square array of cylinders, for which analytical solutions exist, as well as on a real micro-tomographic reconstruction of FiberForm TM , a carbon preform commonly used in thermal protection systems.

36 MATERIALS SCIENCE↗

QUIC-URB and QUIC-fire extension to complex terrain: Development of a terrain-following coordinate system

Ensemble-based approaches to prescribed fire planning cannot be supported by CFD-based models like FIRETEC and WFDS because they are too computationally expensive and cannot leverage LES approaches like CAWFE and WRF-SFIRE because too coarse of resolution. QUIC-Fire was developed to fill this gap but it cannot currently address complex terrain, typical for instance of the Western United States. In this paper, we describe the extension of the diagnostic wind model QUIC-URB, the wind engine of QUIC-Fire, to a terrain-following coordinate system. In particular, the paper presents the mathematical derivation of the wind solver leading to a linear system of equations that are solved through the successive over-relaxation method. The model is validated against a standard test used in previous works (the Askervein Hill) and against a new dataset from measurements in the Socorro Mountains, New Mexico. The terrain-following implementation captured the correct phenomenology for the isolated Askervein Hill, with a wind speed up at the top of the hill. We report the model agreed well with measurements on the upwind side of the peak, but overestimated speed-up on the downwind side of the hill. This is due to the inability of the model to generate flow separation and wake-eddy dynamics. On a common laptop, the divergence-free wind field was obtained in 6 s, making the solver appealing for coupled fire–atmosphere simulations. The Socorro Mountain was highly complex, with many cliff faces, peaks, and valleys. Although the model captures the magnitude and direction of inlet and outlet areas of the domain, it performs rather poorly in the valley region and in the regions near the steep cliffs. Hence, the model shows good agreement with data in areas of open sloped terrain but lacks in areas where flow separation and thermally driven effects may be present (neither effect was addressed in this work). Results highlight that future work should focus on the implementation of parameterizations of wake-eddies, similar to QUIC-URB’s building parameterizations, and on thermodynamic-driven flow.

54 ENVIRONMENTAL SCIENCES↗

Accelerating eigenvalue computation for nuclear structure calculations via perturbative corrections

Subspace projection methods utilizing perturbative corrections have been proposed for computing the lowest few eigenvalues and corresponding eigenvectors of large Hamiltonian matrices. In this paper, we build upon these methods and introduce the term Subspace Projection with Perturbative Corrections (SPPC) method to refer to this approach. We tailor the SPPC for nuclear many-body Hamiltonians represented in a truncated configuration interaction subspace, i.e., the no-core shell model (NCSM). We use the hierarchical structure of the NCSM Hamiltonian to partition the Hamiltonian as the sum of two matrices. The first matrix corresponds to the Hamiltonian represented in a small configuration space, whereas the second is viewed as the perturbation to the first matrix. Eigenvalues and eigenvectors of the first matrix can be computed efficiently. Because of the split, perturbative corrections to the eigenvectors of the first matrix can be obtained efficiently from the solutions of a sequence of linear systems of equations defined in the small configuration space. These correction vectors can be combined with the approximate eigenvectors of the first matrix to construct a subspace from which more accurate approximations of the desired eigenpairs can be obtained. We show by numerical examples that the SPPC method can be more efficient than conventional iterative methods for solving large-scale eigenvalue problems such as the Lanczos, block Lanczos and the locally optimal block preconditioned conjugate gradient (LOBPCG) method. The method can also be combined with other methods to avoid convergence stagnation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Extended Lagrangian Born–Oppenheimer molecular dynamics for orbital-free density-functional theory and polarizable charge equilibration models

We report extended Lagrangian Born–Oppenheimer molecular dynamics (XL-BOMD) is formulated for orbital-free Hohenberg–Kohn density-functional theory and for charge equilibration and polarizable force-field models that can be derived from the same orbital-free framework. The purpose is to introduce the most recent features of orbital-based XL-BOMD to molecular dynamics simulations based on charge equilibration and polarizable force-field models. These features include a metric tensor generalization of the extended harmonic potential, preconditioners, and the ability to use only a single Coulomb summation to determine the fully equilibrated charges and the interatomic forces in each time step for the shadow Born–Oppenheimer potential energy surface. The orbital-free formulation has a charge-dependent, short-range energy term that is separate from long-range Coulomb interactions. This enables local parameterizations of the short-range energy term, while the long-range electrostatic interactions can be treated separately. The theory is illustrated for molecular dynamics simulations of an atomistic system described by a charge equilibration model with periodic boundary conditions. The system of linear equations that determines the equilibrated charges and the forces is diagonal, and only a single Ewald summation is needed in each time step. The simulations exhibit the same features in accuracy, convergence, and stability as are expected from orbital-based XL-BOMD.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Mixed-precision iterative refinement using tensor cores on GPUs to accelerate solution of linear systems

Double-precision floating-point arithmetic (FP64) has been the de facto standard for engineering and scientific simulations for several decades. Problem complexity and the sheer volume of data coming from various instruments and sensors motivate researchers to mix and match various approaches to optimize compute resources, including different levels of floating-point precision. In recent years, machine learning has motivated hardware support for half-precision floating-point arithmetic. A primary challenge in high-performance computing is to leverage reduced-precision and mixed-precision hardware. We show how the FP16/FP32 Tensor Cores on NVIDIA GPUs can be exploited to accelerate the solution of linear systems of equations Ax = b without sacrificing numerical stability. The techniques we employ include multiprecision LU factorization, the preconditioned generalized minimal residual algorithm (GMRES), and scaling and auto-adaptive rounding to avoid overflow. We also show how to efficiently handle systems with multiple right-hand sides. On the NVIDIA Quadro GV100 (Volta) GPU, we achieve a 4×-5× performance increase and 5× better energy efficiency versus the standard FP64 implementation while maintaining an FP64 level of numerical stability.

GMRES↗

AMG Preconditioners based on parallel hybrid coarsening and multi-objective graph matching

We describe preliminary results from a multi-objective graph matching algorithm, in the coarsening step of an aggregation-based Algebraic MultiGrid (AMG) preconditioner, for solving large and sparse linear systems of equations on high-end parallel computers. We have two objectives. First, we wish to improve the convergence behavior of the AMG method when applied to highly anisotropic problems. Second, we wish to extend the parallel package \texttt{PSCToolkit} to exploit multi-threaded parallelism at the node level on multi-core processors. Our matching proposal balances the need to simultaneously compute high weights and large cardinalities by a new formulation of the weighted matching problem combining both these objectives using a parameter $\lambda$. We compute the matching by a parallel $2/3-\varepsilon$-approximation algorithm for maximum weight matchings. Results with the new matching algorithm show that for a suitable choice of the parameter $\lambda$ we compute effective preconditioners in the presence of anisotropy, i.e., smaller solve times, setup times, iterations counts, and operator complexity.

D'Ambra, Pasqua↗

Benchmarking Optimizers for Qumode State Preparation with Variational Quantum Algorithms

Quantum state preparation involves preparing a target state from an initial system, a process integral to applications such as quantum machine learning and solving systems of linear equations. Recently, there has been a growing interest in qumodes due to advancements in the field and their potential applications. However there is a notable gap in the literature specifically addressing this area. This paper aims to bridge this gap by providing performance benchmarks of various optimizers used in state preparation with Variational Quantum Algorithms. We conducted extensive testing across multiple scenarios, including different target states, both ideal and sampling simulations, and varying numbers of basis gate layers. Our evaluations offer insights into the complexity of learning each type of target state and demonstrate that some optimizers perform better than others in this context. Notably, the Powell optimizer was found to be exceptionally robust against sampling errors, making it a preferred choice in scenarios prone to such inaccuracies. Additionally, the Simultaneous Perturbation Stochastic Approximation optimizer was distinguished for its efficiency and ability to handle increased parameter dimensionality effectively.

Kan, Shuwen [Fordham University]↗

Performance Analysis and Optimal Node-aware Communication for Enlarged Conjugate Gradient Methods

Krylov methods are a key way of solving large sparse linear systems of equations but suffer from poor strong scalability on distributed memory machines. Furthermore, this is due to high synchronization costs from large numbers of collective communication calls alongside a low computational workload. Enlarged Krylov methods address this issue by decreasing the total iterations to convergence, an artifact of splitting the initial residual and resulting in operations on block vectors. In this article, we present a performance study of an enlarged Krylov method, Enlarged Conjugate Gradients (ECG), noting the impact of block vectors on parallel performance at scale. Most notably, we observe the increased overhead of point-to-point communication as a result of denser messages in the sparse matrix-block vector multiplication kernel. Additionally, we present models to analyze expected performance of ECG, as well as motivate design decisions. Most importantly, we introduce a new point-to-point communication approach based on node-aware communication techniques that increases efficiency of the method at scale.

97 MATHEMATICS AND COMPUTING↗

OpenMxP-Opensource Mixed Precision Computing

This is an opensource library for benchmarking the system's GPU mixed precision capabilities. The software calculates solution of the system of linear equation in 64bit accuracy using mixed precision techniques and iterative refinement. Original benchmark designed is done by ICL, and it is name HPL-MxP (HPL-AI)

Lu, Hao↗

Terrain-Influenced Winds and Fire-Fire Interactions in Wildland Fire Simulations [Dissertation]

Ensemble-based approaches to prescribed fire planning cannot be supported by computational fluid dynamics based models like FIRETEC and the Wildland-Urban Interface Fire Dynamics Simulator (WFDS) because they are too computationally expensive and cannot leverage large eddy simulation approaches like CAWFE and WRF-SFIRE because they have too coarse of resolution. QUIC-Fire was developed to fill this gap but it cannot currently address complex terrain, that is typical for instance in the Western United States. This dissertation describes a variety of improvements made to QUIC-Fire and its various incorporated algorithms in an effort to make it a viable tool in simulating wildland and prescribed fires on terrain. The modifications made to QUIC-Fire are described in three chapters. The first chapter describes the extension of the diagnostic wind model QUIC-URB, the wind engine of QUIC-Fire, to a terrain-following coordinate system. The terraininfluenced winds it generates are analyzed and compared. In particular, this chapter presents the mathematical derivation of the wind solver leading to a linear system of equations that are solved through the successive over-relaxation method. The model is validated against a standard test used in previous works (the Askervein Hill) and against a new dataset from measurements in the Socorro Mountains, New Mexico. The terrain-following implementation captures the correct phenomenology for the isolated Askervein Hill, with a wind speed up at the top of the hill. The model agrees well with measurements on the upwind side of the peak, but overestimates speed-up on the downwind side of the hill. This is due to the inability of the model to generate flow separation and wake-eddy dynamics. On a common laptop, the divergence-free wind field is obtained in 6 s, making the solver appealing for coupled fire-atmosphere simulations. The Socorro Mountain is highly complex, with many cliff faces, peaks, and valleys. Although the model captures the magnitude and direction of inlet and outlet areas of the domain, it performs rather poorly in the valley region and in the regions near the steep cliffs. Hence, the model shows good agreement with data in areas of open sloped terrain but lacks in areas where flow separation and thermally driven effects may be present (neither effect is addressed in this work). In the second chapter the implementation of the terrain-following version of QUIC-URB into QUIC-Fire, and the necessary changes needed to include terrain are described. No changes to the underlying fire spread algorithm are made other than what is required to correctly account for the inclusion of terrain. Previously published FIRETEC results that use five different topographies that share the same centerline profile are compared to simulation results from the modified QUIC-Fire that use the same topographies and fuels. QUIC-Fire results show overall similar behaviors in terms of how the topographies affect fire shapes and trends in spread rates. Due to the terrain-following version of QUIC-URB being unable to generate flow separations at the crest of hills, fire spread rates in these regions across all topographies are over-predicted when compared to FIRETEC. Lateral fire growth shows similar trends with FIRETEC between topographies but does not capture the increase in spread due to a diagonal interface between grassland and forested fuel region of the domain. These results suggest that there are three algorithms within QUIC-Fire that could use improvement: how flame tilt angle is accounted for, the incorporation of non-local drag effects, and the inclusion of the wake-eddy parameterizations that are used in QUIC-URB. Lastly, the third chapter describes a modification to the initial guess used for the calculation of the QUIC-URB mass-conserved wind solution during fire simulations. The modification is aimed at improving fire-fire interactions in QUIC-Fire simulations. The modification consists of using the solution from the previous timestep as the starting point for the calculation of the solution for the next timestep. Fire-fire interactions is greatly improved by the change but a new source of error is introduced. Due to how plumes are modelled in QUIC-Fire the new solution contains errors where gaps in the plume structure are present. However, these errors are mostly limited to the upper atmosphere, where they do not affect fire behavior at the surface, and their magnitude isn’t significant enough to discount the amount of new fire phenomenology now captured in QUIC-Fire with the change.

58 GEOSCIENCES↗

Electromagnetic Waves in Cold Plasmas

A collisionless fluid model of a plasma is presented and solved to first order in the cold, high frequency limit. From this model, a homogeneous system of linear equations is obtained for a harmonic electric field perturbation from which the Appleton-Hartree equation is derived. Moreover, the polarization states of electromagnetic waves in cold, collisionless plasmas are examined. Complete sets of equations governing the reflection and transmission of electromagnetic waves in both non magnetized and magnetized plasmas are obtained. In solving the boundary problem, the Booker quartic is derived and the nature of its roots discussed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Basic Research Needs in Quantum Computing and Networking

Employing quantum mechanical resources in computing, information processing, and networking opens the door to potential exponential advantages over classical counterparts. However, quantifying and realizing such advantages poses extensive scientific and engineering challenges. Department of Energy (DOE) investments have driven steady progress in addressing such challenges. Recently developed quantum algorithms offer asymptotic exponential advantages in speed or accuracy for fundamental scientific problems. These problems include simulating physical systems, solving systems of linear equations, differential equations, and optimization problems. Empirical demonstrations on nascent quantum hardware suggest better performance on contrived computational tasks than classical analogs. However, the requirements for a quantum computer or network to demonstrate an end-to-end rigorously quantifiable performance improvement over classical analogs remains a grand challenge, especially for problems of practical value. In particular, what will be required for quantum technology to ultimately exhibit scalable, rigorous, and transformative performance advantages for practical applications? In July 2023, DOE’s Advanced Scientific Computing Research program in the Office of Science convened the Workshop on Basic Research Needs in Quantum Computing and Networking, where major opportunities and grand challenges were identified. The following five priority research directions (PRDs) were identified as a result of the workshop.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗