Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “linear equation systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Probabilistic error estimation for non-intrusive reduced models learned from data of systems governed by linear parabolic partial differential equations

This work derives a residual-based a posteriori error estimator for reduced models learned with non-intrusive model reduction from data of high-dimensional systems governed by linear parabolic partial differential equations with control inputs. It is shown that quantities that are necessary for the error estimator can be either obtained exactly as the solutions of least-squares problems in a non-intrusive way from data such as initial conditions, control inputs, and high-dimensional solution trajectories or bounded in a probabilistic sense. Here, the computational procedure follows an offline/online decomposition. In the offline (training) phase, the high-dimensional system is judiciously solved in a black-box fashion to generate data and to set up the error estimator. In the online phase, the estimator is used to bound the error of the reduced-model predictions for new initial conditions and new control inputs without recourse to the high-dimensional system. Numerical results demonstrate the workflow of the proposed approach from data to reduced models to certified predictions.

97 MATHEMATICS AND COMPUTING↗

Charged particle motion in spherically symmetric distributions of magnetic monopoles

The classical equations of motion of a charged particle in a spherically symmetric distribution of magnetic monopoles can be transformed into a system of linear equations, thereby providing a type of integrability. In the case of a single monopole, the solution was given long ago by Poincaré. In the case of a uniform distribution of monopoles, the solution can be expressed in terms of parabolic cylinder functions (essentially the eigenfunctions of an inverted harmonic oscillator). Further, this solution is relevant to recent studies of nonassociative star products, symplectic lifts of twisted Poisson structures, and fluids and plasmas of electric and magnetic charges.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Multirate linearly-implicit GARK schemes

Many complex applications require the solution of initial-value problems where some components change fast, while others vary slowly. Multirate schemes apply different step sizes to resolve different components of the system, according to their dynamics, in order to achieve increased computational efficiency. The stiff components of the system, fast or slow, are best discretized with implicit base methods in order to ensure numerical stability. To this end, linearly implicit methods are particularly attractive as they solve only linear systems of equations at each step. This paper develops the Multirate GARK-ROS/ROW (MR-GARK-ROS/ROW) framework for linearly-implicit multirate time integration. The order conditions theory considers both exact and approximative Jacobians. The effectiveness of implicit multirate methods depends on the coupling between the slow and fast computations; an array of efficient coupling strategies and the resulting numerical schemes are analyzed. Multirate infinitesimal step linearly-implicit methods, that allow arbitrarily small micro-steps and offer extreme computational flexibility, are constructed. The new unifying framework includes existing multirate Rosenbrock(-W) methods as particular cases, and opens the possibility to develop new classes of highly effective linearly implicit multirate integrators.

97 MATHEMATICS AND COMPUTING↗

Assessing the Feasibility of Bordered Block Diagonal Reordering in Power System Matrices using Fully Convolutional Network

In electromagnetic transient (EMT) simulations for power systems and inverter-based resources (IBRs), the arrangement of states within the system's linear equations, represented by matrix A in Ax=b, is critical. The state ordering in matrix A can highlight distinct characteristics of the system's graph, and identifying an optimal state ordering is crucial for efficient computation. The choice of state ordering, however, is dependent on the solver used, as each solver may perform optimally with different matrix patterns. With a wide array of matrix reordering algorithms available, selecting the most suitable one becomes challenging without insights into the matrix's ideal configuration. To address this, the paper proposes a fully convolutional network (FCN) to evaluate the reordering potential of the A matrix into a bordered block diagonal (BBD) pattern, which is commonly observed in power system and IBR modeling. The FCN's assessment aims to streamline the solver's operation, which in turn could substantially reduce the computational time required to find a solution.

Xia, Qianxue↗

Report for the ASCR Workshop on Basic Research Needs in Quantum Computing and Networking - 2023

Employing quantum mechanical resources in computing and networking opens the door to new computation and communication models and potential disruptive advantages over classical counterparts. However, quantifying and realizing such advantages face extensive scientific and engineering challenges. Investments by the Department of Energy (DOE) have driven progress toward addressing such challenges. Quantum algorithms have been recently developed, in some cases offering asymptotic exponential advantages in speed or accuracy, for fundamental scientific problems such as simulating physical systems, solving systems of linear equations, or solving differential equations. Empirical demonstrations on nascent quantum hardware suggest better performance than classical analogs on specialized computational tasks favorable to the quantum computing systems. However, demonstration of an end-to-end, substantial and rigorously quantifiable quantum performance advantage over classical analogs remains a grand challenge, especially for problems of practical value. The definition of requirements for quantum technologies to exhibit scalable, rigorous, and transformative performance advantages for practical applications also remains an outstanding open question, namely, what will be required to ultimately demonstrate practical quantum advantage?

97 MATHEMATICS AND COMPUTING↗

Report for the ASCR Workshop on Basic Research Needs in Quantum Computing and Networking - 2023

Employing quantum mechanical resources in computing and networking opens the door to new computation and communication models and potential disruptive advantages over classical counterparts. However, quantifying and realizing such advantages face extensive scientific and engineering challenges. Investments by the Department of Energy (DOE) have driven progress toward addressing such challenges. Quantum algorithms have been recently developed, in some cases offering asymptotic exponential advantages in speed or accuracy, for fundamental scientific problems such as simulating physical systems, solving systems of linear equations, or solving differential equations. Empirical demonstrations on nascent quantum hardware suggest better performance than classical analogs on specialized computational tasks favorable to the quantum computing systems. However, demonstration of an end-to-end, substantial and rigorously quantifiable quantum performance advantage over classical analogs remains a grand challenge, especially for problems of practical value. The definition of requirements for quantum technologies to exhibit scalable, rigorous, and transformative performance advantages for practical applications also remains an outstanding open question, namely, what will be required to ultimately demonstrate practical quantum advantage?

97 MATHEMATICS AND COMPUTING↗

MAPPRAISER: A massively parallel map-making framework for multi-kilo pixel CMB experiments

Forthcoming cosmic microwave background (CMB) polarized anisotropy experiments have the potential to revolutionize our understanding of the Universe and fundamental physics. The sought-after, tale-telling signatures will be however distributed over voluminous data sets which these experiments will collect. These data sets will need to be efficiently processed and unwanted contributions due to astrophysical, environmental, and instrumental effects characterized and efficiently mitigated in order to uncover the signatures. This poses a significant challenge to data analysis methods, techniques, and software tools which will not only have to be able to cope with huge volumes of data but to do so with unprecedented precision driven by the demanding science goals posed for the new experiments. A keystone of efficient CMB data analysis is solvers of very large linear systems of equations. Such systems appear in very diverse contexts throughout CMB data analysis pipelines, however they typically display similar algebraic structures and can therefore be solved using similar numerical techniques. Linear systems arising in the so-called map-making problem are one of the most prominent and common ones. In this work we present a massively parallel, flexible and extensible framework, comprised of a numerical library, MIDAPACK, and a high level code, MAPPRAISER, which provide tools for solving efficiently such systems. Here, the framework implements iterative solvers based on conjugate gradient techniques: enlarged and preconditioned using different preconditioners. We demonstrate the framework on simulated examples reflecting basic characteristics of the forthcoming data sets issued by ground-based and satellite-borne instruments, executing it on as many as 16,384 compute cores. The software is developed as an open source project freely available to the community at: https://github.com/B3Dcmb/midapack.

79 ASTRONOMY AND ASTROPHYSICS↗

HyKKT

HyKKT (pronounced as "hiked") is a package for solving systems of linear equations of Karush-Kuhn-Tucker (KKT) form, which typically arise in optimization problems, such as optimal power flow analysis. HyKKT uses Cholesky instead of LDL^T factorization and solves the general KKT system to a desired numerical precision via block reduction and conjugate gradient on the Schur complement. Such implementation is more suitable for implementation on graphic processing units (GPUs).

Regev, Shaked↗

Batched Sparse Linear Algebra (Final Report for Subcontract B648960)

This report finalizes design specifications for developing batched kernels for small tensor operations for unassembled matrix-free iterative solvers, batched solvers for partially assembled operators, and batched solvers with support for various sparse formats. The outcome of the project milestones is a set of interfaces to Batched Sparse LA solvers running on hardware accelerators for use in ECP Libraries and Applications. It is part of the development of sparse batched kernels, solvers/preconditioners as well as creating interoperability in xSDK libraries with sparse and dense batched functions to benefit ECP applications. The participants included representatives from ECP libraries (not limited to the xSDK project), applications, and vendors (AMD, Intel, and NVIDIA). Batched sparse linear algebra solvers form the new frontier for algorithmic development and performance engineering. Many applications (ECP and non-ECP alike) require simultaneous solutions of small linear systems of equations that are structurally sparse. To move towards high hardware utilization, it is important to provide these applications with appropriate interfaces to efficient batched sparse solvers running on modern hardware accelerators. We present interface designs in use by HPC software libraries supporting batched sparse linear algebra and the development of sparse batched kernel codes for solvers and preconditioners. We also address the potential interoperability opportunities to keep the software portable between the major hardware accelerators from AMD, Intel, and NVIDIA. The presented interface specifications includes batched band, sparse iterative, and sparse direct solvers. This report summarizes progress in Kokkos Kernels and the xSDK libraries MAGMA, Ginkgo, hypre, SUNDIALS, and SuperLU_dist.

97 MATHEMATICS AND COMPUTING↗

Early Exploration of a Flexible Framework for Efficient Quantum Linear Solvers in Power Systems

The rapid integration of renewable energy resources presents formidable challenges in managing power grids. While advanced computing and machine learning techniques offer some solutions for accelerating grid modeling and simulation, there remain complex problems that classical computers cannot effectively address. Quantum computing, a promising technology, has the potential to fundamentally transform how we manage power systems, especially in scenarios with a higher proportion of renewable energy sources. One critical aspect is solving linear systems of equations, crucial for power system applications like power flow analysis, for which the Harrow-Hassidim-Lloyd (HHL) algorithm is a well-known quantum solution. However, HHL quantum circuits often exhibit excessive depth, making them impractical for current Noisy-Intermediate-Scale-Quantum (NISQ) devices. In this paper, we introduce a versatile framework, powered by NWQSim, that bridges the gap between power system applications and quantum linear solvers available in Qiskit. This framework empowers researchers to efficiently explore power system applications using quantum linear solvers. Through innovative gate fusion strategies, reduced circuit depth, and GPU acceleration, our simulator significantly enhances resource efficiency. Power flow case studies have demonstrated up to a eight-fold speedup compared to Qiskit Aer, all while maintaining comparable levels of accuracy.

quantum computing, Harrow-Hassidim-Lloyd, high-per↗

Milestone 49 Report: Batched Sparse LA Phase 5 Implementation

Batched sparse linear algebra operations in general, and solvers in particular, have become the major algorithmic development activity and foremost performance engineering effort in the numerical software libraries work on modern hardware with accelerators such as GPUs. Many applications, ECP and non-ECP alike, require simultaneous solutions of many small linear systems of equations that are structurally sparse in one form or another. In order to move towards high hardware utilization levels, it is important to provide these applications with appropriate interface designs to be both functionally efficient and performance portable and give full access to the appropriate batched sparse solvers running on modern hardware accelerators prevalent across DOE supercomputing sites since the inception of ECP. To this end, we present here a summary of recent advances on the interface designs in use by HPC software libraries supporting batched sparse linear algebra and the development of sparse batched kernel codes for solvers and preconditioners. We also address the potential interoperability opportunities to keep the corresponding software portable between the major hardware accelerators from AMD, Intel, and NVIDIA, while maintaining the appropriate disclosure levels conforming to the active NDA agreements. The presented interface specifications include a mix of batched band, sparse iterative, and sparse direct solvers with their accompanying functionality that is already required by the application codes or we anticipated to be needed in the near future. This report summarizes progress in Kokkos Kernels and the xSDK libraries MAGMA, Ginkgo, hypre, PETSc, and SuperLU.

97 MATHEMATICS AND COMPUTING↗

Massively parallel modeling and inversion of electrical resistivity tomography data using PFLOTRAN

Abstract. Electrical resistivity tomography (ERT) is a broadly accepted geophysical method for subsurface investigations. Interpretation of field ERT data usually requires the application of computationally intensive forward modeling and inversion algorithms. For large-scale ERT data, the efficiency of these algorithms depends on the robustness, accuracy, and scalability on high-performance computing resources. In this regard, we present a robust and highly scalable implementation of forward modeling and inversion algorithms for ERT data. The implementation is publicly available and developed within the framework of PFLOTRAN, an open-source, state-of-the-art massively parallel subsurface flow and transport simulation code. The forward modeling is based on a finite-volume discretization of the governing differential equations, and the inversion uses a Gauss–Newton optimization scheme. To evaluate the accuracy of the forward modeling, two examples are first presented by considering layered (1D) and 3D earth conductivity models. The computed numerical results show good agreement with the analytical solutions for the layered earth model and results from a well-established code for the 3D model. Inversion of ERT data, simulated for a 3D model, is then performed to demonstrate the inversion capability by recovering the conductivity of the model. To demonstrate the parallel performance of PFLOTRAN's ERT process model and inversion capabilities, large-scale scalability tests are performed by using up to 131 072 processes on a leadership class supercomputer. These tests are performed for the two most computationally intensive steps of the ERT inversion: forward modeling and Jacobian computation. For the forward modeling, we consider models with up to 122 ×106 degrees of freedom (DOFs) in the resulting system of linear equations and demonstrate that the code exhibits almost linear scalability on up to 10 000 DOFs per process. On the other hand, the code shows superlinear scalability for the Jacobian computation, mainly because all computations are fairly evenly distributed over each process with no parallel communication.

58 GEOSCIENCES↗

Toward efficient polynomial preconditioning for GMRES

Here, we present a polynomial preconditioner for solving large systems of linear equations. The polynomial is derived from the minimum residual polynomial (the GMRES polynomial) and is more straightforward to compute and implement than many previous polynomial preconditioners. Our current implementation of this polynomial using its roots is naturally more stable than previous methods of computing the same polynomial. We implement further stability control using added roots, and this allows for high degree polynomials. We discuss the effectiveness and challenges of root-adding and give an additional check for stability. In this article, we study the polynomial preconditioner applied to GMRES; however it could be used with any Krylov solver. This polynomial preconditioning algorithm can dramatically improve convergence for some problems, especially for difficult problems, and can reduce dot products by an even greater margin.

97 MATHEMATICS AND COMPUTING↗

A smeared crack modeling framework accommodating multi-directional fracture at finite strains

A generic smeared crack modeling framework predicated on the deformation gradient decomposition (DGD) approach is proposed for use in dynamic fracture problems at finite strains, accommodating failure along multiple mutually orthogonal fracture planes embedded within an independently defined bulk material model. Within this constitutive framework, the traction equilibrium conditions imposed at each failure surface are used to determine the associated crack displacements stored as internal state variables. In general, the enforcement of interfacial equilibrium entails the implicit solution of a non-linear system of equations within the constitutive update procedure. However, if inertial effects arising due to the relative motion of the fractured material are incorporated within the model, the traction equilibrium conditions are shown to give rise to corresponding dynamic equations of motion governing the time-evolution of the crack opening displacements. For dynamic problems, an explicit time-integration procedure is devised to efficiently update the material state, subject to a set of internal frictionless contact constraints to prevent material inversion. Finally, the efficacy of the proposed modeling framework is investigated through several benchmark dynamic fracture problems run within the explicit finite element code DYNA3D.

42 ENGINEERING↗

Leveraging operator learning to accelerate convergence of the preconditioned conjugate gradient method

We propose a new deflation strategy to accelerate the convergence of the preconditioned conjugate gradient (PCG) method for solving parametric large-scale linear systems of equations. Unlike traditional deflation techniques that rely on eigenvector approximations or recycled Krylov subspaces, we generate the deflation subspaces using operator learning, specifically the Deep Operator Network (DeepONet). To this aim, we introduce two complementary approaches for assembling the deflation operators. The first approach approximates near-null space vectors of the discrete PDE operator using the basis functions learned by the DeepONet. The second approach directly leverages solutions predicted by the DeepONet. To further enhance convergence, we also propose several strategies for prescribing the sparsity pattern of the deflation operator. Here, a comprehensive set of numerical experiments encompassing steady-state, time-dependent, scalar, and vector-valued problems posed on both structured and unstructured geometries is presented and demonstrates the effectiveness of the proposed DeepONet-based deflated PCG method, as well as its generalization across a wide range of model parameters and problem resolutions.

Deflation↗

Simulation toolkit for digital material characterization of large image-based microstructures

In this paper, an efficient image-based simulation toolkit for material characterization is presented, which is scalable to work from personal computers to workstations. The effective thermal conductivity, elasticity, and permeability are evaluated employing a computational homogenization framework based on the Finite Element Method (FEM). Two complementary open-source packages are presented: one developed in Python, which can convert digital images into voxel meshes (pyTomoviewer); the other developed in Julia, that can run numerical simulations to compute effective material properties (chpack). Also, a CUDA C version of chpack is provided (chfem_gpu). They were designed to deal with large multi-phase models, so strategies were devised to minimize their memory footprint, while avoiding a high toll on execution time. The voxel-based approach significantly simplifies the FEM meshes and allows efficient matrix-free implementations. In that sense, to handle large linear systems of equations, the element-by-element (EBE) technique is adopted, in conjunction with a low-memory implementation of the Preconditioned Conjugate Gradient (PCG) method. Finally, the code was thoroughly tested on an artificial geometry made of a square array of cylinders, for which analytical solutions exist, as well as on a real micro-tomographic reconstruction of FiberForm TM , a carbon preform commonly used in thermal protection systems.

36 MATERIALS SCIENCE↗

QUIC-URB and QUIC-fire extension to complex terrain: Development of a terrain-following coordinate system

Ensemble-based approaches to prescribed fire planning cannot be supported by CFD-based models like FIRETEC and WFDS because they are too computationally expensive and cannot leverage LES approaches like CAWFE and WRF-SFIRE because too coarse of resolution. QUIC-Fire was developed to fill this gap but it cannot currently address complex terrain, typical for instance of the Western United States. In this paper, we describe the extension of the diagnostic wind model QUIC-URB, the wind engine of QUIC-Fire, to a terrain-following coordinate system. In particular, the paper presents the mathematical derivation of the wind solver leading to a linear system of equations that are solved through the successive over-relaxation method. The model is validated against a standard test used in previous works (the Askervein Hill) and against a new dataset from measurements in the Socorro Mountains, New Mexico. The terrain-following implementation captured the correct phenomenology for the isolated Askervein Hill, with a wind speed up at the top of the hill. We report the model agreed well with measurements on the upwind side of the peak, but overestimated speed-up on the downwind side of the hill. This is due to the inability of the model to generate flow separation and wake-eddy dynamics. On a common laptop, the divergence-free wind field was obtained in 6 s, making the solver appealing for coupled fire–atmosphere simulations. The Socorro Mountain was highly complex, with many cliff faces, peaks, and valleys. Although the model captures the magnitude and direction of inlet and outlet areas of the domain, it performs rather poorly in the valley region and in the regions near the steep cliffs. Hence, the model shows good agreement with data in areas of open sloped terrain but lacks in areas where flow separation and thermally driven effects may be present (neither effect was addressed in this work). Results highlight that future work should focus on the implementation of parameterizations of wake-eddies, similar to QUIC-URB’s building parameterizations, and on thermodynamic-driven flow.

54 ENVIRONMENTAL SCIENCES↗

Accelerating eigenvalue computation for nuclear structure calculations via perturbative corrections

Subspace projection methods utilizing perturbative corrections have been proposed for computing the lowest few eigenvalues and corresponding eigenvectors of large Hamiltonian matrices. In this paper, we build upon these methods and introduce the term Subspace Projection with Perturbative Corrections (SPPC) method to refer to this approach. We tailor the SPPC for nuclear many-body Hamiltonians represented in a truncated configuration interaction subspace, i.e., the no-core shell model (NCSM). We use the hierarchical structure of the NCSM Hamiltonian to partition the Hamiltonian as the sum of two matrices. The first matrix corresponds to the Hamiltonian represented in a small configuration space, whereas the second is viewed as the perturbation to the first matrix. Eigenvalues and eigenvectors of the first matrix can be computed efficiently. Because of the split, perturbative corrections to the eigenvectors of the first matrix can be obtained efficiently from the solutions of a sequence of linear systems of equations defined in the small configuration space. These correction vectors can be combined with the approximate eigenvectors of the first matrix to construct a subspace from which more accurate approximations of the desired eigenpairs can be obtained. We show by numerical examples that the SPPC method can be more efficient than conventional iterative methods for solving large-scale eigenvalue problems such as the Lanczos, block Lanczos and the locally optimal block preconditioned conjugate gradient (LOBPCG) method. The method can also be combined with other methods to avoid convergence stagnation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗