Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “linear systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A Scalable Multigrid Reduction Framework for Multiphase Poromechanics of Heterogeneous Media

Simulation of multiphase poromechanics involves solving a multiphysics problem in which multiphase flow and transport are tightly coupled with the porous medium deformation. To capture this dynamic interplay, fully implicit methods, also known as monolithic approaches, are usually preferred. The main bottleneck of a monolithic approach is that it requires solution of large linear systems that result from the discretization and linearization of the governing balance equations. Because such systems are nonsymmetric, indefinite, and highly ill-conditioned, preconditioning is critical for fast convergence. Recently, most efforts in designing efficient preconditioners for multiphase poromechanics have been dominated by physics-based strategies. Current state-of-the-art “black-box” solvers such as algebraic multigrid (AMG) are ineffective because they cannot effectively capture the strong coupling between the mechanics and the flow subproblems, as well as the coupling inherent in the multiphase flow and transport process. In this work, we develop an algebraic framework based on multigrid reduction (MGR) that is suited for tightly coupled systems of PDEs. Using this framework, the decoupling between the equations is done algebraically through defining appropriate interpolation and restriction operators. One can then employ existing solvers for each of the decoupled blocks or design a new solver based on knowledge of the physics. We demonstrate the applicability of our framework when used as a “black-box” solver for multiphase poromechanics. Here, we show that the framework is flexible to accommodate a wide range of scenarios, as well as efficient and scalable for large problems.

97 MATHEMATICS AND COMPUTING↗

WeakIdent: Weak formulation for identifying differential equation using narrow-fit and trimming

Data-driven identification of differential equations is an interesting but challenging problem, especially when the given data are corrupted by noise. When the governing differential equation is a linear combination of various differential terms, the identification problem can be formulated as solving a linear system, with the feature matrix consisting of linear and nonlinear terms multiplied by a coefficient vector. This product is equal to the time derivative term, and thus generates dynamical behaviors. The goal is to identify the correct terms that form the equation to capture the dynamics of the given data. We propose a general and robust framework to recover differential equations using a weak formulation with two new mechanisms, narrow-fit and trimming, for both ordinary and partial differential equations (ODEs and PDEs). The weak formulation facilitates an efficient and robust way to handle noise, and two new mechanisms, narrow-fit and trimming, improve the coefficient support and value recoveries respectively. For each sparsity level, Subspace Pursuit is utilized to find an initial set of support from the large dictionary. Then, we focus on highly dynamic regions (rows of the feature matrix), and error normalize the feature matrix in the narrow-fit step. The support is further updated via trimming the terms that contribute the least. Finally, the support set of features with the smallest Cross-Validation error is chosen as the result. A comprehensive set of numerical experiments are presented for both systems of ODEs and PDEs with various noise levels. The proposed method gives a robust recovery of the coefficients, and a significant denoising effect which can handle up to 100% noise-to-signal ratio for some equations. We compare the proposed method with several state-of-the-art algorithms for the recovery of differential equations.

97 MATHEMATICS AND COMPUTING↗

Application of the locally self-consistent embedding approach to the Anderson model with non-uniform random distributions

Highlights: • Typical Medium Theory (TMT) for the Anderson Localization. • Locally Self-Consistent Multiple Scattering Method (LSMS) for Random Disordered Systems. • Linear Scaling Computational Method for Random Disordered Systems. We apply the recently developed embedding scheme for the locally self-consistent method to random disorder electrons systems. The method is based on the locally self-consistent multiple scattering theory and the typical medium theory. The locally self-consistent multiple scattering theory divides a system into many small designated local interaction zones. The subsystem within each local interaction zone is embedded in a self-consistent field from the typical medium theory. This approximation allows the study of random systems with large numbers of sites. We present results for the three dimensional Anderson model with different random disorder potential distributions. Using the typical density of states as an indicator of Anderson localization, we find that the method can capture the localization for commonly studied disorder potentials. These include the uniform distribution, the Gaussian distribution, and even the unbounded Cauchy distribution.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Differences in Tropical Rainfall in Aquaplanet Simulations With Resolved or Parameterized Deep Convection

Abstract This study investigates the effects of resolved deep convection on tropical rainfall and its multi‐scale variability. A series of aquaplanet simulations are analyzed using the Model for Prediction Across Scales‐Atmosphere with horizontal cell spacings from 120 to 3 km. The 3‐km experiment uses a novel configuration with 3‐km cell spacing between 20°S and 20°N and 15‐km cell spacing poleward of 30°N/S. A comparison of those experiments shows that resolved deep convection yields a narrower, stronger, and more equatorward intertropical convergence zone, which is supported by stronger nonlinear horizontal momentum advection in the boundary layer. There is also twice as much tropical rainfall variance in the experiment with resolved deep convection than in the experiments with parameterized convection. All experiments show comparable precipitation variance associated with Kelvin waves; however, the experiment with resolved deep convection shows higher precipitation variance associated with westward propagating systems. Resolved deep convection also yields at least two orders of magnitude more frequent heavy rainfall rates (>2 mm hr −1 ) than the experiments with parameterized convection. A comparison of organized precipitation systems demonstrates that tropical convection organizes into linear systems that are associated with stronger and deeper cold pools and upgradient convective momentum fluxes when convection is resolved. In contrast, parameterized convection results in more circular systems, weaker cold pools, and downgradient convective momentum fluxes. These results suggest that simulations with parameterized convection are missing an important feedback loop between the mean state, convective organization, and meridional gradients of moisture and momentum.

Rios‐Berrios, Rosimar↗

Many-body reactive force field development for carbon condensation in C/O systems under extreme conditions

In this paper, we describe the development of a reactive force field for C/O systems under extreme temperatures and pressures, based on the many-body Chebyshev Interaction Model for Efficient Simulation (ChIMES). The resulting model, which targets carbon condensation under thermodynamic conditions of 6500 K and 2.5 g cm –3 , affords a balance between model accuracy, complexity, and training set generation expense. We show that the model recovers much of the accuracy of density functional theory for the prediction of structure, dynamics, and chemistry when applied to dissociative condensed phase systems at 1:1 and 1:2 C:O ratios, as well as molten carbon. Our C/O modeling approach exhibits a 104 increase in efficiency for the same system size (i.e., 128 atoms) and a linear system size scalability over standard quantum molecular dynamics methods, allowing the simulation of significantly larger systems than previously possible. We find that the model captures the condensed-phase reaction-coupled formation of carbon clusters implied by recent experiments, and that this process is susceptible to strong finite size effects. Overall, we find the present ChIMES model to be well suited for studying chemical processes and cluster formation at pressures and temperatures typical of shock waves. We expect that the present C/O modeling paradigm can serve as a template for the development of a broader high pressure–high temperature force-field for condensed phase chemistry in organic materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Low Precision and Efficient Programming Languages for Sustainable AI: Final Report for the Summer Project of 2024

This document contains all relevant material generated during the authors' summer internship at NREL in 2024. This report shows how to improve energy efficiency of a few code samples by using low-precision data types combined with mixed-precision algorithms. The main applications considered here are (i) linear system solvers using mixed precision, and (ii) neural networks using mixed precision. This report also discusses how programming languages affect energy consumption of algorithms, energy metrics for a code and tools, and the available current software and hardware infrastructure.

97 MATHEMATICS AND COMPUTING↗

Compressed basis GMRES on high-performance graphics processing units

Krylov methods provide a fast and highly parallel numerical tool for the iterative solution of many large-scale sparse linear systems. To a large extent, the performance of practical realizations of these methods is constrained by the communication bandwidth in current computer architectures, motivating the investigation of sophisticated techniques to avoid, reduce, and/or hide the message-passing costs (in distributed platforms) and the memory accesses (in all architectures). This article leverages Ginkgo’s memory accessor in order to integrate a communication-reduction strategy into the (Krylov) GMRES solver that decouples the storage format (i.e., the data representation in memory) of the orthogonal basis from the arithmetic precision that is employed during the operations with that basis. Given that the execution time of the GMRES solver is largely determined by the memory accesses, the cost of the datatype transforms can be mostly hidden, resulting in the acceleration of the iterative step via a decrease in the volume of bits being retrieved from memory. Together with the special properties of the orthonormal basis (whose elements are all bounded by 1), this paves the road toward the aggressive customization of the storage format, which includes some floating-point as well as fixed-point formats with mild impact on the convergence of the iterative process. We develop a high-performance implementation of the “compressed basis GMRES” solver in the Ginkgo sparse linear algebra library using a large set of test problems from the SuiteSparse Matrix Collection. We demonstrate robustness and performance advantages on a modern NVIDIA V100 graphics processing unit (GPU) of up to 50% over the standard GMRES solver that stores all data in IEEE double-precision.

97 MATHEMATICS AND COMPUTING↗

Combining Sparse Approximate Factorizations with Mixed-precision Iterative Refinement

The standard LU factorization-based solution process for linear systems can be enhanced in speed or accuracy by employing mixed-precision iterative refinement. Most recent work has focused on dense systems. We investigate the potential of mixed-precision iterative refinement to enhance methods for sparse systems based on approximate sparse factorizations. In doing so, we first develop a new error analysis for LU- and GMRES-based iterative refinement under a general model of LU factorization that accounts for the approximation methods typically used by modern sparse solvers, such as low-rank approximations or relaxed pivoting strategies. We then provide a detailed performance analysis of both the execution time and memory consumption of different algorithms, based on a selected set of iterative refinement variants and approximate sparse factorizations. Our performance study uses the multifrontal solver MUMPS, which can exploit block low-rank factorization and static pivoting. We evaluate the performance of the algorithms on large, sparse problems coming from a variety of real-life and industrial applications showing that mixed-precision iterative refinement combined with approximate sparse factorization can lead to considerable reductions of both the time and memory consumption.

97 MATHEMATICS AND COMPUTING↗

Customizable wave tailoring nonlinear materials enabled by bilevel inverse design

Abstract Passive wave transformation via nonlinearity is ubiquitous in settings from acoustics to optics and electromagnetics. It is well known that different nonlinearities yield different effects on propagating signals, which raises the question of “what precise nonlinearity is the best for a given wave tailoring application?” In this work, considering a one-dimensional spring-mass chain connected by polynomial springs (a variant of the Fermi-Pasta-Ulam-Tsingou system), we introduce a bilevel inverse design method which couples the shape optimization of structures for tailored constitutive responses with reduced-order nonlinear dynamical inverse design. We apply it to two qualitatively distinct problems—minimization of peak transmitted kinetic energy from impact, and pulse shape transformation—demonstrating our method’s breadth of applicability. For the impact problem, we obtain two fundamental insights. First, small differences in nonlinearity can drastically change the dynamic response of the system, from severely under- to outperforming a comparative linear system. Second, the oft-used strategy of impact mitigation via “energy locking” bistability can be significantly outperformed by our optimal nonlinearity. We validate this case with impact experiments and find excellent agreement. This study establishes a framework for broader passive nonlinear mechanical wave tailoring material design, with applications to computing, signal processing, shock mitigation, and autonomous materials.

Science & Technology - Other Topics↗

Extremum seeking for optimal control problems with unknown time-varying systems and unknown objective functions

We consider the problem of optimal feedback control of an unknown, noisy, time-varying, dynamic system that is initialized repeatedly. Examples include a robotic manipulator which must perform the same motion, such as assisting a human, repeatedly and accelerating cavities in particle accelerators which are turned on for a fraction of a second with given initial conditions and vary slowly due to temperature fluctuations. In this paper, we present an approach that applies to systems of practical interest. The method presented here is model independent; does not require knowledge of the objective function; is robust to measurement noise; is applicable for any set of initial conditions; is applicable to simultaneously controlling an arbitrary number of parameters; and may be implemented with a broad range of continuous or discontinuous functions such as sine or square waves. For systems with convex cost functions we prove that our algorithm will produce controllers that approach the minimal cost. For linear systems we reproduce the cost minimizing linear quadratic regulator optimal controller that could have been designed analytically had the system and cost function been known. We demonstrate the effectiveness of the algorithm with simulation studies of noisy and time-varying systems.

42 ENGINEERING↗

A supernodal all-pairs shortest path algorithm

We show how to exploit graph sparsity in the Floyd-Warshall algorithm for the all-pairs shortest path (Apsp) problem. Floyd-Warshall is an attractive choice for Apsp on high-performing systems due to its structural similarity to solving dense linear systems and matrix multiplication. However, if sparsity of the input graph is not properly exploited, Floyd-Warshall will perform unnecessary asymptotic work and thus may not be a suitable choice for many input graphs. To overcome this limitation, the key idea in our approach is to use the known algebraic relationship between Floyd-Warshall and Gaussian elimination, and import several algorithmic techniques from sparse Cholesky factorization, namely, fill-in reducing ordering, symbolic analysis, supernodal traversal, and elimination tree parallelism. When combined, these techniques reduce computation, improve locality and enhance parallelism. We implement these ideas in an efficient shared memory parallel prototype that is orders of magnitude faster than an efficient multi-threaded baseline Floyd-Warshall that does not exploit sparsity. Our experiments suggest that the Floyd-Warshall algorithm can compete with Dijkstra's algorithm (the algorithmic core of Johnson's algorithm) for several classes sparse graphs.

Sao, Piyush↗

Linear embedding of nonlinear dynamical systems and prospects for efficient quantum algorithms

The simulation of large nonlinear dynamical systems, including systems generated by discretization of hyperbolic partial differential equations, can be computationally demanding. Such systems are important in both fluid and kinetic computational plasma physics. This motivates exploring whether a future error-corrected quantum computer could perform these simulations more efficiently than any classical computer. In this work, we describe a method for mapping any finite nonlinear dynamical system to an infinite linear dynamical system (embedding) and detail three specific cases of this method that correspond to previously studied mappings. Then we explore an approach for approximating the resulting infinite linear system with finite linear systems (truncation). Using a number of qubits only logarithmic in the number of variables of the nonlinear system, a quantum computer could simulate truncated systems to approximate output quantities if the nonlinearity is sufficiently weak. Other aspects of the computational efficiency of the three detailed embedding strategies are also discussed.

97 MATHEMATICS AND COMPUTING↗

A low-rank solver for the stochastic unsteady Navier–Stokes problem

Here we study a low-rank iterative solver for the unsteady Navier–Stokes equations for incompressible flows with a stochastic viscosity. The equations are discretized using the stochastic Galerkin method, and we consider an all-at-once formulation where the algebraic systems at all the time steps are collected and solved simultaneously. The problem is linearized with Picard’s method. To efficiently solve the linear systems at each step, we use low-rank tensor representations within the Krylov subspace method, which leads to significant reductions in storage requirements and computational costs. Combined with effective mean-based preconditioners and the idea of inexact solve, we show that only a small number of linear iterations are needed at each Picard step. The proposed algorithm is tested with a model of flow in a two-dimensional symmetric step domain with different settings to demonstrate the computational efficiency.

97 MATHEMATICS AND COMPUTING↗

Juqbox.jl

The Juqbox.jl package implements functionality for solving the quantum optimal control problem for realizing logical gates in closed quantum systems. The dynamics of the quantum system is modeled by Schroedinger's equation, which takes to form of a linear system of ordinary differential equations (ODE). Juqbox.jl solves this ODE by numerical time stepping and applies a gradient-based optimization technique to determine control pulses for driving an initial state to a final state, according to the desired logical gate transformation. To evaluate the gradient, Juqbox.jl applies the ``first discretize, then optimize'' approach based on a discrete adjoint time stepping technique. The actual optimization is performed by the open source Ipopt libaray. Juqbox.jl is written in the Julia programming language which, among many other features, provides a convenient interface to the Ipopt library.

PETERSSON, NILSA.↗

Anderson acceleration with approximate calculations: Applications to scientific computing

Here we provide rigorous theoretical bounds for Anderson acceleration (AA) that allow for approximate calculations when applied to solve linear problems. We show that, when the approximate calculations satisfy the provided error bounds, the convergence of AA is maintained while the computational time could be reduced. We also provide computable heuristic quantities, guided by the theoretical error bounds, which can be used to automate the tuning of accuracy while performing approximate calculations. For linear problems, the use of heuristics to monitor the error introduced by approximate calculations, combined with the check on monotonicity of the residual, ensures the convergence of the numerical scheme within a prescribed residual tolerance. Motivated by the theoretical studies, we propose a reduced variant of AA, which consists in projecting the least-squares used to compute the Anderson mixing onto a subspace of reduced dimension. The dimensionality of this subspace adapts dynamically at each iteration as prescribed by the computable heuristic quantities. We numerically show and assess the performance of AA with approximate calculations on: (i) linear deterministic fixed-point iterations arising from the Richardson's scheme to solve linear systems with open-source benchmark matrices with various preconditioners and (ii) non-linear deterministic fixed-point iterations arising from non-linear time-dependent Boltzmann equations.

97 MATHEMATICS AND COMPUTING↗

Toward a scalable robust security-constrained optimal power flow using a proximal projection bundle method

Robust security-constrained optimal power flow (rSCOPF) aims to find the worst-case contingencies of alternating current optimal power flow (ACOPF) in power systems. With the rise of GPU architectures on the upcoming supercomputer architectures, optimization algorithms that rely on sparse linear algebra and indefinite linear systems are becoming increasingly hard to solve efficiently (e.g. interior-point method). To address this we revisit a maximin optimization formulation of the rSCOPF and the single-level mixed-integer semidefinite programming (MISDP) reformulation, which is obtained by taking the Lagrangian relaxation of the inner minimization ACOPF problem. In this paper, we focus on the development of a proximal projection bundle method (PPBM) for solving continuous relaxation node subproblems of the MISDP problem, based primarily on the well-known alternating direction method of multipliers. Cutting planes reminiscent of bundle method ideas are also applied in coordination with updates of the proximal parameter. The cutting-plane method can generate a large number of linear inequalities, leading to a large scale but decomposable quadratic programming (QP) subproblem that is amenable to GPUs. We present the numerical results on the IEEE 30, 57, 118, and 300-bus systems by using our PBMM method. We discuss the main computational bottleneck of our method, which is the time taken to solve each iteration of a QP subproblem instance of the PPBM, and how GPU architectures can accelerate this solution process.

bundle method↗

QUIC-URB and QUIC-fire extension to complex terrain: Development of a terrain-following coordinate system

Ensemble-based approaches to prescribed fire planning cannot be supported by CFD-based models like FIRETEC and WFDS because they are too computationally expensive and cannot leverage LES approaches like CAWFE and WRF-SFIRE because too coarse of resolution. QUIC-Fire was developed to fill this gap but it cannot currently address complex terrain, typical for instance of the Western United States. In this paper, we describe the extension of the diagnostic wind model QUIC-URB, the wind engine of QUIC-Fire, to a terrain-following coordinate system. In particular, the paper presents the mathematical derivation of the wind solver leading to a linear system of equations that are solved through the successive over-relaxation method. The model is validated against a standard test used in previous works (the Askervein Hill) and against a new dataset from measurements in the Socorro Mountains, New Mexico. The terrain-following implementation captured the correct phenomenology for the isolated Askervein Hill, with a wind speed up at the top of the hill. We report the model agreed well with measurements on the upwind side of the peak, but overestimated speed-up on the downwind side of the hill. This is due to the inability of the model to generate flow separation and wake-eddy dynamics. On a common laptop, the divergence-free wind field was obtained in 6 s, making the solver appealing for coupled fire–atmosphere simulations. The Socorro Mountain was highly complex, with many cliff faces, peaks, and valleys. Although the model captures the magnitude and direction of inlet and outlet areas of the domain, it performs rather poorly in the valley region and in the regions near the steep cliffs. Hence, the model shows good agreement with data in areas of open sloped terrain but lacks in areas where flow separation and thermally driven effects may be present (neither effect was addressed in this work). Results highlight that future work should focus on the implementation of parameterizations of wake-eddies, similar to QUIC-URB’s building parameterizations, and on thermodynamic-driven flow.

54 ENVIRONMENTAL SCIENCES↗

Xyce(™) Parallel Electronic Simulator v.7.5

The Xyce Parallel Electronic Simulator simulates electronic circuit behavior in DC, AC, HB, MPDE and transient mode using standard analog (DAE) and/or device (PDE) device models including several age and radiation aware devices. It supports a variety of computing platforms (both serial and parallel) computers. Lastly, it uses a variety of modern solution algorithms dynamic parallel load-balancing and iterative solvers.! ! Xyce is primarily used to simulate the voltage and current behavior of a circuit network (a network of electronic devices connected via a conductive network). As a tool, it is mainly used for the design and analysis of electronic circuits.! ! Kirchoff's conservation laws are enforced over a network using modified nodal analysis. This results in a set of differential algebraic equations (DAEs). The resulting nonlinear problem is solved iteratively using a fully coupled Newton method, which in turn results in a linear system that is solved by either a standard sparse-direct solver or iteratively using Trilinos linear solver packages, also developed at Sandia National Laboratories.

Source record↗