Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “linear systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

GPU-resident sparse direct linear solvers for alternating current optimal power flow analysis

Integrating renewable resources within the transmission grid at a wide scale poses significant challenges for economic dispatch as it requires analysis with more optimization parameters, constraints, and sources of uncertainty. This motivates the investigation of more efficient computational methods, especially those for solving the underlying linear systems, which typically take more than half of the overall computation time. In this paper, we present our work on sparse linear solvers that take advantage of hardware accelerators, such as graphical processing units (GPUs), and improve the overall performance when used within economic dispatch computations. We treat the problems as sparse, which allows for faster execution but also makes the implementation of numerical methods more challenging. We present the first GPU-native sparse direct solver that can execute on both AMD and NVIDIA GPUs. We demonstrate significant performance improvements when using high-performance linear solvers within alternating current optimal power flow (ACOPF) analysis. Furthermore, we demonstrate the feasibility of getting significant performance improvements by executing the entire computation on GPU-based hardware. Finally, we identify outstanding research issues and opportunities for even better utilization of heterogeneous systems, including those equipped with GPUs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

GenMod: A generative modeling approach for spectral representation of PDEs with random inputs

Here, we propose a method for quantifying uncertainty in high-dimensional PDE systems with random parameters, where the number of solution evaluations is small. Parametric PDE solutions are often approximated using a spectral decomposition based on polynomial chaos expansions. For the class of systems we consider (i.e., high dimensional with limited solution evaluations) the coefficients are given by an underdetermined linear system in a regression formulation. This implies additional assumptions, such as sparsity of the coefficient vector, are needed to approximate the solution. Here, we present an approach where we assume the coefficients are close to the range of a generative model that maps from a low to a high dimensional space of coefficients. Our approach is inspired be recent work examining how generative models can be used for compressed sensing in systems with random Gaussian measurement matrices. Using results from PDE theory on coefficient decay rates, we construct an explicit generative model that predicts the polynomial chaos coefficient magnitudes. The algorithm we developed to find the coefficients, which we call GenMod, is composed of two main steps. First, we predict the coefficient signs using Orthogonal Matching Pursuit. Then, we assume the coefficients are within a sparse deviation from the range of a sign-adjusted generative model. This allows us to find the coefficients by solving a nonconvex optimization problem, over the input space of the generative model and the space of sparse vectors. We obtain theoretical recovery results for a Lipschitz continuous generative model and for a more specific generative model, based on coefficient decay rate bounds. We examine three high-dimensional problems and show that, for all three examples, the generative model approach outperforms sparsity promoting methods at small sample sizes.

97 MATHEMATICS AND COMPUTING↗

Hierarchical off-diagonal low-rank approximation of Hessians in inverse problems, with application to ice sheet model initialization

Obtaining lightweight and accurate approximations of discretized objective functional Hessians in inverse problems governed by partial differential equations (PDEs) is essential to make both deterministic and Bayesian statistical large-scale inverse problems computationally tractable. The cubic computational complexity of dense linear algebraic tasks, such as Cholesky factorization, that provide a means to sample Gaussian distributions and determine solutions of Newton linear systems is a computational bottleneck at large-scale. These tasks can be reduced to log-linear complexity by utilizing hierarchical off-diagonal low-rank (HODLR) matrix approximations. In this work, we show that a class of Hessians that arise from inverse problems governed by PDEs are well approximated by the HODLR matrix format. In particular, we study inverse problems governed by PDEs that model the instantaneous viscous flow of ice sheets. In these problems, we seek a spatially distributed basal sliding parameter field such that the flow predicted by the ice sheet model is consistent with ice sheet surface velocity observations. Here, we demonstrate the use of HODLR Hessian approximation to efficiently sample the Laplace approximation of the posterior distribution with covariance further approximated by HODLR matrix compression. Computational studies are performed which illustrate ice sheet problem regimes for which the Gauss–Newton data-misfit Hessian is more efficiently approximated by the HODLR matrix format than the low-rank (LR) format. We then demonstrate that HODLR approximations can be favorable, when compared to global LR approximations, for large-scale problems by studying the data-misfit Hessian associated with inverse problems governed by the first-order Stokes flow model on the Humboldt glacier and Greenland ice sheet.

97 MATHEMATICS AND COMPUTING↗

Phase-space entropy cascade and irreversibility of stochastic heating in nearly collisionless plasma turbulence

We consider a nearly collisionless plasma consisting of a species of “test particles” in one spatial and one velocity dimension, stirred by an externally imposed stochastic electric field—a kinetic analog of the Kraichnan model of passive advection. The mean effect on the particle distribution function is turbulent diffusion in velocity space—known as stochastic heating. Accompanying this heating is the generation of fine-scale structure in the distribution function, which we characterize with the collisionless (Casimir) invariant C 2 ∝ ∫ ∫ d x d v 〈 f 2 〉 —a quantity that here plays the role of (negative) entropy of the distribution function. We find that C 2 is transferred from large scales to small scales in both position and velocity space via a phase-space cascade enabled by both particle streaming and nonlinear interactions between particles and the stochastic electric field. We compute the steady-state fluxes and spectrum of C 2 in Fourier space, with k and s denoting spatial and velocity wave numbers, respectively. In our model, the nonlinearity in the evolution equation for the spectrum turns into a fractional Laplacian operator in k space, leading to anomalous diffusion. Whereas even the linear phase mixing alone would lead to a constant flux of C 2 to high s (towards the collisional dissipation range) at every k , the nonlinearity accelerates this cascade by intertwining velocity and position space so that the flux of C 2 is to both high k and high s simultaneously. Integrating over velocity (spatial) wave numbers, the k -space ( s -space) flux of C 2 is constant down to a dissipation length (velocity) scale that tends to zero as the collision frequency does, even though the rate of collisional dissipation remains finite. The resulting spectrum in the inertial range is a self-similar function in the ( k , s ) plane, with power-law asymptotics at large k and s . Our model is fully analytically solvable, but the asymptotic scalings of the spectrum can also be found via a simple phenomenological theory whose key assumption is that the cascade is governed by a “critical balance” in phase space between the linear and nonlinear timescales. We argue that stochastic heating is made irreversible by this entropy cascade and that, while collisional dissipation accessed via phase mixing occurs only at small spatial scales rather than at every scale as it would in a linear system, the cascade makes phase mixing even more effective overall in the nonlinear regime than in the linear one. Published by the American Physical Society 2024

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

NCCS High Performance GMRES Mixed Precision

HPG-MxP is a software package that performs a fixed number of multigrid preconditioned (using a Gauss-Seidel smoother) Generalized minimal residual (PGMRES) iterations in order to solve a possibly nonsymmetric large sparse linear system of equations. It is designed to be a benchmark to measure a computer's performance for sparse linear algebra workloads typical in scientific computing while allowing the use of mixed precision methods. The solution is required to have convergence characteristics and accuracy similar to double precision GMRES. It is based on the High Performance Conjugate Gradient Benchmark (HPCG) which restricts all implementations to use only the IEEE double precision format (FP64). The original implementation (https://github.com/hpg-mxp/hpg-mxp) was written by Ichitaro Yamazaki, Jennifer Loe, Christian Glusa, Sivasankaran Rajamanickam, Piotr Luszczek, and Jack Dongarra. Please refer to that repository for documentation on the original implementation. This version is maintained by the National Center for Computational Sciences at Oak Ridge National Laboratory. It is highly scalable and optimized for Oak Ridge Leadership Computing Facility (OLCF) systems, particularly Frontier.

Kashi, Aditya [Oak Ridge National Laboratory (ORNL↗

Efficient shallow Ritz method for 1D diffusion problems

This paper studies the shallow Ritz method for solving the one-dimensional diffusion problem. It is shown that the shallow Ritz method improves the order of approximation dramatically for non-smooth problems. To realize this optimal or nearly optimal order of the shallow Ritz approximation, we develop a damped block Newton (dBN) method that alternates between updates of the linear and non-linear parameters. Per each iteration, the linear and the non-linear parameters are updated by exact inversion and one step of a modified, damped Newton method applied to a reduced non-linear system, respectively. The computational cost of each dBN iteration is $\mathcal{O}$(n). Starting with the non-linear parameters as a uniform partition of the interval, numerical experiments show that the dBN is capable of efficiently moving mesh points to nearly optimal locations. In conclusion, to improve the efficiency of the dBN further, we propose an adaptive damped block Newton (AdBN) method by combining the dBN with the adaptive neuron enhancement (ANE) method [28].

Diffusion problems↗

SLAC microresonator RF(SMuRF) electronics: A tone-tracking readout system for superconducting microwave resonator arrays

Here, we describe the newest generation of the SLAC Microresonator RF (SMuRF) electronics, a warm digital control and readout system for microwave-frequency resonator-based cryogenic detector and multiplexer systems, such as microwave superconducting quantum interference device multiplexers (μmux) or microwave kinetic inductance detectors. Ultra-sensitive measurements in particle physics and astronomy increasingly rely on large arrays of cryogenic sensors, which in turn necessitate highly multiplexed readout and accompanying room-temperature electronics. Microwave-frequency resonators are a popular tool for cryogenic multiplexing, with the potential to multiplex thousands of detector channels on one readout line. The SMuRF system provides the capability for reading out up to 3328 channels across a 4–8 GHz bandwidth. Notably, the SMuRF system is unique in its implementation of a closed-loop tone-tracking algorithm that minimizes RF power transmitted to the cold amplifier, substantially relaxing system linearity requirements and effective noise from intermodulation products. Here, we present a description of the hardware, firmware, and software systems of the SMuRF electronics, comparing achieved performance with science-driven design requirements. In particular, we focus on the case of large-channel-count, low-bandwidth applications, but the system has been easily reconfigured for high-bandwidth applications. The system described here has been successfully deployed in lab settings and field sites around the world and is baselined for use on upcoming large-scale observatories.

47 OTHER INSTRUMENTATION↗

Resilience and fault tolerance in high-performance computing for numerical weather and climate prediction

Progress in numerical weather and climate prediction accuracy greatly depends on the growth of the available computing power. As the number of cores in top computing facilities pushes into the millions, increased average frequency of hardware and software failures forces users to review their algorithms and systems in order to protect simulations from breakdown. This report surveys hardware, application-level and algorithm-level resilience approaches of particular relevance to time-critical numerical weather and climate prediction systems. A selection of applicable existing strategies is analysed, featuring interpolation-restart and compressed checkpointing for the numerical schemes, in-memory checkpointing, user-level failure mitigation and backup-based methods for the systems. Numerical examples showcase the performance of the techniques in addressing faults, with particular emphasis on iterative solvers for linear systems, a staple of atmospheric fluid flow solvers. The potential impact of these strategies is discussed in relation to current development of numerical weather prediction algorithms and systems towards the exascale. Trade-offs between performance, efficiency and effectiveness of resiliency strategies are analysed and some recommendations outlined for future developments.

54 ENVIRONMENTAL SCIENCES↗

Report on local data recovery approaches suitable for weather and climate prediction (Deliverable 1.3) (V.1.0)

Numerical weather and climate prediction rates as one of the scientific applications whose accuracy improvements greatly depend on the growth of the available computing power. As the number of cores in top computing facilities pushes into the millions, increasing average frequency of hardware and software failures forces users to review their algorithms and systems in order to protect simulations from breakdown. This report surveys approaches for fault-tolerance in numerical algorithms and system resilience in parallel simulations from the perspective of numerical weather and climate prediction systems. A selection of existing strategies is analyzed, featuring interpolation-restart and compressed checkpointing for the numerics, in-memory checkpointing, user-level failure mitigation-based and backup-based methods for the systems. Numerical examples showcase the performance of the techniques in addressing faults, with particular emphasis on iterative solvers for linear systems, a staple of atmospheric fluid flow solvers. The potential impact of these strategies is discussed in relation to current development of numerical weather prediction algorithms and systems towards the exascale. Trade-offs between performance, efficiency and effectiveness of resiliency strategies are analyzed and some recommendations outlined for future developments.

97 MATHEMATICS AND COMPUTING↗

Explicit model predictive control through robust optimization

A strategy that calculates an explicit state feedback policy to regulate constrained uncertain discrete‐time uncertain linear systems is presented. We consider uncertain processes, affected by box‐bounded multiplicative uncertainty as well as bounded additive uncertainty with linear state and inputs constraints. The proposed method includes (i) the calculation of a terminal set constraint and (ii) the robust reformulation of state constraints in the prediction horizon. These features allow the derivation of the desired policy by solving a single multiparametric quadratic programming problem that guarantees feasible operation in the presence of uncertainty. Additionally, we employ variable and constraint elimination approaches to enhance the computational performance of the strategy. We demonstrate the steps and benefits of these developments with a numerical example and a chemical engineering case study.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Estimation of Participation Factors Using the Synchrosqueezed Wavelet Transform

This paper proposes a data-driven approach for estimating participation factors for a power system using only simulation results on selected disturbances. The approach is purely response-based and does not need a linearized system model for eigen-analysis, which makes it applicable to systems whose detailed, complete mathematical models are not available. Considering the unavoidable nonlinearity as exhibited in the transient period of a system response, the Synchrosqueezed Wavelet Transform is applied to simulated responses for modal analysis to obtain participation factors. Based on simulations of Kundur's two-area system using both the electromagnetic transient model and phasor model, the participation factors estimated by the proposed approach are compared with two other signal processing tools, the Prony analysis and continuous wavelet transform, and are also benchmarked with conventional model-based participation factors.

participation factor↗

A Comparative Study of Joint Modeling Methods and Analysis of Fasteners [Slides]

Motivation: Crucial aspect of mechanical design is joining methodology of parts. Ability to analyze joint and fasteners in system for structural integrity is fundamental. Different modeling representations of fasteners include spring, beam, and solid elements. Various methods compared for linear system to decide method appropriate for design study. New method for modeling fastener joint is explored from full system perspective. Analysis results match well with published experimental data for new method.

42 ENGINEERING↗

How to Partition a Quantum Observable

We present a partition of quantum observables in an open quantum system that is inherited from the division of the underlying Hilbert space or configuration space. It is shown that this partition leads to the definition of an inhomogeneous continuity equation for generic, non-local observables. This formalism is employed to describe the local evolution of the von Neumann entropy of a system of independent quantum particles out of equilibrium. Crucially, we find that all local fluctuations in the entropy are governed by an entropy current operator, implying that the production of entanglement entropy is not measured by this partitioned entropy. For systems linearly perturbed from equilibrium, it is shown that this entropy current is equivalent to a heat current, provided that the system-reservoir coupling is partitioned symmetrically. Finally, we show that any other partition of the coupling leads directly to a divergence of the von Neumann entropy. Thus, we conclude that Hilbert-space partitioning is the only partition of the von Neumann entropy that is consistent with the laws of thermodynamics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Convergence of Weak-SINDy Surrogate Models

In this paper, we give an in-depth error analysis for surrogate models generated by a variant of the Sparse Identification of Nonlinear Dynamics (SINDy) method. We start with an overview of a variety of nonlinear system identification techniques, namely SINDy, weak-SINDy, and the occupation kernel method. Under the assumption that the dynamics are a finite linear combination of a set of basis functions, these methods establish a linear system to recover coefficients. We illuminate the structural similarities between these techniques and establish a projection property for the weak-SINDy technique. Following the overview, we analyze the error of surrogate models generated by a simplified version of weak-SINDy. In particular, under the assumption of boundedness of a composition operator given by the solution, we show that (i) the surrogate dynamics converges towards the true dynamics and (ii) the solution of the surrogate model is reasonably close to the true solution. Finally, as an application, we discuss the use of a combination of weak-SINDy surrogate modeling and proper orthogonal decomposition (POD) to build a surrogate model for partial differential equations (PDEs).

97 MATHEMATICS AND COMPUTING↗

Neural Ordinary Differential Equations for Nonlinear System Identification

Neural ordinary differential equations (NODE) have been recently proposed as a promising approach for nonlinear system identification tasks. In this work, we systematically compare their predictive performance with current state-of-the-art nonlinear and classical linear methods. In particular, we present a quantitative study comparing NODE's performance against neural state-space models and classical linear system identification methods and evaluate their inference speed and prediction performance on open-loop errors across eight different dynamical systems. The experiments show that NODEs can consistently improve the prediction accuracy by order of magnitude compared to benchmark methods. Besides improved accuracy, we also observed that NODEs are less sensitive to hyperparameters compared to neural state-space models by paying the cost of increased computation at the inference time.

machine leaning, system identification, physics in↗

symPACK: A GPU-Capable Fan-Out Sparse Cholesky Solver

Sparse symmetric positive definite systems of equations are ubiquitous in scientific workloads and applications. Parallel sparse Cholesky factorization is the method of choice for solving such linear systems. Therefore, the development of parallel sparse Cholesky codes that can efficiently run on today’s large-scale heterogeneous distributed-memory platforms is of vital importance. Modern supercomputers offer nodes that contain a mix of CPUs and GPUs. To fully utilize the computing power of these nodes, scientific codes must be adapted to offload expensive computations to GPUs. We present symPACK, a GPU-capable parallel sparse Cholesky solver that uses one-sided communication primitives and remote procedure calls provided by the UPC++ library. We also utilize the UPC++ "memory kinds" feature to enable efficient communication of GPU-resident data. We show that on a number of large problems, symPACK outperforms comparable state-of-the-art GPU-capable Cholesky factorization codes by up to 14x on the NERSC Perlmutter supercomputer.

Bellavita, Julian↗

Linear solvers for power grid optimization problems: A review of GPU-accelerated linear solvers

The linear equations that arise in interior methods for constrained optimization are sparse symmetric indefinite, and they become extremely ill-conditioned as the interior method converges. These linear systems present a challenge for existing solver frameworks based on sparse LU or LDL T decompositions. Here, we benchmark five well known direct linear solver packages on CPU- and GPU-based hardware, using matrices extracted from power grid optimization problems. The achieved solution accuracy varies greatly among the packages. None of the tested packages delivers significant GPU acceleration for our test cases. For completeness of the comparison we include results for MA57, which is one of the most efficient and reliable CPU solvers for this class of problem.

97 MATHEMATICS AND COMPUTING↗

High-precision quantum algorithms for partial differential equations

Quantum computers can produce a quantum encoding of the solution of a system of differential equations exponentially faster than a classical algorithm can produce an explicit description. However, while high-precision quantum algorithms for linear ordinary differential equations are well established, the best previous quantum algorithms for linear partial differential equations (PDEs) have complexity poly(1/ϵ), where ϵ is the error tolerance. By developing quantum algorithms based on adaptive-order finite difference methods and spectral methods, we improve the complexity of quantum algorithms for linear PDEs to be poly(d,log(1/ϵ)), where d is the spatial dimension. Our algorithms apply high-precision quantum linear system algorithms to systems whose condition numbers and approximation errors we bound. We develop a finite difference algorithm for the Poisson equation and a spectral algorithm for more general second-order elliptic equations.

97 MATHEMATICS AND COMPUTING↗