Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Linear Solvers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Iterative Linearization for Phasor-Defined Optimal Power Dispatch

Optimal power flow (OPF) problems, which dispatch power targets to controllable generating units across a network, must generally account for non-convex constraints on power flow. Furthermore, adapting those problems so as to make them solvable with convex optimization techniques is an area of much academic and operational interest. In this paper, we present a method for solving OPF as a quadratic program by iteratively refining and re-initializing a linearized model of power flow based on the outputs of an associated nonlinear solver. The linear model on which we demonstrate this method is an adapted version of an approximation designed for use with unbalanced distribution networks. As an important benefit, the model allows for the explicit inclusion of nodal voltage phasor values in both the OPF problem's objective and its constraints, which opens the door to the idea of phasor-based control (PBC) design. We show in simulations on the IEEE 13-node test feeder that our method quickly converges to a set of phasor targets that are sufficiently precise for use in operations at the distribution level.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Low Precision and Efficient Programming Languages for Sustainable AI: Final Report for the Summer Project of 2024

This document contains all relevant material generated during the authors' summer internship at NREL in 2024. This report shows how to improve energy efficiency of a few code samples by using low-precision data types combined with mixed-precision algorithms. The main applications considered here are (i) linear system solvers using mixed precision, and (ii) neural networks using mixed precision. This report also discusses how programming languages affect energy consumption of algorithms, energy metrics for a code and tools, and the available current software and hardware infrastructure.

97 MATHEMATICS AND COMPUTING↗

High resolution numerical simulation of the linearized Euler equations in conservation law form

A linearized Euler solver based on a high resolution numerical scheme is presented. The approach is to linearize the flux vector as opposed to carrying through the complete linearization analysis with the dependent variable vector written as a sum of the mean and the perturbed flow. This allows the linearized equations to be maintained in conservation law form. The linearized equations are used to compute unsteady flows in turbomachinery blade rows arising due to blade vibrations. Numerical solutions are compared to theoretical results (where available) and to numerical solutions of the nonlinear Euler equations.

Sreenivas, Kidambi↗

Hardware acceleration for HPS algorithms in two and three dimensions

We provide a flexible, open-source framework for hardware acceleration, namely massively-parallel execution on general-purpose graphics processing units (GPUs), applied to the hierarchical Poincaré–Steklov (HPS) family of algorithms for building fast direct solvers for linear elliptic partial differential equations. To take full advantage of the power of hardware acceleration, we propose two variants of HPS algorithms to improve performance on two- and three-dimensional problems. In the two-dimensional setting, we introduce a novel recomputation strategy that minimizes costly data transfers to and from the GPU; in three dimensions, we modify and extend the adaptive discretization technique of Geldermans and Gillman [1] to greatly reduce peak memory usage. We provide an open-source implementation of these methods written in JAX, a high-level accelerated linear algebra package, which allows for the first integration of a high-order fast direct solver with automatic differentiation tools. We conclude with extensive numerical examples showing our methods are fast and accurate on two- and three-dimensional problems.

Fast direct solvers↗

A Tutorial for Using an Open-Source Solver for the Regional Energy Deployment System (ReEDS) Model

ReEDS is a publicly available model developed at NREL that can be used to analyze the potential evolution of the U.S. electric power system into the future. ReEDS is formulated as a linear program, written in GAMS, and solved using a linear programming solver (e.g., CPLEX). This tutorial provides context and understanding for the implications of using an open-source solver for ReEDS.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

kynema-fmb [SWR-23-07]

Kynema-FMB (FKA: Kynema) is an open-source performance portable flexible multibody (FMB) dynamics solver designed for time-domain simulations. While originally tailored for wind turbine structural dynamics, the formulation and implementation are those of a general flexible-multidbody dynamics solver that can readily be applied to a wide range of systems. Kynema was designed with a narrow focus, namely to provide a lightweight, fast, accurate FMD solver for coupling to computational-fluid-dynamics (CFD) codes, especially the CFD codes in the Kynema suite, for fluid-structure-interaction (FSI) simulations. Kynema-FMB is equipped to model systems that can be represented as a collection of beams and rigid bodies that are connected through constraints. Degrees of freedom are defined in the inertial/global frame of reference and include displacements and rotations (formally as rotation matrices, but stored as quaternions). The underlying formulation is built on a Lie-group time integrator designed for index-3 differential-algebraic equations, which is second-order accurate in time (Bruls et al., 2012). Beam models are based on geometrically exact beam theory and are discretized as high-order spectral finite elements similar to those in BeamDyn (Wang et al., 2017). The governing equations for a FMD system like a wind turbine constitute a highly nonlinear system of constrained partial-differential equations. Kynema-FMB uses analytical Jacobians in the nonlinear-system solves in each time step. Linear systems use sparse storage and several third-party sparse-linear-system solvers are enabled. Ill conditioning of linear systems is mitigated with preconditioning described in Bottasso et al, 2008. Kynema-FMB is integrated with a simple open-source controller (ROSCO). There is an application programming interface (API) for coupling to geometry-resolved CFD (like that in Sharma et al., 2023) and actuator-force CFD (like that in Kuhn et al., 2025). In the latter, for actuator-line models, Kynema-FMB includes an internal blade-element solver that depends on user-provided lookup tables for coefficients of lift and drag, i.e., aerodynamic polars. Kynema-FMB is written in C++ and leverages Kokkos and Kokkos-Kernels (KokkosEcosystem) as its performance portability layer enabling simulations on both CPU and GPU systems. The repository is equipped with extensive automated testing at the unit and regression/system levels. The following describes the high-level development objectives conceived for Kynema: *Kynema will follow modern software development best practices, including test-driven development (TDD), version control, hierarchical automated testing, and continuous integration (CI) for a robust development environment. *The core data structures are memory efficient and enable vectorization and parallelization at multiple levels. *Data structures are data-oriented to exploit methods for accelerated computing including high utilization of chip resources (e.g., single instruction multiple data (SIMD) instruction sets) and parallelization using GP-GPUs. *The computational algorithms incorporate robust open-source libraries for mathematical operations, resource allocation, and data management. *The API design considers multiple stakeholder needs and ensure integration with existing and future ecosystems for data science, machine learning, and AI. *Kynema-FMB is written in modern C++ and leverages Kokkos as its performance-portability library with inspiration from the kynema stack.

Sprague, MichaelA.↗

Prediction of unsteady aerodynamic loads in cascades using the linearized Euler equations on deforming grids

A linearized Euler solver for calculating unsteady flows in turbomachinery blade rows due to both incident gusts and blade motion is presented. Using the linearized Euler technique, one decomposes the flow into a mean (or steady) flow plus an unsteady, harmonically varying, small disturbance flow. Linear variable coefficient equations describe the small disturbance behavior of the flow, and are solved using a pseudo-time marching Lax-Wendroff scheme. For the blade motion problem, a harmonically deforming computational rid that conforms to the motion of vibrating blades eliminates large error producing mean flow gradient terms that would otherwise appear in the unsteady flow tangency boundary condition. The paper also presents a new, numerically exact, nonreflecting far-field boundary condition based on an eigenanalysis of the discretized equations. Computed flow solutions demonstrate the computational accuracy and efficiency of the present method. The solution of the linearized Euler equations requires one to two orders of magnitude less computer time than solution of the nonlinear Euler equations using traditional time-accurate time-marching techniques. In addition, the deformable grid significantly improves the accuracy of the solution.

Hall, Kenneth C.↗

A Parallel Symmetric Successive Overrelaxation Method for OVERFLOW

The block Jacobi symmetric successive overrelaxation (SSOR) algorithm has been reformulated as a parallelized algorithm for the OVERFLOWstructured, overset grid, computational fluid dynamics flow solver. Simple changes to the flow solver required to implement the algorithm are discussed. A series of test cases are presented that demonstrate how the addition of implicit overset boundaries has improved the robustness and nonlinear convergence characteristics of the flow solver.

Computational Fluid Dynamics↗

Novel Solver Algorithms for Nearly Singular Linear Systems Arising in Combustion Modelling

Direct Numerical Simulations of realistic combustion devices are extremely challenging due to the wide separation of scales in the simulation, for example an internal combustion (IC) engine chamber, and the flame thickness of a high-pressure flame. The PeleLMeX solver uses adaptive mesh refinement (AMR) to evolve multi-species reacting flows in the low Mach number limit at the Exascale and relies on an embedded boundary (EB) approach to represent complex geometries. In that framework, the EB geometries often give rise to very small cut-cells along the boundary, which translate into extreme ill-conditioning of the pressure-projection, with eigenvalues that span 15-16 orders of magnitude. In this talk, we focus on the case of a typical IC piston bowl geometry for which we present on a novel approach towards solving these nearly singular linear systems with ILU-based, C-AMG smoothers on massively parallel architectures. In particular, we use scaling and equilibration algorithms to handle the non-normality of the upper triangular factors. This enables us to approximate the highly sequential triangular solve algorithm, embedded in the AMG smoothing-solve phase, with Jacobi iterations. This approximation can be written as a convergent Neumann series whose terms are composed of highly parallel sparse matrix vector multiplications. The result is an algorithm that substantially decreases setup and solve time, compared to state-of-the-art, for these challenging linear systems.

combustion modelling↗

Prediction of Undsteady Flows in Turbomachinery Using the Linearized Euler Equations on Deforming Grids

A linearized Euler solver for calculating unsteady flows in turbomachinery blade rows due to both incident gusts and blade motion is presented. The model accounts for blade loading, blade geometry, shock motion, and wake motion. Assuming that the unsteadiness in the flow is small relative to the nonlinear mean solution, the unsteady Euler equations can be linearized about the mean flow. This yields a set of linear variable coefficient equations that describe the small amplitude harmonic motion of the fluid. These linear equations are then discretized on a computational grid and solved using standard numerical techniques. For transonic flows, however, one must use a linear discretization which is a conservative linearization of the non-linear discretized Euler equations to ensure that shock impulse loads are accurately captured. Other important features of this analysis include a continuously deforming grid which eliminates extrapolation errors and hence, increases accuracy, and a new numerically exact, nonreflecting far-field boundary condition treatment based on an eigenanalysis of the discretized equations. Computational results are presented which demonstrate the computational accuracy and efficiency of the method and demonstrate the effectiveness of the deforming grid, far-field nonreflecting boundary conditions, and shock capturing techniques. A comparison of the present unsteady flow predictions to other numerical, semi-analytical, and experimental methods shows excellent agreement. In addition, the linearized Euler method presented requires one or two orders-of-magnitude less computational time than traditional time marching techniques making the present method a viable design tool for aeroelastic analyses.

Clark, William S.↗

A Mixed Integer Linear Programming-basedDistributed Energy Management for Three-phaseUnbalanced Active Distribution Network

A mixed integer linear programming (MILP)–baseddistributed energy management for three-phase unbalancedactive distribution network is proposed. Modern distributionnetworks have becoming more and more active with increasingdeployment of microgrids, distributed energy resources (DERs)as well as controllable loads. Considering various ownership andcontrol models of microgrids, DERs and controllable loads, adistributed energy management was formulated using the alternatingdirection method of multipliers (ADMM) algorithm. ByADMM, the distribution management system (DMS) and theseactive components are coordinated through price signals, whichare adjusted according to the generation-load mismatch per nodeper phase. To enable resolution of the ADMM-based distributedoptimization using more accessible and popular MILP solver,different linearization techniques were proposed to linearize theaugmented Lagrangian terms and other nonlinear terms. Resultsof case studies on a three-phase active distribution network withthree microgrids and several DERs and controllable loads validatedthe effectiveness of proposed MILP-based distributed energymanagement. In addition, the capability of proposed method inmitigating phase power unbalance has been demonstrated.

Liu, Guodong↗

Multigrid deflation for Lattice QCD

Computing the trace of the inverse of large matrices is typically addressed through statistical methods. Deflating out the lowest eigenvectors or singular vectors of the matrix reduces the variance of the trace estimator. This work summarizes our efforts to reduce the computational cost of computing the deflation space while achieving the desired variance reduction for Lattice QCD applications. Previous efforts computed the lower part of the singular spectrum of the Dirac operator by using an eigensolver preconditioned with a multigrid linear system solver. Despite the improvement in performance in those applications, as the problem size grows the runtime and storage demands of this approach will eventually dominate the stochastic estimation part of the computation. In this work, we propose to compute the deflation space in one of the following two ways. First, by using an inexact eigensolver on the Hermitian, but maximally indefinite, operator. Second, by exploiting the fact that the multigrid prolongator for this operator is rich in components toward the lower part of the singular spectrum. We show experimentally that the inexact eigensolver can approximate the lower part of the spectrum even for ill-conditioned operators. Also, the deflation based on the multigrid prolongator is more efficient to compute and apply, and, despite its limited ability to approximate the fine level spectrum, it obtains similar variance reduction on the trace estimator as deflating with approximate eigenvectors from the fine level operator.

97 MATHEMATICS AND COMPUTING↗

High-order dimensionally-split Cartesian embedded boundary method for non-dissipative schemes

Centered finite-difference schemes are commonly used for high-fidelity turbulent flow simulations in canonical configurations because of their non-dissipative property and computational efficiency. However, their use in flow simulations over complex geometries is limited by the requirements of a structured grid and a stable boundary treatment in the absence of artificial (numerical) dissipation. Cartesian embedded boundary (EB) approaches provide an efficient structured-grid framework to apply difference schemes over complex domains. However, they are often restricted to low orders of accuracy because of numerical instabilities at the embedded boundaries and the issues of small-cell problem that are difficult to address with high-order accuracy. The present work discusses a systematic approach to obtain high-order EB methods with non-dissipative centered schemes in the interior. This approach, based on satisfying the primary and secondary conservation conditions, is employed to derive EB schemes that are up to sixth-order accurate in the interior and fourth-order accurate globally for hyperbolic, parabolic as well as incompletely parabolic problems. The proposed finite-difference discretization is, by construction, dimensionally split and addresses the small-cell problem without any cell/geometry transformations, thus, highly simplifying implementation in a flow solver. Various linear and non-linear numerical tests are performed to evaluate the stability and the accuracy of the proposed EB schemes.

97 MATHEMATICS AND COMPUTING↗

Solving differential‐algebraic equations in power system dynamic analysis with quantum computing

Abstract Power system dynamics are generally modeled by high dimensional non‐linear differential‐algebraic equations (DAEs) given a large number of components forming the network. These DAEs' complexity can grow exponentially due to the increasing penetration of distributed energy resources, whereas their computation time becomes sensitive due to the increasing interconnection of the power grid with other energy systems. This paper demonstrates the use of quantum computing algorithms to solve DAEs for power system dynamic analysis. We leverage a symbolic programming framework to equivalently convert the power system's DAEs into ordinary differential equations (ODEs) using index reduction methods and then encode their data into qubits using amplitude encoding. The system non‐linearity is captured by Hamiltonian simulation with truncated Taylor expansion so that state variables can be updated by a quantum linear equation solver. Our results show that quantum computing can solve the power system's DAEs accurately with a computational complexity polynomial in the logarithm of the system dimension. We also illustrate the use of recent advanced tools in scientific machine learning for implementing complex computing concepts, that is, Taylor expansion, DAEs/ODEs transformation, and quantum computing solver with abstract representation for power engineering applications.

computational complexity↗

Validation of the GFS model for gyrokinetic stability of NSTX pedestal data

This study presents a large database validation of the gyro fluid system (GFS) model for linear gyrokinetic stability for high-mode (H-mode) edge transport barrier conditions in the national spherical torus experiment (NSTX) tokamak. The database of linear stability calculations with the CGYRO gyrokinetic code was produced using plasma profile measurements from NSTX discharges to identify kinetic ballooning modes (KBM), trapped electron modes (TEM), and micro-tearing modes (MTM) that limit the pressure profile gradient in the H-mode barrier. A novel Bayesian optimization approach determines optimal resolution parameters for GFS specifically for spherical tokamak pedestal conditions. Our results demonstrate that GFS, with optimized resolution, can achieve accurate linear stability analysis in NSTX pedestal conditions for reduced resolution compared to CGYRO. GFS can accurately find the KBM, TEM, and MTM instability branches. Parametric analysis reveals that GFS accuracy in this extreme pedestal parameter range is degraded for low magnetic shear and near the separatrix conditions. These findings establish GFS as a fast linear eigenmode solver for spherical tokamak pedestal gyrokinetic stability and demonstrate a systematic methodology for determining the optimum resolution settings.

Yang, Minglei [Oak Ridge National Laboratory (ORNL↗

Co-design for Particle Applications at Exascale

Co-design across the Exascale Computing Project (ECP) has been critical for both enabling science applications and bringing disparate communities together. Developing and porting applications to the various high-performance computing (HPC) architectures on pre-exascale and exascale computers has been quite challenging due to the diversity of hardware features and software stacks. The Co-design Center for Particle Applications (CoPA) has developed and enhanced the Cabana and PROGRESS/BML libraries to facilitate the creation of new particle applications, make existing particle applications exascale capable, and allow teams to explore new capabilities. Particle methods from atomistic, mesoscale, continuum, through cosmological scales have been built with Cabana, along with new possibilities for application coupling. Similarly, the PROGRESS/BML library has enabled quantum particle applications with linear algebra solvers to use advanced hardware. Across these CoPA-developed libraries, the co-design abstraction layer combines performance portability with math library support to facilitate separation of concerns and directly support science runs.

97 MATHEMATICS AND COMPUTING↗

GSoFa: Scalable Sparse Symbolic LU Factorization on GPUs

Decomposing a matrix $\mathbf {A}$ into a lower matrix $\mathbf {L}$ and an upper matrix $\mathbf {U}$, which is also known as LU decomposition, is an essential operation in numerical linear algebra. For a sparse matrix, LU decomposition often introduces more nonzero entries in the $\mathbf {L}$ and $\mathbf {U}$ factors than in the original matrix. A symbolic factorization step is needed to identify the nonzero structures of $\mathbf {L}$ and $\mathbf {U}$ matrices. Attracted by the enormous potentials of the Graphics Processing Units (GPUs), an array of efforts have surged to deploy various LU factorization steps except for the symbolic factorization, to the best of our knowledge, on GPUs. This article introduces gSoFa, the first GPU-based symbolic factorization design with the following three optimizations to enable scalable LU symbolic factorization for nonsymmetric pattern sparse matrices on GPUs. First, here we introduce a novel fine-grained parallel symbolic factorization algorithm that is well suited for the Single Instruction Multiple Thread (SIMT) architecture of GPUs. Second, we tailor supernode detection into a SIMT friendly process and strive to balance the workload, minimize the communication and saturate the GPU computing resources during supernode detection. Third, we introduce a three-pronged optimization to reduce the excessive space consumption problem faced by multi-source concurrent symbolic factorization. Taken together, gSoFa achieves up to 31× speedup from 1 to 44 Summit nodes (6 to 264 GPUs) and outperforms the state-of-the-art CPU project, on average, by 5×. Notably, gSoFa also achieves up to 47 percent of the peak memory throughput of a V100 GPU in the Summit Supercomputer.

97 MATHEMATICS AND COMPUTING↗