Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Linear systems solvers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Optimal charging scheduling and management for a fast-charging battery electric bus system

Herein we discuss how battery electric buses (BEBs) are rapidly being embraced by public transit agencies because of their environmental and economic benefits. To address the problems of limited driving range and time-consuming charging for BEBs, manufacturers have developed rapid on-route charging technology that utilizes typical layovers at terminals to charge buses in operation using high power. With on-route fast-charging, BEBs are as capable as their diesel counterparts in terms of range and operating time. However, on-route fast-charging makes it more challenging to schedule and manage charging events for a BEB system. First, on-route fast-charging may lead to high electricity power demand charges. Second, it may increase electricity energy charges because of charging that occurs during on-peak hours. Without careful charging scheduling and management, on-route fast-charging may significantly increase fuel costs and reduce the economic attractiveness of BEBs. The present study proposes a network modeling framework to optimize the charging scheduling and management for a fast-charging BEB system, effectively minimizing total charging costs. The charging schedule determines when to charge a BEB, while the charging management strategically controls the actual charging power. Charging costs include both electricity demand charges and energy charges. The charging scheduling and management problem is first formulated as a nonlinear nonconvex program with time-continuous variables. A discretizing method and a linear reformulation technique are then adopted to reformulate the model as a linear program, which can be easily solved using off-the-shelf solvers, even for large-scale problems. Finally, the model is demonstrated with extensive numerical studies based on two real-world bus networks. The results demonstrate that the proposed model can effectively determine the optimal charging scheduling and management for a fast-charging BEB system, which carries the potential for use in large-scale real-world bus networks.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Iterative methods in GPU-resident linear solvers for nonlinear constrained optimization

Linear solvers are major computational bottlenecks in a wide range of decision support and optimization computations. The challenges become even more pronounced on heterogeneous hardware, where traditional sparse numerical linear algebra methods are often inefficient. For example, methods for solving ill-conditioned linear systems have relied on conditional branching, which degrades performance on hardware accelerators such as graphical processing units (GPUs). To improve the efficiency of solving ill-conditioned systems, our computational strategy separates computations that are efficient on GPUs from those that need to run on traditional central processing units (CPUs). Our strategy maximizes the reuse of expensive CPU computations. Iterative methods, which thus far have not been broadly used for ill-conditioned linear systems, play an important role in our approach. In particular, we extend ideas from Arioli et al., (2007) to implement iterative refinement using inexact LU factors and flexible generalized minimal residual (FGMRES), with the aim of efficient performance on GPUs. In conclusion, we focus on solutions that are effective within broader application contexts, and discuss how early performance tests could be improved to be more predictive of the performance in a realistic environment.

97 MATHEMATICS AND COMPUTING↗

Gravitational Self-force Errors of Poisson Solvers on Adaptively Refined Meshes

An error in the gravitational force that the source of gravity induces on itself (a self-force error) violates both the conservation of linear momentum and the conservation of energy. If such errors are present in a self-gravitating system and are not sufficiently random to average out, the obtained numerical solution will become progressively more unphysical with time: the system will acquire or lose momentum and energy due to numerical effects. In this paper, we demonstrate how self-force errors can arise in the case where self-gravity is solved on an adaptively refined mesh when the refinement is nonuniform. Here, we provide the analytical expression for the self-force error and numerical examples that demonstrate such self-force errors in idealized settings. We also show how these errors can be corrected to an arbitrary order by straightforward addition of correction terms at the refinement boundaries.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A Fast Algebraic Multigrid Solver and Accurate Discretization for Highly Anisotropic Heat Flux I: Open Field Lines

We present a novel solver technique for the anisotropic heat flux equation, aimed at the high level of anisotropy seen in magnetic confinement fusion plasmas. Such problems pose two major challenges: (i) discretization accuracy and (ii) efficient implicit linear solvers. We simultaneously address each of these challenges by constructing a new finite element discretization with excellent accuracy properties, tailored to a novel solver approach based on algebraic multigrid (AMG) methods designed for advective operators. We pose the problem in a mixed formulation, introducing the directional temperature gradient as an auxiliary variable. The temperature and auxiliary fields are discretized in a scalar discontinuous Galerkin space with upwinding principles used for discretizations of advection. We demonstrate the proposed discretization’s superior accuracy over other discretizations of anisotropic heat flux, achieving error 1000x smaller for anisotropy ratio of 10 9 , for closed field lines. The block matrix system is reordered and solved in an approach where the two advection operators are inverted using AMG solvers based on approximate ideal restriction, which is particularly efficient for upwind discontinuous Galerkin discretizations of advection. To ensure that the advection operators are nonsingular, in this paper we restrict ourselves to considering open (acyclic) magnetic field lines for the linear solvers. We demonstrate fast convergence of the proposed iterative solver in highly anisotropic regimes where other diffusion-based AMG methods fail.

97 MATHEMATICS AND COMPUTING↗

Nonlinear Elimination Applied to Radiation Diffusion

We apply a nonlinearly preconditioned, quasi-Newton framework to accelerate the numerical solution of the thermal radiative transfer (TRT) equations. This framework was inspired by the unpublished method that has existed for years in Teton, Lawrence Livermore National Laboratory’s deterministic TRT code. In this paper, we cast this iteration scheme within a formal nonlinear preconditioning framework and compare its performance against other iteration schemes in the framework. With proper choices of iteration controls for the various levels of the solver, we can recover the standard linearized one-step method, a full nonlinear Newton scheme, as well as the method in Teton. In brief, the nonlinear preconditioning TRT scheme formally eliminates the material temperature equation from the nonlinear system in a nonlinear analog of a Schur complement. This nonlinear elimination step involves solving a decoupled nonlinear equation for each spatial degree of freedom and is therefore inexpensive. By applying a quasi-Newton iteration scheme on the new system, we obtain a three-level iteration scheme that is at least as efficient as commonly used TRT schemes. The new method allows full convergence to the nonlinear backward Euler time-discretized system, increasing accuracy and robustness, while using a similar number of linear iterations as the more common linearized one-step methods Eq. (4).

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Failure Probability Constrained AC Optimal Power Flow

Despite cascading failures being the central cause of blackouts in power transmission systems, existing operational and planning decisions are made largely by ignoring their underlying cascade potential. This paper posits a reliability-aware AC Optimal Power Flow formulation that seeks to design a dispatch point which has a low operator-specified likelihood of triggering a cascade starting from any single component outage. By exploiting a recently developed analytical model of the probability of component failure, our Failure Probability-constrained ACOPF (FP-ACOPF) utilizes the system's expected first failure time as a smoothly tunable and interpretable signature of cascade risk. Here, we use techniques from bilevel optimization and numerical linear algebra to efficiently formulate and solve the FP-ACOPF using off-the-shelf solvers. Extensive simulations on the IEEE 118-bus case show that, when compared to the unconstrained and N-1 security-constrained ACOPF, our probability-constrained dispatch points can significantly lower the probabilities of long severe cascades and of large demand losses, while incurring only minor increases in total generation costs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Asynchronous Iterative Solvers for Extreme-Scale Computing

The Asynchronous Iterative Solvers for Extreme-Scale Computing (AsyncIS) project aims to explore more efficient numerical algorithms by decreasing their overhead. AsyncIS does this by replacing the outer Krylov subspace solver with an asynchronous optimized Schwarz method, thereby removing the global synchronization and bulk synchronous operations typically used in numerical codes. AsyncIS—a U.S. Department of Energy (DOE)-funded collaboration between Georgia Tech, the University of Tennessee, Knoxville, Temple University, and Sandia National Laboratories—also focuses on the development and optimization of asynchronous preconditioners (i.e., preconditioners that are generated and/or applied in an asynchronous fashion). The novel preconditioning algorithms that provide fine-grained parallelism enable preconditioned Krylov solvers to run efficiently on large-scale distributed systems and manycore accelerators like GPUs.

97 MATHEMATICS AND COMPUTING↗

Constrained Local Approximate Ideal Restriction for Advection-Diffusion Problems

Herein this paper focuses on developing a reduction-based algebraic multigrid (AMG) method that is suitable for solving general (non)symmetric linear systems and is naturally robust from pure advection to pure diffusion. Initial motivation comes from a new reduction-based AMG approach, $\ell \text{AIR}$ (local approximate ideal restriction), that was developed for solving advection-dominated problems. Though this new solver is very effective in the advection-dominated regime, its performance degrades in cases where diffusion becomes dominant. This is consistent with the fact that in general, reduction-based AMG methods tend to suffer from growth in complexity and/or convergence rates as the problem size is increased, especially for diffusion-dominated problems in two or three dimensions. Motivated by the success of $\ell \text{AIR}$ in the advective regime, our aim in this paper is to generalize the AIR framework with the goal of improving the performance of the solver in diffusion-dominated regimes. To do so, we propose a novel way to combine mode constraints as used commonly in energy-minimization AMG methods with the local approximation of ideal operators used in $\ell \text{AIR}$. The resulting constrained $\ell \text{AIR}$ algorithm is able to achieve fast scalable convergence on advective and diffusive problems. In addition, it is able to achieve standard low complexity hierarchies in the diffusive regime through aggressive coarsening, something that was previously difficult for reduction-based methods.

97 MATHEMATICS AND COMPUTING↗

An efficient high-order numerical solver for diffusion equations with strong anisotropy

In this paper, we present an interior penalty discontinuous Galerkin finite element scheme for solving diffusion problems with strong anisotropy arising in magnetized plasmas for fusion applications. Additionally, we demonstrate the accuracy produced by the high-order scheme and develop an efficient preconditioning technique to solve the corresponding linear system, which is robust to the mesh size and anisotropy of the problem. Several numerical tests are provided to validate the accuracy and efficiency of the proposed algorithm.

97 MATHEMATICS AND COMPUTING↗

Assessing the Feasibility of Bordered Block Diagonal Reordering in Power System Matrices using Fully Convolutional Network

In electromagnetic transient (EMT) simulations for power systems and inverter-based resources (IBRs), the arrangement of states within the system's linear equations, represented by matrix A in Ax=b, is critical. The state ordering in matrix A can highlight distinct characteristics of the system's graph, and identifying an optimal state ordering is crucial for efficient computation. The choice of state ordering, however, is dependent on the solver used, as each solver may perform optimally with different matrix patterns. With a wide array of matrix reordering algorithms available, selecting the most suitable one becomes challenging without insights into the matrix's ideal configuration. To address this, the paper proposes a fully convolutional network (FCN) to evaluate the reordering potential of the A matrix into a bordered block diagonal (BBD) pattern, which is commonly observed in power system and IBR modeling. The FCN's assessment aims to streamline the solver's operation, which in turn could substantially reduce the computational time required to find a solution.

Xia, Qianxue↗

Linear complexity

We present factorization and solution phases for a new linear complexity direct solver designed for concurrent batch operations on fine-grained parallel architectures, for matrices amenable to hierarchical representation. We focus on the strong-admissibility-based $\mathscr{H}^{2}$ format, where strong recursive skeletonization factorization compresses remote interactions. We build upon previous implementations of $\mathscr{H}^{2}$ matrix construction for efficient factorization and solution algorithm design, which are illustrated graphically in stepwise detail. The algorithms are ‘blackbox’ in the sense that the only inputs are the matrix and right-hand side, without analytical or geometrical information about the origin of the system. We demonstrate linear complexity scaling in both time and memory on four representative families of dense matrices up to one million in size. Parallel scaling up to 16 threads is enabled by a multi-level matrix graph coloring and avoidance of dynamic memory allocations thanks to prefix-sum memory management. An experimental backward error analysis is included. We break down the timings of different phases, identify phases that are memory-bandwidth limited, and discuss alternatives for phases that may be sensitive to the trend to employ lower precisions for performance.

Boukaram, Wajih↗

Impact of Reordering on the LU Factorization Performance of Bordered Block-Diagonal Sparse Matrix

Power engineers rely on computer-based simulation tools to assess grid performance and ensure security. At the core of these tools are solvers for sparse linear equations. When transformed into a bordered block-diagonal (BBD) structure, part of the sparse linear equation solving can be parallelized. This work focuses on using the Schur-complement-based method for LU factorization on BBD matrices, specifically, Jacobian matrices from large-scale systems. Our findings show that the natural ordering method outperforms the default ordering method in computational performance for each block of the BBD matrix. This observation is validated using synthetic 25k-bus and 70k-bus cases, showing a speedup of up to 38% when using natural ordering without permutation. Additionally, the impact of the number of partitions is studied, and the result shows that computational performance improves with more, smaller partitions in the BBD matrices.

BBD matrix↗

A low-rank solver for the stochastic unsteady Navier–Stokes problem

Here we study a low-rank iterative solver for the unsteady Navier–Stokes equations for incompressible flows with a stochastic viscosity. The equations are discretized using the stochastic Galerkin method, and we consider an all-at-once formulation where the algebraic systems at all the time steps are collected and solved simultaneously. The problem is linearized with Picard’s method. To efficiently solve the linear systems at each step, we use low-rank tensor representations within the Krylov subspace method, which leads to significant reductions in storage requirements and computational costs. Combined with effective mean-based preconditioners and the idea of inexact solve, we show that only a small number of linear iterations are needed at each Picard step. The proposed algorithm is tested with a model of flow in a two-dimensional symmetric step domain with different settings to demonstrate the computational efficiency.

97 MATHEMATICS AND COMPUTING↗

Design Considerations for GPU-based Mixed Integer Programming on Parallel Computing Platforms

Mixed Integer Programming (MIP) is a powerful abstraction in combinatorial optimization that finds real-life application across many significant sectors. The recent proliferation of graphical processing unit (GPU)-based accelerated computing architectures in large-scale parallel computing or supercomputing presents new opportunities as well as challenges in the advancement of MIP solver technology to effectively use the new accelerated computing platforms and scale to large parallel systems. Here, we recount the conventional processor-based strategies and focus on configurations where the most promising intersection lies between parallel MIP solver approaches and the specific strengths of accelerated parallel platforms. We note that the best potential lies in solving problems whose individual matrix sizes (of the linear program relaxation) fit entirely within one accelerator's memory and whose branch-and-bound (or branch-and-cut) trees cannot be fully contained within a small number of computational nodes. Additionally, we identify ideal features of computational linear algebra support on GPU accelerators that would help advance this direction of scalable parallel solution of MIP problems on GPU-based accelerated computing architectures.

Perumalla, Kalyan↗

A High-Throughput Solver for Marginalized Graph Kernels on GPU

Here, we present the design and optimization of a solver for efficient and high-throughput computation of the marginalized graph kernel on General Purpose GPUs. The graph kernel is computed using the conjugate gradient method to solve a generalized Laplacian of the tensor product between a pair of graphs. To cope with the large gap between the instruction throughput and the memory bandwidth of the GPUs, our solver forms the graph tensor product on-the-fly without storing it in memory. This is achieved by using threads in a warp cooperatively to stream the adjacency and edge label matrices of individual graphs by small square matrix blocks called tiles, which are then staged in registers and the shared memory for later reuse. Warps across a thread block can further share tiles via the shared memory to increase data reuse. We exploit the sparsity of the graphs hierarchically by storing only non-empty tiles using a coordinate format and nonzero elements within each tile using bitmaps. We propose a new partition-based reordering algorithm for aggregating nonzero elements of the graphs into fewer but denser tiles to further exploit sparsity. We carry out extensive theoretical analyses on the graph tensor product primitives for tiles of various density and evaluate their performance on synthetic and real-world datasets. Our solver delivers three to four orders of magnitude speedup over existing CPU-based solvers such as GraKeL and GraphKernels. The capability of the solver enables kernel-based learning tasks at unprecedented scales.

97 MATHEMATICS AND COMPUTING↗

A Scalable Multigrid Reduction Framework for Multiphase Poromechanics of Heterogeneous Media

Simulation of multiphase poromechanics involves solving a multiphysics problem in which multiphase flow and transport are tightly coupled with the porous medium deformation. To capture this dynamic interplay, fully implicit methods, also known as monolithic approaches, are usually preferred. The main bottleneck of a monolithic approach is that it requires solution of large linear systems that result from the discretization and linearization of the governing balance equations. Because such systems are nonsymmetric, indefinite, and highly ill-conditioned, preconditioning is critical for fast convergence. Recently, most efforts in designing efficient preconditioners for multiphase poromechanics have been dominated by physics-based strategies. Current state-of-the-art “black-box” solvers such as algebraic multigrid (AMG) are ineffective because they cannot effectively capture the strong coupling between the mechanics and the flow subproblems, as well as the coupling inherent in the multiphase flow and transport process. In this work, we develop an algebraic framework based on multigrid reduction (MGR) that is suited for tightly coupled systems of PDEs. Using this framework, the decoupling between the equations is done algebraically through defining appropriate interpolation and restriction operators. One can then employ existing solvers for each of the decoupled blocks or design a new solver based on knowledge of the physics. We demonstrate the applicability of our framework when used as a “black-box” solver for multiphase poromechanics. Here, we show that the framework is flexible to accommodate a wide range of scenarios, as well as efficient and scalable for large problems.

97 MATHEMATICS AND COMPUTING↗

Sensor enabled data-driven predictive analytics for modeling and control with high penetration of DERs in distribution systems

The electric power grid is undergoing a tremendous transformation due to the increasing penetration of renewable energy resources beginning with wind and more recently with the distributed energy resources (DERs) such as solar and battery storage. DERs have dramatically changed the role of the distribution systems in the overall power grid, and they are expected to contribute a significant portion of power generation in the future. If current trends for DERs continue, system operation and control will need to change dramatically for improved grid reliability and resiliency. As renewable resources increase in penetration, new and challenging operational, planning, and design problems are expected to emerge. Some of the key challenges that arise in the planning and operation of the future grid are: 1) Quantifying the impact of high DER penetration in distribution systems on bulk grid behavior over multiple time scales. 2) Identifying whether a particular DER configuration/settings have a large impact on the overall grid behavior. These challenges can be addressed in an offline manner using detailed T&D grid models and they can also be addressed in an online manner using sensor measurements. In particular, the advancement and planned growth in sensor technology in power grid over various voltage levels provide us with a unique opportunity to tackle these challenges from a data analytic perspective without needing detailed T&D grid models. A few questions that naturally arise when addressing the challenges from DERs using sensor data are: 1) How can we use limited sensor measurements to monitor & control voltage stability and small signal stability of the bulk system? 2) How can we ensure that the developed data analytic methods are robust to data availability and quality issues? 3) How can we compute the developed analytics in a scalable manner using streaming measurements? In this project, we addressed the aforementioned challenges arising from DERs and answered the questions raised above on how to effectively use the sensor measurements to enhance the reliability and performance of the electric grid. Thus, the overarching goal of this project is to develop effective reduced/representative system models from data that make the computational complexity sufficiently manageable so as to be useful to simulate, analyze, and even control complex non-linear power systems dynamics with large penetrations of DERs. In order to achieve the objective, the project team established a four-fold technical approach 1) Formulated a combined transmission-distribution co-simulation framework for data generation and validation, 2) Derived reduced/representative models of power systems based on data-driven methods for efficient computation and appropriate representation of system behavior, 3) Developed data driven characterization of power system behavior based on transfer operator theory, machine learning and optimization for model estimation, 4) Incorporated a scalable data management and processing architecture using distributed Kafka streaming applications that coordinate input data streams to the developed data analytics. The key accomplishments of the project are: 1) Development of a scalable multi-timescale T&D co-simulation framework (both for steady state and for dynamic co-simulation) using commercial solvers (PSSE and GridLAB-D). The steady-state T&D co-simulation interface is shared with our industry partner (PJM). 2) A structured reduced order dynamic model of distribution systems that can represent partial motor stalling along with a systematic procedure to derive the model parameters. 3) A PMU based online method to monitor, localize and mitigate fault-induced delayed voltage recovery using DER reactive support and load control in distribution systems. 4) Development of linear operator based robust methodologies for dynamic state estimation, uncertainty quantification, system identification and trajectory prediction for power system dynamics. 5) An adaptive damping control for utilizing wind energy resources to provide oscillation damping and system stability. 6) Implementation of Kafka-based framework for efficient processing of streaming data using Linux-based local virtual environment.

DER integration↗

Mathematical solutions in internal dose assessment: A comparison of Python-based differential equation solvers in biokinetic modeling

Abstract In biokinetic modeling systems employed for radiation protection, biological retention and excretion have been modeled as a series of discretized compartments representing the organs and tissues of the human body. Fractional retention and excretion in these organ and tissue systems have been mathematically governed by a series of coupled first-order ordinary differential equations (ODEs). The coupled ODE systems comprising the biokinetic models are usually stiff due to the severe difference between rapid and slow transfers between compartments. In this study, the capabilities of solving a complex coupled system of ODEs for biokinetic modeling were evaluated by comparing different Python programming language solvers and solving methods with the motivation of establishing a framework that enables multi-level analysis. The stability of the solvers was analyzed to select the best performers for solving the biokinetic problems. A Python-based linear algebraic method was also explored to examine how the numerical methods deviated from an analytical or semi-analytical method. Results demonstrated that customized implicit methods resulted in an enhanced stable solution for the inhaled 60 Co (Type M) and 131 I (Type F) exposure scenarios for the inhalation pathway of the International Commission on Radiological Protection (ICRP) Publication 130 Human Respiratory Tract Model (HRTM). The customized implementation of the Python-based implicit solvers resulted in approximately consistent solutions with the Python-based matrix exponential method ( expm ). The differences generally observed between the implicit solvers and expm are attributable to numerical precision and the order of numerical approximation of the numerical solvers. This study provides the first analysis of a list of Python ODE solvers and methods by comparing their usage for solving biokinetic models using the ICRP Publication 130 HRTM and provides a framework for the selection of the most appropriate ODE solvers and methods in Python language to implement for modeling the distribution of internal radioactivity.

61 RADIATION PROTECTION AND DOSIMETRY↗