Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Linear systems solvers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

An investigation of Newton-Sketch and subsampled Newton methods

Sketching, a dimensionality reduction technique, has received much attention in the statistics community. In this paper, we study sketching in the context of Newton's method for solving finite-sum optimization problems in which the number of variables and data points are both large. In this work, we study two forms of sketching that perform dimensionality reduction in data space: Hessian subsampling and randomized Hadamard transformations. Each has its own advantages, and their relative tradeoffs have not been investigated in the optimization literature. Additionally, our study focuses on practical versions of the two methods in which the resulting linear systems of equations are solved approximately, at every iteration, using an iterative solver. The advantages of using the conjugate gradient method vs. a stochastic gradient iteration are revealed through a set of numerical experiments, and a complexity analysis of the Hessian subsampling method is presented.

97 MATHEMATICS AND COMPUTING↗

Towards Efficient Alternating Current Optimal Power Flow Analysis on Graphical Processing Units

We present a solution of sparse ACOPF analysis on GPU. In particular, we discuss the performance bottlenecks and detail our efforts to accelerate the linear solver, a core component of ACOPF that dominates the computational time. ACOPF solutions of two large-scale systems, synthetic Northeast (25,000 buses) and Eastern (70,000 buses) \cite{birchfield2017tamu-cases} on GPU show promising speed-up compared to CPU based solution using a state-of-the-art solver. To our knowledge, this is the first result demonstrating acceleration of sparse ACOPF on GPUs.

Power grid analysis, GPU↗

Predicting Flow in Fracture Networks With Quantum Algorithms

Uncertainty quantification plays a crucial role in the modeling of subsurface flow. For instance, uncertainties in the properties of geologic fracture networks significantly impact flow, requiring numerous simulations to accurately estimate quantities of interest. However, each simulation is computationally expensive because it requires solving a large linear system to capture features that involve both small and large fractures. An example is in percolation, where the interaction of many small fractures (which cumulatively can have a large surface area) with the rock matrix must be modeled precisely. Quantum computing is an emerging tool with the potential to address this issue. Quantum algorithms offer a significant speedup in solving linear systems, achieving efficiencies that are challenging to match with classical approaches. These classical approaches include direct solvers, such as LU decomposition, and iterative methods, notably preconditioned conjugate gradient, commonly used in subsurface modeling to solve large sparse systems. However, applying quantum algorithms to geologic fracture flow requires careful attention to algorithmic and problem-specific constraints to fully realize this quantum advantage. In this work we describe a quantum algorithm for generalized Monte Carlo applications with a quadratic speedup over the classical approaches which can be combined with the quantum speedup, currently under investigation, for solving quantum linear systems for subsurface flow. We show that for quantum algorithms the computational cost of estimating a quantity of interest for a statistical ensemble of networks is roughly the same as that of a single realization, essentially implying that one can get uncertainty quantification for free.

58 GEOSCIENCES↗

Demonstration and performance testing of extreme-resolution simulations with static meshes on Summit (CPU & GPU) for a parked-turbine configuration and an actuator-line (mid-fidelity model) wind farm configuration (ECP-Q4 FY2020 Milestone Report)

The goal of the ExaWind project is to enable predictive simulations of wind farms comprised of many megawatt-scale turbines situated in complex terrain. Predictive simulations will require computational fluid dynamics (CFD) simulations for which the mesh resolves the geometry of the turbines and captures the rotation and large deflections of blades. Whereas such simulations for a single turbine are arguably petascale class, multi-turbine wind farm simulations will require exascale-class resources. The primary physics codes in the ExaWind simulation environment are Nalu-Wind, an unstructured-grid solver for the acoustically incompressible Navier-Stokes equations, AMR-Wind, a block-structured-grid solver with adaptive mesh refinement capabilities, and OpenFAST, a wind-turbine structural dynamics solver. The Nalu-Wind model consists of the mass-continuity Poisson-type equation for pressure and Helmholtz-type equations for transport of momentum and other scalars. For such modeling approaches, simulation times are dominated by linear-system setup and solution for the continuity and momentum systems. For the ExaWind challenge problem, the moving meshes greatly affect overall solver costs as reinitialization of matrices and recomputation of preconditioners is required at every time step. The choice of overset-mesh methodology to model the moving and non-moving parts of the computational domain introduces constraint equations in the elliptic pressure-Poisson solver. The presence of constraints greatly affects the performance of algebraic multigrid preconditioners.

17 WIND ENERGY↗

Hydrodynamic characterization of the coastal pioneer array ocean observing system

Ocean observation buoys require relatively small amounts of power, yet traditionally necessitate costly resupply trips for battery replacement. With the offshore location of the buoys and small power requirements, wave energy may be an effective solution for providing consistent and reliable power to support the buoy instrumentation. The US National Science Foundation Ocean Observatories Initiative (OOI) includes arrays of point absorber-like buoy systems used for ocean observation that have been deployed at multiple locations including the Southern Mid-Atlantic Bight. A study is currently underway to design a pitch resonator wave energy converter to supplement existing renewable energy generation for powering observation instrumentation. This paper details field measurements from surface moorings of the OOI Coastal Pioneer Array, which informs the subsequent development of a numerical model for the moored observation system. The model is developed in Wave Energy Converter Simulator (WEC-Sim), which leverages the Simscape multibody solver within the MATLAB/Simulink framework and linear potential flow theory to simulate the hydrodynamic interactions and multibody dynamics in 6 degrees of freedom. Multiple tuning variables are considered to produce a model for the system that matches well with empirical data (about 8% error). In conclusion, the WEC-Sim model will serve as a platform for integrating the pitch resonator wave energy converter concept and deployment preparation (detailed design including power take-off and control systems, response evaluation, etc.).

hydrodynamic modeling↗

Noisy-Intermediate-Scale Quantum Electromagnetic Transients Program

Quantum-empowered electromagnetic transients program (QEMTP) is a promising paradigm for tackling EMTP's computational burdens. Nevertheless, no existing studies truly achieve a practical and scalable QEMTP operable on today's noisy-intermediate-scale quantum (NISQ) computers. The strong reliance on noise-free and fault-tolerant quantum devices--which appears to be decades away--hinder practical applications of current QEMTP methods. Here, we devise a NISQ-QEMTP methodology which for the first time transitions the QEMTP operations from ideal, noise-free quantum simulators to real, noisy quantum computers. The main contributions lie in: (1) a shallow-depth QEMTP quantum circuit for mitigating noises on NISQ quantum devices; (2) practical QEMTP linear solvers incorporating executable quantum state preparation and measurements for nodal voltage computations; (3) a noise-resilient QEMTP algorithm leveraging quantum resources logarithmically scaled with power system dimension; (4) a quantum shifted frequency analysis (QSFA) for accelerating QEMTP by exploiting dynamic phasor simulations with larger time steps; (5) a systematical analysis on QEMTPs performance under various noisy quantum environments. Extensive experiments systematically verify the accuracy, efficacy, universality and noise-resilience of QEMTP on both noise-free simulators and IBM real quantum computers.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

ExaWind: Predictive Wind Energy Simulations

This presentation describes the ExaWind project and the team's progress in creating a suite of performance-portable codes designed for predictive simulations of wind farms on next-generation exascale-class supercomputers. Such simulations will require the resolution of scales spanning many orders of magnitude, from blade boundary layers to wind farm flow structures. In the U.S., the first exascale systems will be GPU accelerated, and different GPU manufacturers have been chosen for the different systems. At the heart of the ExaWind software is a hybrid-solver approach based on the codes Nalu-Wind and AMR-Wind, which are computational fluid dynamics solvers for the incompressible Navier-Stokes equations. Nalu-Wind is an unstructured-grid code used to resolve wind turbine geometry and blade boundary layers, whereas AMR-Wind is a structured-grid background solver for atmospheric turbulent flow and turbine wake propagation. The models are coupled with overset meshes and global linear systems are approximated through a loose-coupling algorithm. Results will include validation-quality high-fidelity simulations and strong/weak scaling results from the Summit supercomputer.

49 EE - Wind and Water Power Program - Wind (EE-4W↗

Quantum computing and preconditioners for hydrological linear systems

Modeling hydrological fracture networks is a hallmark challenge in computational earth sciences. Accurately predicting critical features of fracture systems, e.g. percolation, can require solving large linear systems far beyond current or future high performance capabilities. Quantum computers can theoretically bypass the memory and speed constraints faced by classical approaches, however several technical issues must first be addressed. Chief amongst these difficulties is that such systems are often ill-conditioned, i.e. small changes in the system can produce large changes in the solution, which can slow down the performance of linear solving algorithms. We test several existing quantum techniques to improve the condition number, but find they are insufficient. We then introduce the inverse Laplacian preconditioner, which improves the scaling of the condition number of the system from O(N) to O($\sqrt{N}$) and admits a quantum implementation. These results are a critical first step in developing a quantum solver for fracture systems, both advancing the state of hydrological modeling and providing a novel real-world application for quantum linear systems algorithms.

97 MATHEMATICS AND COMPUTING↗

Finite Element Analysis of the TRUST Nonlinear Dynamics Testbed

This paper builds on prior work conducted within the Los Alamos National Laboratory (LANL) Testbeds to Reduce Uncertainty in Simulations and Tests (TRUST) program. Specifically, it builds on finite element (FE) modeling efforts for the TRUST program’s Nonlinear Dynamics (ND) testbed. Historically, the FE model for the ND testbed has exclusively utilized a Lanczos eigensolver that linearly extracts the system’s natural frequencies. This paper investigates the Abaqus 2024’s explicit dynamic solver, which implements a central difference explicit solver. The central difference method used in Abaqus can capture nonlinear material responses in structural dynamic simulations making it suitable for the ND FE model.

42 ENGINEERING↗

A Mixed integer linear programming‐based distributed energy management for networked microgrids considering network operational objectives and constraints

Abstract Mixed integer linear programming (MILP)–based distributed energy management for networked microgrids embedded modern distribution systems is proposed. Considering the diverse ownership of microgrids, distributed energy resources (DERs) that interface directly with utilities and responsive loads, an alternating direction method of multipliers–based distributed framework was formulated for the scheduling of networked microgrids embedded modern distribution systems by adjusting nodal price signals iteratively. In addition, to make the formulated optimization problems resolvable through more accessible and popular MILP solvers, different linearisation techniques were employed to transform the nonlinear terms into linear or mixed integer linear formats. The proposed MILP‐based distributed method preserves all participants' autonomy (e.g., microgrids, DERs that interface directly with utilities and responsive loads), while incentivising them to actively participate in the distribution system operation with price signals. The proposed method is validated with results of numerical simulation using a modern distribution system consisting of multiple networked microgrids, DERs that interface directly with utilities, as well as responsive loads.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Algebraic Multigrid with Filtering: An Efficient Preconditioner for Interior Point Methods in Large-Scale Contact Mechanics Optimization

Large-scale contact mechanics simulations are crucial in many engineering fields such as structural design and manufacturing. In the frictionless case, contact can be modeled by minimizing an energy functional; however, these problems are often nonlinear, nonconvex, and increasingly difficult to solve as mesh resolution increases. In this work, we employ a Newton-based interior-point (IP) filter line-search method, an effective approach for large-scale constrained optimization. While this method converges rapidly, each iteration requires solving a large saddle-point linear system that becomes ill-conditioned as the optimization process converges, largely due to IP treatment of the contact constraints. Such ill-conditioning can hinder solver scalability and increase iteration counts with mesh refinement. Here, to address this, we introduce a novel preconditioner, algebraic multigrid with filtering (AMGF), tailored to the Schur complement of the saddle-point system. Building on the classical AMG solver, commonly used for elasticity, we augment it with a specialized subspace correction that filters near null space components introduced by contact interface constraints. Through theoretical analysis and numerical experiments on a range of linear and nonlinear contact problems, we demonstrate that the proposed solver achieves mesh independent convergence and maintains robustness against the ill-conditioning that notoriously plagues IP methods. These results indicate that AMGF makes contact mechanics simulations more tractable and broadens the applicability of Newton-based IP methods in challenging engineering scenarios. More broadly, AMGF is well suited for problems, optimization or otherwise, where solver performance is limited by a low-dimensional subspace, such as those arising from localized constraints, interface conditions, or model heterogeneities. This makes the method widely applicable beyond contact mechanics and constrained optimization.

Mathematics and Computing↗

symPACK: A GPU-Capable Fan-Out Sparse Cholesky Solver

Sparse symmetric positive definite systems of equations are ubiquitous in scientific workloads and applications. Parallel sparse Cholesky factorization is the method of choice for solving such linear systems. Therefore, the development of parallel sparse Cholesky codes that can efficiently run on today’s large-scale heterogeneous distributed-memory platforms is of vital importance. Modern supercomputers offer nodes that contain a mix of CPUs and GPUs. To fully utilize the computing power of these nodes, scientific codes must be adapted to offload expensive computations to GPUs. We present symPACK, a GPU-capable parallel sparse Cholesky solver that uses one-sided communication primitives and remote procedure calls provided by the UPC++ library. We also utilize the UPC++ "memory kinds" feature to enable efficient communication of GPU-resident data. We show that on a number of large problems, symPACK outperforms comparable state-of-the-art GPU-capable Cholesky factorization codes by up to 14x on the NERSC Perlmutter supercomputer.

Bellavita, Julian↗

Potential quantum advantage for simulation of fluid dynamics

Numerical simulation of turbulent fluid dynamics needs to either parametrize turbulence—which introduces large uncertainties—or explicitly resolve the smallest scales—which is prohibitively expensive. Here, we provide evidence through analytic bounds and numerical studies that a potential quantum speedup can be achieved to simulate fluid dynamics using quantum computing. Specifically, we provide a lattice Boltzmann formulation of fluid dynamics for which we give evidence that low-order Carleman linearization is much more accurate than previously believed for these systems. This is achieved via a combination of reformulating the Navier-Stokes nonlinearity (u·$\triangledown$u) to lattice-Boltzmann nonlinearity (u 2 ) and accurately linearizing the dynamical equations, which effectively trades nonlinearity for additional degrees of freedom that add negligible expense in the quantum solver. Based on this, we apply a quantum algorithm for simulating the Carleman-linearized lattice Boltzmann equation and provide evidence that its cost scales logarithmically with system size compared with polynomial scaling in the best known classical algorithms. In this paper, we suggest that a quantum advantage may exist for simulating fluid dynamics, paving the way for simulating nonlinear multiscale transport phenomena in a wide range of disciplines using quantum computing.

42 ENGINEERING↗

A fast implicit solver for semiconductor models in one space dimension

Several different approaches are proposed for solving fully implicit discretizations of a simplified Boltzmann-Poisson system with a linear relaxation-type collision kernel. This system models the evolution of free electrons in semiconductor devices under a low-density assumption. At each implicit time step, the discretized system is formulated as a fixed-point problem, which can then be solved with a variety of methods. A key algorithmic component in all the approaches considered here is a recently developed sweeping algorithm for Vlasov-Poisson systems. A synthetic acceleration scheme has been implemented to accelerate the convergence of iterative solvers by using the solution to a drift-diffusion equation as a preconditioner. The performance of four iterative solvers and their accelerated variants has been compared on problems modeling semiconductor devices with various electron mean-free-path.

97 MATHEMATICS AND COMPUTING↗

Development of a continuous synthesis process for carbamazepine using validated in-line Raman spectroscopy and kinetic modelling for disturbance simulation

Mitigation of failure modes in the continuous synthesis (CS) of a drug substance (DS) has the potential to widen the adoption of continuous manufacturing (CM) technologies by the pharmaceutical industry. Here, this work demonstrates the development of a robust continuous process for the synthesis of carbamazepine (CBZ), an essential medicine as per the World Health Organization (WHO), facilitated by kinetic modelling and monitored by in-line Raman spectroscopy. Accurate kinetic modelling and the use of validated process analytical technology (PAT) models for quantitative measurement were found to play an important role in developing CS of drug substances. Kinetic data for the formation of CBZ from iminostilbene (ISB) were collected by batch reaction sampling and high-performance liquid chromatography (HPLC) analysis. A non-linear solver and iterative method was applied to determine two sets of Arrhenius parameters simultaneously for the reaction system by minimizing the standard error of the model fit. The start-up and dynamic equilibrium stages for the CS of CBZ using a continuous stirred tank reactor (CSTR) were modelled based on the batch kinetic data and employed to optimize conversion and simulate process disturbances. An in-line Raman spectroscopy method was successfully developed, validated, and integrated to determine the concentrations of CBZ and ISB within the operating range for the CS. The CS kinetic model was evaluated experimentally from startup to dynamic equilibrium over 10 residence times with monitoring by HPLC and in-line Raman spectroscopy. The developed kinetic model in tandem with in-line Raman spectroscopy successfully predicted disturbances due to changes in process variables and can serve as a useful tool in the future design of advanced process control strategies for the continuous synthesis of CBZ.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Combining Sparse Approximate Factorizations with Mixed-precision Iterative Refinement

The standard LU factorization-based solution process for linear systems can be enhanced in speed or accuracy by employing mixed-precision iterative refinement. Most recent work has focused on dense systems. We investigate the potential of mixed-precision iterative refinement to enhance methods for sparse systems based on approximate sparse factorizations. In doing so, we first develop a new error analysis for LU- and GMRES-based iterative refinement under a general model of LU factorization that accounts for the approximation methods typically used by modern sparse solvers, such as low-rank approximations or relaxed pivoting strategies. We then provide a detailed performance analysis of both the execution time and memory consumption of different algorithms, based on a selected set of iterative refinement variants and approximate sparse factorizations. Our performance study uses the multifrontal solver MUMPS, which can exploit block low-rank factorization and static pivoting. We evaluate the performance of the algorithms on large, sparse problems coming from a variety of real-life and industrial applications showing that mixed-precision iterative refinement combined with approximate sparse factorization can lead to considerable reductions of both the time and memory consumption.

97 MATHEMATICS AND COMPUTING↗

Milestone 49 Report: Batched Sparse LA Phase 5 Implementation

Batched sparse linear algebra operations in general, and solvers in particular, have become the major algorithmic development activity and foremost performance engineering effort in the numerical software libraries work on modern hardware with accelerators such as GPUs. Many applications, ECP and non-ECP alike, require simultaneous solutions of many small linear systems of equations that are structurally sparse in one form or another. In order to move towards high hardware utilization levels, it is important to provide these applications with appropriate interface designs to be both functionally efficient and performance portable and give full access to the appropriate batched sparse solvers running on modern hardware accelerators prevalent across DOE supercomputing sites since the inception of ECP. To this end, we present here a summary of recent advances on the interface designs in use by HPC software libraries supporting batched sparse linear algebra and the development of sparse batched kernel codes for solvers and preconditioners. We also address the potential interoperability opportunities to keep the corresponding software portable between the major hardware accelerators from AMD, Intel, and NVIDIA, while maintaining the appropriate disclosure levels conforming to the active NDA agreements. The presented interface specifications include a mix of batched band, sparse iterative, and sparse direct solvers with their accompanying functionality that is already required by the application codes or we anticipated to be needed in the near future. This report summarizes progress in Kokkos Kernels and the xSDK libraries MAGMA, Ginkgo, hypre, PETSc, and SuperLU.

97 MATHEMATICS AND COMPUTING↗