Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Linear systems solvers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Aerodynamic shape optimization via sensitivity analysis on decomposed computational domains

Direct and iterative method considered to be most applicable to large systems of linear equations arising in discrete sensitivity analysis are assessed. Based on a single-domain grid, computations are performed using a banded matrix solver and an iterative solver, the generalized minimum residual (GMRES) method. The banded matrix solver is found to be generally the most economical method for those applications where the number of right-hand sides is large (i.e., a large number of design variables or a large number of adjoint vectors). For systems of equations that are too large to be solved by direct methods, an approach is proposed whereby the computational domain is divided into small subdomains, and each subdomain is solved separately.

Eleshaky, Mohamed E.↗

Multigrid approaches to non-linear diffusion problems on unstructured meshes

The efficiency of three multigrid methods for solving highly non-linear diffusion problems on two-dimensional unstructured meshes is examined. The three multigrid methods differ mainly in the manner in which the nonlinearities of the governing equations are handled. These comprise a non-linear full approximation storage (FAS) multigrid method which is used to solve the non-linear equations directly, a linear multigrid method which is used to solve the linear system arising from a Newton linearization of the non-linear system, and a hybrid scheme which is based on a non-linear FAS multigrid scheme, but employs a linear solver on each level as a smoother. Results indicate that all methods are equally effective at converging the non-linear residual in a given number of grid sweeps, but that the linear solver is more efficient in cpu time due to the lower cost of linear versus non-linear grid sweeps.

Mavriplis, Dimitri J.↗

Combining Sparse Approximate Factorizations with Mixed-precision Iterative Refinement

The standard LU factorization-based solution process for linear systems can be enhanced in speed or accuracy by employing mixed-precision iterative refinement. Most recent work has focused on dense systems. We investigate the potential of mixed-precision iterative refinement to enhance methods for sparse systems based on approximate sparse factorizations. In doing so, we first develop a new error analysis for LU- and GMRES-based iterative refinement under a general model of LU factorization that accounts for the approximation methods typically used by modern sparse solvers, such as low-rank approximations or relaxed pivoting strategies. We then provide a detailed performance analysis of both the execution time and memory consumption of different algorithms, based on a selected set of iterative refinement variants and approximate sparse factorizations. Our performance study uses the multifrontal solver MUMPS, which can exploit block low-rank factorization and static pivoting. We evaluate the performance of the algorithms on large, sparse problems coming from a variety of real-life and industrial applications showing that mixed-precision iterative refinement combined with approximate sparse factorization can lead to considerable reductions of both the time and memory consumption.

97 MATHEMATICS AND COMPUTING↗

Milestone 49 Report: Batched Sparse LA Phase 5 Implementation

Batched sparse linear algebra operations in general, and solvers in particular, have become the major algorithmic development activity and foremost performance engineering effort in the numerical software libraries work on modern hardware with accelerators such as GPUs. Many applications, ECP and non-ECP alike, require simultaneous solutions of many small linear systems of equations that are structurally sparse in one form or another. In order to move towards high hardware utilization levels, it is important to provide these applications with appropriate interface designs to be both functionally efficient and performance portable and give full access to the appropriate batched sparse solvers running on modern hardware accelerators prevalent across DOE supercomputing sites since the inception of ECP. To this end, we present here a summary of recent advances on the interface designs in use by HPC software libraries supporting batched sparse linear algebra and the development of sparse batched kernel codes for solvers and preconditioners. We also address the potential interoperability opportunities to keep the corresponding software portable between the major hardware accelerators from AMD, Intel, and NVIDIA, while maintaining the appropriate disclosure levels conforming to the active NDA agreements. The presented interface specifications include a mix of batched band, sparse iterative, and sparse direct solvers with their accompanying functionality that is already required by the application codes or we anticipated to be needed in the near future. This report summarizes progress in Kokkos Kernels and the xSDK libraries MAGMA, Ginkgo, hypre, PETSc, and SuperLU.

97 MATHEMATICS AND COMPUTING↗

Scalable Predictive Control and Optimization for Grid Integration of Large-Scale Distributed Energy Resources

Integrating a large number of distributed energy resources (DERs) into the power grid needs a scalable power balancing method. We formulate the power balancing problem as a look-ahead optimization problem to be solved sequentially by a power distribution system aggregator based on a model predictive control (MPC) framework. Solving large-scale look-ahead control problems requires proper configuration of the control steps. In this paper, to solve large-scale control problems, we propose a variable time granularity where control time steps nearby the current control step have finer resolutions. The aggregator objective includes maximization of power production revenue and minimization of power purchasing expense, renewable power curtailment, and mileage costs for energy storage and electric vehicle (EV) charging stations while satisfying system capacity and operational constraints. The control problem is formulated as a mixed-integer linear program (MILP) and solved using the XpressMP solver. We perform simulations considering a copper plate representation of a large distribution network consisting of 2507 devices (controllable DERs), including curtailable photovoltaics (PVs), energy storage batteries, EV charging stations, and buildings with heating, ventilation, and air conditioning units (HVACs). We show the effectiveness of the proposed approach in managing DERs interactively for maximum energy trading profit and local supply-demand power balancing. Finally, we demonstrate that the proposed method outperforms other benchmark controllers regarding computation time without compromising operational performance.

DER↗

Newton solution of inviscid and viscous problems

The application of Newton iteration to inviscid and viscous airfoil calculations is examined. Spatial discretization is performed using upwind differences with split fluxes. The system of linear equations which arises as a result of linearization in time is solved directly using either a banded matrix solver or a sparse matrix solver. In the latter case, the solver is used in conjunction with the nested dissection strategy, whose implementation for airfoil calculations is discussed. The boundary conditions are also implemented in a fully implicit manner, thus yielding quadratic convergence. Complexities such as the ordering of cell nodes and the use of a far field vortex to correct freestream for a lifting airfoil are addressed. Various methods to accelerate convergence and improve computational efficiency while using Newton iteration are discussed. Results are presented for inviscid, transonic nonlifting and lifting airfoils and also for laminar viscous cases.

Venkatakrishnan, V.↗

Sparse Linear Solvers for Large-scale Electromagnetic Transient Simulations

Linear solvers form the basis for electromagnetic transient (EMT) simulations. There is a need to speed up EMT simulations as larger regions are analyzed using EMT simulations. For the same, the performance of linear solvers plays an important role. Exploiting the sparsity of the matrices generated in EMT simulations could assist with speed-up. Scalability is also crucial as power grids expand, demanding solutions capable of accommodating the increasing system size. Recent studies from the North American Electric Reliability Corporation (NERC) increasingly emphasize that EMT simulation models of the power grid will grow larger with the inclusion of power electronics components. Parallelisms in sparsity patterns exploit modern central processing units (CPUs), multi-core CPUs, and graphics processing units (GPUs) architectures in sparse solver designs. Therefore, this paper explores publicly available existing linear solvers and investigates their efficiency in large-scale power grid simulations. A large-scale power grid is developed by increasing the size of the IEEE 39 bus test system to up to 39000 bus systems.

Hsu, Kuan-Chieh↗

Performance issues for iterative solvers in device simulation

Due to memory limitations, iterative methods have become the method of choice for large scale semiconductor device simulation. However, it is well known that these methods still suffer from reliability problems. The linear systems which appear in numerical simulation of semiconductor devices are notoriously ill-conditioned. In order to produce robust algorithms for practical problems, careful attention must be given to many implementation issues. This paper concentrates on strategies for developing robust preconditioners. In addition, effective data structures and convergence check issues are also discussed. These algorithms are compared with a standard direct sparse matrix solver on a variety of problems.

Fan, Qing↗

Pyomo.GDP: an ecosystem for logic based modeling and optimization development

We present three core principles for engineering-oriented integrated modeling and optimization tool sets—intuitive modeling contexts, systematic computer-aided reformulations, and flexible solution strategies—and describe how new developments in Pyomo.GDP for Generalized Disjunctive Programming (GDP) advance this vision. We describe a new logical expression system implementation for Pyomo.GDP allowing for a more intuitive description of logical propositions. The logical expression system supports automated reformulation of these logical constraints to linear constraints. We also describe two new logic-based global optimization solver implementations built on Pyomo.GDP that exploit logical structure to avoid “zero-flow” numerical difficulties that arise in nonlinear network design problems when nodes or streams disappear. These new solvers also demonstrate the capability to link to external libraries for expanded functionality within an integrated implementation. We present these new solvers in the context of a flexible array of solution paths available to GDP models. Finally, we present results on a new library of GDP models demonstrating the value of multiple solution approaches.

42 ENGINEERING↗

Towards exascale for wind energy simulations

We examine large-eddy-simulation modeling approaches and computational performance of two open-source computational fluid dynamics codes for the simulation of atmospheric boundary layer flows that are of direct relevance to wind energy production. The first code, NekRS, is a high-order, unstructured-grid, spectral element code. The second code, AMR-Wind, is a second-order, block-structured, finite-volume code with adaptive mesh refinement capabilities. The objective of this study is to co-develop these codes in order to improve model fidelity and performance for each. These features will be critical for running ABL-based applications such as wind farm analysis on advanced computing architectures. To this end, we investigate the performance of NekRS and AMR-Wind on the Oak Ridge Leadership Facility supercomputers Summit, using 4 to 800 nodes (24 to 4,800 NVIDIA V100 GPUs), and Crusher, the testbed for the Frontier exascale system, using 18 to 384 Graphics Compute Dies on AMD MI250X GPUs. We compare strong- and weak-scaling capabilities, linear solver performance, and time to solution. We also identify leading inhibitors to parallel scaling.

17 WIND ENERGY↗

Optimization to Generate Equations of State for Hydrogen Production

On a high level, the larger project in question, HydroGEN, aims to develop software used for finding equations of state (EOS) to optimize catalyst configuration for H 2 production through water splitting. In particular, this summer project focused on solving the nonlinear equations used in fitting the equations. This problem involved using Python to solve a linear system with nonlinear constraints. In order for this to be achieved, Pyomo was used to build a model and the solver Ipopt, interior point optimizer, was used. Pyomo is a Python-based language developed at Sandia; it is an optimization modeling language. Rather than solving the entire problem at once, a toy problem was created, simplifying the problem down to the most important focus. This problem had a known solution, comparable to the calculated solution to assess accuracy and as progress was made towards finding solutions, complexity was gradually added to the problem. After building and solving the toy problem, it was found that it gave reasonably accurate solutions, better compared to the two existing solvers previously used with this project in terms of functionality. The solver is now ready for implementation into the project’s main software.

08 HYDROGEN↗

Component-Based Development of CFD Software FUN3D

FUN3D, a suite of Computational Fluid Dynamics simulation and design tools developed at the NASA Langley Research Center, has undergone continuous development since the late1980s. It contains a large portion of legacy code. Extending it with new capabilities becomes increasingly difficult. To improve the extensibility and reusability, FUN3D is moving toward component-based development. New features, such as Stabilized Finite Elements, Yoga, and Sparse Linear Algebra Toolkit, are integrated into the system as components. Some existing features such as the Node-Centered Finite Volume Solver, are also being refactored to components. The integration of these components poses new requirements on the development workflow. In this paper, we describe the Continuous Integration of FUN3D to support component-based development, and discuss the tools used, the practices followed, and lessons learned during the transition from the traditional approach.

computational fluid dynamics software↗

Reduced-Order Aerodynamic Modeling Based on CFD Frequency Responses from Multisine Inputs

A system identification analysis was performed to determine a reduced-order model (ROM) of a computational fluid dynamics (CFD) solver in support of linear aeroservoelastic model development and feedback control design. The approach was applied to the FUN3D code for the half-span wind tunnel test article used in the NASA-Boeing collaboration called the Integrated Adaptive Wing Technology Maturation (IAWTM) project. In a transonic flow condition, multiple inputs (11 structural mode displacements and 3 control surface deflections) were simultaneously excited with orthogonal phase-optimized multisines while multiple outputs (the corresponding 14 generalized aerodynamic forces) were recorded. From these recorded times series, the matrix of frequency responses was computed and subsequently fit using rational function approximations (RFAs). It was found that the entire (14 x 14) matrix of frequency responses could be determined from a single CFD run and that results generally followed trends predicted using other methods. Differences were attributed to the modeling fidelity and nonlinearities from structural mode and control surface interactions at higher reduced frequencies. More accurate fits of the RFAs to the frequency response data were obtained by making two CFD runs, one with only structural mode excitations and one with only control surface excitations, which reduced the degree of nonlinearity in the modeling data.

Aeroservoelasticity↗

Consistent Second Moment Methods with Scalable Linear Solvers for Radiation Transport

Second moment methods (SMMs) are developed that are consistent with the discontinuous Galerkin spatial discretization of the discrete ordinates (or S\(_N\)) transport equations. The low-order (LO) diffusion system of equations is discretized with fully consistent P\(_1\), local discontinuous Galerkin (LDG), and interior penalty (IP) methods. A discrete residual approach is used to derive SMM correction terms that make each of the LO systems consistent with the high-order discretization. We show that the consistent methods are more accurate and have better solution quality than independently discretized LO systems, that they preserve the diffusion limit, and that the LDG and IP consistent SMMs can be scalably solved in parallel on a challenging, multimaterial benchmark problem.

97 MATHEMATICS AND COMPUTING↗

Domain decomposition methods for the parallel computation of reacting flows

Domain decomposition is a natural route to parallel computing for partial differential equation solvers. Subdomains of which the original domain of definition is comprised are assigned to independent processors at the price of periodic coordination between processors to compute global parameters and maintain the requisite degree of continuity of the solution at the subdomain interfaces. In the domain-decomposed solution of steady multidimensional systems of PDEs by finite difference methods using a pseudo-transient version of Newton iteration, the only portion of the computation which generally stands in the way of efficient parallelization is the solution of the large, sparse linear systems arising at each Newton step. For some Jacobian matrices drawn from an actual two-dimensional reacting flow problem, comparisons are made between relaxation-based linear solvers and also preconditioned iterative methods of Conjugate Gradient and Chebyshev type, focusing attention on both iteration count and global inner product count. The generalized minimum residual method with block-ILU preconditioning is judged the best serial method among those considered, and parallel numerical experiments on the Encore Multimax demonstrate for it approximately 10-fold speedup on 16 processors.

Keyes, David E.↗

NREL Price Series Developed for the ARPA-E FLECCS Program

The price data for four regions (CAISO, ERCOT, MISO-W, and PJM-W) are developed using the ReEDS to PLEXOS conversion as described in (Gagnon et al. 2020). The reference ReEDS case chosen is based on the 2020 Standard Scenario Mid-case, which uses the 2020 ReEDS model version (Cole et al. 2020; Ho et al. 2021). All ReEDS model inputs use 2020 Standard Scenarios Mid-case assumptions except for CO2 prices, which are implemented as linearly increasing CO2 price trajectories beginning at $0/tCO2 in 2020 and ending at either $100/tCO2 or $150/tCO2 in 2035 to dive capacity expansion towards a low-carbon system that could support CCS deployment. However, these scenarios also prohibit CCS deployment in this time frame so that resulting price data are not influenced by the deployment and operation of CCS itself. Implementing the ReEDS to PLEXOS conversion tool, PLEXOS is then simulated using the 2035 ReEDS infrastructure for both CO2 price scenarios, with the following model version and setup: PLEXOS Version: 8.2 Solver: Xpress-MP 35.01.01 Mixed integer optimization relative gap 1% System configuration: • Total number of nodes: 134 (consistent with ReEDS balancing areas) • Line losses enforced using piecewise linear approximation • Energy dump was enabled Price data is aggregated to the ISO/RTO level using load-weighted averages. References: Cole, Wesley, Sean Corcoran, Nathaniel Gates, Daniel Mai, Trieu, and Paritosh Das. 2020. “2020 Standard Scenarios Report: A U.S. Electricity Sector Outlook.” NREL/TP-6A20-77442. Golden, CO: National Renewable Energy Laboratory. https://www.nrel.gov/docs/fy21osti/77442.pdf. Gagnon, Pieter, Will Frazier, Elaine Hale, and Wesley Cole. 2020. “Cambium Documentation: Version 2020.” NREL/TP-6A20-78239. National Renewable Energy Lab. (NREL), Golden, CO (United States). https://doi.org/10.2172/1734551. Ho, Jonathan, Jonathon Becker, Maxwell Brown, Patrick Brown, Ilya (ORCID:0000000284917814) Chernyakhovskiy, Stuart Cohen, Wesley (ORCID:000000029194065X) Cole, et al. 2021. “Regional Energy Deployment System (ReEDS) Model Documentation: Version 2020.” NREL/TP-6A20-78195. Golden, CO: National Renewable Energy Laboratory. https://doi.org/10.2172/1788425.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Design of a Modular Monolithic Implicit Solver for Multi-Physics Applications

The design of a modular multi-physics high-order space-time finite-element framework is presented together with its extension to allow monolithic coupling of different physics. One of the main objectives of the framework is to perform efficient high- fidelity simulations of capsule/parachute systems. This problem requires simulating multiple physics including, but not limited to, the compressible Navier-Stokes equations, the dynamics of a moving body with mesh deformations and adaptation, the linear shell equations, non-re effective boundary conditions and wall modeling. The solver is based on high-order space-time - finite element methods. Continuous, discontinuous and C1-discontinuous Galerkin methods are implemented, allowing one to discretize various physical models. Tangent and adjoint sensitivity analysis are also targeted in order to conduct gradient-based optimization, error estimation, mesh adaptation, and flow control, adding another layer of complexity to the framework. The decisions made to tackle these challenges are presented. The discussion focuses first on the "single-physics" solver and later on its extension to the monolithic coupling of different physics. The implementation of different physics modules, relevant to the capsule/parachute system, are also presented. Finally, examples of coupled computations are presented, paving the way to the simulation of the full capsule/parachute system.

Carton De Wiart, Corentin↗