Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Linear systems solvers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Optimization of a Solver for Computational Materials and Structures Problems on NVIDIA Volta and AMD Instinct GPUs

The Scalable Implementation of Finite Elements by NASA (ScIFEN) is a software package developed to solve complex computational materials and structures problems using the finite element method (FEM). In this paper, we describe optimization techniques to speed up the linear solver computation that occurs within the ScIFEN application. We consider GPUs from two different vendors, NVIDIA and AMD as our target platforms for optimization and highlight differences in performance and optimization techniques. The NVIDIA GPU Volta V100 is used in the Summit system deployed at Oak Ridge National Laboratory, and the new exascale system, Frontier, will be using AMD Radeon Instinct GPU. We evaluated the performance of various optimization techniques on test matrices, ranging in size from100K to 4M, that are representative of ScIFEN applications. The linear solver computation is memory-bound on both GPUs. Our experiments show that on the NVIDIA GPU we obtained up to79%of the theoretical peak bandwidth, while the AMD GPU achieved 59%. Overall, the NVIDIA V100 GPU outperforms the AMD MI 25 GPU1. We observed an overall speedup of up to37X on an NVIDIA V100 compared to an Intel Skylake 12-coremachine. The solver for a 4M degree of freedom system took under 2.5 seconds.

Mohammad Zubair↗

Evaluation of out-of-core computer programs for the solution of symmetric banded linear equations

FORTRAN coded out-of-core equation solvers that solve using direct methods symmetric banded systems of simultaneous algebraic equations. Banded, frontal and column (skyline) solvers were studied as well as solvers that can partition the working area and thus could fit into any available core. Comparison timings are presented for several typical two dimensional and three dimensional continuum type grids of elements with and without midside nodes. Extensive conclusions are also given.

Dunham, R. S.↗

Parallel variable-band Choleski solvers for computational structural analysis applications on vector multiprocessor supercomputers

A Choleski method used to solve linear systems of equations that arise in large scale structural analyses is described. The method uses a novel variable-band storage scheme and is structured to exploit fast local memory caches while minimizing data access delays between main memory and vector registers. Several parallel implementations of this method are described for the CRAY-2 and CRAY Y-MP computers demonstrating the use of microtasking and autotasking directives. A portable parallel language, FORCE, is also used for two different parallel implementations, demonstrating the use of CRAY macrotasking. Results are presented comparing the matrix factorization times for three representative structural analysis problems from runs made in both dedicated and multi-user modes on both the CRAY-2 and CRAY Y-MP computers. CPU and wall clock timings are given for the various parallel methods and are compared to single processor timings of the same algorithm. Computation rates over 1 GIGAFLOP (1 billion floating point operations per second) on a four processor CRAY-2 and over 2 GIGAFLOPS on an eight processor CRAY Y-MP are demonstrated as measured by wall clock time in a dedicated environment. Reduced wall clock times for the parallel methods relative to the single processor implementation of the same Choleski algorithm are also demonstrated for runs made in multi-user mode.

Poole, E. L.↗

Structural Analysis Using NX Nastran 9.0

NX Nastran is a powerful Finite Element Analysis (FEA) software package used to solve linear and non-linear models for structural and thermal systems. The software, which consists of both a solver and user interface, breaks down analysis into four files, each of which are important to the end results of the analysis. The software offers capabilities for a variety of types of analysis, and also contains a respectable modeling program. Over the course of ten weeks, I was trained to effectively implement NX Nastran into structural analysis and refinement for parts of two missions at NASA's Kennedy Space Center, the Restore mission and the Orion mission.

NX Nastran↗

Predicting Flow in Fracture Networks With Quantum Algorithms

Uncertainty quantification plays a crucial role in the modeling of subsurface flow. For instance, uncertainties in the properties of geologic fracture networks significantly impact flow, requiring numerous simulations to accurately estimate quantities of interest. However, each simulation is computationally expensive because it requires solving a large linear system to capture features that involve both small and large fractures. An example is in percolation, where the interaction of many small fractures (which cumulatively can have a large surface area) with the rock matrix must be modeled precisely. Quantum computing is an emerging tool with the potential to address this issue. Quantum algorithms offer a significant speedup in solving linear systems, achieving efficiencies that are challenging to match with classical approaches. These classical approaches include direct solvers, such as LU decomposition, and iterative methods, notably preconditioned conjugate gradient, commonly used in subsurface modeling to solve large sparse systems. However, applying quantum algorithms to geologic fracture flow requires careful attention to algorithmic and problem-specific constraints to fully realize this quantum advantage. In this work we describe a quantum algorithm for generalized Monte Carlo applications with a quadratic speedup over the classical approaches which can be combined with the quantum speedup, currently under investigation, for solving quantum linear systems for subsurface flow. We show that for quantum algorithms the computational cost of estimating a quantity of interest for a statistical ensemble of networks is roughly the same as that of a single realization, essentially implying that one can get uncertainty quantification for free.

58 GEOSCIENCES↗

Hydrodynamic characterization of the coastal pioneer array ocean observing system

Ocean observation buoys require relatively small amounts of power, yet traditionally necessitate costly resupply trips for battery replacement. With the offshore location of the buoys and small power requirements, wave energy may be an effective solution for providing consistent and reliable power to support the buoy instrumentation. The US National Science Foundation Ocean Observatories Initiative (OOI) includes arrays of point absorber-like buoy systems used for ocean observation that have been deployed at multiple locations including the Southern Mid-Atlantic Bight. A study is currently underway to design a pitch resonator wave energy converter to supplement existing renewable energy generation for powering observation instrumentation. This paper details field measurements from surface moorings of the OOI Coastal Pioneer Array, which informs the subsequent development of a numerical model for the moored observation system. The model is developed in Wave Energy Converter Simulator (WEC-Sim), which leverages the Simscape multibody solver within the MATLAB/Simulink framework and linear potential flow theory to simulate the hydrodynamic interactions and multibody dynamics in 6 degrees of freedom. Multiple tuning variables are considered to produce a model for the system that matches well with empirical data (about 8% error). In conclusion, the WEC-Sim model will serve as a platform for integrating the pitch resonator wave energy converter concept and deployment preparation (detailed design including power take-off and control systems, response evaluation, etc.).

hydrodynamic modeling↗

ExaWind: Predictive Wind Energy Simulations

This presentation describes the ExaWind project and the team's progress in creating a suite of performance-portable codes designed for predictive simulations of wind farms on next-generation exascale-class supercomputers. Such simulations will require the resolution of scales spanning many orders of magnitude, from blade boundary layers to wind farm flow structures. In the U.S., the first exascale systems will be GPU accelerated, and different GPU manufacturers have been chosen for the different systems. At the heart of the ExaWind software is a hybrid-solver approach based on the codes Nalu-Wind and AMR-Wind, which are computational fluid dynamics solvers for the incompressible Navier-Stokes equations. Nalu-Wind is an unstructured-grid code used to resolve wind turbine geometry and blade boundary layers, whereas AMR-Wind is a structured-grid background solver for atmospheric turbulent flow and turbine wake propagation. The models are coupled with overset meshes and global linear systems are approximated through a loose-coupling algorithm. Results will include validation-quality high-fidelity simulations and strong/weak scaling results from the Summit supercomputer.

49 EE - Wind and Water Power Program - Wind (EE-4W↗

Quantum computing and preconditioners for hydrological linear systems

Modeling hydrological fracture networks is a hallmark challenge in computational earth sciences. Accurately predicting critical features of fracture systems, e.g. percolation, can require solving large linear systems far beyond current or future high performance capabilities. Quantum computers can theoretically bypass the memory and speed constraints faced by classical approaches, however several technical issues must first be addressed. Chief amongst these difficulties is that such systems are often ill-conditioned, i.e. small changes in the system can produce large changes in the solution, which can slow down the performance of linear solving algorithms. We test several existing quantum techniques to improve the condition number, but find they are insufficient. We then introduce the inverse Laplacian preconditioner, which improves the scaling of the condition number of the system from O(N) to O($\sqrt{N}$) and admits a quantum implementation. These results are a critical first step in developing a quantum solver for fracture systems, both advancing the state of hydrological modeling and providing a novel real-world application for quantum linear systems algorithms.

97 MATHEMATICS AND COMPUTING↗

Finite Element Analysis of the TRUST Nonlinear Dynamics Testbed

This paper builds on prior work conducted within the Los Alamos National Laboratory (LANL) Testbeds to Reduce Uncertainty in Simulations and Tests (TRUST) program. Specifically, it builds on finite element (FE) modeling efforts for the TRUST program’s Nonlinear Dynamics (ND) testbed. Historically, the FE model for the ND testbed has exclusively utilized a Lanczos eigensolver that linearly extracts the system’s natural frequencies. This paper investigates the Abaqus 2024’s explicit dynamic solver, which implements a central difference explicit solver. The central difference method used in Abaqus can capture nonlinear material responses in structural dynamic simulations making it suitable for the ND FE model.

42 ENGINEERING↗

A Mixed integer linear programming‐based distributed energy management for networked microgrids considering network operational objectives and constraints

Abstract Mixed integer linear programming (MILP)–based distributed energy management for networked microgrids embedded modern distribution systems is proposed. Considering the diverse ownership of microgrids, distributed energy resources (DERs) that interface directly with utilities and responsive loads, an alternating direction method of multipliers–based distributed framework was formulated for the scheduling of networked microgrids embedded modern distribution systems by adjusting nodal price signals iteratively. In addition, to make the formulated optimization problems resolvable through more accessible and popular MILP solvers, different linearisation techniques were employed to transform the nonlinear terms into linear or mixed integer linear formats. The proposed MILP‐based distributed method preserves all participants' autonomy (e.g., microgrids, DERs that interface directly with utilities and responsive loads), while incentivising them to actively participate in the distribution system operation with price signals. The proposed method is validated with results of numerical simulation using a modern distribution system consisting of multiple networked microgrids, DERs that interface directly with utilities, as well as responsive loads.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A study of equation solvers for linear and non-linear finite element analysis on parallel processing computers

Concurrent computing environments provide the means to achieve very high performance for finite element analysis of systems, provided the algorithms take advantage of multiple processors. The authors have examined several algorithms for both linear and nonlinear finite element analysis. The performance of these algorithms on an Alliant FX/80 parallel supercomputer has been studied. For single load case linear analysis, the optimal solution algorithm is strongly problem dependent. For multiple load cases or nonlinear analysis through a modified Newton-Raphson method, decomposition algorithms are shown to have a decided advantage over element-by-element preconditioned conjugate gradient algorithms.

Watson, Brian C.↗

Complex generalized minimal residual algorithm for iterative solution of quantum-mechanical reactive scattering equations

Complex dense matrices corresponding to the D + H2 and O + HD reactions were solved using a complex generalized minimal residual (GMRes) algorithm described by Saad and Schultz (1986) and Saad (1990). To provide a test case with a different structure, the H + H2 system was also considered. It is shown that the computational effort for solutions with the GMRes algorithm depends on the dimension of the linear system, the total energy of the scattering problem, and the accuracy criterion. In several cases with dimensions in the range 1110-5632, the GMRes algorithm outperformed the LAPACK direct solver, with speedups for the linear equation solution as large as a factor of 23.

Chatfield, David C.↗

Algebraic Multigrid with Filtering: An Efficient Preconditioner for Interior Point Methods in Large-Scale Contact Mechanics Optimization

Large-scale contact mechanics simulations are crucial in many engineering fields such as structural design and manufacturing. In the frictionless case, contact can be modeled by minimizing an energy functional; however, these problems are often nonlinear, nonconvex, and increasingly difficult to solve as mesh resolution increases. In this work, we employ a Newton-based interior-point (IP) filter line-search method, an effective approach for large-scale constrained optimization. While this method converges rapidly, each iteration requires solving a large saddle-point linear system that becomes ill-conditioned as the optimization process converges, largely due to IP treatment of the contact constraints. Such ill-conditioning can hinder solver scalability and increase iteration counts with mesh refinement. Here, to address this, we introduce a novel preconditioner, algebraic multigrid with filtering (AMGF), tailored to the Schur complement of the saddle-point system. Building on the classical AMG solver, commonly used for elasticity, we augment it with a specialized subspace correction that filters near null space components introduced by contact interface constraints. Through theoretical analysis and numerical experiments on a range of linear and nonlinear contact problems, we demonstrate that the proposed solver achieves mesh independent convergence and maintains robustness against the ill-conditioning that notoriously plagues IP methods. These results indicate that AMGF makes contact mechanics simulations more tractable and broadens the applicability of Newton-based IP methods in challenging engineering scenarios. More broadly, AMGF is well suited for problems, optimization or otherwise, where solver performance is limited by a low-dimensional subspace, such as those arising from localized constraints, interface conditions, or model heterogeneities. This makes the method widely applicable beyond contact mechanics and constrained optimization.

Mathematics and Computing↗

symPACK: A GPU-Capable Fan-Out Sparse Cholesky Solver

Sparse symmetric positive definite systems of equations are ubiquitous in scientific workloads and applications. Parallel sparse Cholesky factorization is the method of choice for solving such linear systems. Therefore, the development of parallel sparse Cholesky codes that can efficiently run on today’s large-scale heterogeneous distributed-memory platforms is of vital importance. Modern supercomputers offer nodes that contain a mix of CPUs and GPUs. To fully utilize the computing power of these nodes, scientific codes must be adapted to offload expensive computations to GPUs. We present symPACK, a GPU-capable parallel sparse Cholesky solver that uses one-sided communication primitives and remote procedure calls provided by the UPC++ library. We also utilize the UPC++ "memory kinds" feature to enable efficient communication of GPU-resident data. We show that on a number of large problems, symPACK outperforms comparable state-of-the-art GPU-capable Cholesky factorization codes by up to 14x on the NERSC Perlmutter supercomputer.

Bellavita, Julian↗

Reduced-Order Aerodynamic Modeling Based on CFD Frequency Responses from Multisine Inputs

A system identification analysis was performed to determine reduced-order models of a computational fluid dynamics (CFD) solver for linear aeroelastic analysis and control design. The application was to the FUN3D code and the flexible half-span wind tunnel test article, in transonic flow conditions, used in the NASA-Boeing collaboration called the Integrated Adaptive Wing Technology Maturation (IAWTM) project. Multiple inputs (structural mode displacements and control surface deflections) were simultaneously excited with orthogonal phase-optimized multisines and multiple outputs (generalized aerodynamic forces) were recorded, from which the matrix of frequency responses were computed using a single CFD run. A state-space model was then fit to the frequency response data using a maximum-likelihood estimator.

System identification↗

Potential quantum advantage for simulation of fluid dynamics

Numerical simulation of turbulent fluid dynamics needs to either parametrize turbulence—which introduces large uncertainties—or explicitly resolve the smallest scales—which is prohibitively expensive. Here, we provide evidence through analytic bounds and numerical studies that a potential quantum speedup can be achieved to simulate fluid dynamics using quantum computing. Specifically, we provide a lattice Boltzmann formulation of fluid dynamics for which we give evidence that low-order Carleman linearization is much more accurate than previously believed for these systems. This is achieved via a combination of reformulating the Navier-Stokes nonlinearity (u·$\triangledown$u) to lattice-Boltzmann nonlinearity (u 2 ) and accurately linearizing the dynamical equations, which effectively trades nonlinearity for additional degrees of freedom that add negligible expense in the quantum solver. Based on this, we apply a quantum algorithm for simulating the Carleman-linearized lattice Boltzmann equation and provide evidence that its cost scales logarithmically with system size compared with polynomial scaling in the best known classical algorithms. In this paper, we suggest that a quantum advantage may exist for simulating fluid dynamics, paving the way for simulating nonlinear multiscale transport phenomena in a wide range of disciplines using quantum computing.

42 ENGINEERING↗

Implicit solvers for unstructured meshes

Implicit methods were developed and tested for unstructured mesh computations. The approximate system which arises from the Newton linearization of the nonlinear evolution operator is solved by using the preconditioned GMRES (Generalized Minimum Residual) technique. Three different preconditioners were studied, namely, the incomplete LU factorization (ILU), block diagonal factorization, and the symmetric successive over relaxation (SSOR). The preconditioners were optimized to have good vectorization properties. SSOR and ILU were also studied as iterative schemes. The various methods are compared over a wide range of problems. Ordering of the unknowns, which affects the convergence of these sparse matrix iterative methods, is also studied. Results are presented for inviscid and turbulent viscous calculations on single and multielement airfoil configurations using globally and adaptively generated meshes.

Venkatakrishnan, V.↗

Development of a continuous synthesis process for carbamazepine using validated in-line Raman spectroscopy and kinetic modelling for disturbance simulation

Mitigation of failure modes in the continuous synthesis (CS) of a drug substance (DS) has the potential to widen the adoption of continuous manufacturing (CM) technologies by the pharmaceutical industry. Here, this work demonstrates the development of a robust continuous process for the synthesis of carbamazepine (CBZ), an essential medicine as per the World Health Organization (WHO), facilitated by kinetic modelling and monitored by in-line Raman spectroscopy. Accurate kinetic modelling and the use of validated process analytical technology (PAT) models for quantitative measurement were found to play an important role in developing CS of drug substances. Kinetic data for the formation of CBZ from iminostilbene (ISB) were collected by batch reaction sampling and high-performance liquid chromatography (HPLC) analysis. A non-linear solver and iterative method was applied to determine two sets of Arrhenius parameters simultaneously for the reaction system by minimizing the standard error of the model fit. The start-up and dynamic equilibrium stages for the CS of CBZ using a continuous stirred tank reactor (CSTR) were modelled based on the batch kinetic data and employed to optimize conversion and simulate process disturbances. An in-line Raman spectroscopy method was successfully developed, validated, and integrated to determine the concentrations of CBZ and ISB within the operating range for the CS. The CS kinetic model was evaluated experimentally from startup to dynamic equilibrium over 10 residence times with monitoring by HPLC and in-line Raman spectroscopy. The developed kinetic model in tandem with in-line Raman spectroscopy successfully predicted disturbances due to changes in process variables and can serve as a useful tool in the future design of advanced process control strategies for the continuous synthesis of CBZ.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗