Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “fast solver”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Fast structural design and analysis via hybrid domain decomposition on massively parallel processors

A hybrid domain decomposition framework for static, transient and eigen finite element analyses of structural mechanics problems is presented. Its basic ingredients include physical substructuring and /or automatic mesh partitioning, mapping algorithms, 'gluing' approximations for fast design modifications and evaluations, and fast direct and preconditioned iterative solvers for local and interface subproblems. The overall methodology is illustrated with the structural design of a solar viewing payload that is scheduled to fly in March 1993. This payload has been entirely designed and validated by a group of undergraduate students at the University of Colorado using the proposed hybrid domain decomposition approach on a massively parallel processor. Performance results are reported on the CRAY Y-MP/8 and the iPSC-860/64 Touchstone systems, which represent both extreme parallel architectures. The hybrid domain decomposition methodology is shown to outperform leading solution algorithms and to exhibit an excellent parallel scalability.

Farhat, Charbel↗

Numerical simulation of steady transonic flow about airfoils

A computer code has been developed that couples a fast transonic full-potential AF2 solver with both an efficient integral boundary-layer method and a viscous wedge approximation of the shock/boundary-layer interaction. The efficiency of the coupled analysis methods and the method of coupling has resulted in a uniquely efficient analysis tool. The airfoil geometry is modified by the displacement thickness before the shock and the displacement thickness plus the viscous wedge thickness after the shock by considering the viscous effects as an equivalent transpiration boundary condition. The flow about conventional and supercritical airfoils under moderately strong shock situations has been calculated. Comparisons with experimental data indicate that this viscous correction method has improved the accuracy of the full-potential analysis. Furthermore, the computer time required to obtain a converged solution has been reduced.

Lee, S. C.↗

Final Technical Report - Center for Simulation of Fusion Relevant RF Actuators

We have developed a suite of 3D electromagnetic field solvers, both FEM and FDTD based, that account for the RF antenna and vacuum vessel geometries with unprecedented accuracy. Workflows were developed that make it possible to translate CAD models for the antenna and vacuum vessel to physics meshes for RF wave simulation. Nonlinear RF sheath formation has been incorporated self-consistently as a boundary condition in these solvers. We have also carried out extensive studies of the impact of RF sheaths on the ion energy angle distribution at plasma-material interfaces, using high fidelity particle-in-cell codes. Comprehensive simulation models were developed to assess the impact of blob-like edge turbulence on RF wave propagation and the impact of the RF ponderomotive force on the plasma scrape-off layer (SOL). A fluid transport solver for the far-SOL was also developed which accounts for the high parallel to perpendicular heat anisotropy on an unstructured mesh, thus making it possible to precisely represent an antenna structure in the presence of edge transport. Finally we have developed a hierarchy of core wave propagation and absorption models that self-consistently combine continuum Fokker Planck and Monte Carlo treatments of fast ion evolution with ICRF full-wave field solvers and continuum Fokker Planck treatments of fast electron evolution with both full-wave and ray tracing models.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

The three-dimensional Multi-Block Advanced Grid Generation System (3DMAGGS)

As the size and complexity of three dimensional volume grids increases, there is a growing need for fast and efficient 3D volumetric elliptic grid solvers. Present day solvers are limited by computational speed and do not have all the capabilities such as interior volume grid clustering control, viscous grid clustering at the wall of a configuration, truncation error limiters, and convergence optimization residing in one code. A new volume grid generator, 3DMAGGS (Three-Dimensional Multi-Block Advanced Grid Generation System), which is based on the 3DGRAPE code, has evolved to meet these needs. This is a manual for the usage of 3DMAGGS and contains five sections, including the motivations and usage, a GRIDGEN interface, a grid quality analysis tool, a sample case for verifying correct operation of the code, and a comparison to both 3DGRAPE and GRIDGEN3D. Since it was derived from 3DGRAPE, this technical memorandum should be used in conjunction with the 3DGRAPE manual (NASA TM-102224).

Alter, Stephen J.↗

Application of fast Fourier transforms to the direct solution of a class of two-dimensional separable elliptic equations on the sphere

An efficient, direct, second-order solver for the discrete solution of a class of two-dimensional separable elliptic equations on the sphere (which generally arise in implicit and semi-implicit atmospheric models) is presented. The method involves a Fourier transformation in longitude and a direct solution of the resulting coupled second-order finite-difference equations in latitude. The solver is made efficient by vectorizing over longitudinal wave-number and by using a vectorized fast Fourier transform routine. It is evaluated using a prescribed solution method and compared with a multigrid solver and the standard direct solver from FISHPAK.

Moorthi, Shrinivas↗

Linearized Aeroelastic Solver Applied to the Flutter Prediction of Real Configurations

A fast-running unsteady aerodynamics code, LINFLUX, was previously developed for predicting turbomachinery flutter. This linearized code, based on a frequency domain method, models the effects of steady blade loading through a nonlinear steady flow field. The LINFLUX code, which is 6 to 7 times faster than the corresponding nonlinear time domain code, is suitable for use in the initial design phase. Earlier, this code was verified through application to a research fan, and it was shown that the predictions of work per cycle and flutter compared well with those from a nonlinear time-marching aeroelastic code, TURBO-AE. Now, the LINFLUX code has been applied to real configurations: fans developed under the Energy Efficient Engine (E-cubed) Program and the Quiet Aircraft Technology (QAT) project. The LINFLUX code starts with a steady nonlinear aerodynamic flow field and solves the unsteady linearized Euler equations to calculate the unsteady aerodynamic forces on the turbomachinery blades. First, a steady aerodynamic solution is computed for given operating conditions using the nonlinear unsteady aerodynamic code TURBO-AE. A blade vibration analysis is done to determine the frequencies and mode shapes of the vibrating blades, and an interface code is used to convert the steady aerodynamic solution to a form required by LINFLUX. A preprocessor is used to interpolate the mode shapes from the structural dynamics mesh onto the computational fluid dynamics mesh. Then, LINFLUX is used to calculate the unsteady aerodynamic pressure distribution for a given vibration mode, frequency, and interblade phase angle. Finally, a post-processor uses the unsteady pressures to calculate the generalized aerodynamic forces, eigenvalues, an esponse amplitudes. The eigenvalues determine the flutter frequency and damping. Results of flutter calculations from the LINFLUX code are presented for (1) the E-cubed fan developed under the E-cubed program and (2) the Quiet High Speed Fan (QHSF) developed under the Quiet Aircraft Technology project. The results are compared with those obtained from the TURBO-AE code. A graph of the work done per vibration cycle for the first vibration mode of the E-cubed fan is shown. It can be seen that the LINFLUX results show a very good comparison with TURBO-AE results over the entire range of interblade phase angle. The work done per vibration cycle for the first vibration mode of the QHSF fan is shown. Once again, the LINFLUX results compare very well with the results from the TURBOAE code.

Reddy, Tondapu S.↗

Parallel-vector computation for linear structural analysis and non-linear unconstrained optimization problems

Several parallel-vector computational improvements to the unconstrained optimization procedure are described which speed up the structural analysis-synthesis process. A fast parallel-vector Choleski-based equation solver, pvsolve, is incorporated into the well-known SAP-4 general-purpose finite-element code. The new code, denoted PV-SAP, is tested for static structural analysis. Initial results on a four processor CRAY 2 show that using pvsolve reduces the equation solution time by a factor of 14-16 over the original SAP-4 code. In addition, parallel-vector procedures for the Golden Block Search technique and the BFGS method are developed and tested for nonlinear unconstrained optimization. A parallel version of an iterative solver and the pvsolve direct solver are incorporated into the BFGS method. Preliminary results on nonlinear unconstrained optimization test problems, using pvsolve in the analysis, show excellent parallel-vector performance indicating that these parallel-vector algorithms can be used in a new generation of finite-element based structural design/analysis-synthesis codes.

Nguyen, D. T.↗

Randomized Algorithms for Linear Solvers

Recently, randomized algorithms in numerical linear algebra, specifically those centered around random sketching, have gained traction in primarily theoretical research due to their potential to significantly reduce problem dimensionality at the cost of an O(1) multiplicative distortion factor. It has been assumed that this sketching can be done efficiently, but thorough investigation into how precisely to do it has been neglected. Moreover, the theory-based community has argued for sketching’s ability to reduce computational cost via complexity analysis, but has not researched how it affects the stability of the algorithms. At Sandia, efficient linear solvers that scale well on modern HPC architectures while maintaining stability are imperative for practical applications. In this LDRD, we developed a random sketching strategy that is substantially faster than existing ones, and demonstrate its superior performance in practice on a NVIDIA H100 GPU. Moreover, we show how this can be used to significantly outperform existing linear least squares solvers while improving the solver’s stability as well. Additionally, we demonstrate how this sketching strategy can be used to make a fast, stable QR factorization that can subsequently be used in s-step and block Krylov solvers. Finally, we incorporate a sketching-based block orthogonalization scheme into s-step GMRES, which is stable and faster than existing approaches on the Perlmutter supercomputer.

97 MATHEMATICS AND COMPUTING↗

Rethinking materials simulations: Blending direct numerical simulations with neural operators

Abstract Materials simulations based on direct numerical solvers are accurate but computationally expensive for predicting materials evolution across length- and time-scales, due to the complexity of the underlying evolution equations, the nature of multiscale spatiotemporal interactions, and the need to reach long-time integration. We develop a method that blends direct numerical solvers with neural operators to accelerate such simulations. This methodology is based on the integration of a community numerical solver with a U-Net neural operator, enhanced by a temporal-conditioning mechanism to enable accurate extrapolation and efficient time-to-solution predictions of the dynamics. We demonstrate the effectiveness of this hybrid framework on simulations of microstructure evolution via the phase-field method. Such simulations exhibit high spatial gradients and the co-evolution of different material phases with simultaneous slow and fast materials dynamics. We establish accurate extrapolation of the coupled solver with large speed-up compared to DNS depending on the hybrid strategy utilized. This methodology is generalizable to a broad range of materials simulations, from solid mechanics to fluid dynamics, geophysics, climate, and more.

36 MATERIALS SCIENCE↗

The block adaptive multigrid method applied to the solution of the Euler equations

In the present study, a scheme capable of solving very fast and robust complex nonlinear systems of equations is presented. The Block Adaptive Multigrid (BAM) solution method offers multigrid acceleration and adaptive grid refinement based on the prediction of the solution error. The proposed solution method was used with an implicit upwind Euler solver for the solution of complex transonic flows around airfoils. Very fast results were obtained (18-fold acceleration of the solution) using one fourth of the volumes of a global grid with the same solution accuracy for two test cases.

Pantelelis, Nikos↗

A Parallel Multigrid Solver for Viscous Flows on Anisotropic Structured Grids

This paper presents an efficient parallel multigrid solver for speeding up the computation of a 3-D model that treats the flow of a viscous fluid over a flat plate. The main interest of this simulation lies in exhibiting some basic difficulties that prevent optimal multigrid efficiencies from being achieved. As the computing platform, we have used Coral, a Beowulf-class system based on Intel Pentium processors and equipped with GigaNet cLAN and switched Fast Ethernet networks. Our study not only examines the scalability of the solver but also includes a performance evaluation of Coral where the investigated solver has been used to compare several of its design choices, namely, the interconnection network (GigaNet versus switched Fast-Ethernet) and the node configuration (dual nodes versus single nodes). As a reference, the performance results have been compared with those obtained with the NAS-MG benchmark.

Prieto, Manuel↗

Technical report series on global modeling and data assimilation. Volume 2: Direct solution of the implicit formulation of fourth order horizontal diffusion for gridpoint models on the sphere

High order horizontal diffusion of the form K Delta(exp 2m) is widely used in spectral models as a means of preventing energy accumulation at the shortest resolved scales. In the spectral context, an implicit formation of such diffusion is trivial to implement. The present note describes an efficient method of implementing implicit high order diffusion in global finite difference models. The method expresses the high order diffusion equation as a sequence of equations involving Delta(exp 2). The solution is obtained by combining fast Fourier transforms in longitude with a finite difference solver for the second order ordinary differential equation in latitude. The implicit diffusion routine is suitable for use in any finite difference global model that uses a regular latitude/longitude grid. The absence of a restriction on the timestep makes it particularly suitable for use in semi-Lagrangian models. The scale selectivity of the high order diffusion gives it an advantage over the uncentering method that has been used to control computational noise in two-time-level semi-Lagrangian models.

Max J. Suarez↗

An Adaptive Unstructured Grid Method by Grid Subdivision, Local Remeshing, and Grid Movement

An unstructured grid adaptation technique has been developed and successfully applied to several three dimensional inviscid flow test cases. The approach is based on a combination of grid subdivision, local remeshing, and grid movement. For solution adaptive grids, the surface triangulation is locally refined by grid subdivision, and the tetrahedral grid in the field is partially remeshed at locations of dominant flow features. A grid redistribution strategy is employed for geometric adaptation of volume grids to moving or deforming surfaces. The method is automatic and fast and is designed for modular coupling with different solvers. Several steady state test cases with different inviscid flow features were tested for grid/solution adaptation. In all cases, the dominant flow features, such as shocks and vortices, were accurately and efficiently predicted with the present approach. A new and robust method of moving tetrahedral "viscous" grids is also presented and demonstrated on a three-dimensional example.

Pirzadeh, Shahyar Z.↗

A physics-informed deep learning description of Knudsen layer reactivity reduction

A physics-informed neural network (PINN) is used to evaluate the fast ion distribution in the hot spot of an inertial confinement fusion target. The use of tailored input and output layers to the neural network is shown to enable a PINN to learn the parametric solution to the Vlasov–Fokker–Planck equation in the absence of any synthetic or experimental data. As an explicit demonstration of the approach, the specific problem of Knudsen layer fusion yield reduction is treated. Here, the predictions from the Vlasov–Fokker–Planck PINN are used to provide a non-perturbative solution of the fast ion tail in the vicinity of the hot spot, thus allowing the spatial profile of the fusion reactivity to be evaluated for a range of collisionalities and hot spot conditions. Excellent agreement is found between the predictions of the Vlasov–Fokker–Planck PINN and the results from traditional numerical solvers with respect to both the energy and spatial distribution of fast ions and the fusion reactivity profile, demonstrating that the Vlasov–Fokker–Planck PINN provides an accurate and efficient means of determining the impact of Knudsen layer yield reduction across a broad range of plasma conditions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Physics-Informed Deep Learning Description of Knudsen Layer Reactivity Reduction

A physics-informed neural network (PINN) is used to evaluate the fast ion distribution in the hot spot of an inertial confinement fusion target. The use of tailored input and output layers to the neural network is shown to enable a PINN to learn the parametric solution to the Vlasov–Fokker–Planck equation in the absence of any synthetic or experimental data. As an explicit demonstration of the approach, the specific problem of Knudsen layer fusion yield reduction is treated. Here, the predictions from the Vlasov–Fokker–Planck PINN are used to provide a non-perturbative solution of the fast ion tail in the vicinity of the hot spot, thus allowing the spatial profile of the fusion reactivity to be evaluated for a range of collisionalities and hot spot conditions. Excellent agreement is found between the predictions of the Vlasov–Fokker–Planck PINN and the results from traditional numerical solvers with respect to both the energy and spatial distribution of fast ions and the fusion reactivity profile, demonstrating that the Vlasov–Fokker–Planck PINN provides an accurate and efficient means of determining the impact of Knudsen layer yield reduction across a broad range of plasma conditions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Finite Difference informed Random Walk solver for simulating radiation defect evolution in polycrystalline structures with strongly inhomogeneous diffusivity

Diffusivity of species and defects on grain boundaries is usually several orders of magnitude larger than that inside grains. Such strongly inhomogeneous diffusivity requires prohibitively high computational demands for modeling microstructural evolution. Here, this paper presents a highly-efficient numerical solver, combining the Finite Difference method and Random Walk model, designed for accurately modeling strongly inhomogeneous diffusion within polycrystalline structures. The proposed solver, termed Finite Difference informed Random Walk (FDiRW), integrates a customized Finite Difference (cFD) scheme tailored for fast diffusion along thin grain boundaries represented by a single-layer of nodes. Numerical experiments demonstrate that the FDiRW solver achieves an impressive efficiency gain of 1560x compared to traditional Finite Difference methods while maintaining accuracy, making it feasible for personal computer machines to handle diffusional systems with strongly inhomogeneous diffusivity across static polycrystalline microstructures. The model has been successfully applied to simulate radiation defect evolution, showcasing its scalability to engineering scales in both length and time dimensions.

36 MATERIALS SCIENCE↗

Implementing Ordinary Differential Equation Solvers in Rust Programming Language for Modeling Vehicle Powertrain Systems: Preprint

Efficient and accurate ordinary differential equation (ODE) solvers are necessary for powertrain and vehicle dynamics modeling. However, current commercial ODE solvers can be financially prohibitive, leading to a need for accessible, effective, open-source ODE solvers designed for powertrain modeling. Rust is a compiled programming language that has the potential to be used for fast and easy-to-use powertrain models, given its exceptional computational performance, robust package ecosystem, and short time required for modelers to become proficient. However, of the three commonly used (>3,000 downloads) packages in Rust with ODE solver capabilities, only one has more than four numerical methods implemented, and none are designed specifically for modeling physical systems. Therefore, the goal of the Differential Equation System Solver (DESS) was to implement accurate ODE solvers in Rust designed for the component-based problems often seen in powertrain modeling. DESS is a text-based software package that provides a flexible framework for building and solving systems of ODEs. This allows DESS to be included as a dependency for automotive powertrain models that require a variety of solvers and solver configurations. Seven explicit ODE solver methods have been implemented in DESS: Euler’s, Heun’s, midpoint, Ralston’s, classic Runge-Kutta, Bogacki-Shampine, and Cash-Karp. These represent five fixed-step methods and two adaptive-step methods. This paper shows that the solver implementations increase accuracy and computational efficiency compared to Euler's method when modeling a system of three thermal masses in Rust. DESS also includes features designed for modeling component-based physical systems. Users can define relationships between nodes in their system, which the package then translates into a system of equations, leading to simpler and more intuitive code. In the case of a three-thermal-mass system, the user can specify node thermal properties (e.g., thermal capacitance), how nodes are interconnected, and thermal conductance between nodes rather than providing a system of equations. The core contribution from this work is an open-source, text-based Rust package with ODE solvers for automotive powertrain modeling to support cost-free, fast, and accurate simulation.

ADVANCED PROPULSION SYSTEMS↗

Porting OVERFLOW CFD Code to GPUs: To Hackathons and Beyond!

OVERFLOW is an overset, structured computational fluid dynamics (CFD) code written in Fortran which is widely used in the government, industry, and academia. Over the last several years the OVERFLOW developers have been working to port miniapps based on computationally expensive parts of OVERFLOW to run on GPUs, primarily using OpenACC. This effort started at our first hackathon in 2019 and since then the OVERFLOW team has attended two additional hackathons (virtually). These hackathon environments have provided a great place to collaborate with others and learn from experts. These learning experiences enabled porting two miniapps to run effectively on NVIDIA GPUs using OpenACC. The first miniapp focused on motifs found in the solver itself and the final ported version runs three times fast ona single V100 compared to a 40 core, dual-socket Intel Skylake node. The speed up in this solverminiapp required multiple design changes including increasing the amount of parallelism available and the amount of work performed in each kernel. The second miniapp focused on overset MPI communication, also saw significant speedups over the CPU implementation using a CUDA-aware MPI implementation through OpenACC. This presentation will discuss our experience at the hackathons, our process of porting the miniapps to run on the GPUs, and several lessons learned throughout.

OpenACC↗