Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “direct solver”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Software issues in three dimensional continuum shape optimization employing boundary formulations

The paper addresses the issue of how individual computational techniques for the accurate and economical calculation of information required for large-scale 3D continuum structural shape optimization via boundary element analysis (BEA) formulations can be prudently selected and incorporated in large-scale BEA programs, employing either direct or iterative equation solvers. A series of techniques that allows for the incorporation of economical shape design sensitivity analysis capability with a minimal amount of software development is described. The use of reanalysis is shown to be the key ingredient associated with each of these techniques. It is concluded that, from a software engineering perspective, capabilities that facilitate 3D shape optimization can be implemented in BEA programs with only modest investments in program development.

Kane, James H.↗

Fast structural design and analysis via hybrid domain decomposition on massively parallel processors

A hybrid domain decomposition framework for static, transient and eigen finite element analyses of structural mechanics problems is presented. Its basic ingredients include physical substructuring and /or automatic mesh partitioning, mapping algorithms, 'gluing' approximations for fast design modifications and evaluations, and fast direct and preconditioned iterative solvers for local and interface subproblems. The overall methodology is illustrated with the structural design of a solar viewing payload that is scheduled to fly in March 1993. This payload has been entirely designed and validated by a group of undergraduate students at the University of Colorado using the proposed hybrid domain decomposition approach on a massively parallel processor. Performance results are reported on the CRAY Y-MP/8 and the iPSC-860/64 Touchstone systems, which represent both extreme parallel architectures. The hybrid domain decomposition methodology is shown to outperform leading solution algorithms and to exhibit an excellent parallel scalability.

Farhat, Charbel↗

Optimization of Wing-Body Configurations by the Euler Equations

This paper describes a new wing-body design procedure which is based on the Euler equations and a constrained numerical optimization technique. The geometry modification is based on a set of fundamental modes defined on the unit interval. A design example involving a generic wing-body model is presented to demonstrate the usefulness of the design program. It is shown that the use of an Euler solver coupled with a direct numerical optimization procedure is affordable on the current generation of supercomputers.

Chang, I.-Chung↗

RANS-MP: A Portable Parallel Navier-Stokes Solver

RANS-MP, a new implementation of a single-grid Navier-Stokes solver using the diagonalized Beam-Warming approximate-factorization scheme, is presented. This first release of the completely rewritten solver employs the following optimizations: (1) Bi-directional multi-partition method for the ADI solver part; this improves granularity and load balance; (2) Improved cache usage through elimination of non-unit-stride array access (possible in part due to multi-partitioning); (3) Preprocessing of communicating boundary conditions to streamline logic during time stepping; (4) Truly parallel, high-performance I/O using the newly-developed MPI-IO library; (5) Elimination of large amounts of redundant operations through efficient use of workspace. Results of some realistic wing computations on the IBM SP2 computer will be presented. We will demonstrate that excellent absolute performance and scalability are obtained with RANS-MP, even for relatively small grid sizes. Besides high performance, an outstanding feature of RANS-MP is its true portability, due to the use of the portable message passing and I/O libraries MPI and MPI-IO.

VanderWijngaart, Rob F.↗

Parallel Domain Decomposition Formulation and Software for Large-Scale Sparse Symmetrical/Unsymmetrical Aeroacoustic Applications

The overall objectives of this research work are to formulate and validate efficient parallel algorithms, and to efficiently design/implement computer software for solving large-scale acoustic problems, arised from the unified frameworks of the finite element procedures. The adopted parallel Finite Element (FE) Domain Decomposition (DD) procedures should fully take advantages of multiple processing capabilities offered by most modern high performance computing platforms for efficient parallel computation. To achieve this objective. the formulation needs to integrate efficient sparse (and dense) assembly techniques, hybrid (or mixed) direct and iterative equation solvers, proper pre-conditioned strategies, unrolling strategies, and effective processors' communicating schemes. Finally, the numerical performance of the developed parallel finite element procedures will be evaluated by solving series of structural, and acoustic (symmetrical and un-symmetrical) problems (in different computing platforms). Comparisons with existing "commercialized" and/or "public domain" software are also included, whenever possible.

Nguyen, D. T.↗

Parallel Finite Element Domain Decomposition for Structural/Acoustic Analysis

A domain decomposition (DD) formulation for solving sparse linear systems of equations resulting from finite element analysis is presented. The formulation incorporates mixed direct and iterative equation solving strategics and other novel algorithmic ideas that are optimized to take advantage of sparsity and exploit modern computer architecture, such as memory and parallel computing. The most time consuming part of the formulation is identified and the critical roles of direct sparse and iterative solvers within the framework of the formulation are discussed. Experiments on several computer platforms using several complex test matrices are conducted using software based on the formulation. Small-scale structural examples are used to validate thc steps in the formulation and large-scale (l,000,000+ unknowns) duct acoustic examples are used to evaluate the ORIGIN 2000 processors, and a duster of 6 PCs (running under the Windows environment). Statistics show that the formulation is efficient in both sequential and parallel computing environmental and that the formulation is significantly faster and consumes less memory than that based on one of the best available commercialized parallel sparse solvers.

Nguyen, Duc T.↗

Issue Summary of INL Phase IV Transient Results for IAEA CRP on HTGR UAM Benchmark

This report details the Parallel and Highly Innovative Simulation for Idaho National Laboratory (INL) Code System (PHISICS)/Reactor Excursions and Leak Analysis Program (RELAP5)-3D results obtained for the transient core exercises defined for Phase IV of the International Atomic Energy Agency (IAEA) Coordinated Research Project (CRP) on high-temperature gas cooled reactor (HTGR) uncertainty analysis in modeling (UAM). The Phase III models and results are linked to the earlier Standardized Computer Analyses for Licensing Evaluation (SCALE)/Sampler/New ESC-based Weighting Transport (NEWT) data generated for the lattice physics (lattice) stage Phase I of the CRP. The focus of this report is the Uncertainty/Sensitivity Assessment (U/SA) of the prismatic modular high-temperature gas cooled reactor (MHTGR)-350 design, and specifically for Exercises IV-1 and IV-2 of the benchmark: the Control Rod Withdrawal (CRW) and Pressurised Loss of Cooling (PLOFC) events. The statistical U/SA methodology is implemented and demonstrated using the RAVEN code, based on perturbed cross-section libraries obtained from the SCALE/Sampler sequence. Uncertainties in nuclear data (cross-sections and the average number of neutrons produced per fission, 235U[¯v ]) lead to standard deviations (uncertainties of one s) of approximately 0.5% in the core eigenvalues of the MHTGR-350 and core models. For the coupled neutronics/thermal fluid model, local power density uncertainties up to 3.6% were observed in the colder regions of the core, while the local maximum fuel temperature uncertainties reached 1.5% for the models that included thermal fluid uncertainties. The addition of thermal fluid uncertainties dominated the impacts of nuclear data uncertainties in all cases. The main contributors to uncertainties in the power density and fuel temperatures during the transients were uncertainties in the reactor operating conditions (total power, inlet mass flow rate and inlet gas temperature). Variations in the bypass flows did not have significant impact on any of the output variables. For the nuclear data uncertainties it was found that the 235U(¯v ) / 235U(¯v ) covariance produced the largest sensitivities in terms of its impact on the eigenvalue and peak reactor power. It was also observed that the impact of any nuclear data uncertainties on the maximum fuel temperature was much less significant that the impact on eigenvalue and power. Another important finding was that although the use of eight or more energy groups is recommended for best-estimate HTGR simulation, two-group models produced acceptable uncertainty and sensitivity results for most FOMs. Since the statistical U/SA methodology is computationally expensive, and most transient solver requirements will scale directly with the number of energy groups, two energy groups could be used by HTGR developers during the early stages of design when larger uncertainty margins can be tolerated.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Advanced System Thermal Fluids Solver Development for SAM

This work summarizes a feasibility study on testing numerical algorithms that are suitable and efficient for advanced system analysis code development under the mutli-physics framework, MOOSE. The key is the implementation of a high-order one-dimensional staggered-grid finite volume method (SG-FVM), and its direct interaction with the linear/nonlinear solver, PETSc. Leveraging the existing capabilities of the SAM code, significant code coverages were established in the finite volume method code. This in turn allows for a suite of test problems with different problem sizes and levels of complexity to be used to quantify the performance improvement of the finite volume method code. As evidently shown in this study, the implemented SG-FVM demonstrated superior performance improvement against a direct finite element method implementation through MOOSE for the wide range of selected problems. On two computer systems, the speedup was observed to be significant, with at least one order of magnitude of solving time reduction. In addition, for a complex reactor model, transient simulation was performed using the finite volume method code, the results of which agree very well with the reference results from the finite element method code. Overall, this study demonstrates a successful feasibility study on the proposed numerical algorithms and software structure to support advanced system analysis tool development. In this work, short-term priority development and testing items were identified, and long-term code adoption and integration plans were made for the eventual deployment of the finite volume method in the SAM code.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Numerical Simulations for Landing Gear Noise Generation and Radiation

Aerodynamic noise from a landing gear in a uniform flow is computed using the Ffowcs Williams -Hawkings (FW-H) equation. The time accurate flow data on the surface is obtained using a finite volume flow solver on an unstructured and. The Ffowcs Williams-Hawkings equation is solved using surface integrals over the landing gear surface and over a permeable surface away from the landing gear. Two geometric configurations are tested in order to assess the impact of two lateral struts on the sound level and directivity in the far-field. Predictions from the Ffowcs Williams-Hawkings code are compared with direct calculations by the flow solver at several observer locations inside the computational domain. The permeable Ffowcs Williams-Hawkings surface predictions match those of the flow solver in the near-field. Far-field noise calculations coincide for both integration surfaces. The increase in drag observed between the two landing gear configurations is reflected in the sound pressure level and directivity mainly in the streamwise direction.

Morris, Philip J.↗

PeleLMeX [SWR-22-48]

PeleLMeX is a solver for high fidelity reactive flow simulations, namely direct numerical simulation (DNS) and large eddy simulation (LES). The solver combines a low Mach number approach, adaptive mesh refinement (AMR), embedded boundary (EB) geometry treatment and high performance computing (HPC) to provide a flexible tool to address research questions on platforms ranging from small workstations to the world's largest GPU-accelerated supercomputers. PeleLMeX has been used to study complex flame/turbulence interactions in RCCI engines and hydrogen combustion or the effect of sustainable aviation fuel on gas turbine combustion. PeleLMeX is part of the Pele combustion Suite (https://amrex-combustion.github.io/)

Day, Marcus↗

An Adaptively-Refined, Cartesian, Cell-Based Scheme for the Euler and Navier-Stokes Equations

A Cartesian, cell-based scheme for solving the Euler and Navier-Stokes equations in two dimensions is developed and tested. Grids about geometrically complicated bodies are generated automatically, by recursive subdivision of a single Cartesian cell encompassing the entire flow domain. Where the resulting cells intersect bodies, polygonal 'cut' cells are created. The geometry of the cut cells is computed using polygon-clipping algorithms. The grid is stored in a binary-tree data structure which provides a natural means of obtaining cell-to-cell connectivity and of carrying out solution-adaptive refinement. The Euler and Navier-Stokes equations are solved on the resulting grids using a finite-volume formulation. The convective terms are upwinded, with a limited linear reconstruction of the primitive variables used to provide input states to an approximate Riemann solver for computing the fluxes between neighboring cells. A multi-stage time-stepping scheme is used to reach a steady-state solution. Validation of the Euler solver with benchmark numerical and exact solutions is presented. An assessment of the accuracy of the approach is made by uniform and adaptive grid refinements for a steady, transonic, exact solution to the Euler equations. The error of the approach is directly compared to a structured solver formulation. A non smooth flow is also assessed for grid convergence, comparing uniform and adaptively refined results. Several formulations of the viscous terms are assessed analytically, both for accuracy and positivity. The two best formulations are used to compute adaptively refined solutions of the Navier-Stokes equations. These solutions are compared to each other, to experimental results and/or theory for a series of low and moderate Reynolds numbers flow fields. The most suitable viscous discretization is demonstrated for geometrically-complicated internal flows. For flows at high Reynolds numbers, both an altered grid-generation procedure and a different formulation of the viscous terms are shown to be necessary. A hybrid Cartesian/body-fitted grid generation approach is demonstrated. In addition, a grid-generation procedure based on body-aligned cell cutting coupled with a viscous stensil-construction procedure based on quadratic programming is presented.

Coirier, William John↗

Modeling MTS pyrolysis and SiC deposition kinetics using principal component analysis and neural networks

Accurate chemical kinetics modeling is crucial for improving the efficiency of chemical processing and synthesis of ceramic matrix composites. Detailed kinetic models are computationally expensive due to the large number of transported chemical species, while the simplified physics-based models, such as single-step global mechanisms, are efficient but often overlook key chemical intermediates and pathways. Recent deep learning approaches promise accurate and cost-effective models. Yet, they require additional closures for the transported nonlinear latent variables, complicating integration with existing solvers. In this work, we develop a hybrid linear—nonlinear reduced model for silicon carbide deposition from methyltrichlorosilane precursor by combining principal component analysis (PCA) and autoencoder (AE) neural network (NN) approaches. PCA is used to identify a smaller set of linear transport variables, enabling direct reuse of conventional transport solvers. NNs then reconstruct the full chemical state from these reduced variables. We demonstrate the method on a chemical vapor deposition reactor—comprising a gas-phase pyrolysis plug flow reactor and a heterogeneous surface reactor—over a wide range of temperatures, pressures, and residence times. Our PCA–AE model achieves high accuracy with only five transported scalars, achieving an eightfold cost reduction compared to detailed mechanisms, in both a priori (using data from the test set only) and a posteriori (coupled with a differential equation solver). In conclusion, notable errors arise primarily near training domain boundaries and for long residence times, indicating the need for domain shift indicators and better long-horizon predictions in future reduced chemistry model development.

autoencoder neural networks↗

GCAM Regional Tuning: A framework to tune GCAM parameters

GCAM assumptions typically generate scenarios that are designed to be internally consistent and globally coherent. The gcamdata tool which facilitates the compilation of data sets and user assumptions is not well suited to tailoring to specific country or regional realities, sponsor requirements, or perform harmonization for model intercomparison needs. As described in this report, the GCAM Regional Tuning project develops a computational framework that enables users to adjust GCAM parameters, so model outputs match targeted outcomes at user-defined spatial, temporal, and sectoral resolutions. The framework integrates GCAM, gcamdata, and gcamwrapper with a set of flexible “tuning directives” and an iterative numerical solver. Users can define targets (e.g., technology shares in power generation, BEV uptake, sectoral service demands), select tuners that manipulate relevant GCAM parameters (e.g., share weights, cost adders, elasticities), and export tuned parameters as reusable GCAM XML inputs for future runs. We demonstrate the approach and document usage, diagnostics, and known limitations, and we outline potential future directions.

97 MATHEMATICS AND COMPUTING↗

Some fast elliptic solvers on parallel architectures and their complexities

The discretization of separable elliptic partial differential equations leads to linear systems with special block triangular matrices. Several methods are known to solve these systems, the most general of which is the Block Cyclic Reduction (BCR) algorithm which handles equations with nonconsistant coefficients. A method was recently proposed to parallelize and vectorize BCR. Here, the mapping of BCR on distributed memory architectures is discussed, and its complexity is compared with that of other approaches, including the Alternating-Direction method. A fast parallel solver is also described, based on an explicit formula for the solution, which has parallel computational complexity lower than that of parallel BCR.

Gallopoulos, E.↗

Some fast elliptic solvers on parallel architectures and their complexities

The discretization of separable elliptic partial differential equations leads to linear systems with special block tridiagonal matrices. Several methods are known to solve these systems, the most general of which is the Block Cyclic Reduction (BCR) algorithm which handles equations with nonconstant coefficients. A method was recently proposed to parallelize and vectorize BCR. In this paper, the mapping of BCR on distributed memory architectures is discussed, and its complexity is compared with that of other approaches including the Alternating-Direction method. A fast parallel solver is also described, based on an explicit formula for the solution, which has parallel computational compelxity lower than that of parallel BCR.

Gallopoulos, E.↗

A Formalization of Core Why3 in Coq

Intermediate verification languages like Why3 and Boogie have made it much easier to build program verifiers, transforming the process into a logic compilation problem rather than a proof automation one. Why3 in particular implements a rich logic for program specification with polymorphism, algebraic data types, recursive functions and predicates, and inductive predicates; it translates this logic to over a dozen solvers and proof assistants. Accordingly, it serves as a backend for many tools, including Frama-C, EasyCrypt, and GNATProve for Ada SPARK. But how can we be sure that these tools are correct? The alternate foundational approach, taken by tools like VST and CakeML, provides strong guarantees by implementing the entire toolchain in a proof assistant, but these tools are harder to build and cannot directly take advantage of SMT solver automation. As a first step toward enabling automated tools with similar foundational guarantees, we give a formal semantics in Coq for the logic fragment of Why3. We show that our semantics are useful by giving a correct-by-construction natural deduction proof system for this logic, using this proof system to verify parts of Why3's standard library, and proving sound two of Why3's transformations used to convert terms and formulas into the simpler logics supported by the backend solvers.

97 MATHEMATICS AND COMPUTING↗

Lattice Green’s Functions for High-Order Finite Difference Stencils

Lattice Green's Functions (LGFs) are fundamental solutions to discretized linear operators, and as such they are a useful tool for solving discretized elliptic PDEs on domains that are unbounded in one or more directions. The majority of existing numerical solvers that make use of LGFs rely on a second-order discretization and operate on domains with free-space boundary conditions in all directions. Under these conditions, fast expansion methods are available that enable precomputation of 2D or 3D LGFs in linear time, avoiding the need for brute-force multi-dimensional quadrature of numerically unstable integrals. Here we focus on higher-order discretizations of the Laplace operator on domains with more general boundary conditions, by (1) providing an algorithm for fast and accurate evaluation of the LGFs associated with high-order dimension-split centered finite differences on unbounded domains, and (2) deriving closed-form expressions for the LGFs associated with both dimension-split and Mehrstellen discretizations on domains with one unbounded dimension. Through numerical experiments we demonstrate that these techniques provide LGF evaluations with near machine-precision accuracy, and that the resulting LGFs allow for numerically consistent solutions to high-order discretizations of the Poisson's equation on fully or partially unbounded 3D domains.

97 MATHEMATICS AND COMPUTING↗

On the role of artificial viscosity in Navier-Stokes solvers

A method is proposed to determine directly the amount of artificial viscosity needed for stability using an eigenvalue analysis for a finite difference representation of the Navier-Stokes equations. The stability and growth of small perturbations about a steady flow over the airfoils are analyzed for various amounts of artificial viscosity. The eigenvalues were determined for a small perturbation about a steady inviscid flow over a NACA 0012 airfoil at a Mach number of 0.8 and angle of attack of 0 degrees. The movement of the eigenvalue constellation with respect to the amount of artificial viscosity is studied. The stability boundries as a function of the amount of artificial viscosity from both the eigenvalue analysis and the time marching scheme are also presented. This procedure not only allows for determining the effect of varying amounts of artificial viscosity, but also for the effects of different forms of terms for artificial viscosity.

Mahajan, Aparajit J.↗