Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “solver”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

A Comparison of GPU-Accelerated Multiphase CFD Solvers on the Polaris Supercomputer: Part 1

This report is in support of the Innovative and Novel Computational Impact on Theory and Experiment (INCITE) program sponsored by the U.S. Department of Energy (USDOE). With INCITE-level resources, one project, titled BubblyFlow, was granted computational resources for the 2025 calendar year on the Polaris supercomputer at the Argonne Leadership Computing Facility (ALCF). The project aims to conduct simulations to understand the fundamental characteristics of turbulent bubbly flow phenomena in nature. Staff at the ALCF and Argonne’s Computational Science division, along with collaborators at the City College of New York and University of Illinois at Chicago, helped a summer student to assess the accuracy and performance of two high performance computing (HPC) codes. Both codes, ImExLBM and FluTAS, are fundamentally different in their mathematical and numerical modeling. However, both may be used to solve the same physical problem. The collaboration sought to better understand the differences between both codes in terms of accuracy and efficiency. This would ultimately help the BubblyFlow project better utilize resources and establish a knowledge-base of code capabilities in future simulation campaigns. We compare ImExLBM and FluTAS, two high-performance multiphase computational fluid dynamics (CFD) solvers, in terms of physical fidelity, time-to-solution, and parallel efficiency. We validate ImExLBM (Implicit-Explicit Lattice Boltzmann Method) against a canonical benchmark and assess it’s performance relative to FluTAS (Fluid Transport Accelerated Solver), a well-established open-source CFD code.

97 MATHEMATICS AND COMPUTING↗

Quantification of LEU Holdup using gamma ray imaging and inverse transport solver

Holdup is the residual amount of special nuclear material (SNM) remaining in a processing facility after the bulk materials have been cleaned out. In commercial uranium processing facilities, quantification of holdup is a major challenge because of the highly variable shapes and sizes of the deposits. Any method that attempts to generalize and calibrate deposit shapes in order to quantify holdup will be prone to high uncertainties. Uncertainties on the order of ±50% are typical in holdup results. In international safeguards applications, a ±50% uncertainty can result in a large amount of material unaccounted for (MUF) thereby increasing the difficulty of detecting material diversion and facility misuse. An imaging-based methodology has been developed with the objective of significantly reducing this uncertainty by using the true deposit shape, instead of relying on oversimplified geometric assumptions. The project is a collaboration between ORNL, Y-12, and the University of Tennessee, Knoxville, TN. Uranium sources of known masses were measured using the Germanium Gamma-ray Imager (GeGI), a high-resolution imaging spectrometer, creating a pixelated map for each spectral bin. Two different gamma imaging methods are employed in this work: coded aperture imaging and Compton imaging. A validated MonteCarlo model of the detector has been developed using the GEANT4 code for determining the intrinsic response of the detector, its enclosure, and the coded aperture mask. An inverse transport solver based on the Markov Chain Monte-Carlo approach known as Differential Evolution Adaptive Metropolis (DREAM) is employed to use the measurement data from the image pixels (coded aperture or Compton) to solve for the mass of 235 U in the deposit. A reliable method based on the DREAM solver has been developed to flag the infinite thickness condition of a uranium deposit. The project team is working towards improving the image reconstruction for Compton imaging so that a better localization of the source can be achieved. Besides treating the coded aperture and Compton imaging methods independently, the project is also evaluating a combined method that uses the Compton scatter data from a coded aperture measurement. GEANT4 simulations are being performed to evaluate the combined approach. The impact on the DREAM optimization as the source thickness progressively approaches infinite thickness is being evaluated. A number of uranium sources available at ORNL have been measured, and the DREAM results have been tested and validated for the coded aperture imaging. A similar effort will be carried out to validate the Compton based method once the development of algorithms for better localization are complete. The imaging based quantification is very amenable to unattended monitoring of holdup accumulation at key measurement points. A proof of concept measurement has been completed to demonstrate this capability The current work used the high energy resolution imager GeGI. However, the approach and methodologies are applicable to other imagers such as the cadmium zin telluride (CZT) based imager manufactured by H3D, Inc.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A Fluid Solver for Studying Torsional Galloping in Solar-Tracking PV Panel Arrays

Solar-tracking photovoltaic arrays are susceptible to aeroelastic fluttering during high-wind events. This dynamic fluttering behavior can grow in amplitude until the panels enter an unstable mode known as torsional galloping which can lead to panel failure or total array destruction. To better understand the physics of the torsional galloping phenomenon, and to inform the discussion around panel design and recommended panel stow positions during high wind events, a fluid-structure interaction solver composed of a simulated atmospheric boundary layer with simplified panel structural responses was designed. The simulation choices and features of this solver were informed by the geometry and physical properties of an experimental panel array known to exhibit torsional galloping behavior during hind-wind events. These simulations revealed that the torsional galloping instability is driven by a combination of cyclic vortex shedding from the sun-facing side of the panel and the elastic properties of the torque tube linking the panel assemblies. Testing different stow angles across a range of wind speeds indicates that panels are generally more stable when stowed at negative angles where the leading edge is closer to the ground, hypothesized to be due to ground-blocking effects.

41 EE - Solar Energy Technologies Office (EE-4S)↗

Two-Stage Gauss-Seidel Preconditioners and Smoothers for Krylov Solvers on a GPU Cluster: Preprint

Gauss-Seidel (GS) relaxation is often employed as a preconditioner for a Krylov solver or as a smoother for Algebraic Multigrid (AMG). However, the requisite sparse triangular solve is difficult to parallelize on many-core architectures such as graphics processing units (GPUs). In the present study, the performance of the sequential GS relaxation based on a triangular solve is compared with two-stage variants, replacing the direct triangular solve with a fixed number of inner Jacobi-Richardson (JR) iterations. When a small number of inner iterations is sufficient to maintain the Krylov convergence rate, the two-stage GS (GS2) often outperforms the sequential algorithm on many-core architectures. The GS2 algorithm is also compared with JR. When they perform the same number of ops for SpMV (e.g. three JR sweeps compared to two GS sweeps with one inner JR sweep), the GS2 iterations, and the Krylov solver preconditioned with GS2, may converge faster than the JR iterations. Moreover, for some problems (e.g. elasticity), it was found that JR may diverge with a damping factor of one, whereas two-stage GS may improve the convergence with more inner iterations. Finally, to study the performance of the two-stage smoother and preconditioner for a practical problem, these were applied to incompressible uid ow simulations on GPUs.

algebraic multigrid↗

Discrete-Element and Material-Point Method (DEM and MPM) Based Solvers for Sustainable Technologies

We present the use of discrete element method (DEM) and material point method (MPM) in three relevant green technology applications that include biomass feedstock handling, lithium-ion battery manufacturing, and high-pressure reverse osmosis. Our open-source DEM and MPM solvers are developed using performance portable grid and particle management library, AMReX, thus enabling superior performance on NVIDIA and AMD GPUs with > 100 million particles. Our DEM solver resolves the motion of individual particles in a granular system and includes a bonded sphere method for modeling non-spherical particles along with Hertzian and liquid bridge-based contact models. We simulate highly variable biomass feedstock flows in large-scale hoppers for biofuel production and electrode calendering in battery manufacturing using DEM. Our simulations predict flow blockage in large scale biomass hoppers and electrode microstructure variations, thus providing valuable information for biofuel and battery manufacturers, respectively. The second half of the talk will be on MPM and its application towards pore resolved simulations of reverse osmosis membranes under compressive loads. We present a validation study of our MPM simulations with membrane microscopy imaging thus providing useful insights on membrane stability under high pressure conditions. We also present a spectral stability analysis of using linear hat, quadratic and cubic spline basis in MPM indicating regions of numerical stability.

BIOMASS FUELS,MATHEMATICS AND COMPUTING↗

AMR-Wind: A Performance-Portable, High-Fidelity Flow Solver for Wind Farm Simulations

We present AMR-Wind, a verified and validated high-fidelity computational-fluid-dynamics code for wind farm flows. AMR-Wind is a block-structured, adaptive-mesh, incompressible-flow solver that enables predictive simulations of the atmospheric boundary layer and wind plants. It is a highly scalable code designed for parallel high-performance computing with a specific focus on performance portability for current and future computing architectures, including graphical processing units (GPUs). In this paper, we detail the governing equations, the numerical methods, and the turbine models. Establishing a foundation for the correctness of the code, we present the results of formal verification and validation. The verification studies, which include a novel actuator line test case, indicate that AMR-Wind is spatially and temporally second-order accurate. The validation studies demonstrate that the key physics capabilities implemented in the code, including actuator disk models, actuator line models, turbulence models, and large eddy simulation (LES) models for atmospheric boundary layers, perform well in comparison to reference data from established computational tools and theory. We conclude with a demonstration simulation of a 12-turbine wind farm operating in a turbulent atmospheric boundary layer, detailing computational performance and realistic wake interactions.

17 WIND ENERGY↗

Investigation of grid-based vorticity-velocity large eddy simulation off-body solvers for application to overset CFD

Accurately predicting unsteady wakes and vortex-dominated flows is essential to a wide range of engineering applications, including aircraft, rotorcraft, shipboard operations, bio-inspired unsteady flight and propulsion, wind turbines, and urban flows. While current CFD software can model the complete flow field and wake system, the computational costs incurred in high Reynolds number unsteady turbulent flow simulations often remain prohibitive for routine engineering use, particularly for applications involving moving components. Prior work has demonstrated that by adopting a vorticity-velocity formulation in a grid-based off-body flow solver (VorTran-M and VorTran-M2) one can lower these costs by several orders of magnitude when compared to conventional approaches. This paper describes the extensions made to VorTran-M2 to support turbulent flows, and associated benchmarking activity to assess its performance for problems involving strong stretching and diffusion processes, whose competing contributions to the vorticity field are core drivers of turbulent flow evolution. Predictions are presented for: (i) the Kida-Pelz problem whose inviscid form is of mathematical interest due to its apparent formation of singular flow in finite time; and (ii) the Taylor Green vortex arrangement, which has been extensively studied as a fundamental simulation challenge in the turbulent modeling community. Here, the results are used to evaluate the overall predictive ability and performance of two sub-grid scale models incorporated into VorTran-M2. Results indicate that the computational cost savings seen previously for inviscid and convection dominated problems extend to turbulent flow simulations supporting the viability of VorTran-M2 as a low cost means for accurately modeling the far-field and background flow, particularly when long duration vorticity evolution is of interest.

42 ENGINEERING↗

Benchmark of the KGMf with a coupled Boltzmann equation solver

The Kinetic Global Model framework (KGMf) is an open-source general-purpose global model (spatially averaged) simulation code developed to explore the reaction kinetics and pathways in plasma discharge systems. It contains species continuity and electron energy balance equations, with a time-dependent evaluated electron energy distribution function (EEDF) for electron impact reactions. The EEDF is utilized to determine the rate coefficients for electron impact reactions, which can have profound impact on the temporal evolutions of plasma parameters. Previously, the EEDF was commonly assumed as Maxwellian or an analytical function of the effective electron temperature. In this work, the KGMf is coupled with a Boltzmann equation (BE) solver to self-consistently compute the EEDF. The EEDF evolution frequency is determined based on relative changes of the reduced electric field. The KGMf is benchmarked with the ZDPlasKin code based on high-pressure low-temperature argon plasma discharge cases. Additionally, the temporal evolutions of reduced electric field, electron temperature, EEDF, reaction rates, and species densities, are obtained and compared under different discharge conditions, showing good agreement between the KGMf and the ZDPlasKin simulations. The application of the KGMf for predicting breakdown times in high power microwave discharges is also presented, which shows qualitative agreement with particle-in-cell simulations. The KGMf can be further applied for more complicated plasma discharge systems, where the reaction kinetics are more intricate (e.g., plasma-assisted combustion systems).

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Accelerated impurity solver for DMFT and its diagrammatic extensions

Here, we present ComCTQMC, a GPU accelerated quantum impurity solver. It uses the continuous-time quantum Monte Carlo (CTQMC) algorithm wherein the partition function is expanded in terms of the hybridisation function (CT-HYB). ComCTQMC supports both partition and worm-space measurements, and it uses improved estimators and the reduced density matrix to improve observable measurements whenever possible. ComCTQMC efficiently measures all one and two-particle Green's functions, all static observables which commute with the local Hamiltonian, and the occupation of each impurity orbital. ComCTQMC can solve complex-valued impurities with crystal fields that are hybridized to both fermionic and bosonic baths. Most importantly, ComCTQMC utilizes graphical processing units (GPUs), if available, to dramatically accelerate the CTQMC algorithm when the Hilbert space is sufficiently large. We demonstrate acceleration by a factor of over 600 (100) in a simulation of δ-Pu at 600 K with (without) crystal fields. In easier problems, the GPU offers less impressive acceleration or even decelerates the CTQMC. Here we describe the theory, algorithms, and structure used by ComCTQMC in order to achieve this set of features and level of acceleration.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Iterative methods in GPU-resident linear solvers for nonlinear constrained optimization

Linear solvers are major computational bottlenecks in a wide range of decision support and optimization computations. The challenges become even more pronounced on heterogeneous hardware, where traditional sparse numerical linear algebra methods are often inefficient. For example, methods for solving ill-conditioned linear systems have relied on conditional branching, which degrades performance on hardware accelerators such as graphical processing units (GPUs). To improve the efficiency of solving ill-conditioned systems, our computational strategy separates computations that are efficient on GPUs from those that need to run on traditional central processing units (CPUs). Our strategy maximizes the reuse of expensive CPU computations. Iterative methods, which thus far have not been broadly used for ill-conditioned linear systems, play an important role in our approach. In particular, we extend ideas from Arioli et al., (2007) to implement iterative refinement using inexact LU factors and flexible generalized minimal residual (FGMRES), with the aim of efficient performance on GPUs. In conclusion, we focus on solutions that are effective within broader application contexts, and discuss how early performance tests could be improved to be more predictive of the performance in a realistic environment.

97 MATHEMATICS AND COMPUTING↗

A numerical Poisson solver with improved radial solutions for a self-consistent locally scaled self-interaction correction method

Abstract The universal applicability of density functional approximations is limited by self-interaction error made by these functionals. Recently, a novel one-electron self-interaction-correction (SIC) method that uses an iso-orbital indicator to apply the SIC at each point in space by scaling the exchange-correlation and Coulomb energy densities was proposed. The locally scaled SIC (LSIC) method is exact for the one-electron densities, and unlike the well-known Perdew–Zunger SIC (PZSIC) method recovers the uniform electron gas limit of the uncorrected density functional approximation, and reduces to PZSIC method as a special case when isoorbital indicator is set to the unity. Here, we present a numerical scheme that we have adopted to evaluate the Coulomb potential of the electron density scaled by the iso-orbital indicator required for the self-consistent LSIC calculations. After analyzing the behavior of the finite difference method (FDM) and the green function solution to the radial part of the Poisson equation, we adopt a hybrid approach that uses the FDM for the Coulomb potential due to the monopole and the GF for all higher-order terms. The performance of the resultant hybrid method is assessed using a variety of systems. The results show improved accuracy than earlier numerical schemes. We also find that, even with a generic set of radial grid parameters, accurate energy differences can be obtained using a numerical Coulomb solver in standard density functional studies.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A Performance Portable, Fully Implicit Landau Collision Operator with Batched Linear Solvers

Modern accelerators use hierarchical parallel programming models that enable massive multithreading within a processing element (PE), with multiple PEs per device driven by traditional processes. Batching is a technique for exposing PE-level parallelism in algorithms that have traditionally run on MPI processes or multiple threads within a single process. Opportunities for batching arise in, for example, kinetic discretizations of magnetized plasmas where collisions are advanced in velocity space at each spatial point independently. This paper builds on previous work on a high-performance, fully nonlinear, Landau collision operator by batching the linear solver, as well as batching the spatial point problems and adding new support for multiple grids for multiscale, multispecies problems. An anisotropic relaxation verification test that agrees well with previously published results and analytical models is presented. The performance results from NVIDIA A100 and AMD MI250X nodes are presented with hardware utilization analysis for each architecture. Finally, the entire implicit Landau operator time advance is implemented in Kokkos for performance portability, running entirely on the device and is available in the PETSc numerical library.

97 MATHEMATICS AND COMPUTING↗

A Semi-Algebraic Two Level Solver

We develop a simple semi-algebraic 2-level solver built on traditional multigrid ideas. It is designed to be easily incorporated into existing simulation software. It exhibits good convergence for many classes of challenging problems including discontinuous diffusion, convection- diffusion, and Helmholtz equations. It has built-in structure that makes it simple to generalize in several interesting directions.

97 MATHEMATICS AND COMPUTING↗

A multipurpose lifting-line flow solver for arbitrary wind energy concepts

Abstract. In this work, we extend the AeroDyn module of OpenFAST to support arbitrary collections of wings, rotors, and towers. The new standalone AeroDyn driver supports arbitrary motions of the lifting surfaces and complex turbulent inflows. Aerodynamics and inflow are assembled into one module that can be readily coupled with an elastic solver. We describe the features and updates necessary for the implementation of the new AeroDyn driver. We present different case studies of the driver to illustrate its application to concepts such as multirotors, kites, or vertical-axis wind turbines. We perform verification and validation of some of the new features using the following test cases: elliptical wings, horizontal-axis wind turbines, and 2D and 3D vertical-axis wind turbines. The wind turbine simulations are compared to existing tools and field measurements. We use this opportunity to describe some limitations of current models and to highlight areas that we think should be the focus of future research in wind turbine aerodynamics.

17 WIND ENERGY↗

Using Griffin's Transmutation Solver to Calculate Radiation Damage

Displacement Radiation Damage originates from all nuclides, not just those that are naturally occurring. Currently only damage from naturally-occurring nuclides, or sometimes damage from one transmutation product is considered. It is proposed that the transmutation solvers implemented in many codes be used to calculate this radiation damage. This would explicitly treat damage from all sources without additionally burdening the user. A proof-of-concept implementation was created in Griffin. The implementation showed that only minor code modifications are necessary to add this feature. When compared against an analytical benchmark the results Griffin could calculate were very accurate.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Multipoint Aerostructural Optimization of Wind Turbine Rotors Using a Coupled Blade‐Resolved Aerostructural Solver

Physics‐based design optimization workflows thread the needle between computational cost limitations and simulation complexity, often compromising between modeling detail and the range of operating design conditions. Multipoint aerostructural optimization of wind turbine rotors has so far been confined to low‐fidelity analyses or to high‐fidelity studies with simplified structural models, leaving the most complex design trade‐offs unexplored. We close this gap by performing the first tightly coupled gradient‐based multipoint aerostructural rotor optimization using 3D aerodynamic and structural solvers with discrete coupled adjoints. The optimizer simultaneously varies blade planform, airfoil shapes, and structural thickness through more than 270 design variables, minimizing a weighted combination of rotor mass and power across multiple wind speeds. Applied to a modified DTU 10‐MW benchmark under conservative structural and aerodynamic constraints, our multipoint optimization reduces rotor mass by up to 36% and increases power by 12%–15% across the main operating conditions; biasing the objective toward power yields power gains up to 18% and a 17% mass reduction. For a nominal wind distribution, 3‐point rotor designs accounting for low RPM and high thrust conditions capture dominant trade‐offs and outperform single‐point designs. Adding two off‐design points changes individual‐condition power by less than 3% but leaves the weighted average within 0.5%, and the mass‐power bias has a stronger effect on the final design than the operating‐point weighting itself. Our framework extends naturally to richer load cases and site‐specific wind distributions, providing a basis for high‐fidelity multipoint design earlier in industrial workflows.

17 WIND ENERGY↗

A stochastic solver based on the residence time algorithm for crystal plasticity models

Abstract The deformation of crystalline materials by dislocation motion takes place in discrete amounts determined by the Burgers vector. Dislocations may move individually or in bundles, potentially giving rise to intermittent slip. This confers plastic deformation with a certain degree of variability that can be interpreted as being caused by stochastic fluctuations in dislocation behavior. However, crystal plasticity (CP) models are almost always formulated in a continuum sense, assuming that fluctuations average out over large material volumes and/or cancel out due to multi-slip contributions. Nevertheless, plastic fluctuations are known to be important in confined volumes at or below the micron scale, at high temperatures, and under low strain rate/stress deformation conditions. Here, we develop a stochastic solver for CP models based on the residence-time algorithm that naturally captures plastic fluctuations by sampling among the set of active slip systems in the crystal. The method solves the evolution equations of explicit CP formulations, which are recast as stochastic ordinary differential equations and integrated discretely in time. The stochastic CP model is numerically stable by design and naturally breaks the symmetry of plastic slip by sampling among the active plastic shear rates with the correct probability. This can lead to phenomena such as intermittent slip or plastic localization without adding external symmetry-breaking operations to the model. The method is applied to body-centered cubic tungsten single crystals under a variety of temperatures, loading orientations, and imposed strain rates.

42 ENGINEERING↗

PyAlbany: A Python interface to the C++ multiphysics solver Albany

Albany is a parallel C++ finite element library for solving forward and inverse problems involving partial differential equations (PDEs). In this paper we introduce PyAlbany, a newly developed Python interface to the Albany library. PyAlbany can be used to effectively drive Albany enabling fast and easy analysis and post-processing of applications based on PDEs that are pre-implemented in Albany. PyAlbany relies on the library PyBind11 to bind Python with C++ Albany code. Here we detail the implementation of PyAlbany and showcase its capabilities through a number of examples targeting a heat-diffusion problem. In particular we consider the following: (1) the generation of samples for a Monte Carlo application, (2) a scalability study, (3) a study of parameters on the performance of a linear solver, and finally (4) a tool for performing eigenvalue decompositions of matrix-free operators for a Bayesian inference application.

97 MATHEMATICS AND COMPUTING↗