Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Mesh Optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

316 records · Page 18

A New Semistructured Algebraic Multigrid Method

Multigrid methods are well suited to large massively parallel computer architectures because they are mathematically optimal and display good parallelization properties. Since current architecture trends are favoring regular compute patterns to achieve high performance, the ability to express structure has become much more important. The hypre software library provides high-performance multigrid preconditioners and solvers through conceptual interfaces, including a semistructured interface that describes matrices primarily in terms of stencils and logically structured grids. This paper presents a new semistructured algebraic multigrid (SSAMG) method built on this interface. The numerical convergence and performance of a CPU implementation of this method are evaluated for a set of semistructured problems. In conclusion, SSAMG achieves significantly better setup times than hypre’s unstructured AMG solvers and comparable convergence. In addition, the new method is capable of solving more complex problems than hypre’s structured solvers.

97 MATHEMATICS AND COMPUTING↗

Compare linear-system solver and preconditioner stacks with emphasis on GPU performance and propose phase-2 NGP solver development pathway

The goal of the ExaWind project is to enable predictive simulations of wind farms comprised of many megawatt-scale turbines situated in complex terrain. Predictive simulations will require computational fluid dynamics (CFD) simulations for which the mesh resolves the geometry of the turbines and captures the rotation and large deflections of blades. Whereas such simulations for a single turbine are arguably petascale class, multi-turbine wind farm simulations will require exascale-class resources. The primary physics codes in the ExaWind project are Nalu-Wind, which is an unstructured-grid solver for the acoustically incompressible Navier-Stokes equations, and OpenFAST, which is a whole-turbine simulation code. The Nalu-Wind model consists of the mass-continuity Poisson-type equation for pressure and a momentum equation for the velocity. For such modeling approaches, simulation times are dominated by linear-system setup and solution for the continuity and momentum systems. For the ExaWind challenge problem, the moving meshes greatly affect overall solver costs as reinitialization of matrices and recomputation of preconditioners is required at every time step. In this report we evaluated GPU-performance baselines for the linear solvers in the Trilinos and hypre solver stacks using two representative Nalu-Wind simulations: an atmospheric boundary layer precursor simulation on a structured mesh, and a fixed-wing simulation using unstructured overset meshes. Both strong-scaling and weak-scaling experiments were conducted on the OLCF supercomputer Summit and similar proxy clusters. We focused on the performance of multi-threaded Gauss-Seidel and two-stage Gauss-Seidel that are extensions of classical Gauss-Seidel; of one-reduce GMRES, a communication-reducing variant of the Krylov GMRES; and algebraic multigrid methods that incorporate the afore-mentioned methods. The team has established that AMG methods are capable of solving linear systems arising from the fixed-wing overset meshes on CPU, a critical intermediate result for ExaWind FY20 Q3 and Q4 milestones. For the fixed-wing strong-scaling study (model with 3M grid-points), the team identified that Nalu-Wind simulations with the new Trilinos and hypre solvers scale to modest GPU counts, maintaining above 70% efficiency up to 6 GPUs. However, there still remain significant bottlenecks to performance: matrix assembly (hypre), AMG setup (hypre and Trilinos) In the weak-scaling experiments (going from 0.4M to 211M gridpoints), it's shown that the solver apply phases are faster on GPUs, but that Nalu-Wind simulation times grow, primarily due to the multigrid-setup process. Finally, based on the report outcomes, we propose a linear solver path-forward for the remainder of the ExaWind project. Near term, the NREL team will continue their work on GPU-based linear-system assembly. They will also investigate how the use of alternatives to the NVIDIA UVM (unified virtual memory) paradigm affects performance. Longer term, the NREL team will evaluate algorithmic performance on other types of accelerators and merge their improvements back to the main hypre repository branch. Near term, the Trilinos team will address performance bottlenecks identified in this milestone, such as implementing a GPU-based segregated momentum solve and reusing matrix graphs across linear-system assembly phases. Longer term, the Trilinos team will do detailed analysis and optimization of multigrid setup.

17 WIND ENERGY↗

CG-Kit: Code Generation Toolkit for performant and maintainable variants of source code applied to Flash-X hydrodynamics simulations

CG-Kit is a new Code Generation tool-Kit that we have developed as a part of the solution for portability and maintainability for multiphysics computing applications. The development of CG-Kit is rooted in the urgent need created by the shifting landscape of high-performance computing platforms and the algorithmic complexities of a particular large-scale multiphysics application: Flash-X. To efficiently use computing resources on a heterogeneous node, an application must have a map of computation to resources and a mechanism to move the data and computation to the resources according to the map. Most existing performance portability solutions are focussed on abstracting the expression of computations so that a unified source code can be specialized to run on different resources. However, such an approach is insufficient for a code like Flash-X, which has a multitude of code components that can be assembled in various permutations and combinations to form different instances of applications. Similar challenges apply to any code that has composability, where a single specified way of apportioning work among devices may not be optimal. Additionally, use cases arise where the optimal control flow of computation may differ for different devices while the underlying numerics remain identical. This combination leads to unique challenges including handling an existing large code base in Fortran and/or C/C++, subdivision of code into a great variety of units supporting a wide range of physics and numerical methods, different parallelization techniques for distributed and shared memory systems and accelerator devices, and heterogeneity of computing platforms requiring coexisting variants of parallel algorithms. All of these challenges demand that scientific software developers apply existing knowledge about domain applications, algorithms, and computing platforms to determine custom abstractions and granularity for code generation. There is a critical lack of tools to tackle those problems. CG-Kit is designed to fill this gap by providing a user with the ability to express their desired control flow and computation-to-resource map in the form a pseudocode-like recipe. It consists of standalone tools that can be combined into highly specific and, we argue, highly effective portability and maintainability toolchains. Here we present the design of our new tools: parametrized source trees, control flow graphs, and recipes. The tools are implemented in Python. They are agnostic to the programming language of the source code targeted for code generation. In conclusion, we demonstrate the capabilities of the toolkit with two examples, first, multithreaded variants of the basic AXPY operation, and second, variants of parallel algorithms within a hydrodynamics solver, called Spark, from Flash-X that operates on block-structured adaptive meshes.

Algorithmic portability↗

Enabling the Broader Use of MOOSE for Nuclear Energy and Other Simulation

This Final Scientific and Technical Report summarizes work performed under the Phase IIA SBIR project “Enabling the Broader Use of MOOSE for Nuclear Energy and Other Simulation” (DE-SC0020906) from August 2023 through August 2025. The objective of the Phase IIA effort was to mature and harden capabilities developed during Phase II, with the goal of enabling practical interoperability between Coreform’s isogeometric analysis (IGA) technologies and the Multiphysics Object-Oriented Simulation Environment (MOOSE), while improving robustness, performance, and scalability for complex, nuclear-relevant geometries. Over the course of Phase IIA, the project established and validated an extraction-based interoperability pathway between Coreform tools and MOOSE. A combined mesh and matrix format was defined collaboratively with MOOSE developers and integrated into the solver, enabling standard MOOSE workflows to operate on data exported from Coreform’s IGA and Flex Representation Method (FRM) pipelines. Early demonstrations validated architectural compatibility using linear solid mechanics problems, while later efforts focused on benchmark testing and external use. By the end of the project period, engineers at BWXT were able to independently set up and execute a simulation using the Coreform–MOOSE workflow and provide direct feedback that informed further refinement. In parallel, substantial effort was devoted to improving the robustness of trimmed U-spline construction for complex CAD geometries. A growing test suite of nuclear-relevant models was compiled through collaboration with multiple stakeholders and used to drive extensive bug fixing and reliability improvements. These efforts resulted in improved robustness and performance, including the addition of fallback capabilities that enhance reliability when the underlying commercial CAD kernel fails. Performance-oriented work progressed later in the project, with the development and demonstration of methods to decompose complex geometries into structured subregions and updated data representations to support more efficient solver processing. Additionally, extensive enhancements to threadsafe parallel data structures and trimming operations established a foundation for scalable processing of large assemblies. Collaboration with Sandia National Laboratories on the SGM geometric modeling kernel advanced to a functioning interface test case, positioning the workflow for future kernel integration. Overall, the Phase IIA effort successfully transitioned the project from architectural proof-of-concept to externally exercised, solver-integrated capability, while clarifying remaining technical challenges related to standardization, performance optimization, and kernel integration.

42 ENGINEERING↗

A Systematic Interpretation of Subsurface Proppant Concentration from Drilling Mud Returns: Case Study from Hydraulic Fracturing Test Site (HFTS-2) in Delaware Basin

The aim of this study is generation and validation of a proppant log using analysis of drilling mud returns for child wells. Proppant log provides qualitative as well as quantitative insights into spatial distribution of proppant sand particles from prior stimulation of parent wells. While the basic methodology was developed and formalized during analysis of material collected from through fracture cores at Hydraulic Fracturing Test Site in Midland Basin (HFTS – 1), the test wells at HFTS – 2 in the neighboring Delaware Basin allowed the opportunity to validate the workflow on actual mud return samples from subsurface. As a child well is being drilled, periodic mud return samples are collected at the rig site and preserved for analysis. The workflow involves systematic cleaning of the samples including various steps such as washing, drying and segregation of samples into relevant size fractions of interest (< Mesh 20) based on specifications of pumped sand during stimulation of the parent well. Clean samples are imaged using high resolution transparency scanning. Scan images are then systematically analyzed for particles of interest using computer vision techniques. Sample counts are further validated using elemental analysis of smaller sub-samples at various depths of interest. This step is necessary to isolate proppant versus other naturally occurring minerals such as sulphates and carbonates which show similar optical properties. We successfully correlated proppant distribution against the existing parent well and validated propped versus relatively un-propped zones for a child well at the test site. The advantage of testing the proppant log concept at the HFTS – 2 site is the plethora of additional diagnostic data that is available to validate our primary observations. We can correlate spatial proppant distribution against variability in stimulation response based on independent observations such as image logs, microseismic attributes as well as DAS response, all of which tend to corroborate one another. One of our significant successes was being able to describe varying degrees of impact of the parent well along the lateral length of a stimulated child well. Our workflow represents a systematic and one-of-a-kind interpretation of spatial proppant distribution while drilling child wells. This provides unique opportunities to better understand the current state of the Downloaded from http://onepetro.org/URTECONF/proceedings-pdf/21URTC/2-21URTC/D021S031R003/2477415/urtec-2021-5189-ms.pdf/1 by Carol Worster on 28 February 2022 URTeC 5189 2 reservoir being targeted including zones which are likely more drained relative to others and how the planned completion of the child well can be improved. Lastly, this log can be useful is validating optimal well spacing in relatively new fields under development.

58 GEOSCIENCES↗

NSFnets (Navier-Stokes flow nets): Physics-informed neural networks for the incompressible Navier-Stokes equations

In the last 50 years there has been a tremendous progress in solving numerically the Navier-Stokes equations using finite differences, finite elements, spectral, and even meshless methods. Yet, in many real cases, we still cannot incorporate seamlessly (multi-fidelity) data into existing algorithms, and for industrial-complexity applications the mesh generation is time consuming and still an art. Moreover, solving ill-posed problems (e.g., lacking boundary conditions) or inverse problems is often prohibitively expensive and requires different formulations and new computer codes. Here, we employ physics-informed neural networks (PINNs), encoding the governing equations directly into the deep neural network via automatic differentiation, to overcome some of the aforementioned limitations for simulating incompressible laminar and turbulent flows. We develop the Navier-Stokes flow nets (NSFnets) by considering two different mathematical formulations of the Navier-Stokes equations: the velocity-pressure (VP) formulation and the vorticity-velocity (VV) formulation. Since this is a new approach, we first select some standard benchmark problems to assess the accuracy, convergence rate, computational cost and flexibility of NSFnets; analytical solutions and direct numerical simulation (DNS) databases provide proper initial and boundary conditions for the NSFnet simulations. The spatial and temporal coordinates are the inputs of the NSFnets, while the instantaneous velocity and pressure fields are the outputs for the VP-NSFnet, and the instantaneous velocity and vorticity fields are the outputs for the VV-NSFnet. This is unsupervised learning and, hence, no labeled data are required beyond boundary and initial conditions and the fluid properties. The residuals of the VP or VV governing equations, together with the initial and boundary conditions, are embedded into the loss function of the NSFnets. No data is provided for the pressure to the VP-NSFnet, which is a hidden state and is obtained via the incompressibility constraint without extra computational cost. Unlike the traditional numerical methods, NSFnets inherit the properties of neural networks (NNs), hence the total error is composed of the approximation, the optimization, and the generalization errors. Here, we empirically attempt to quantify these errors by varying the sampling (“residual”) points, the iterative solvers, and the size of the NN architecture. For the laminar flow solutions, we show that both the VP and the VV formulations are comparable in accuracy but their best performance corresponds to different NN architectures. The initial convergence rate is fast but the error eventually saturates to a plateau due to the dominance of the optimization error. For the turbulent channel flow, we show that NSFnets can sustain turbulence at , but due to expensive training we only consider part of the channel domain and enforce velocity boundary conditions on the subdomain boundaries provided by the DNS data base. We also perform a systematic study on the weights used in the loss function for balancing the data and physics components, and investigate a new way of computing the weights dynamically to accelerate training and enhance accuracy. In the last part, we demonstrate how NSFnets should be used in practice, namely for ill-posed problems with incomplete or noisy boundary conditions as well as for inverse problems. We obtain reasonably accurate solutions for such cases as well without the need to change the NSFnets and at the same computational cost as in the forward well-posed problems. As a result, we also present a simple example of transfer learning that will aid in accelerating the training of NSFnets for different parameter settings.

97 MATHEMATICS AND COMPUTING↗

Improving I/O Performance for Exascale Applications through Online Data Layout Reorganization

The applications being developed within the U.S. Exascale Computing Project (ECP) to run on imminent Exascale computers will generate scientific results with unprecedented fidelity and record turn-around time. Many of these codes are based on particle-mesh methods and use advanced algorithms, especially dynamic load-balancing and mesh-refinement, to achieve high performance on Exascale machines. Yet, as such algorithms improve parallel application efficiency, they raise new challenges for I/O logic due to their irregular and dynamic data distributions. Thus, while the enormous data rates of Exascale simulations already challenge existing file system write strategies, the need for efficient read and processing of generated data introduces additional constraints on the data layout strategies that can be used when writing data to secondary storage. We review these I/O challenges and introduce two online data layout reorganization approaches for achieving good tradeoffs between read and write performance. We demonstrate the benefits of using these two approaches for the ECP particle-in-cell simulation WarpX, which serves as a motif for a large class of important Exascale applications. Here, we show that by understanding application I/O patterns and carefully designing data layouts we can increase read performance by more than 80 percent.

97 MATHEMATICS AND COMPUTING↗

Measurement of Transport Properties of Woody Biomass Feedstock Particles Before and After Pyrolysis by Numerical Analysis of X-Ray Tomographic Reconstructions

Lignocellulosic biomass has a complex, species-specific microstructure that governs heat and mass transport during conversion processes. A quantitative understanding of the evolution of pore size and structure is critical to optimize conversion processes for biofuel and bio-based chemical production. Further, improving our understanding of the microstructure of biochar coproduct will accelerate development of its myriad applications. This work quantitatively compares the microstructural features and the anisotropic permeabilities of two woody feedstocks, red oak and Douglas fir, using X-ray computed tomography (XCT) before and after the feedstocks are subjected to pyrolysis. Quantitative analysis of the three-dimensional (3D) reconstructions allows for direct calculations of void fractions, pore size distributions and tortuosity factors. Next, 3D images are imported into an immersed boundary based finite volume solver to simulate gas flow through the porous structure and to directly calculate the principal permeabilities along longitudinal, radial, and tangential directions. The permeabilities of native biomass are seen to differ by three to four orders of magnitude in the different principal directions, but we find that this anisotropy is substantially reduced in the biochar formed during pyrolysis. The quantitative transport properties reported here enhance the ability of pyrolysis simulations to account for feedstock-specific effects and thereby provide a useful touchstone for the biorefining community.

09 BIOMASS FUELS↗

An Adaptive Newton-Based Free-Boundary Grad–Shafranov Solver

Equilibria in magnetic confinement devices result from force balancing between the Lorentz force and the plasma pressure gradient. In an axisymmetric configuration like a tokamak, such an equilibrium is described by an elliptic equation for the poloidal magnetic flux, commonly known as the Grad–Shafranov equation. It is challenging to develop a scalable and accurate free-boundary Grad–Shafranov solver, since it is a fully nonlinear optimization problem that simultaneously solves for the magnetic field coil current outside the plasma to control the plasma shape. In this work, we develop a Newton-based free-boundary Grad–Shafranov solver using adaptive finite elements and preconditioning strategies. The free-boundary interaction leads to the evaluation of a domain-dependent nonlinear form of which its contribution to the Jacobian matrix is achieved through shape calculus. The optimization problem aims to minimize the distance between the plasma boundary and specified control points while satisfying two nontrivial constraints, which correspond to the nonlinear finite element discretization of the Grad–Shafranov equation and a constraint on the total plasma current involving a nonlocal coupling term. The linear system is solved by a block factorization, and AMG is called for subblock elliptic operators. The unique contributions of this work include the treatment of a global constraint, preconditioning strategies, nonlocal reformulation, and the implementation of adaptive finite elements. Furthermore, it is found that the resulting Newton solver is robust, successfully reducing the nonlinear residual to 1e-6 and lower in a small handful of iterations while addressing the challenging case to find a Taylor state equilibrium where conventional Picard-based solvers fail to converge.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Breakup dynamics in a pressure-swirl injector for urea-water solution applications: A computational study

The co-optimization of in-cylinder combustion and after-treatment technology has become a major aspect in engine design and development, with the goal of meeting the increasingly restrictive emission regulations in the transportation industry. Selective Catalytic Reduction is a robust technology to control the emission of NO x , and the injection of urea in water solution is the exhaust tailpipe is a key aspect of its operation. The proposed work uses high-fidelity Computational Fluid Dynamics to characterize the atomization dynamics of the liquid jet in relevant cross-flow conditions. The study focuses on a commercial low-pressure (9 bar) pressure-swirl injector which is characterized in its internal geometry through high-resolution X-ray micro-computational tomography. The internal two-phase flow has been modeled according to the volume-of-fluid approach in a large eddy simulation framework and validated against near-nozzle X-ray radiography measurement. Moreover, characterizing the breakup dynamics for the swirling hollow cone formation, and assessing the influence of the cross-flow in the breakup dynamics was completed. The results have been reported proposing Re-Oh maps and probability density functions of the spray kinematics. Higher cross-flow momentum generates an increase in the jet intact length and a reduction of the liquid droplet diameters. The axial momentum of the jet is affected by the cross-flow already in the near-nozzle region, determining a relevant deviation of the spray velocities. In conclusion, this work aims to inform the initialization of Eulerian-Lagrangian spray models through the assignment of droplet kinematics and static one-way coupling between volume-of-fluid results and Lagrangian spray parcels, to be used for system-size domain simulations.

33 ADVANCED PROPULSION SYSTEMS↗