Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Domain decomposition”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Implementation of a multiblock sensitivity analysis method in numerical aerodynamic shape optimization

A multiblock sensitivity analysis method is applied in a numerical aerodynamic shape optimization technique. The Sensitivity Analysis Domain Decomposition (SADD) scheme which is implemented in this study was developed to reduce the computer memory requirements resulting from the aerodynamic sensitivity analysis equations. Discrete sensitivity analysis offers the ability to compute quasi-analytical derivatives in a more efficient manner than traditional finite-difference methods, which tend to be computationally expensive and prone to inaccuracies. The direct optimization procedure couples CFD analysis based on the two-dimensional thin-layer Navier-Stokes equations with a gradient-based numerical optimization technique. The linking mechanism is the sensitivity equation derived from the CFD discretized flow equations, recast in adjoint form, and solved using direct matrix inversion techniques. This investigation is performed to demonstrate an aerodynamic shape optimization technique on a multiblock domain and its applicability to complex geometries. The objectives are accomplished by shape optimizing two aerodynamic configurations. First, the shape optimization of a transonic airfoil is performed to investigate the behavior of the method in highly nonlinear flows and the effect of different grid blocking strategies on the procedure. Secondly, shape optimization of a two-element configuration in subsonic flow is completed. Cases are presented for this configuration to demonstrate the effect of simultaneously reshaping interfering elements. The aerodynamic shape optimization is shown to produce supercritical type airfoils in the transonic flow from an initially symmetric airfoil. Multiblocking effects the path of optimization while providing similar results at the conclusion. Simultaneous reshaping of elements is shown to be more effective than individual element reshaping due to the inclusion of mutual interference effects.

Lacasse, James M.↗

A Multi-Level Parallelization Concept for High-Fidelity Multi-Block Solvers

The integration of high-fidelity Computational Fluid Dynamics (CFD) analysis tools with the industrial design process benefits greatly from the robust implementations that are transportable across a wide range of computer architectures. In the present work, a hybrid domain-decomposition and parallelization concept was developed and implemented into the widely-used NASA multi-block Computational Fluid Dynamics (CFD) packages implemented in ENSAERO and OVERFLOW. The new parallel solver concept, PENS (Parallel Euler Navier-Stokes Solver), employs both fine and coarse granularity in data partitioning as well as data coalescing to obtain the desired load-balance characteristics on the available computer platforms. This multi-level parallelism implementation itself introduces no changes to the numerical results, hence the original fidelity of the packages are identically preserved. The present implementation uses the Message Passing Interface (MPI) library for interprocessor message passing and memory accessing. By choosing an appropriate combination of the available partitioning and coalescing capabilities only during the execution stage, the PENS solver becomes adaptable to different computer architectures from shared-memory to distributed-memory platforms with varying degrees of parallelism. The PENS implementation on the IBM SP2 distributed memory environment at the NASA Ames Research Center obtains 85 percent scalable parallel performance using fine-grain partitioning of single-block CFD domains using up to 128 wide computational nodes. Multi-block CFD simulations of complete aircraft simulations achieve 75 percent perfect load-balanced executions using data coalescing and the two levels of parallelism. SGI PowerChallenge, SGI Origin 2000, and a cluster of workstations are the other platforms where the robustness of the implementation is tested. The performance behavior on the other computer platforms with a variety of realistic problems will be included as this on-going study progresses.

Hatay, Ferhat F.↗

From PINNs to PIKANs: recent advances in physics-informed machine learning

Physics-Informed Neural Networks (PINNs) have emerged as a key tool in Scientific Machine Learning since their introduction in 2017, enabling the efficient solution of ordinary and partial differential equations using sparse measurements. Over the past few years, significant advancements have been made in the training and optimization of PINNs, covering aspects such as network architectures, adaptive refinement, domain decomposition, and the use of adaptive weights and activation functions. A notable recent development is the Physics-Informed Kolmogorov-Arnold Networks (PIKANS), which leverage a representation model originally proposed by Kolmogorov in 1957, offering a promising alternative to traditional PINNs. In this review, we provide a comprehensive overview of the latest advancements in PINNs, focusing on improvements in network design, feature expansion, optimization techniques, uncertainty quantification, and theoretical insights. We also survey key applications across a range of fields, including biomedicine, fluid and solid mechanics, geophysics, dynamical systems, heat transfer, chemical engineering, and beyond. Lastly, we review computational frameworks and software tools developed by both academia and industry to support PINN research and applications.

Kolmogorov-Arnold networks↗

Linear and Nonlinear Solvers for Simulating Multiphase Flow within Large-Scale Engineered Subsurface Systems

Simulation of multiphase flow in the subsurface is well-known to be computationally challenging. While there have been many studies that have explored approaches to overcoming these challenges, they often utilize relatively simple case studies. In this paper, we focus on the unique numerical challenges posed by modeling large-scale engineered subsurface systems, characterized by discrete features embedded in a heterogeneous natural subsurface setting. The man-made features such as shafts, tunnels, and barriers often cause multiple challenges in modeling the domain for multiphase porous media flow. This flow scenario can have a wide range of applications such as nuclear waste repositories, enhanced recovery of a petroleum reservoir, geothermal engineering, and carbon sequestration. An example of these severe numerical challenges is the case of performance assessment (PA) for Waste Isolation Pilot Plant (WIPP), the only operating deep geological repository in the US, which simulates extreme material properties of bedded salt rock formation and extreme contrast due to open excavation next to the formation. The models have extremes not only of permeability and porosity but also of the constitutive models needed for multiphase flow; additionally, they have process models like salt creep closure reducing porosity over time, fracturing in clay and anhydrite interbeds of the bedded salt, gas generation from the waste materials, and unintentional human borehole intrusions in some scenarios. Numerical simulations require the solution of coupled systems of nonlinear PDEs; in our work, we use the open-source simulator PFLOTRAN which is based on Finite Volume discretization. The solution of the nonlinear equations requires use of the Newton-Raphson iteration at each time step, which entails the solution of the linearized Jacobian system at each iteration. The effects of all the processes (i.e., large number of unknowns, highly nonlinear constitutive relations, large contrasts in material properties in short distances) lead to an ill-conditioned Jacobian matrix that severely challenges traditional linear solver, i.e., stabilized biconjugate gradient with block Jacobi incomplete LU preconditioner (BCGS-ILU) leading to non-convergence for traditional Newton-Raphson nonlinear solver causing unacceptably long computation time for each model. This paper presents linear solvers such as constrained pressure residual (CPR) two-stage preconditioner with alternate-block-factorization (ABF) and quasi- implicit pressure and explicit saturation (QIMPES) decouplers and flexible generalized residual solver (FGMRES). The new general-purpose nonlinear solver, Newton trust-region dogleg Cauchy (NTRDC), is also introduced to resolve extreme nonlinearities in the models. We demonstrate the effectiveness of each method relative to the default BCGS-Newton solver. The two best cases had nearly 50 times speed-up and achieved completion of a simulation in 14 hours that never completed due to non-convergence with the default solver. We also investigate the strong scalability of each method and discuss some of the deficiencies found for Block Jacobi preconditioner using parallel domain decomposition, and node packing effects of modern processor architecture.

Preconditioner, Nonlinear, Porous media, Multiphas↗

Model-parallel Fourier neural operators as learned surrogates for large-scale parametric PDEs

Fourier neural operators (FNOs) are a recently introduced neural network architecture for learning solution operators of partial differential equations (PDEs), which have been shown to perform significantly better than comparable deep learning approaches. Once trained, FNOs can achieve speed-ups of multiple orders of magnitude over conventional numerical PDE solvers. However, due to the high dimensionality of their input data and network weights, FNOs have so far only been applied to two-dimensional or small three-dimensional problems. To remove this limited problem-size barrier, we propose a model-parallel version of FNOs based on domain-decomposition of both the input data and network weights. Here, we demonstrate that our model-parallel FNO is able to predict time-varying PDE solutions of over 2.6 billion variables on Perlmutter using up to 512 A100 GPUs and show an example of training a distributed FNO on the Azure cloud for simulating multiphase CO 2 dynamics in the Earth’s subsurface.

58 GEOSCIENCES↗

HiPACE++: A portable, 3D quasi-static particle-in-cell code

Modeling plasma accelerators is a computationally challenging task and the quasi-static particle-in-cell algorithm is a method of choice in a wide range of situations. In this work, we present the first performance-portable, quasi-static, three-dimensional particle-in-cell code HiPACE++. By decomposing all the computation of a 3D domain in successive 2D transverse operations and choosing appropriate memory management, HiPACE++ demonstrates orders-of-magnitude speedups on modern scientific GPUs over CPU-only implementations. The 2D transverse operations are performed on a single GPU, avoiding time-consuming communications. The longitudinal parallelization is done through temporal domain decomposition, enabling near-optimal strong scaling from 1 to 512 GPUs. HiPACE++ is a modular, open-source code enabling efficient modeling of plasma accelerators from laptops to state-of-the-art supercomputers.

97 MATHEMATICS AND COMPUTING↗

Code modernization strategies for short-range non-bonded molecular dynamics simulations

Modern HPC systems are increasingly relying on greater core counts and wider vector registers. Thus, applications need to be adapted to fully utilize these hardware capabilities. One class of applications that can benefit from this increase in parallelism are molecular dynamics simulations. In this paper, we describe our efforts at modernizing the ESPResSo++ simulation package for molecular dynamics by restructuring its particle data layout for efficient memory accesses and applying vectorization techniques to benefit the calculation of short-range non-bonded forces, which results in an overall three times speedup and serves as a baseline for further optimizations. We also implement fine-grained parallelism for multi-core CPUs through HPX, a C++ runtime system which uses lightweight threads and an asynchronous many-task approach to maximize concurrency. Our goal is to evaluate the performance of an HPX-based approach compared to the bulk-synchronous MPI-based implementation. This requires the introduction of an additional layer to the domain decomposition scheme that defines the task granularity. On spatially inhomogeneous systems, which impose a corresponding load-imbalance in traditional MPI-based approaches, we demonstrate that by choosing an optimal task size, the efficient work-stealing mechanisms of HPX can overcome the overhead of communication resulting in an overall 1.4 times speedup compared to the baseline MPI version.

97 MATHEMATICS AND COMPUTING↗

EUTERPE: A global gyrokinetic code for stellarator geometry

The current state of the EUTERPE code is described with emphasis on the implemented models and their numerical implementation. The code solves the multi-species electromagnetic gyrokinetic equations in the full volume of a three-dimensional domain. Noise reduction of the particle-in-cell method is achieved by using a δf-method and Fourier filters. The field equations are discretized with B-splines and the resulting system of equations is solved iteratively. For linear simulations a phase-factor transformation is applied in order to strongly reduce the necessary grid resolution. Apart from the full gyrokinetic model, other numerically less expensive hybrid models are also implemented. They are mainly tailored for comparison with fluid theory and for studying the interaction of the bulk plasma with fast particles. The code is parallelized for CPUs by particle and domain decomposition. Good scalability up to several thousand nodes is demonstrated.

97 MATHEMATICS AND COMPUTING↗

Long-time integration of parametric evolution equations with physics-informed DeepONets

Ordinary and partial differential equations (ODEs/PDEs) play a paramount role in analyzing and simulating complex dynamic processes across all corners of science and engineering. In recent years machine learning tools are aspiring to introduce new effective ways of simulating such equations, however existing approaches are not able to reliably return stable and accurate predictions across long temporal horizons. We aim to address this challenge by introducing an effective framework for learning evolution operators that map random initial conditions to associated ODE/PDE solutions within a short time interval. Such operators can be parametrized by deep neural networks that are trained in an entirely self-supervised manner without requiring one to generate any paired input-output observations. Global long-time predictions across a range of initial conditions can be then obtained by iteratively evaluating the trained model using each prediction as the initial condition for the next evaluation step. Here, this introduces a new approach to temporal domain decomposition that is shown to be effective in performing accurate long-time simulations for a wide range of parametric ODE and PDE systems, from wave propagation, to reaction-diffusion dynamics and stiff chemical kinetics, introducing a new way of rapidly emulating non-equilibrium processes in science and engineering.

97 MATHEMATICS AND COMPUTING↗

Evaluation of the spatial self-shielding impact for TRISO-based nuclear fuel depletion

Reactor physics analyses of nuclear cores with nuclear fuel concepts containing tristructural isotropic (TRISO) particles, such as pebbles or compact fuel elements, rely on various degrees of simplification to keep these highly heterogeneous problems computationally tractable. One such limitation regards the level of spatial discretization employed during burnup calculations, where traditionally only a limited number of spatial zones are modeled at the full core level and assume that the spectrum is constant within the fuel elements and TRISO particles in this depletion zone. This type of assumption neglects the impact of spatial self-shielding effect within the kernels (microscale level) as well as within the compact or pebbles (mesoscale level). Furthermore, the Monte Carlo code Serpent 2 contains many relevant features for efficiently modeling this type of geometry, including a collision-based domain decomposition intended for very large burnup calculations, which we leveraged for this work to quantify the impact of capturing neutron flux variations occurring at the micro- and mesoscale level on a series of high-temperature gas-cooled reactor fuel element depletion problems. While spatial self-shielding is observed at both scales, with differences from a volume-averaged burnup of ±7% within the kernels and ±2% between TRISO particles within the fuel element, the conjugated effect on nuclide inventories and multiplication factor are negligible, hence confirming that assuming a single average spectrum value may be sufficient for most applications.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Novel strategies for modal-based structural material identification

Here, we present modal-based methods for model calibration in structural dynamics, and address several key challenges in the solution of gradient-based optimization problems with eigenvalues and eigenvectors, including the solution of singular Helmholtz problems encountered in sensitivity calculations, non-differentiable objective functions caused by mode swapping during optimization, and cases with repeated eigenvalues. Unlike previous literature that relied on direct solution of the eigenvector adjoint equations, we present a parallel iterative domain decomposition strategy (Adjoint Computation via Modal Superposition with Truncation Augmentation) for the solution of the singular Helmholtz problems. For problems with repeated eigenvalues we present a novel Mode Separation via Projection algorithm, and in order to address mode swapping between inverse iterations we present a novel Injective mode ordering metric. We present the implementation of these methods in a massively parallel finite element framework with the ability to use measured modal data to extract unknown structural model parameters from large complex problems. A series of increasingly complex numerical examples are presented that demonstrate the implementation and performance of the methods in a massively parallel finite element framework [7], [5], using gradient-based optimization techniques in the Rapid Optimization Library (ROL) [21].

36 MATERIALS SCIENCE↗

Scalable freeform optimization of wide-aperture 3D metalenses by zoned discrete axisymmetry

We introduce a novel framework for design and optimization of 3D freeform metalenses that attains nearly linear scaling of computational cost with diameter, by breaking the lens into a sequence of radial “zones” with 𝑛-fold discrete axisymmetry, where 𝑛 increases with radius. This allows vastly more design freedom than imposing continuous axisymmetry, while avoiding the compromises of the locally periodic approximation (LPA) or scalar diffraction theory. Using a GPU-accelerated finite-difference time-domain (FDTD) solver in cylindrical coordinates, we perform full-wave simulation and topology optimization within each supra-wavelength zone. We validate our approach by designing millimeter and centimeter-scale, poly-achromatic, 3D freeform metalenses which outperform the state of the art. By demonstrating the scalability and resulting optical performance enabled by our “zoned discrete axisymmetry” (ZDA) and supra-wavelength domain decomposition, we highlight the potential of our framework to advance large-scale meta-optics and next-generation photonic technologies.

Sun, Mengdi [Wesleyan University]↗

Status of GPU capabilities within the Shift Monte Carlo radiation transport code

Shift is a general-purpose Monte Carlo (MC) radiation transport code for fission, fusion, and national security applications. Shift has been adapted to efficiently run on GPUs in order to leverage leadership-class supercomputers. This work presents Shift’s current GPU capabilities. These include core radiation transport capabilities for eigenvalue and fixed-source simulations, and support for non-uniform domain decomposition, Doppler broadening, free-gas elastic scattering, general-purpose geometry, hybrid MC/deterministic transport, and depletion. Transport results demonstrate a 2–5× GPU-to-CPU speedup on a per-node basis for an eigenvalue problem on the Frontier supercomputer and a 28× speedup for a fixed-source problem on the Summit supercomputer.

Biondo, Elliott [ORNL] (ORCID:0000000290881360)↗

Learning the structure of wind: A data-driven nonlocal turbulence model for the atmospheric boundary layer

In this work, we develop a novel data-driven approach to modeling the atmospheric boundary layer. This approach leads to a nonlocal, anisotropic synthetic turbulence model which we refer to as the deep rapid distortion (DRD) model. Our approach relies on an operator regression problem that characterizes the best fitting candidate in a general family of nonlocal covariance kernels parameterized in part by a neural network. This family of covariance kernels is expressed in Fourier space and is obtained from approximate solutions to the Navier–Stokes equations at very high Reynolds numbers. Each member of the family incorporates important physical properties such as mass conservation and a realistic energy cascade. The DRD model can be calibrated with noisy data from field experiments. After calibration, the model can be used to generate synthetic turbulent velocity fields. To this end, we provide a new numerical method based on domain decomposition which delivers scalable, memory-efficient turbulence generation with the DRD model as well as others. We demonstrate the robustness of our approach with both filtered and noisy data coming from the 1968 Air Force Cambridge Research Laboratory Kansas experiments. Using these data, we witness exceptional accuracy with the DRD model, especially when compared to the International Electrotechnical Commission standard.

17 WIND ENERGY↗

Preparation of the SU(3) lattice Yang-Mills vacuum with variational quantum methods

Studying QCD and other gauge theories on quantum hardware requires the preparation of physically interesting states. The variational quantum eigensolver provides a way of performing vacuum state preparation on quantum hardware. Here in this work, variational quantum eigensolver is applied to pure SU(3) lattice Yang-Mills on a single plaquette and one dimensional plaquette chains. Bayesian optimization and gradient descent were investigated for performing the classical optimization. Ansatz states for plaquette chains are constructed in a scalable manner from smaller systems using domain decomposition and a stitching procedure analogous to the density matrix renormalization group. Small examples are performed on IBM’s superconducting Manila processor.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Interface learning of multiphysics and multiscale systems

Complex natural or engineered systems comprise multiple characteristic scales, multiple spatiotemporal domains, and even multiple physical closure laws. To address such challenges, we introduce an interface learning paradigm and put forth a data-driven closure approach based on memory embedding to provide physically correct boundary conditions at the interface. To enable the interface learning for hyperbolic systems by considering the domain of influence and wave structures into account, we put forth the concept of upwind learning toward a physics-informed domain decomposition. The promise of the proposed approach is shown for a set of canonical illustrative problems. Here, we highlight that high-performance computing environments can benefit from this methodology to reduce communication costs among processing units in emerging machine-learning-ready heterogeneous platforms toward exascale era.

42 ENGINEERING↗

An Algorithmic and Software Pipeline for Very Large Scale Scientific Data Compression with Error Guarantees

Efficient data compression is becoming increasingly critical for storing scientific data because many scientific applications produce vast amounts of data. This paper presents an end-to-end algorithmic and software pipeline for data compression that guarantees both error bounds on primary data (PD) and derived data, known as Quantities of Interest (QoI).We demonstrate the effectiveness of the pipeline by compressing fusion data generated by a large-scale fusion code, XGC, which produces tens of petabytes of data in a single day. We demonstrate that the compression is conducted by setting aside computational resources known as staging nodes, and does not impact the simulation performance. For efficient parallel I/O, the pipeline uses ADIOS2, which many codes such as XGC already use for their parallel I/O. We show that our approach can compress the data by two orders of magnitude while guaranteeing high accuracy on both the PD and the QoIs. Further, the amount of resources required by compression is a few percent of the resources required by simulation while ensuring that the compression time for each stage is less than the corresponding simulation time.This pipeline consists of three main steps. The first step decomposes the data using domain decomposition into small subdomains. Each subdomain is then compressed independently to achieve a high level of parallelism. The second step uses existing techniques that guarantee error bounds on the primary data for each subdomain. The third step uses a post-processing optimization technique based on Lagrange multipliers to reduce the QoI errors for data corresponding to each subdomain. The Lagrange multipliers generated can be further quantized or truncated to increase the compression level. All of the above characteristics of our approach make it highly practical to apply on-the-fly compression while guaranteeing errors on QoIs that are critical to the scientists.

Banerjee, Tania↗

Tuning Multigrid Methods with Robust Optimization and Local Fourier Analysis

Local Fourier analysis is a useful tool for predicting and analyzing the performance of many efficient algorithms for the solution of discretized PDEs, such as multigrid and domain decomposition methods. The crucial aspect of local Fourier analysis is that it can be used to minimize an estimate of the spectral radius of a stationary iteration, or the condition number of a preconditioned system, in terms of a symbol representation of the algorithm. In practice, this is a “minimax” problem, minimizing with respect to solver parameters the appropriate measure of work, which involves maximizing over the Fourier frequency. Often, several algorithmic parameters may be determined by local Fourier analysis in order to obtain efficient algorithms. Analytical solutions to minimax problems are rarely possible beyond simple problems; the status quo in local Fourier analysis involves grid sampling, which is prohibitively expensive in high dimensions. In this paper, we propose and explore optimization algorithms to solve these problems efficiently. Finally, several examples, with known and unknown analytical solutions, are presented to show the effectiveness of these approaches.

97 MATHEMATICS AND COMPUTING↗