Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “mesh data structure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Discrete global grid system-based flow routing datasets in the Amazon and Yukon basins

Abstract. Discrete global grid systems (DGGS) are emerging spatial data structures widely used to organize geospatial datasets across scales. While DGGS have found applications in various scientific disciplines, including atmospheric science and ecology, their integration into physically based hydrological models and Earth system models (ESMs) has been hindered by the lack of flow routing datasets based on DGGS. In response to this gap, this study pioneers the development of new flow routing datasets using icosahedral Snyder equal-area (ISEA) DGGS and a novel mesh-independent flow direction model. We present flow routing datasets for two large basins, the tropical Amazon River basin and the Arctic Yukon River basin. These datasets (1) facilitate the adoption of DGGS for hydrological models and (2) provide flow routing inputs for evaluation of DGGS-based flow routing in the Amazon and Yukon river basins. The data are available at https://doi.org/10.5281/zenodo.8377765 (Liao, 2023).

54 ENVIRONMENTAL SCIENCES↗

Aeroelastic Analysis Using Deforming Cartesian Grids

Ongoing work in air-vehicle design illustrates the potential of advanced concepts to provide significant improvements in efficiency; but with their incorporation of lightweight flexible structures, such configurations may require active control systems to ensure reliability and safety. However, many contemporary analysis methods are inefficient for aeroelastic analysis and design of such configurations. This paper describes the development of a new approach that automates the geometry setup, mesh generation, and assembly of fluid–structural coupling interfaces to enable efficient aeroelastic and aeroservoelastic analysis of advanced concepts. The core elements for this approach are a cut-cell Cartesian grid-based computational fluid dynamics solver, a nonlinear beam element structural model, a conservative fluid–structural interface treatment, and the formulation and implementation of a new deforming grid capability within the cut-cell Cartesian grid solver. In this paper, emphasis is on this latter component with detailed description given of the mesh motion strategy, evaluation of fluxes and structural loads at the surface, and computation of geometrical properties such as cell volume, directed face areas, centroids, and motion-induced fluxes for deforming Cartesian grids required to advance the flow states. Aeroelastic simulations exercising the capability show favorable agreement with data and predictions in the literature for subsonic and supersonic applications.

97 MATHEMATICS AND COMPUTING↗

Reduced order modeling for flow and transport problems with Barlow Twins self-supervised learning

Abstract We propose a unified data-driven reduced order model (ROM) that bridges the performance gap between linear and nonlinear manifold approaches. Deep learning ROM (DL-ROM) using deep-convolutional autoencoders (DC–AE) has been shown to capture nonlinear solution manifolds but fails to perform adequately when linear subspace approaches such as proper orthogonal decomposition (POD) would be optimal. Besides, most DL-ROM models rely on convolutional layers, which might limit its application to only a structured mesh. The proposed framework in this study relies on the combination of an autoencoder (AE) and Barlow Twins (BT) self-supervised learning, where BT maximizes the information content of the embedding with the latent space through a joint embedding architecture. Through a series of benchmark problems of natural convection in porous media, BT–AE performs better than the previous DL-ROM framework by providing comparable results to POD-based approaches for problems where the solution lies within a linear subspace as well as DL-ROM autoencoder-based techniques where the solution lies on a nonlinear manifold; consequently, bridges the gap between linear and nonlinear reduced manifolds. We illustrate that a proficient construction of the latent space is key to achieving these results, enabling us to map these latent spaces using regression models. The proposed framework achieves a relative error of 2% on average and 12% in the worst-case scenario (i.e., the training data is small, but the parameter space is large.). We also show that our framework provides a speed-up of $$7 \times 10^{6}$$ 7 × 10 6 times, in the best case, and $$7 \times 10^{3}$$ 7 × 10 3 times on average compared to a finite element solver. Furthermore, this BT–AE framework can operate on unstructured meshes, which provides flexibility in its application to standard numerical solvers, on-site measurements, experimental data, or a combination of these sources.

97 MATHEMATICS AND COMPUTING↗

On the Response of a Herschel–Bulkley Fluid Due to a Moving Plate

In this paper, we study the boundary-layer flow of a Herschel–Bulkley fluid due to a moving plate; this problem has been experimentally investigated by others, where the fluid was assumed to be Carbopol, which has similar properties to cement. The computational fluid dynamics finite volume method from the open-source toolbox/library OpenFOAM is used on structured quad grids to solve the mass and the linear momentum conservation equations using the solver “overInterDyMFoam” customized with non-Newtonian viscosity libraries. The governing equations are solved numerically by using regularization methods in the context of the overset meshing technique. The results indicate that there is a good comparison between the experimental data and the simulations. The boundary layer thicknesses are predicted within the uncertainties of the measurements. The simulations indicate strong sensitivities to the rheological properties of the fluid.

36 MATERIALS SCIENCE↗

DS-GL: Advancing Graph Learning via Harnessing the Power of Nature within Dynamic Systems

With the rapid digitization of the world, an increasing number of real-world applications are turning to nonEuclidean data, modeled as graphs. Due to their intrinsic high complexity and irregularity, learning from graph data demands tremendous computational power. Recently, CMOS-compatible Ising machines, i.e., dynamic systems composed of CMOS components, have emerged as a new approach that harnesses the inherent power of natural annealing within dynamic systems to efficiently resolve binary optimization problems and have been adopted for traditional graph computation, such as max-cut. However, when performing complex Graph Learning (GL) tasks, Ising machines face significant hurdles: (i) they are inherently binary and thus ill-suited for real-valued problems; (ii) their expensive all-to-all coupling network that guarantees effective natural annealing poses daunting scalability concerns. To address these challenges, this paper proposes a nature-powered graph learning framework dubbed DS-GL, which is the first effort to transform the process of solving graph learning problems into the natural annealing process within a parameterized dynamic system embodied as a CMOS chip. To tackle the two major hurdles, DS-GL first augments the Ising machine architecture to modify the self-reaction term of its Hamiltonian function from linear to quadratic, effectively serving as an energy regulator. This adjustment maintains the system’s original physical interpretation while enabling it to process continuous, real-valued data. Second, to address the scaling issue, DS-GL further upgrades the real-valued dense Ising machine by decomposing it into a mesh-based multi-PE dynamic system that supports efficient distributed spatial-temporal co-annealing across different PEs through sparse interconnects. By exploiting the inherent sparsity and component structures in real-world graphs, DS-GL is able to map complex graph learning tasks onto the scalable dynamic system while maintaining high accuracy. Evaluations with three diverse GL applications across six real-world datasets, including traffic flow and COVID-19 prediction, show that DS-GL can deliver from 102× to 106× speedups and 500× energy reduction over Graph Neural Networks on GPUs, with 5% - 20% accuracy enhancement.

Song, Ruibing↗

Deciphering supramolecular and polymer-like behavior in metallogels: real-time insights into temperature-modulated gelation and rapid self-assembly dynamics

Bis(pyridyl) urea-based gelators, namely L2 and its isomeric mixture ( L1 + L2 ), are known to self-assemble into 1D architectures capable of inducing supramolecular gelation. Coordination with metal ions such as Ag( I ), Cu( II ), and Fe( III ) introduces structural reinforcement, enabling the formation of distinct 3D networks governed by metal-specific coordination geometries. Here, we present a comprehensive investigation into the temperature-responsive behavior (20–60 °C) of L2 and L1 + L2 , both in the absence and presence of Ag( I ), Dy( III ), Fe( III ), Cu( II ), and Ho( III ), using real-time small-angle neutron scattering (SANS). To probe long-term structural evolution/kinetics of self-assembly, real-time small-angle X-ray scattering (SAXS) was employed on L2 + Ag gels, complemented by differential scanning calorimetry (DSC) to evaluate thermal transitions. Our results reveal strikingly divergent gelation behaviors: L2 forms a highly rigid, covalent polymer-like network, while L1 + L2 exhibits remarkable thermal adaptability. Upon metal coordination, the assemblies exhibit pronounced crystallinity and exceptional thermal stability, as evidenced by persistent Bragg reflections and invariant d-spacings. Intriguingly, L2 : Fe (2 : 1) and L1 : L2 : Fe (0.5 : 0.5 : 1) in acetonitrile-d 3 (ACN-d 3 ) deviate from this trend, forming thermally labile amorphous gels. These systems show a complete loss of crystalline order, reduced Porod exponents—indicative of collapsed or branched fiber morphologies—and prominent melting and glass transition events in DSC. Fitting SANS and SAXS data to the correlation length model unveiled insightful nanostructural features. While most systems displayed minimal temperature-induced variation in mesh size or surface morphology, L2 : Ag in dimethyl sulfoxide-d 6 (DMSO-d 6 )/D 2 O and L2 : Fe (1 : 1) in ACN-d 3 exhibited a rare combination of thermally stable correlation lengths and increasing high- q exponents—strongly suggesting progressive fiber densification or surface smoothing within a robust gel framework. These findings highlight the tunability and structural resilience of supramolecular gels through precise control of ligand architecture, metal coordination, and temperature, offering valuable design principles for functional soft materials.

Pajoubpong, Jinnipha [Univ. of Cincinnati, OH (Uni↗

SAVY-4000 Finite-Element Drop Test Analysis

PFE Auxiliary Systems conducted drop testing on SAVY-4000 containers to evaluate structural response under 12-foot drop conditions. In support of that effort, a finite-element modeling capability was developed to simulate drop response across multiple container sizes and impact orientations. The purpose of this work was to provide a consistent analysis framework that could support interpretation of testing, compare response trends across multiple configurations, and generate quantities of interest for later comparison with experimental data. More broadly, the analysis and testing were intended to assess whether the containers continued to perform their primary function after a 12-foot drop, namely maintaining structural integrity and containment of the contents. The modeling approach combined an implicit preload analysis with an explicit drop simulation so that each drop event began from a mechanically realistic assembled condition, including compression of the silicone O-ring. Separate models were developed for 2-quart, 5-quart, 12-quart, and 10-gallon containers. The results were evaluated in terms of strain-gauge response, collar-lid gap behavior, and accumulated plastic strain. In addition, parametric studies were performed on the 2-quart container to assess sensitivity to O-ring stiffness, friction, canister thickness, geometry tolerance, and mesh density. The simulations showed that predicted drop responses depended strongly on both container size and drop orientation. Gap metrics identified cases in which the predicted collar-lid opening exceeded the nominal O-ring cross-section threshold, while plastic strain metrics identified localized regions of elevated permanent deformation. Parametric studies showed that the predicted response was especially sensitive to the assumed O-ring stiffness and contact friction, while the geometry tolerance study produced smaller changes in the cases examined. The main value of this work was that it established a repeatable modeling and simulation workflow to support drop-test implementation, evaluate effects of future configuration changes, and understand modeling assumptions that most influenced predicted response. At the current stage, the results were viewed as preliminary model predictions rather than validated predictions. The next step would be to compare drop-test data to the model so that predictive values of the workflow could be refined and used with greater confidence to assess whether the containers maintained structural integrity and containment of the contents after a 12-foot drop.

42 ENGINEERING↗

A hybrid architecture for volt-var control in active distribution grids

Modern active distribution grids are characterized by the increasing penetration of distributed energy resources (DERs). The proper coordination and scheduling of a large numbers of these small-scale and spatially distributed DERs is necessary, and warrants the use of novel distributed approaches. In this paper, we propose a hybrid volt-var control architecture for the distribution grid, which leverages existing centralized and local approaches to planning, decision making, and control, and augments it with distributed optimization and distributed control for DER management. First, we propose a convex model to describe the power physics of distribution grids of meshed topology and unbalanced structure, based on current injection and McCormick Envelopes. Second, we employ the distributed proximal atomic coordination (PAC) algorithm to coordinate DERs to provide voltage support. We implement volt-var optimization by optimally coordinating DERs including PV smart inverters and demand response. We present results using the IEEE-34 bus network, using real data from a distribution feeder in Hawaii, to model load and PV generation. Different levels of DER penetration and objective functions are simulated. Finally, our results show the need for the coordination of DERs to improve voltage profiles, even in networks with existing voltage control devices. Further, we show the need for flexible reactive power capabilities to achieve desired grid performance.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Development of Numerical Model of Metal Foam with PCM for the Estimation of Effective Thermal Conductivity

Global warming due to climate change is a threat to humankind. Nuclear energy is one of the promising solutions to reduce fossil fuel usage. Nuclear energy can handle the base load, compensating for the volatility of renewable energy. If nuclear energy could achieve load following capability, the combination with renewable energy would be more suitable. Thermal energy storage (TES) is one of the options for enabling load following of nuclear reactors. The TES makes it possible to store surplus nuclear thermal energy and release it later as needed. In Idaho National Laboratory (INL), a new concept of latent heat TES integrated with high-temperature heat pipe has been proposed and is under development, which is called Heat pipe-Integrated Thermal Battery (HITB). HITB exchanges thermal energy between the reactor system and TES via heat pipe. The heat transferred to TES medium, made of phase change material (PCM), stores energy as sensible heat and/or latent heat. As PCM typically has poor thermal conductivity, however, various heat transfer enhancement techniques are required to achieve a rapid charging cycle. There are many techniques to enhance the heat transfer ability of TES medium such as disk, fin, and metal foam. Among them, metal foam is an appropriate option to enhance the heat transfer because it maximizes the heat transfer area through metal wicks. Metal foam is a lightweight metal structure that has a high porosity of over 0.9. The typical materials for metal foam are Aluminum, Copper, Nickel, and Silicon Carbide (SiC). Metal foam not only enhances heat transfer via conduction but also increases contact surface area. In the HITB design , the metal foam is being considered as one of the options to enhance the heat transfer of TES medium (PCM) [1]. To predict the enhanced thermal performance of TES, one should properly estimate the effective thermal conductivity of metal foam combined with PCM material or calculate heat transfer in distributed model. There are many experimental works that provides effective thermal conductivity of metal foam with various PCM [2,3]. Also, many theoretical models were developed based on the unit cell model of metal foam [4,5]. With a distributed model, on the other hand, detail heat transfer characteristics between metal foam and PCM material can be analyzed considering the geometry or buoyancy effect. However, due to the complex geometry of metal foam pores, the computational cost for three-dimensional modeling highly increases. Therefore, if metal foam structure can be modeled in simple and repetitive design, the computational cost would decrease Among the various metal foam models [2], lattice model is one of the simple and extendable design. The porosity and pores per inch (PPI) can be characterized by the size and spatial distance of lattice structure. If the three-dimensional metal foam model consists of lattice structure could properly estimate the heat transfer, which is characterized by effective thermal conductivity, it would be a good option to assess the thermal performance of metal foam with PCM. In this study, a three-dimensional numerical model was developed to simulate conductive heat transfer between metal foam and PCM. The three-dimensional lattice structure of square pillars was selected as a basic structure of the metal foam. The calculation result was characterized by the effective thermal conductivity of the whole domain. A sensitivity study was conducted for mesh size, domain size, and PPI to check whether the calculation result gives a converged result or not. Lastly, the effective thermal conductivity from the lattice model was compared with existing experimental data to validate the model result

25 ENERGY STORAGE↗

A projection method for particle resampling

Particle discretizations of partial differential equations are advantageous for high-dimensional kinetic models in phase-space due to their better scalability than continuum approaches with respect to dimension. Complex processes collectively referred to as particle noise hamper long time simulations with particle methods. One approach to address this problem is particle mesh adaptivity, or remapping, known as particle resampling and remeshing. Here, this work introduces a resampling method that projects particles to and from a (finite element) function space. The method is simple, using standard sparse linear algebra and finite element techniques, and it preserves all moments up to the order of a polynomial represented exactly by the continuum function space. It is distinguished from most other mesh-based methods in that new particle positions and number are decoupled from the mesh, allowing particle and continuum meshes to be adapted relatively independently. While this work is developed with structured particle and continuum phase-space grids on 1X + 1V Vlasov-Poisson models of Landau damping and two-stream instability, the method is well-suited to unstructured grids. Stable long time dynamics are demonstrated up to time T = 500. Reproducibility artifacts and data are publicly available.

Kinetic methods↗

Domain architecture and catalysis of the Staphylococcus aureus fatty acid kinase

Fatty acid kinase (Fak) is a two-component enzyme that generates acyl-phosphate for phospholipid synthesis. Fak consists of a kinase domain protein (FakA) that phosphorylates a fatty acid enveloped by a fatty acid binding protein (FakB). The structural basis for FakB function has been established, but little is known about FakA. Here, we used limited proteolysis to define three separate FakA domains: the amino terminal FakA_N, the central FakA_L, and the carboxy terminal FakA_C. The isolated domains lack kinase activity, but activity is restored when FakA_N and FakA_L are present individually or connected as FakA_NL. The X-ray structure of the monomeric FakA_N captures the product complex with ADP and two Mg 2+ ions bound at the nucleotide site. The FakA_L domain encodes the dimerization interface along with conserved catalytic residues Cys240, His282, and His284. AlphaFold analysis of FakA_L predicts the catalytic residues are spatially clustered and pointing away from the dimerization surface. Furthermore, the X-ray structure of FakA_C shows that it consists of two subdomains that are structurally related to FakB. Analytical ultracentrifugation demonstrates that FakA_C binds FakB, and site-directed mutagenesis confirms that a positively charged wedge on FakB meshes with a negatively charged groove on FakA_C. Finally, small angle X-ray scattering analysis is consistent with freely rotating FakA_N and FakA_C domains tethered by flexible linkers to FakA_L. These data reveal specific roles for the three independently folded FakA protein domains in substrate binding and catalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Comparative Study of the Perceptual Sensitivity of Topological Visualizations to Feature Variations

Color maps are a commonly used visualization technique in which data are mapped to optical properties, e.g., color or opacity. Color maps, however, do not explicitly convey structures (e.g., positions and scale of features) within data. Topology-based visualizations reveal and explicitly communicate structures underlying data. Although our understanding of what types of features are captured by topological visualizations is good, our understanding of people's perception of those features is not. Further, this paper evaluates the sensitivity of topology-based isocontour, Reeb graph, and persistence diagram visualizations compared to a reference color map visualization for synthetically generated scalar fields on 2-manifold triangular meshes embedded in 3D. In particular, we built and ran a human-subject study that evaluated the perception of data features characterized by Gaussian signals and measured how effectively each visualization technique portrays variations of data features arising from the position and amplitude variation of a mixture of Gaussians. For positional feature variations, the results showed that only the Reeb graph visualization had high sensitivity. For amplitude feature variations, persistence diagrams and color maps demonstrated the highest sensitivity, whereas isocontours showed only weak sensitivity. These results take an important step toward understanding which topology-based tools are best for various data and task scenarios and their effectiveness in conveying topological variations as compared to conventional color mapping.

97 MATHEMATICS AND COMPUTING↗

Uncertainty Quantification and Sensitivity Analysis for Simulation of Hostile Blast Events [Slides]

A mesh convergence study shows that shock arrival times converge for basic flow shots. Sensitivity study reveals a surprisingly high sensitivity to the placement of the HE package. Bayesian inference tools used on the experimental datasets suggest this is a real effect. Sensitivity analysis of ambient conditions such as temperature and pressure reveal minimal effect. Porting of pressure data from QUINOA simulations to perform structural analysis on the aeroshell geometry was successful, and shows a strong dependence of payload response to the angle of attack.

42 ENGINEERING↗

HfZr_BCC_SolidSolution_128atoms_VASP6

We performed density functional theory (DFT) calculations for body-centered-cubic (BCC) structures with 128 lattices sites of solid solution binary alloys hafnium-zirconium (Hf-Zr). The electronic structures of alloys have been calculated using Vienna Ab initio Simulation Package (VASP). Within this package the DFT approach is used to reduce many-body Schrodinger equation to set of single particle Kohn-Sham (KS) equations. The generalized electronic exchange-correlation functional is described by generalized gradient approximation with the Perdew-Burke-Ernzerhof parametrization. The electron-ion interactions is described by pseudopotentials developed within the plane-wave basis projector augmented-wave (PAW) approach. These pseudopotentials are available at the VASP portal (http://cms.mpi.univie.ac.at/vasp/). Our calculations have been run with the pseudopotentials treating s and p semi-core states as valence in case for the elements Hf and Zr. The electronic densities and potentials are expanded over plane-waves with energy cutoff of 350 eV. 2x2x2 k-mesh and normal precision were used. The alloys were modeled by supercell containing 128 randomly distributed atoms. At initial step the atoms occupy perfect bcc lattice cites. This initial structure was optimized until energy changes less than 1e-6 eV, while forces acting on atoms don't exceed 1e-2 eV/angstrom. The electron-ion interaction is described by PAW pseudopotentials. The calculations have been collected by sampling chemical compositions across the entire compositional range. The chemical compositions have been sampled by progressively changing the number of atoms per constituent by 4. For each chemical composition of binaries and ternaries, the first-principle calculations have been run for 100 randomized arrangements of the constituents on the BCC lattice sites. We collected data for a total of 3,100 randomized atomic structures over 31 chemical compositions.

36 MATERIALS SCIENCE↗

AMReX v2024

The software framework, AMReX, supports the development of block-structured adaptive mesh refinement (AMR) algorithms for solving systems of partial differential equations. AMR reduces the computational cost and memory footprint compared to a uniform mesh while preserving the essential local descriptions of different physical processes in complex multiphysics algorithms. AMR uses a hierarchical representation of the solution at multiple levels of resolution where the solution on each level is defined on the union of data containers at that resolution. These data containers, which represent the solution over a logically rectangular subregion of the domain, can contain field data defined on a mesh, Lagrangian particles or combinations of both. In addition to these basic data types, AMReX supports a multilevel embedded boundary representation of complex geometry; linear solvers for cell-centered and nodal data; asynchronous I/O in a native format readable by ParaView, VisIt and yt; and interfaces to hypre and PETSc solvers. AMReX enables applications to run on distributed memory architectures with multicore CPUs and with GPU accelerators. AMReX uses a lightweight abstraction layer that effectively hides the details of the architecture from the application. The framework currently supports CUDA, HIP and SYCL for GPU acceleration and OpenMP for multi-core CPU architectures.

Almgren, Ann↗

HfTa_BCC_SolidSolution_128atoms_VASP6

We performed density functional theory (DFT) calculations for body-centered-cubic (BCC) structures with 128 lattices sites of solid solution binary alloys hafnium-tantalum (Hf-Ta). The electronic structures of alloys have been calculated using Vienna Ab initio Simulation Package (VASP). Within this package the DFT approach is used to reduce many-body Schrodinger equation to set of single particle Kohn-Sham (KS) equations. The generalized electronic exchange-correlation functional is described by generalized gradient approximation with the Perdew-Burke-Ernzerhof parametrization. The electron-ion interactions is described by pseudopotentials developed within the plane-wave basis projector augmented-wave (PAW) approach. These pseudopotentials are available at the VASP portal (http://cms.mpi.univie.ac.at/vasp/). Our calculations have been run with the pseudopotentials treating s and p semi-core states as valence in case for the elements Hf and Ta. The electronic densities and potentials are expanded over plane-waves with energy cutoff of 350 eV. 2x2x2 k-mesh and normal precision were used. The alloys were modeled by supercell containing 128 randomly distributed atoms. At initial step the atoms occupy perfect BCC lattice sites. This initial structure was optimized until energy changes less than 1e-6 eV, while forces acting on atoms don't exceed 1e-2 eV/angstrom. The electron-ion interaction is described by PAW pseudopotentials. The calculations have been collected by sampling chemical compositions across the entire compositional range. The chemical compositions have been sampled by progressively changing the number of atoms per constituent by 4. For each chemical composition of binaries and ternaries, the first-principle calculations have been run for 100 randomized arrangements of the constituents on the BCC lattice sites. We collected data for a total of 3,040 randomized atomic structures over 31 chemical compositions. Further methodological and structural information is contained in the dataset README.txt file.

36 MATERIALS SCIENCE↗

Advancing SiC clad fuel performance model: bridging micro- and macro-scale models and experimental validation

This report presents a workflow for advancing fuel performance modeling of SiC composite cladding for light-water reactors by linking microscale, experimental data-informed finite element analysis with rod-scale fuel performance codes such as BISON. The workflow uses X-ray computed tomography (XCT) to capture the actual geometry and processing-induced defects of as-fabricated SiC composite tube specimens, particularly porosity and wall-thickness variations, and converts the segmented XCT volumes into image-based finite element meshes for high fidelity structural analysis.

Koyanagi, Takaaki [Oak Ridge National Laboratory (↗

Parthenon—a performance portable block-structured adaptive mesh refinement framework

On the path to exascale the landscape of computer device architectures and corresponding programming models has become much more diverse. While various low-level performance portable programming models are available, support at the application level lacks behind. To address this issue, we present the performance portable block-structured adaptive mesh refinement (AMR) framework Parthenon, derived from the well-tested and widely used Athena++ astrophysical magnetohydrodynamics code, but generalized to serve as the foundation for a variety of downstream multi-physics codes. Parthenon adopts the Kokkos programming model, and provides various levels of abstractions from multidimensional variables, to packages defining and separating components, to launching of parallel compute kernels. Parthenon allocates all data in device memory to reduce data movement, supports the logical packing of variables and mesh blocks to reduce kernel launch overhead, and employs one-sided, asynchronous MPI calls to reduce communication overhead in multi-node simulations. Using a hydrodynamics miniapp, we demonstrate weak and strong scaling on various architectures including AMD and NVIDIA GPUs, Intel and AMD x86 CPUs, IBM Power9 CPUs, as well as Fujitsu A64FX CPUs. At the largest scale on Frontier (the first TOP500 exascale machine), the miniapp reaches a total of 1.7 × 10 13 zone-cycles/s on 9216 nodes (73,728 logical GPUs) at [Formula: see text] weak scaling parallel efficiency (starting from a single node). In combination with being an open, collaborative project, this makes Parthenon an ideal framework to target exascale simulations in which the downstream developers can focus on their specific application rather than on the complexity of handling massively-parallel, device-accelerated AMR.

97 MATHEMATICS AND COMPUTING↗