Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “performance portable algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Speeding Up Hartree–Fock in JuliaChem with Density Fitting

In this work, the density fitting (DF) approximation is added to the restricted Hartree–Fock (RHF) implementation in the JuliaChem computational chemistry code. Utilizing a DF algorithm that uses symmetry and integral screening, a significant reduction in time to compute the Fock matrix is achieved. The symmetry and screening DF-RHF techniques were adapted to be performed on graphics processing units (GPUs), which are well suited to perform the matrix multiplications that comprise the bulk of the Fock build time in DF-RHF. The JuliaChem DF-RHF GPU algorithm employs a novel approach that automatically switches between two DF-RHF algorithms depending on the number of basis functions in the calculation. The JuliaChem GPU DF-RHF implementation demonstrates up to 2× speedup for Fock build times compared to the existing best-in-class GPU DF-RHF implementation by operating directly on screened intermediate matrices. Due to the high portability of the Julia language code, the JuliaChem CPU and GPU DF-RHF implementations could be benchmarked on a variety of CPU and GPU architectures from multiple hardware vendors.

Hayes, John J. [Ames Laboratory, and Iowa State Un↗

2025 Advances in NekRS: Supporting improved performance for nuclear applications

This report presents several 2025 advancements in NekRS, a high-fidelity spectral element CFD code developed at Argonne National Laboratory to support the NEAMS thermal-hydraulics program. The forthcoming v25 release consolidates several of these advances, adding new features for portability across heterogeneous GPU architectures, real-time in situ visualization, improved turbulence modeling, and conjugate heat transfer coupling. Over the past year, NekRS has demonstrated strong scalability and performance on DOE’s leading exascale platforms, including Aurora and Frontier, confirming its readiness for some of the largest and most complex simulations attempted to date. These achievements provide a powerful new platform for high-fidelity data generation, which in turn supports the development and validation of advanced closure models critical for reactor safety and design. Significant algorithmic innovations have also been introduced. A new global runtime h-refinement capability simplifies workflows by reducing mesh preparation burdens and enabling coarse-to-fine restarts. Building on this, a novel multigrid strategy was implemented to accelerate pressure and transport solves at scale, addressing long-standing bottlenecks in exascale CFD. Together, these developments improve both the efficiency and accessibility of high-fidelity simulations for reactor-relevant problems. Collectively, these enhancements represent a major step forward in simulation technology, positioning NekRS as a cornerstone of NEAMS efforts to enable accurate, efficient, and scalable high-fidelity analysis of advanced nuclear systems.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Automated Subpixel Snow Parameter Mapping with AVIRIS Data

We describe an automated algorithm (MEMSCAG) for mapping subpixel snow covered area (SCA) and snow grain size with AVIRIS data. The algorithm is based on the multiple endmember approach to spectral mixture analysis in which the spectral endmembers and the number of endmembers can vary on a pixel-by-pixel basis. This approach accounts for surface cover heterogeneity within a scene. The mixture analysis runs on endmembers from a spectral library of snow, vegetation, rock, soil, and lake ice spectra. Snow endmembers of varying grain size were produced with a radiative transfer model. All non-snow endmembers were collected with a portable field spectrometer. Mapping is performed through sequential 2-endmember, 3-endmember, and 4- endmember mixture model runs, each subject to constraints on RMS, residuals, fractions and priority. Grain size is determined by the grain size of the snow endmember used in the optimal mixture model. We apply MEMSCAG to AVIRIS data collected over Mammoth Mountain, CA and the northern site of the BOREAS in Manitoba, Canada. MEMSCAG produces appropriate snow covered area estimates in all regions. A preliminary comparison of grain size estimates from MEMSCAG with field measurements demonstrates high accuracy.

Painter, Thomas H.↗

TROJID: A portable software package for upper-stage trajectory optimization

Performance optimization for upper-stage exoatmospheric vehicles often is performed within the framework of a full capability trajectory simulation package requiring either a large mainframe computer or powerful work-station. Since these software packages tend to include capabilities providing for high-fidelity boost and reentry simulations, the programs usually are quite large and not very portable. The program TROJID is an attempt to provide an environment for the optimization of upper-stage trajectories within a small package capable of being run on a standard desktop microcomputer. Utilizing a state-of-the-art nonlinear programming algorithm and a trajectory simulator implementing impulsive burns and an analytic coast phase propagator, TROJID is capable of producing trajectories for optimal multi-burn upper-stage orbit transfers. The package has been designed to allow full generality in definition of both the trajectory simulator and the parameter optimization problem.

Hammes, Steven M.↗

Batched Sparse Linear Algebra (Final Report for Subcontract B648960)

This report finalizes design specifications for developing batched kernels for small tensor operations for unassembled matrix-free iterative solvers, batched solvers for partially assembled operators, and batched solvers with support for various sparse formats. The outcome of the project milestones is a set of interfaces to Batched Sparse LA solvers running on hardware accelerators for use in ECP Libraries and Applications. It is part of the development of sparse batched kernels, solvers/preconditioners as well as creating interoperability in xSDK libraries with sparse and dense batched functions to benefit ECP applications. The participants included representatives from ECP libraries (not limited to the xSDK project), applications, and vendors (AMD, Intel, and NVIDIA). Batched sparse linear algebra solvers form the new frontier for algorithmic development and performance engineering. Many applications (ECP and non-ECP alike) require simultaneous solutions of small linear systems of equations that are structurally sparse. To move towards high hardware utilization, it is important to provide these applications with appropriate interfaces to efficient batched sparse solvers running on modern hardware accelerators. We present interface designs in use by HPC software libraries supporting batched sparse linear algebra and the development of sparse batched kernel codes for solvers and preconditioners. We also address the potential interoperability opportunities to keep the software portable between the major hardware accelerators from AMD, Intel, and NVIDIA. The presented interface specifications includes batched band, sparse iterative, and sparse direct solvers. This report summarizes progress in Kokkos Kernels and the xSDK libraries MAGMA, Ginkgo, hypre, SUNDIALS, and SuperLU_dist.

97 MATHEMATICS AND COMPUTING↗

A Novel 24 GHz One-Shot, Rapid and Portable Microwave Imaging System

Development of microwave and millimeter wave imaging systems has received significant attention in the past decade. Signals at these frequencies penetrate inside of dielectric materials and have relatively small wavelengths. Thus. imaging systems at these frequencies can produce images of the dielectric and geometrical distributions of objects. Although there are many different approaches for imaging at these frequencies. they each have their respective advantageous and limiting features (hardware. reconstruction algorithms). One method involves electronically scanning a given spatial domain while recording the coherent scattered field distribution from an object. Consequently. different reconstruction or imaging techniques may be used to produce an image (dielectric distribution and geometrical features) of the object. The ability to perform this accurate~v and fast can lead to the development of a rapid imaging system that can be used in the same manner as a video camera. This paper describes the design of such a system. operating at 2-1 GHz. using modulated scatterer technique applied to 30 resonant slots in a prescribed measurement domain.

Ghasr, M. T.↗

Evaluation of Portable Programming Models to Accelerate LArTPC Detector Simulations

The Liquid Argon Time Projection Chamber (LArTPC) technology is widely used in high energy physics experiments, including the upcoming Deep Underground Neutrino Experiment (DUNE). Accurately simulating LArTPC detector responses is essential for analysis algorithm development and physics model interpretations. Accurate LArTPC detector response simulations are computationally demanding, and can become a bottleneck in the analysis workflow. Compute devices such as General-Purpose Graphics Processing Units (GPGPUs) have the potential to substantially accelerate simulations compared to traditional CPU-only processing. The software development that requires often carries the cost of specialized code refactorization and porting to match the target hardware architecture. With the rapid evolution and increased diversity of the computer architecture landscape, it is highly desirable to have a portable solution that also maintains reasonable performance. We report our ongoing effort in evaluating Kokkos as a basis for this portable programming model using LArTPC simulations in the context of the Wire-Cell Toolkit, a C++ library for LArTPC simulations, data analysis, reconstruction and visualization.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Evaluation of Portable Programming Models to Accelerate LArTPC Detector Simulations

The Liquid Argon Time Projection Chamber (LArTPC) technology is widely used in high energy physics experiments, including the upcoming Deep Underground Neutrino Experiment (DUNE). Accurately simulating LArTPC detector responses is essential for analysis algorithm development and physics model interpretations. Accurate LArTPC detector response simulations are computationally demanding, and can become a bottleneck in the analysis workflow. Compute devices such as General-Purpose Graphics Processing Units (GPGPUs) have the potential to substantially accelerate simulations compared to traditional CPU-only processing. The software development for these compute accelerators often carries the cost of specialized code refactorization and porting to match the target hardware architecture. With the rapid evolution and increased diversity of the computer architecture landscape, it is highly desirable to have a portable solution that also maintains reasonable performance. We report our ongoing effort in evaluating Kokkos as a basis for this portable programming model using LArTPC simulations in the context of the Wire-Cell Toolkit, a C++ library for LArTPC simulations, data analysis, reconstruction and visualization.

47 OTHER INSTRUMENTATION↗

Module-OT: A Turnkey Solution for Securing Energy Systems

The Modular Security Apparatus for Managing Distributed Cryptography for Command-and-Control Messages on Operational Technology Networks (Module-OT) is a flexible and lightweight solution for grid-edge devices focusing on end-to-end security. It is a bump- in- the-wire solution acting as a secure conduit for data between devices or systems across a network. It improves the cybersecurity posture of DER systems by providing authentication, authorization, and data integrity to secure DER communications. Additionally, it performs key management, provides data security through whitelisting Internet Protocol addresses and ports, blocks unauthorized connections, controls user access, and allows serial or Ethernet connections for added flexibility. The core software is portable to various Linux-based operating systems and is developed to be customized by the developer and researcher communities. Module-OT has been validated in the lab, has been demonstrated at a 500-KW PV-plus-storage site, and has been proven ready to secure operational technology devices. Its core functionality meets current standards, including validation procedures of the NIST Cryptographic Algorithm Validation Program (CAVP) and the Federal Information Processing Standard (FIPS 140-2). Because of its capability to provide an accessible and affordable option for stepping up security across modern energy systems, Module-OT can serve as an effective technological option to standardize cybersecurity moving forward.

cryptography↗

ExaWind: Predictive Wind Energy Simulations

This presentation describes the ExaWind project and the team's progress in creating a suite of performance-portable codes designed for predictive simulations of wind farms on next-generation exascale-class supercomputers. Such simulations will require the resolution of scales spanning many orders of magnitude, from blade boundary layers to wind farm flow structures. In the U.S., the first exascale systems will be GPU accelerated, and different GPU manufacturers have been chosen for the different systems. At the heart of the ExaWind software is a hybrid-solver approach based on the codes Nalu-Wind and AMR-Wind, which are computational fluid dynamics solvers for the incompressible Navier-Stokes equations. Nalu-Wind is an unstructured-grid code used to resolve wind turbine geometry and blade boundary layers, whereas AMR-Wind is a structured-grid background solver for atmospheric turbulent flow and turbine wake propagation. The models are coupled with overset meshes and global linear systems are approximated through a loose-coupling algorithm. Results will include validation-quality high-fidelity simulations and strong/weak scaling results from the Summit supercomputer.

49 EE - Wind and Water Power Program - Wind (EE-4W↗

Statistical results from the Virginia Tech propagation experiment using the Olympus 12, 20, and 30 GHz satellite beacons

Virginia Tech has performed a comprehensive propagation experiment using the Olympus satellite beacons at 12.5, 19.77, and 29.66 GHz (which we refer to as 12, 20, and 30 GHz). Four receive terminals were designed and constructed, one terminal at each frequency plus a portable one with 20 and 30 GHz receivers for microscale and scintillation studies. Total power radiometers were included in each terminal in order to set the clear air reference level for each beacon and also to predict path attenuation. More details on the equipment and the experiment design are found elsewhere. Statistical results for one year of data collection were analyzed. In addition, the following studies were performed: a microdiversity experiment in which two closely spaced 20 GHz receivers were used; a comparison of total power and Dicke switched radiometer measurements, frequency scaling of scintillations, and adaptive power control algorithm development. Statistical results are reported.

Stutzman, Warren L.↗

Data from: 'Abiotic influences on continuous conifer forest structure across a subalpine watershed'

This package archives the core data used for analysis and inference in 'Abiotic influences on continuous conifer forest structure across a subalpine watershed' (Worsham et al., 2025). All data were collected in the East River, Washington Gulch, Slate River, and Coal Creek watersheds of Colorado. In the paper, we quantified the relative influence of climate, topographic, edaphic, and geologic factors on conifer stand structure and composition, and their functional relationships, at the watershed scale. We used waveform LiDAR data to derive spatially continuous stand structure metrics. We fused these with a species-level classification map to estimate tree species abundance. We applied generalized additive and generalized boosted models to evaluate the covariability of structural and compositional metrics with abiotic variables. The package contains the essential products required for reproducing our analysis and the tables and figures reported in the publication. The products comprise four classes: (1) geospatial data, (2) tabular data used for inferential analysis, (3) tabular data describing analytical results and performance statistics, and (4) a data user guide. (1) includes discretized waveform LiDAR data, locations and attributes of individual tree crowns, sampling locations and domain boundaries, a canopy height model, and raster files of estimated forest structural and compositional metrics at 100 m grid scale. (2) includes all response and explanatory variable values applied in inferential models. Response variables include conifer forest stand density, basal area, 95th percentile height, quadratic mean diameter, and others. Explanatory variables include climatic water deficit, actual evapotranspiration, elevation, heat load, soil available water content, and others. (3) includes results of training and testing several individual tree detection (ITD) algorithms, as well as inferential modeling results. (4) is a PDF user guide for this data package, including detailed descriptions and data dictionaries for all files. The data package root contains 17 assets: 8 compressed tape archive (.tar.gz) files, 5 comma-separated values (.csv) files, 3 Geographic Tagged Image File Format (GeoTIFF) (.tif) files, and 1 Portable Document Format (.pdf) file. The compressed .tar.gz archives contain ESRI shapefiles (.shp) .tif, compressed LASer (.laz), and .csv files. The archives must first be decompressed using the widely distributed command-line software utility TAR. All other files, including constituent files within the .tar.gz archives, can be opened in the open-source R statistical computing environment. Alternatively, .csv files may also be read in any simple text editor software or Microsoft Excel. Geospatial files including .shp and .tif files can also be opened in GIS software, such as QGIS (open-source) or ESRI ArcGIS (proprietary). The .pdf Data User Guide can be read with Adobe Acrobat Reader or other compatible readers.

2018 NEON and 2025 CHESS Campaigns↗

Productive Programming of Distributed Systems with the SHAD C++ Library

High-performance computing (HPC) is often perceived as a matter of making large-scale systems (e.g., clusters) run as fast as possible, regardless the required programming effort. However, the idea of "bringing HPC to the masses" has recently emerged. Inspired by this vision, we have designed SHAD, the Scalable High-performance Algorithms and Data-structures library. SHAD is open source software, written in C++, for C++ developers. Unlike other HPC libraries for distributed systems, which rely on SPMD models, SHAD adopts a shared-memory programming abstraction, to make C++ programmers feel at home. Underneath, SHAD manages tasking and data-movements, moving the computation where data resides and taking advantage of asynchrony to tolerate network latency. At the bottom of his stack, SHAD can interface with multiple runtime systems: this not only improves developer’s productivity, by hiding the complexity of such software and of the underlying hardware, but also greatly enhance code portability. Thanks to its abstraction layers, SHAD can indeed target different systems, ranging from laptops to HPC clusters, without any need for modifying the user-level code. We have prototyped and open-sourced the implementation of (a subset of) the C++ standard library (STL) targeting multi-node HPC clusters. Our work allows plain STL-based C++ code to scale on HPC systems, with no need for rewriting the code to exploit the complex hardware. SHAD is available under Apache v2 License at https://github.com/pnnl/SHAD. In this paper we overview the design of the SHAD library, depicting its main components: runtime systems abstractions for tasking; parallel and distributed data-structures; STL-compliant interfaces and algorithms.

Castellana, Vito G.↗

A scalable matrix-free spectral element approach for unsteady PDE constrained optimization using PETSc/TAO

In this work, we provide a new approach for the efficient matrix-free application of the transpose of the Jacobian for the spectral element method for the adjoint-based solution of partial differential equation (PDE) constrained optimization. This results in optimizations of nonlinear PDEs using explicit integrators where the integration of the adjoint problem is not more expensive than the forward simulation. Solving PDE constrained optimization problems entails combining expertise from multiple areas, including simulation, computation of derivatives, and optimization. The Portable, Extensible Toolkit for Scientific computation (PETSc) together with its companion package, the Toolkit for Advanced Optimization (TAO), is an integrated numerical software library that contains an algorithmic/software stack for solving linear systems, nonlinear systems, ordinary differential equations, differential algebraic equations, and large-scale optimization problems and, as such, is an ideal tool for performing PDE-constrained optimization. This paper describes an efficient approach in which the software stack provided by PETSc/TAO can be used for large-scale nonlinear time-dependent problems. Time integration can involve a range of high-order methods, both implicit and explicit. The PDE-constrained optimization algorithm used is gradient-based and seamlessly integrated with the simulation of the physical problem.

97 MATHEMATICS AND COMPUTING↗

Development of NDE/NDT Tools for High-Volume & High-Speed Inspection of CFRP Structures in Automotive Manufacturing

Main advantages of the air-coupled ultrasound testing (ACUT) and electromagnetic testing (EMT) techniques for NDE of CFRP composites were non-contact sensing, scalability for high-speed inspection, cost-effectiveness, and non-hazardous operation. Despite these advantages, no systems that would satisfy the project requirements were commercially available. Hence, one of the major efforts of the Michigan State University (MSU) team at the initial stage of the project was to close this technological gap by developing, optimizing, and validating array sensors that would provide sufficient sensitivity, spatial coverage, and resolution for robust defect detection. Optimization of the ACUT and EMT sensor designs was performed using experimentally validated finite element models. Initial experiments using array probes were conducted on relatively flat CFRP samples. In parallel, the MSU team designed and assembled a portable platform with two robotic arms. The robots were equipped with newly designed sensors that enabled high-speed NDE of curved CFRP parts. Presently, the developed robotic platform can be used as a demo/template NDE system, which is easily adaptable to manufacturing environments and in-line NDE. The ACUT NDE system developed by the MSU team used a high-power 4-channel pulser receiver for parallel data acquisition. The array probes were designed by stacking commercially available ACUT transducers, which operated in the frequency range between 100 kHz and 500 kHz. MSU optimized the excitation procedure and developed wave focusing cones so as to reduce the crosstalk between the transducers and to provide higher pulse repletion frequency (PRF). The through-transmission (TT) and single-side access (SSA) inspection modes were successfully implemented. In the TT-ACUT, structural defects in CFRP were detected by passing ultrasonic waves through the test part. Hence, the ACUT transmitters and receivers needed to be placed on the opposite sides of the test part. In the SSA-ACUT, guided waves (GW) were excited in the test part using the transmitters and were sensed by the receivers from the same side. Multi-channel TT-ACUT and SSA-ACUT provided high-speed NDE, and were successfully validated on CFRP test samples with interlaminar delaminations and other embedded defects The EM techniques developed by the MSU team included: 1) eddy current testing (ECT), 2) capacitive imaging (CI) and hybrid dual-mode imaging. In ECT, structural damage was detected in CFRP using coils sensor arrays. In ECT, the excitation magnetic field is generated by passing an alternating current through a coil, which is placed above the test sample. The excitation field penetrates the conductive sample and induces the eddy currents in its transect. In turn, the eddy currents generate the reaction field, which affects the total field sensed by a coil. Hence, the presence of structural flaws will alter the eddy current flow and the picked-up signal. ECT is mostly sensitive to local changes of the electric conductivity of the test sample, and CFRPs are mostly conductive in the direction of carbon fibers. Hence, ECT was well suited for the detection of fiber damage/fiber irregularities. The MSU team developed printed circuit boards (PCB) with coil sensor arrays optimized for NDE of CFRP. Unlike most commercial probes designed for ECT of metallic structures, the MSU array probes were designed for operation in [1-10] MHz frequency range, which was optimal for low-conductive CFRP. Multiple sensing topologies (coil groups excitation/sensing arrangements) were implemented and successfully validated. Capacitive Imaging (CI) technique developed by MSU was complementary to ECT. In contrast to ECT, which was sensitive to local changes of the electrical conductivity, the CI was sensitive to local changes of the dielectric constant. Therefore, CI could provide information about matrix damage/matrix irregularities in CFRP. The MSU CI sensor arrays were made of multiple circular or rectangular open-plate capacitors printed on PCB. Sensors of this type are not commercially available. In addition to ECT and CI, the MSU team developed a hybrid (dual-mode) inductive/capacitive measurement technique that synergistically combined the benefits of inductive and capacitive sensing for rapid NDE of fiber reinforced polymer (FRP) composite structures. Fiber damage and fiber irregularities in FRPs were detected by configuring hybrid sensors as coil sensors. Similarly, matrix damage, matrix irregularities and interlaminar delaminations were detected by configuring hybrid sensors as capacitive sensors. ECT and CI were performed sequentially by means of electronic switching. Hence, eliminating the need for mounting two separate sensor arrays on the probe. Portable robotic platform was developed by MSU for multi-technique high-speed NDE of CFRP test parts. The platform had two 6-axis robots, which enabled inspection of curved parts in approximately a 6×6×6 ft 3 active scan area. On the software side, the MSU team integrated scripts for NDE hardware control with scripts for robot motion control. MSU also implemented automated path planning for the robots, reconstruction of part’s surfaces via stereovision, 3D rendering of inspection data, and image processing algorithms for enhanced defect detection. Automotive composite parts manufactured by Plasan Composites from Phase I were used to validate the ACUT and EMT techniques on representative testbeds. Among those parts were three X-braces for a Dodge Viper, one composite calibration plaque with known defects at known locations, and four other test sections, including sections from a front splitter, a corner section from a composite hood, and a high-pressure RTM panel made using non crimp fabric. Other test samples included CFRP and GFRP calibration plates with fiber/matrix defects fabricated at MSU/CVRC.

36 MATERIALS SCIENCE↗

A New Electromagnetic Instrument for Thickness Gauging of Conductive Materials

Eddy current techniques are widely used to measure the thickness of electrically conducting materials. The approach, however, requires an extensive set of calibration standards and can be quite time consuming to set up and perform. Recently, an electromagnetic sensor was developed which eliminates the need for impedance measurements. The ability to monitor the magnitude of a voltage output independent of the phase enables the use of extremely simple instrumentation. Using this new sensor a portable hand-held instrument was developed. The device makes single point measurements of the thickness of nonferromagnetic conductive materials. The technique utilized by this instrument requires calibration with two samples of known thicknesses that are representative of the upper and lower thickness values to be measured. The accuracy of the instrument depends upon the calibration range, with a larger range giving a larger error. The measured thicknesses are typically within 2-3% of the calibration range (the difference between the thin and thick sample) of their actual values. In this paper the design, operational and performance characteristics of the instrument along with a detailed description of the thickness gauging algorithm used in the device are presented.

Fulton, J. P.↗

Extending SEER for Extreme Heterogeneity

Heterogeneous and multi-device nodes are increasingly common in high-performance computing and data centers, yet existing programming models often lack simple, transparent, and portable support for these diverse architectures. The main contribution of this work is the development of novel SEER capabilities to address this challenge by providing a descriptive programming model that allows applications to seamlessly leverage heterogeneous nodes across various device types. SEER uses efficient memory management and can select the proper device[s] depending on the computational cost of the applications. This is completely transparent to the programmer, thereby providing a highly productive programming environment. Integrating extreme heterogeneity into the SEER library as shown with the use of NVIDIA and AMD GPUs simultaneously allows it to expand and exploit the performance possibilities. Our analysis based on the well-known Conjugate Gradient algorithm reports accelerations above 1.5 × on computationally demanding steps of such an algorithm by using both architectures simultaneously.

Teranishi, Keita [ORNL] (ORCID:0000000166472690)↗

Load Balancing Sequences of Unstructured Adaptive Grids

Mesh adaption is a powerful tool for efficient unstructured grid computations but causes load imbalance on multiprocessor systems. To address this problem, we have developed PLUM, an automatic portable framework for performing adaptive large-scale numerical computations in a message-passing environment. This paper makes several important additions to our previous work. First, a new remapping cost model is presented and empirically validated on an SP2. Next, our load balancing strategy is applied to sequences of dynamically adapted unstructured grids. Results indicate that our framework is effective on many processors for both steady and unsteady problems with several levels of adaption. Additionally, we demonstrate that a coarse starting mesh produces high quality load balancing, at a fraction of the cost required for a fine initial mesh. Finally, we show that the data remapping overhead can be significantly reduced by applying our heuristic processor reassignment algorithm.

Biswas, Rupak↗