Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelism”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Intrinsic Toroidal Rotation Driven by Turbulent and Neoclassical Processes in Tokamak Plasmas from Global Gyrokinetic Simulations

Gyrokinetic tokamak plasmas can exhibit intrinsic toroidal rotation driven by the residual stress. While most studies have attributed the residual stress to the parallel-momentum flux from the turbulent E × B motion, the parallel-momentum flux from the drift-orbit motion (denoted $Π^D_\parallel$) and the E × B-momentum flux from the E × B motion (denoted $Π_{E×B}$) are often neglected. Here, we use the global total-f gyrokinetic code XGC to study the residual stress in the core and the edge of a DIII-D H-mode plasma. Numerical results show that both $Π^D_\parallel$ and $Π_{E×B}$ make up a significant portion of the residual stress. In particular, $Π^D_\parallel$ in the core is higher than the collisional neoclassical level in the presence of turbulence, while in the edge it represents an outflux of countercurrent momentum even without turbulence. Using a recently developed “orbit-flux” formulation, we show that the higher-than-neoclassical-level $Π^D_\parallel$ in the core is driven by turbulence, while the outflux of countercurrent momentum from the edge is mainly due to collisional ion orbit loss. In conclusion, these results suggest that $Π^D_\parallel$ and $Π_{E×B}$ can be important for the study of intrinsic toroidal rotation.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Performance Evaluation of Multi-Vendor Grid-Forming Inverters for Grid-Connected Operation Through Hardware Experimentation: Preprint

Existing real-world projects of GFM inverters that operate in parallel to power grids typically are sized between dozens and a few hundreds Megawatt (MW) scale according to a recent NERC GFM inverter white paper. These large systems are often difficult to evaluate prior to deployments because of their large size. The performance of smaller GFM inverters (dozens to a few hundreds MW) that operate parallel with power grids (distribution systems) is even less understood. There is an opportunity to better understand these systems through hardware testing under controlled laboratory conditions. Therefore, this paper presents the functional performance evaluation tests of multiple (three) commercial GFM inverters when they operate in parallel with the grid through hardware experiments. The goal of these tests is to explore the GFM inverters' functionalities and dynamic response when in parallel with power grids to eventually develop universal specifications for GFM inverters. Both steady state (changing the inverter's frequency and voltage droop) and transient (adding step change in grid's frequency/voltage) tests are performed for each GFM inverter with the same testing circuit and testing protocol. The experimental results indicate the bench-marked performance that: 1) the GFM inverters can be dispatched through frequency and voltage droop intercepts to output the target power when paralleled to the grid; 2) the GFM inverters automatically respond to system frequency and voltage events to output the needed power, however, the GFM inverters all show stability issues when absorbing reactive power from the grid.

grid-forming inverters↗

Architecture-Dependent Thin Film Self-Assembly of Star Polystyrene-poly(2-vinylpyridine) Block Copolymers

Controlling the orientation of nanostructured block copolymer (BCP) thin films is essential for their use in templating, transport, and pattern transfer. Conventional efforts mainly focus on adjusting enthalpic interactions between the blocks and interfaces, while entropic contributions are often overlooked. Here, we show that the morphology of BCP thin films can be precisely tuned by the architectural design of star BCPs. Specifically, we synthesized multiarm star BCPs with a polystyrene (PS) core and poly(2-vinylpyridine) (P2VP) corona, which exhibits a lamellar microdomain morphology. The entropic penalty associated with a parallel orientation of the microdomains to the substrate is controlled by varying the number of arms comprising the star BCPs, from 2-arms (triblock) to 3-arms and 4-arms. We systematically investigated the thin film morphology at different depths using grazing incidence small-angle X-ray and neutron scattering (GISAXS and GISANS), atomic force microscopy (AFM), water contact angle (WCA), and interference microscopy. The results show that 2-arm star BCPs show a parallel orientation, the 3-arm star BCPs form a uniform PS film at the air surface with a vertical orientation of the microdomains underneath, and the 4-arm star BCPs exhibit a parallel microdomain orientation at the air surface with mixed parallel and perpendicular microdomain orientation in the bulk. Additionally, we found that the inclination angle of microdomains at the edges of islands and holes, resulting from the incommensurability between film thickness and the characteristic period of the microdomain morphology, increases with a higher number of arms. This suggests a greater grain boundary tilt angle in the microdomains of star-shaped block copolymers (BCPs). When silicon substrates were modified with PS homopolymer, the selective interaction between substrate and core blocks promotes a parallel orientation for the 4-arm star BCPs. In conclusion, this work shows that control of arm number in star BCPs affords diverse BCP thin film morphologies, offering insights into the star BCP conformations in thin films across different depths.

36 MATERIALS SCIENCE↗

Importance of $\delta B_{\|}$ on ETG stability, turbulence, and transport in NSTX

This study employs electron-scale gyrokinetic simulations to investigate the electron temperature gradient (ETG) driven instabilities, turbulence, and transport in the pedestal region of the National Spherical Torus Experiment, comparing non-lithiated (narrow pedestal) and lithiated (wide pedestal) scenarios. Our findings reveal that, in the non-lithiated case, a branch of strongly unstable ETG modes exhibiting finite parallel magnetic field fluctuations ($\delta B_{\parallel} \neq 0$) emerges at the pedestal top and upper density pedestal region. This branch is uncovered only when $\delta B_{\parallel}$ is retained in the simulations and is associated with substantial electrostatic electron heat flux. This region of strong ETG transport corresponds to the only region in the plasma where the pressure gradient is far below the critical gradient for kinetic ballooning modes. We investigated the origin of this finite $\delta B_{\parallel}$ ETG branch by analyzing the gyrokinetic field equations. Nonlinear saturation is also analyzed and contrasted for simulations with and without $\delta B_{\parallel}$. In contrast with the nonlithiated case, ETG modes in the lithiated case produce substantial transport in the steep gradient region, but are negligible at the pedestal top.

ETG↗

Taking control of compressible modes: bulk viscosity and the turbulent dynamo

Many polyatomic astrophysical plasmas are compressible and out of chemical and thermal equilibrium, introducing a bulk viscosity into the plasma via the internal degrees of freedom of the molecular composition, directly impacting the decay of compressible modes, $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$. This is especially important for small-scale, turbulent dynamo processes in the interstellar medium (ISM), which are known to be sensitive to the effects of compression. To control the viscous properties of $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$, we perform trans-sonic, visco-resistive dynamo simulations with additional bulk viscosity $\nu _{\text{bulk}}$, deriving a new $\nu _{\text{bulk}}$ Reynolds number $\text{Re}_{\text{bulk}}$, and viscous Prandtl number $\text{P}\nu \equiv \text{Re}_{\text{bulk}}/ \text{Re}_{\text{shear}}$, where $\text{Re}_{\text{shear}}$ is the shear viscosity Reynolds number. We derive a framework for decomposing $E_{\rm mag}$ growth rates into incompressible and compressible terms via orthogonal tensor decompositions of $\boldsymbol {\nabla }\otimes \mathrm{{\boldsymbol {\mathit {v}}}}$, where $\mathrm{{\boldsymbol {\mathit {v}}}}$ is the fluid velocity. We find that $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$ play a dual role, growing and decaying $E_{\rm mag}$, and that field-line stretching is the main driver of growth, even in compressible dynamos. In the absence of $\nu _{\text{bulk}}$ ($\text{P}\nu \rightarrow \infty$), $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$ pile up on small-scales, creating a spectral bottleneck, which disappears for $\text{P}\nu \approx 1$. As $\text{P}\nu$ decreases, $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$ are dissipated at increasingly larger scales, in turn suppressing incompressible modes through a coupling between high-k modes. We emphasize the importance of further understanding the role of $\nu _{\text{bulk}}$ in compressible astrophysical plasmas, which we estimate could be as strong as the shear viscosity in the cold ISM, and highlight that compressible direct numerical simulations without bulk viscosity have unresolved compressible mode dissipation scales.

MHD↗

Shapiro steps and stability of skyrmions interacting with alternating anisotropy under the influence of ac and dc drives

Here we use atomistic simulations to examine the sliding dynamics of a skyrmion in a two-dimensional system containing a periodic one-dimensional stripe pattern of variations between low and high values of the perpendicular magnetic anisotropy. The skyrmion changes in size as it crosses the interface between two anisotropy regions. On applying combined dc and ac driving in either parallel or perpendicular directions, we observe a wide variety of Shapiro steps, Shapiro spikes, and phase-locking phenomena. The phase-locked orbits have two-dimensional dynamics due to the gyrotropic or Magnus dynamics of the skyrmions and are distinct from the phase-locked orbits found for strictly overdamped systems. Along a given Shapiro step when the ac drive is perpendicular to the dc drive, the velocity parallel to the ac drive is locked while the velocity in the perpendicular direction increases with increasing drive to form Shapiro spikes. At the transition between adjacent Shapiro steps, the parallel velocity jumps up to the next step value, and the perpendicular velocity drops. The skyrmion Hall angle shows a series of spikes as a function of increasing dc drive, where the jumps correspond to the transition between different phase-locked steps. At high drives, the Shapiro steps and Shapiro spikes are lost. When both the ac and dc drives are parallel to the stripe periodicity direction, Shapiro steps appear, while if the dc drive is parallel to the stripe periodicity direction and the ac drive is perpendicular to the stripe periodicity, then there are only two locked phases, and the skyrmion motion consists of a combination of sliding along the interfaces between the two anisotropy values and jumping across the interfaces.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

HydraGNN v4.0

The new version of HydraGNN v4.0 provides additional core capabilities, such as: Inclusion of multi-body atomistic cluster expansion MACE, polarizable atom interaction neural network PAINN, and equivariant principal neighborhood aggregation (PNAEq) among the message passing layers supported -Inclusion of graph transformers to directly model long-range interactions between nodes that are distant in the graph topology Integration of graph transformers with message passing layers by combining the graph embedding generated by the two mechanisms, which allows for an improved expressivity of the HydraGNN architecture Improved re-implementation of multi-task learning (MTL) to allow its use for stabilized training across imbalanced, multi-source, multi-fidelity data Introduction of multi-task parallelism, a newly proposed type of model parallelism specifically for MTL architectures, which allows to dispatch different output decoding heads to different GPU devices Integration of multi-task parallelism with pre-existing distributed data parallelism to enable a 2D parallelization for distributed training Improved portability of the distributed training across Intel GPUs, which has been testes on ALCF exascale supercomputer Aurora Inclusion of 2-level fine-grained energy profilers portable across NVIDIA, AMD, and Intel GPUs to monitor the power and energy consumption associated with different functions executed by the HydraGNN code during data pre-load and training Restructuring of previous examples and inclusion of new sets of examples to illustrate the download, preprocess, and training of HydraGNN models on new large-scale open-source datasets for atomistic materials modeling (e.g., Alexandria, Transition1x, OMat24, OMol25)

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

Multithreaded copy ('cp')

This is a modification to 'cp' and 'mv' commands to make them multi-threaded. Simple benchmarks showed that multi-threading could reduce the time to copy a large Linux source directory by over 2x. The 'cp' and 'mv' utilities are part of the existing Coreutils (https://www.gnu.org/software/coreutils/) software package that get installed on all Linux distros. Changes: * Add '-j|--parallel ' flags to 'cp' and 'mv'. This allows the utilities to recursively copy regular files in directories in parallel. This does NOT parallelize multiple single file copies to a destination (like 'cp file2 file2 file3 dst/'). Along with this, add in new 'CP_NUM_THREADS' and 'MV_NUM_THREADS' environment variables to set the number of threads. This can be useful when you want to enable parallelism by default in /etc/profile. The maximum number of threads is internally capped to the number of CPUs. * Add a '-j' flag to 'sort' to complement its existing '--parallel' flag. This is only done for consistency with 'cp' and 'mv'. * Add test cases for the new flags. Also, run each 'cp' and 'mv' test both in single-threaded and multithreaded modes for extra coverage.

Hutter, AnthonyJ [Lawrence Livermore National Labo↗

Experience with the alpaka performance portability library in the CMS software

ion Library for Parallel Kernel Acceleration) is a header-only C++ library that provides performance portability across different back-ends, abstracting the underlying levels of parallelism. It supports serial and parallel execution on CPUs, and extremely parallel execution on NVIDIA, AMD and Intel GPUs.This contribution will show how alpaka is used in the CMS software to develop and maintain a single code base; to use different toolchains to build the code for each supported back-end, and link them into a single application; to seamlessly select the best backend at runtime, and implement portable reconstruction algorithms that run efficiently on CPUs and GPUs from different vendors. It will describe the validation and deployment of the alpaka-based implementation in the CMS High Level Trigger, and highlight how it achieves near-native performance.

Alawieh, Jaafar [CERN]↗

Radiative divertor detachment and impurity transport with nitrogen and neon seeding in KSTAR H-mode plasmas

Achieving core-edge compatible divertor detachment is a critical requirement for stable operation in future fusion devices. This study compares nitrogen and neon seeding in KSTAR H-mode plasmas with carbon walls, combining experiments and SOLPS-ITER modelling to evaluate their radiative dissipation and core-edge compatibility. Experimentally, N seeding achieved stronger divertor detachment, with larger reductions in target particle and heat fluxes, a higher divertor radiation fraction, and a lower core radiation fraction than Ne. In contrast, Ne seeding triggered a significant rise in core radiation followed by H–L back transitions, limiting the maximum total radiated power fraction to roughly half that of N. SOLPS-ITER simulations reproduced the experimental trends and revealed that the better core-edge compatibility of N arises from its higher divertor retention in addition to its higher cooling factor. The relative positions of the stagnation points of impurity poloidal velocity and ionization sources did not explain the different divertor compression. Instead, in the present modelling, the higher impurity parallel particle flux, resulting from the higher parallel impurity velocity, explains the stronger nitrogen impurity compression in the divertor region. The parallel temperature distribution with N was more favourable for achieving higher impurity parallel velocity than with Ne, because the impurity velocity is governed by modifications of the main ion flow due to friction and thermal forces, both of which strongly depend on the temperature. Ultimately, this behaviour is attributed to the strongly divertor-localized radiation of N. These results demonstrate that N is more effective than Ne in achieving radiative divertor detachment while maintaining low core contamination in KSTAR, consistent with observations in other present tokamaks.

KSTAR↗

Entanglement-enhanced ac magnetometry in the presence of Markovian noise

Entanglement is a resource to improve the sensitivity of quantum sensors. In an ideal case, using an entangled state as a probe to detect target fields, we can beat the standard quantum limit by which all classical sensors are bounded. However, since entanglement is fragile against decoherence, it is unclear whether entanglement-enhanced metrology is useful in a noisy environment. Its benefit is indeed limited when estimating the amplitude of dc magnetic fields under the effect of parallel Markovian decoherence, where the noise operator is parallel to the target field. In this paper, on the contrary, we show an advantage to using an entanglement over the classical strategy under the effect of parallel Markovian decoherence when we try to detect ac magnetic fields. We consider a scenario to induce a Rabi oscillation of the qubits with the target ac magnetic fields. Although we can, in principle, estimate the amplitude of the ac magnetic fields from the Rabi oscillation, the signal becomes weak if the qubit frequency is significantly detuned from the frequency of the ac magnetic field. We show that, by using the Greenberger-Horne-Zeilinger (GHZ) states, we can significantly enhance the signal of the detuned Rabi oscillation even under the effect of parallel Markovian decoherence. Further, our method is based on the fact that the interaction time between the GHZ states and ac magnetic fields scales as 1/L to mitigate the decoherence effect, where L is the number of qubits, which contributes to improving the bandwidth of the detectable frequencies of the ac magnetic fields. Our results pave the way for new applications of entanglement-enhanced ac magnetometry.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A Flexible Forwarding Scheme to Improve Latency-Bound Irregular P2P Communication in MPI

We propose an algorithm to efficiently perform latency-bound communication scenarios that consist of many small messages. In these parallel scenarios, processes typically pass around a lot of small-sized messages of a few KBs of size. Performing communication operations with P2P MPI routines or collective MPI routines (including neighborhood collectives) in such scenarios may not always yield the optimal results and may not resolve the latency bottleneck. To this end, we develop a regular structure called virtual process topology (VPT) on which the messages can be communicated in a structured and controlled manner. Using parameters of this topology, one can tune the rate of aggression in tackling the latency costs. We demonstrate that our communication algorithm is preferable to MPI P2P and collective routines for latency-bound communication and it can easily be adapted only by replacing calls to MPI routines in a parallel application. We show how to adapt existing topology-aware mapping heuristics to address the volume overhead due to communicating messages on the VPT. Moreover, we propose a novel swap-based mapping heuristic to address this overhead by optimizing the maximum volume handled by a process. Experiments on synthetic communication graphs as well as real-world applications such as parallel Canonical Polyadic sparse tensor decomposition and parallel sparse matrix-dense matrix multiplication show that our approach is a powerful way of overcoming the bottlenecks posed by sparse and latency-bound irregular communication.

communication algorithm↗

Integrating ytopt and libEnsemble to autotune OpenMC

Ytopt is a Python machine-learning-based autotuning software package developed within the ECP PROTEAS-TUNE project. The ytopt software adopts an asynchronous search framework that consists of sampling a small number of input parameter configurations and progressively fitting a surrogate model over the input-output space until exhausting the user-defined maximum number of evaluations or the wall-clock time. libEnsemble is a Python toolkit for coordinating workflows of asynchronous and dynamic ensembles of calculations across massively parallel resources developed within the ECP PETSc/TAO project. libEnsemble helps users take advantage of massively parallel resources to solve design, decision, and inference problems and expands the class of problems that can benefit from increased parallelism. In this paper we present our methodology and framework to integrate ytopt and libEnsemble to take advantage of massively parallel resources to accelerate the autotuning process. Specifically, we focus on using the proposed framework to autotune the ECP ExaSMR application OpenMC, an open source Monte Carlo particle transport code. OpenMC has seven tunable parameters some of which have large ranges such as the number of particles in-flight, which is in the range of 100,000 to 8 million, with its default setting of 1 million. Setting the proper combination of these parameter values to achieve the best performance is extremely time-consuming. Therefore, we apply the proposed framework to autotune the MPI/OpenMP offload version of OpenMC based on a user-defined metric such as the figure of merit (FoM) (particles/s) or energy efficiency energy-delay product (EDP) on Crusher at Oak Ridge Leadership Computing Facility. In conclusion, the experimental results show that we achieve the improvement up to 29.49% in FoM and up to 30.44% in EDP.

Autotuning↗

Fermilab's controls development with virtual accelerator

Control Systems development is often the last thing considered when designing and building new equipment, e.g. a new detector or superconducting RF LINAC; however when the new equipment is installed, it is the first thing desired to be operational for testing. Due to frequent delays in building new equipment and project deadlines, control system development and testing is often curtailed. A way to alleviate this problem is to simulate the control system, though this will be challenging for complex systems.The Fermilab PIP-II (proton improvement plan - II) project is being constructed at Fermilab to deliver $800\,MeV$ protons of $>1\,MW$ beam power to replace the present LINAC for the remainder of the existing accelerator complex. The new LINAC consists of a warm front end (WFE), 23 superconducting RF cryomodules (of 5 types), and a beam transfer line (BTL) to the existing complex.The accelerator physics group has a parallel project to create a digital twin (DT) of the PIP-II accelerator. We have coupled the EPICS controls to this DT and are developing both the DT and EPICS software in parallel. This will allow us to develop the EPICS software framework, the HMIs, sequences, high level physics applications, and other services for use in a fully functional control system.This presentation will detail the work that we have performed to date and show demonstrations of controlling and monitoring the status of the accelerator, as well as future plans for this work.

Hanlet, Pierrick [Fermilab]↗

Fiats: Functional inference and training for surrogates

Fiats provides a platform for research on the training and deployment of neural-network surrogate models for computational science. Fiats also supports exploring, advancing, and combining functional, object-oriented, and parallel programming patterns in Fortran 2023. As such, the Fiats name has dual expansions: “Functional Inference And Training for Surrogates” or “Fortran Inference And Training for Science.” Fiats inference and training procedures are pure and therefore satisfy a language constraint imposed on procedure invocations inside Fortran’s parallel loop construct: do concurrent. Furthermore, the Fiats training procedures are built around a do concurrent parallel reduction. Several compilers can automatically parallelize do concurrent on Central Processing Units (CPUs) or Graphics Processing Units (GPUs). Fiats thus aims to achieve performance portability through standard language mechanisms.

Rouson, Damian [Lawrence Berkeley National Laborat↗

MOOSE ProbML: Parallelizable Probabilistic Machine Learning and Uncertainty Quantification Capabilities

The Multiphysics Object Oriented Simulation Environment (MOOSE) is a widely used open- source finite element software for performing multiphysics multiscale simulations in a massively parallel fashion. Recently, the computational team at Idaho National Laboratory (INL) has implemented Probabilistic Machine Learning (ProbML) capabilities in MOOSE—in a parallelized fashion—and enable active learning with large-scale computational models for tasks such as surrogate model development, scale bridging, forward/inverse uncertainty quantification (UQ), Bayesian optimization, etc. This presentation summarizes these developments in MOOSE along with demonstrations on several real applications relevant to nuclear energy. At the fundamental level, samplers like Monte Carlo/Latin Hypercube, variance reduction, parallelized Markov Chain Monte Carlo (MCMC) support uncertainty propagation in both forward and inverse settings. These samplers can be integrated with the Gaussian processes (GP) suite in MOOSE, which offer several variants like scalar GPs, multi-output GPs, and deep GPs, to enable active learning. These GPs can be tuned using gradient-based optimization methods like Adam and its variants or gradient-free methods like the elliptical slice sampler (a variant of MCMC adept under Gaussian settings) for more complex covariance kernels or likelihoods whose gradient computations can be cumbersome. A variety of batch acquisition functions permit parallelized evaluation of the computational model and support different learning objectives with high efficiency like Bayesian inference, global surrogate development, optimization, etc. Furthermore, libtorch integration supports training, evaluation, and re-training of neural networks and other complex machine learning models in active learning settings. The impacts of these developments are shown on several real applications: (1) nuclear fuel inverse UQ and model inadequacy assessment using the Kennedy O’Hagan framework; (2) uncertainty aware surrogate modeling for additive manufacturing to predict field quantities; (3) nuclear reactor rare events analysis; and (4) complex fluid flow prediction using a global surrogate with quantified prediction uncertainty. Finally, the outlook of MOOSE ProbML is discussed for both outer-loop and inner-loop computations in the broad view to accelerate fuels and materials qualification, address gaps in knowledge and data, and assess new reactor/fuel systems.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Leveraging Pre-Built Catalogs and Object-Level Scheduling to Eliminate I/O Bottlenecks in HPC Environments

Modern High-Performance Computing (HPC) environments face mounting challenges due to the shift from large to small file datasets, along with an increasing number of users and parallelized applications. As HPC systems rely on Parallel File Systems (PFS), such as Lustre for data processing, performance bottlenecks stemming from Object Storage Target (OST) contention have become a significant concern. Existing solutions, such as LADS with its object-level scheduling approach, fall short in large-scale HPC environments due to their inability to effectively address metadata I/O bottlenecks and the growing number of I/O processes. This study highlights the pressing need for a comprehensive solution that tackles both OST contention and metadata I/O challenges in diverse HPC workloads. To address these challenges, we propose SwiftLoad, an object-level I/O scheduling framework that leverages a metadata catalog to enhance the performance and efficiency of parallel HPC utilities. The adoption of the metadata catalog mitigates the metadata I/O bottlenecks that commonly occur in HPC utilities, a challenge that is particularly pronounced in object-level I/O scheduling. SwiftLoad addresses OST contention and the uneven distribution of I/O processes across different OSTs through mathematical modeling and incorporates a Loader Configuration Module to regulate the number of I/O processes. Evaluated with two representative utilities—data deduplication profiling and data augmentation—SwiftLoad achieved performance improvements of up to 5.63x and 11.0x, respectively, on a production supercomputer.

HPC↗

Toucan: A performance portable, scalable implementation of the DECA algorithm

In the field of additive manufacturing (AM), cellular automata (CA) is extensively used to simulate microstructural evolution during solidification. However, while traditional CA approaches are relatively fast, they still require a substantial number of time steps, are limited to moderate volumes, and are relatively difficult to improve through parallelism due to the highly localized nature of the solidification front. Here, to address these issues of time to solution and load balancing, we introduce Toucan, a parallel, performance-portable, and scalable code written in C++ with the Kokkos library that leverages the discrete event inspired cellular automata (DECA) algorithm to perform parallel-in-time (PinT) grain growth simulations. Toucan effectively mitigates load balancing issues by distributing the computational workload more evenly across processors, enhancing scalability and efficiency. We conduct both strong and weak scaling studies on up to 64 GPUs on the Frontier supercomputer, demonstrating that Toucan significantly outperforms the current state-of-the-art, time-stepped CA code, ExaCA, on both single and multi-GPU simulations. Even in AM-specific weak scaling scenarios, Toucan maintains near-ideal scaling, in contrast to the linear increase observed with ExaCA due to the moving laser raster pattern. This study highlights Toucan’s potential to transform microstructural simulations in AM by radically improving both efficiency and scalability over existing methods.

36 MATERIALS SCIENCE↗