Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed and parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Sequence length scaling in vision transformers for scientific images on frontier

Vision Transformers (ViTs) are pivotal for foundational models in scientific imagery, including Earth science applications, due to their capability to process large sequence lengths. While transformers for text have inspired scaling sequence lengths in ViTs, adapting these for ViTs introduces unique challenges. We develop distributed sequence parallelism for ViTs, enabling them to handle up to 1M tokens. Our approach, leveraging DeepSpeed-Ulysses and Long-Sequence-Segmentation with model sharding, is the first to apply sequence parallelism in ViT training, achieving a 94% batch scaling efficiency on 2,048 AMD-MI250X GPUs. Evaluating sequence parallelism in ViTs, particularly in models up to 10B parameters, highlighted substantial bottlenecks. We countered these with hybrid sequence, pipeline, and flash attention strategies, to scale beyond single GPU memory limits. Our method significantly enhances climate modeling accuracy by 20% in temperature predictions, marking the first training of a vision transformer model to convergence with a sequence length of 188K tokens, using full self-attention.

Tsaris, Aristeidis (aris) [ORNL] (ORCID:0000000277↗

Current Practices in Distribution Utility Resilience Planning for Winter Storms

This report is part of a series of hazard-focused case studies examining common practices in electric utility resilience planning. We use standard terminology defining resilience as the ability to anticipate, withstand, absorb, and recover from hazards that cause long duration outages. We distinguish between reliability and resilience using Institute of Electrical and Electronics Engineers (IEEE) 1366-2022, which defines major events as an event that exceeds reasonable design and/or operational limits of the electric power system. Resilience planning is focused on major event days and reliability planning is focused on nonmajor event days. Utility resilience plans are assessed according to common resilience components identified in existing resilience frameworks. The focus of this report is on winter storms in which the primary hazards are heavy snowfall, freezing rain, ice, extreme cold, severe wind, and flooding. These hazards can also contribute to generation shortages, resulting in bulk power system impacts that have consequences for the distribution system, such as load shedding. Stand-alone reports focusing on wildfires and nonwinter storms have been published in parallel with this report. This report can be used as a starting point for understanding potential investment prioritization processes and investment options. This report is intended to improve utility resilience planning by supporting constructive dialogue among utilities, regulators, and other stakeholders.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Flexible User-Defined Domain Decomposition in Kilometer-Scale E3SM Land Model Simulation

The Energy Exascale Earth System Model (E3SM) Land Model (ELM) has been extended to kilometer-scale (km-ELM) resolutions, enabling high-fidelity simulations of terrestrial processes at 1 km x 1 km grid spacing. In ELM, domain decomposition partitions the computational domain across processors, ensuring efficient parallel execution. Currently, round-robin decomposition is applied, providing a straightforward way to distribute computational workload. As ELM continues evolving at the kilometer-scale (km-scale), particularly with integrating lateral flow modeling, decomposition strategies must also account for the increased workload and data movement. This paper introduces a flexible user-defined domain decomposition framework, allowing users to customize domain partitioning based on application requirements. The impact of different decomposition strategies is evaluated across various applications concerning computation, communication, and I/O. Results demonstrate that while 1D partitioning yields superior I/O performance, k-nearest neighbors (KNN) clustering effectively reduces inter-process communication overhead. This study lays the groundwork for scalable partitioning in large-scale land surface simulations, enhancing next-generation Earth system modeling.

Wang, Dali [ORNL] (ORCID:0000000168065108)↗

Stable parallel training of Wasserstein conditional generative adversarial neural networks

In this work, we propose a stable, parallel approach to train Wasserstein conditional generative adversarial neural networks (W-CGANs) under the constraint of a fixed computational budget. Differently from previous distributed GANs training techniques, our approach avoids inter-process communications, reduces the risk of mode collapse and enhances scalability by using multiple generators, each one of them concurrently trained on a single data label. The use of the Wasserstein metric also reduces the risk of cycling by stabilizing the training of each generator. We illustrate the approach on the CIFAR10, CIFAR100, and ImageNet1k datasets, three standard benchmark image datasets, maintaining the original resolution of the images for each dataset. Performance is assessed in terms of scalability and final accuracy within a limited fixed computational time and computational resources. To measure accuracy, we use the inception score, the Fréchet inception distance, and image quality. An improvement in inception score and Fréchet inception distance is shown in comparison to previous results obtained by performing the parallel approach on deep convolutional conditional generative adversarial neural networks as well as an improvement of image quality of the new images created by the GANs approach. Weak scaling is attained on both datasets using up to 2000 NVIDIA V100 GPUs on the OLCF supercomputer Summit.

97 MATHEMATICS AND COMPUTING↗

Density and Magnetic Field Asymmetric Kelvin‐Helmholtz Instability

Abstract The Kelvin‐Helmholtz (KH) instability can transport mass, momentum, magnetic flux, and energy between the magnetosheath and magnetosphere, which plays an important role in the solar‐wind‐magnetosphere coupling process for different planets. Meanwhile, strong density and magnetic field asymmetry are often present between the magnetosheath (MSH) and magnetosphere (MSP), which could affect the transport processes driven by the KH instability. Our magnetohydrodynamics simulation shows that the KH growth rate is insensitive to the density ratio between the MSP and the MSH in the compressible regime, which is different than the prediction from linear incompressible theory. When the interplanetary magnetic field (IMF) is parallel to the planet's magnetic field, the nonlinear KH instability can drive a double mid‐latitude reconnection (DMLR) process. The total double reconnected flux depends on the KH wavelength and the strength of the lower magnetic field. When the IMF is anti‐parallel to the planet's magnetic field, the nonlinear interaction between magnetic reconnection and the KH instability leads to fast reconnection (i.e., close to Petschek reconnection even without including kinetic physics). However, the peak value of the reconnection rate still follows the asymmetric reconnection scaling laws. We also demonstrate that the DMLR process driven by the KH instability mixes the plasma from different regions and consequently generates different types of velocity distribution functions. We show that the counter‐streaming beams can be simply generated via the change of the flux tube connection and do not require parallel electric fields.

Astronomy & Astrophysics↗

Laser-induced slip casting as an additive manufacturing approach for silicon carbide

Here, this work presents processing silicon carbide (SiC) with the laser-induced slip casting (LIS) additive manufacturing (AM). SiC was stabilized in water with polyethyleneimine (PEI) dispersant, and SiC slurries were made with rheology for LIS printing. High-density ceramic parts were printed, followed by single-step binder burnout and sintering. The printed parts achieved 93–95 % of theoretical density. X-ray computed tomography (XCT) revealed a small distribution of flaws exceeding 100 microns. The mechanical properties were measured in both parallel and perpendicular to the printing layers, and the orientation with layers perpendicular to the bending moment resulted in higher strength compared to the parallel direction. Porosity resulting from processing and large inclusions of boron carbide (B4C) were the root cause of failure in the measured samples. Despite these defects through this effort, this new approach demonstrates promise for green forming of SiC with densities greater than 95 % theoretical and tensile strengths above 250 MPa.

Additive Manufacturing↗

On the Efficient Evaluation of the Exchange Correlation Potential on Graphics Processing Unit Clusters

The predominance of Kohn–Sham density functional theory (KS-DFT) for the theoretical treatment of large experimentally relevant systems in molecular chemistry and materials science relies primarily on the existence of efficient software implementations which are capable of leveraging the latest advances in modern high-performance computing (HPC). With recent trends in HPC leading toward increasing reliance on heterogeneous accelerator-based architectures such as graphics processing units (GPU), existing code bases must embrace these architectural advances to maintain the high levels of performance that have come to be expected for these methods. In this work, we purpose a three-level parallelism scheme for the distributed numerical integration of the exchange-correlation (XC) potential in the Gaussian basis set discretization of the Kohn–Sham equations on large computing clusters consisting of multiple GPUs per compute node. In addition, we purpose and demonstrate the efficacy of the use of batched kernels, including batched level-3 BLAS operations, in achieving high levels of performance on the GPU. We demonstrate the performance and scalability of the implementation of the purposed method in the NWChemEx software package by comparing to the existing scalable CPU XC integration in NWChem.

97 MATHEMATICS AND COMPUTING↗

Distributed Macroscopic Traffic Simulation with Open Traffic Models

This paper presents OTM-MPI, an extension of the Open Traffic Models platform (OTM) for running macroscopic traffic simulations in high-performance computing environments. OTM-MPI represents the first open-source, distributed-memory, macroscopic simulation model developed for modern high performance parallel machines and large networks. Macroscopic simulations are appropriate for studying regional traffic scenarios when aggregate trends are of interest, rather than individual vehicle traces. They are also appropriate for studying the routing behavior of classes of vehicles, such as app-informed vehicles. The network partitioning was performed with METIS. Inter-process communication was done with MPI (message-passing interface). Results are provided for two networks: one realistic network which was obtained from Open Street Maps for Chattanooga, TN, and another larger synthetic grid network. The software recorded a speedups of 198x using 256 cores for Chattanooga, and 475x with 1,024 cores for the synthetic network.

macro-scopic traffic simulation↗

Record acceleration of the two-dimensional Ising model using a high-performance wafer-scale engine

The versatility and wide-ranging applicability of the Ising model, originally introduced to study phase transitions in magnetic materials, have made it a cornerstone in statistical physics and a valuable tool for evaluating the performance of emerging computer hardware. Here, we present a novel implementation of the two-dimensional Ising model on Cerebras Wafer-Scale Engine (WSE) – a revolutionary processor that is opening new frontiers in computing. In our deployment of the checkerboard algorithm, we optimized the Ising model to take advantage of the unique WSE architecture. Specifically, we employed a compressed bit representation storing 16 spins on each int16 word, and efficiently distributed the spins over the processing units enabling seamless weak scaling and limiting communications to only immediate neighboring units. Our implementation can handle up to 754 simulations in parallel, achieving an aggregate of over 61.8 trillion flip attempts per second for Ising models with up to 200 million spins. This represents a gain of up to 148 times over previously reported single-devices with a highly optimized implementation on NVIDIA V100 and up to 88 times in productivity compared to NVIDIA H100. Our findings highlight the significant potential of the WSE in scientific computing, particularly in the field of materials modeling.

Ising model↗

Heating and acceleration of ions with Kappa distribution functions by low‐frequency Alfvén wave

Abstract Heating and acceleration of ions with Kappa distribution functions (with parameter ) in low‐beta plasmas, by a low‐frequency Alfvén wave, is investigated using test‐particle simulations, yielding interesting new results. As long as the Alfvén wave amplitude is sufficiently large, the computed net heating energy of ions becomes independent of the wave frequency and amplitude, always approaching the same value of . The eventual energy of ions is dictated only by the initial ion energy and the ratio of the magnetic field energy density to the plasma density. The heating effect of the Kappa ions increases with . During the heating process, the ions are picked up by the Alfvén wave and pitch angle scattered, forming a quasi‐isotropic spherical shell velocity distribution. The Kappa ions are accelerated in the parallel direction, reaching a bulk flow speed roughly equal to the local Alfvén speed. Higher ‐value in the initial Kappa distribution leads to faster saturation. The above results may explain certain features of the ion heating and acceleration in the solar wind and corona.

Li, Kehua↗

A Robust Parallel Distributed State Estimation for Large Scale Distribution Systems

The growing need and interest in real-time monitoring of large distribution networks motivated by the rapid population of renewable sources, EVs and etc. demand a computationally efficient state estimation framework. Furthermore, this paper presents an improved computational framework for implementing a robust state estimator using a multi-core processor. The main contribution of the paper is the proposed computational framework along with two partitioning strategies which enable fast and robust state estimation for large scale radial and/or meshed distribution systems. Formulation of the proposed method and its implementation are described in detail. Performance of the estimator is tested by simulations first using a small 84-bus radial distribution system. Then the method’s scalability is demonstrated by simulations on two very large scale distribution networks one configured radially and the other meshed each containing over 12,500 buses.

42 ENGINEERING↗

rustpix

rustpix is a high-performance, open-source Rust library with first-class Python bindings (via PyO3) for processing pixel-detector data in neutron imaging. It targets time-stamping detectors such as Timepix3 (TPX3) at ORNL's Spallation Neutron Source (VENUS beamline), where each detected neutron deposits charge across a cluster of pixels within a very high-rate event stream (96M+ hits/sec). rustpix parses TPX3 event data in parallel using memory-mapped I/O, offers four interchangeable clustering algorithms (ABS adjacency-based search, DBSCAN, graph/union-find connected components, and a parallel grid method), and extracts weighted, super-resolved centroids to produce neutron-event lists. A streaming architecture lets it process files larger than available memory. rustpix is distributed as a pip-installable Python package (with NumPy integration), Rust crates, a command-line tool, and an interactive GUI; it writes HDF5, Apache Arrow, and CSV; and it is designed to extend to TPX4 and other detector types. Released as open-source under the MIT License.

Zhang, Chen [Oak Ridge National Laboratory (ORNL),↗

Enabling Recycling of Composites: Understanding the Impacts of Multiple Thermal Processing Cycles

When considering the utilization of recycled short carbon fiber feedstock materials for advanced manufacturing, understanding the material degradation behavior is essential in determining how many times a composite material can be effectively reprocessed and remanufactured. This study characterizes the degradation behavior of short carbon fiber acrylonitrile butadiene (CF-ABS) that has been reprocessed five times with twin screw extrusion. Parallel plate rheology was completed to observe the degradation in complex viscosity of the recycled feedstock materials. Gel permeation chromatography (GPC) was utilized to characterize the changes in molecular weight distribution of the recycled materials as a result of thermal and mechanical degradation during the re-processing steps. Rheological characterization, GPC, and twin-screw processing data help inform the process optimizations required to process the recycled feedstock material. Successful characterization of the degradation behavior of short fiber composite feedstock materials aids in increased understanding of the lifespan of high value carbon fiber composite materials and aids in process optimization of recycled composite materials.

Walker, Roo↗

TriC: Distributed-memory Triangle Counting by Exploiting the Graph Structure

Graph analytics has emerged as an important tool in the analysis of large scale data from diverse application domains such as social networks, cyber security and bioinformatics. Counting the number of triangles in a graph is a fundamental kernel with several applications such as detecting the community structure of a graph or in identifying important vertices in a graph. The ubiquity of massive datasets is driving the need to scale graph analytics on parallel systems. However, numerous challenges exist in efficiently parallelizing graph algorithms, especially on distributed-memory systems. Irregular memory accesses and communication patterns, low computation to communication ratios, and the need for frequent synchronization are some of the leading challenges. In this paper, we present TriC, our distributed-memory implementation of triangle counting in graphs using the Message Passing Interface (MPI), as a submission to the 2020 GraphChallenge competition. Using a set of synthetic and real-world inputs from the challenge, we demonstrate a speedup of up to 90x relative to previous work on 32 processor-cores of a NERSC Cori node. We also provide details from distributed runs with up to8192 processes along with strong scaling results. The observations presented in this work provide an understanding of the system-level bottlenecks at scale that specifically impact sparse-irregular workloads and will therefore benefit other efforts to parallelize graph algorithms.

Halappanavar, Mahantesh↗

Operating Stresses and Their Effects on Degradation of LSM-Based SOFC Cathodes

The performance of solid oxides fuel cells (SOFCs) with four different Mn excess of lanthanum strontium manganite (LSM) -based cathodes were examined under different temperatures (1,000 °C, 900 °C), current densities (0, 380, and 760 mA cm -2 ), and cathode atmospheres (p O 2 )=0.1,0.15,0.21) for durations ranging from 58 to 1,008 h. Each yttria-stabilized zirconia (YSZ) electrolyte-supported “button” cell had a porous Ni/YSZ composite anode and a porous LSM/YSZ composite cathode. The cells’ output voltage versus time were recorded and electrochemical impedance spectroscopy (EIS) and linear sweep voltammetry (LSV) measurements were performed every 24 hours. The total area specific resistance (ASR) was calculated from these measurements. The values of ASR from both EIS and LSV were comparable between each other but lower than that from durability test. Distribution of relaxation times (DRT) analysis was performed to investigate the electrochemical processes and their corresponding relaxation frequencies as well as their attributions to the total ASR. The series resistance and parallel resistance that obtained from the equivalent circuit fit of the Nyquist plot were higher at low temperature and low p O 2 regardless of LSM compositions. The area of the peaks derived from DRT analysis increased with time during individual tests. The microstructures of the tested cells were examined using scanning electron microscopy and energy-dispersive x-ray spectroscopy (SEM/EDS). Manganese oxide particles were observed near the cathode-electrolyte interface after prolonged (1,008 h) testing in air of a cell with LSM of 11% manganese excess, and after a short test (58 h) under low oxygen (p O 2 = 0.10) in a cell with LSM of 5% manganese excess. However, the role of MnOx in cell degradation is still unknown.

08 HYDROGEN↗

A Scalable Parallel Hypergraph Generator (HyGen)

Graphs are extensively used to model real-world complex systems. An edge in a graph can model pairwise relationships. However, multiway relationships (connections between three or more vertices) are common in many complex systems such as cellular process, image segmentation, and circuit design. A graph edge cannot model multiway relationships. A hypergraph, which can connect more than two vertices, is thus a better option to model multiway relationships. A large-scale hypergraph analysis has the potential to find useful insights from a complex system and assist in knowledge discovery. Currently a limited number of hypergraphs exists that are representative of real-world datasets. Moreover, real-world hypergraph datasets are small in size and inadequate to incorporate future needs. A graph generator that can produce large-scale synthetic hypergraphs can solve the above mentioned problems. In this paper, we present a scalable parallel hypergraph generator (HyGen) based on the Message Passing Interface (MPI) standard. To generate hypergraphs, HyGen takes the following parameter values as inputs: i) number of vertices, ii) number of hyperedges, iii) number of clusters, iv) vertex distribution, v) hyperedge distribution, vi) local cluster cardinality, and vii) global cluster cardinality. We have demonstrated that HyGen can generate hypergraphs of various sizes in a scalable fashion. HyGen takes approximately four minutes to generate a hypergraph with 4.8 million vertices, 1.6 million hyperedges, and 800 clusters using 1,024 processes on a leadership class computing platform. Our strong and weak scaling experiments on supercomputers demonstrate that HyGen can quickly create large-scale hypergraphs in a parallel manner, thus providing a useful capability for hypergraph analysis.

Hasan, S M Shamimul↗

Adaptive Spatially Aware I/O for Multiresolution Particle Data Layouts

Large-scale simulations on nonuniform particle distributions that evolve over time are widely used in cosmology, molecular dynamics, and engineering. Such data are often saved in an unstructured format that neither preserves spatial locality nor provides metadata for accelerating spatial or attribute subset queries, leading to poor performance of visualization tasks. Furthermore, the parallel I/O strategy used typically writes a file per process or a single shared file, neither of which is portable or scalable across different HPC systems. We present a portable technique for scalable, spatially aware adaptive aggregation that preserves spatial locality in the output. We evaluate our approach on two supercomputers, Stampede2 and Summit, and demonstrate that it outperforms prior approaches at scale, achieving up to 2.5× faster writes and reads for nonuniform distributions. Furthermore, the layout written by our method is directly suitable for visual analytics, supporting low-latency reads and attribute-based filtering with little overhead.

Usher, Will↗

Kinetic entropy-based measures of distribution function non-Maxwellianity: theory and simulations

We investigate kinetic entropy-based measures of the non-Maxwellianity of distribution functions in plasmas, i.e. entropy-based measures of the departure of a local distribution function from an associated Maxwellian distribution function with the same density, bulk flow and temperature as the local distribution. First, we consider a form previously employed by Kaufmann & Paterson (J. Geophys. Res., vol. 114, 2009, A00D04), assessing its properties and deriving equivalent forms. To provide a quantitative understanding of it, we derive analytical expressions for three common non-Maxwellian plasma distribution functions. We show that there are undesirable features of this non-Maxwellianity measure including that it can diverge in various physical limits and elucidate the reason for the divergence. We then introduce a new kinetic entropy-based non-Maxwellianity measure based on the velocity-space kinetic entropy density, which has a meaningful physical interpretation and does not diverge. We use collisionless particle-in-cell simulations of two-dimensional anti-parallel magnetic reconnection to assess the kinetic entropy-based non-Maxwellianity measures. We show that regions of non-zero non-Maxwellianity are linked to kinetic processes occurring during magnetic reconnection. We also show the simulated non-Maxwellianity agrees reasonably well with predictions for distributions resembling those calculated analytically. These results can be important for applications, as non-Maxwellianity can be used to identify regions of kinetic-scale physics or increased dissipation in plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗