Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed and parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Algorithms and programming tools for image processing on the MPP, part 2

A number of algorithms were developed for image warping and pyramid image filtering. Techniques were investigated for the parallel processing of a large number of independent irregular shaped regions on the MPP. In addition some utilities for dealing with very long vectors and for sorting were developed. Documentation pages for the algorithms which are available for distribution are given. The performance of the MPP for a number of basic data manipulations was determined. From these results it is possible to predict the efficiency of the MPP for a number of algorithms and applications. The Parallel Pascal development system, which is a portable programming environment for the MPP, was improved and better documentation including a tutorial was written. This environment allows programs for the MPP to be developed on any conventional computer system; it consists of a set of system programs and a library of general purpose Parallel Pascal functions. The algorithms were tested on the MPP and a presentation on the development system was made to the MPP users group. The UNIX version of the Parallel Pascal System was distributed to a number of new sites.

Reeves, Anthony P.↗

Applications of a transonic wing design method

A method for designing wings and airfoils at transonic speeds using a predictor/corrector approach was developed. The procedure iterates between an aerodynamic code, which predicts the flow about a given geometry, and the design module, which compares the calculated and target pressure distributions and modifies the geometry using an algorithm that relates differences in pressure to a change in surface curvature. The modular nature of the design method makes it relatively simple to couple it to any analysis method. The iterative approach allows the design process and aerodynamic analysis to converge in parallel, significantly reducing the time required to reach a final design. Viscous and static aeroelastic effects can also be accounted for during the design or as a post-design correction. Results from several pilot design codes indicated that the method accurately reproduced pressure distributions as well as the coordinates of a given airfoil or wing by modifying an initial contour. The codes were applied to supercritical as well as conventional airfoils, forward- and aft-swept transport wings, and moderate-to-highly swept fighter wings. The design method was found to be robust and efficient, even for cases having fairly strong shocks.

Campbell, Richard L.↗

Proximity Portability and in Transit , M-to-N Data Partitioning and Movement in SENSEI [Book Chapter]

In high-performance parallel in situ processing, the term in transit processing refers to those configurations where data must move from a producer to a consumer that runs on separate resources. In the context of parallel and distributed computing on an HPC platform one of the central challenges is to determine a mapping of data from producer ranks to consumer ranks. This problem is complicated by the heterogeneity that arises in producer-consumer pairs, such as when producer and consumer codes have different levels of concurrency, different scaling characteristics, or different data models. The resulting mapping and movement of data from M producer to N consumer ranks can have a significant impact on aggregate application performance, particularly when the data consumer requires only a subset of the overall data for its task. This chapter focuses on the design considerations that underlie SENSEI’s implementation to this challenging problem. These design considerations extend the core SENSEI architecture and include ideas like the need to accommodate flexibility in the choice of different partitioning methods, the ability for a data consumer to request and receive only the subset of data needed for its particular operation, and the ability to leverage any of several different data transport tools. The idea of proximity portability, being able to use different data transport methods as part of an in transit workflow, is illustrated through the use of three different transport layers where switching from one transport tool to another is accomplished with only a configuration file change. Here, the chapter also includes a performance analysis summary showing the performance gains that are possible in terms of multiple metrics, such as memory footprint, time to solution, and amount of data moved, when using optimized partitioners in an in transit setting, gains that are made possible by the implementation shaped by specific design considerations.

Bethel, E. Wes↗

Sequence length scaling in vision transformers for scientific images on frontier

Vision Transformers (ViTs) are pivotal for foundational models in scientific imagery, including Earth science applications, due to their capability to process large sequence lengths. While transformers for text have inspired scaling sequence lengths in ViTs, adapting these for ViTs introduces unique challenges. We develop distributed sequence parallelism for ViTs, enabling them to handle up to 1M tokens. Our approach, leveraging DeepSpeed-Ulysses and Long-Sequence-Segmentation with model sharding, is the first to apply sequence parallelism in ViT training, achieving a 94% batch scaling efficiency on 2,048 AMD-MI250X GPUs. Evaluating sequence parallelism in ViTs, particularly in models up to 10B parameters, highlighted substantial bottlenecks. We countered these with hybrid sequence, pipeline, and flash attention strategies, to scale beyond single GPU memory limits. Our method significantly enhances climate modeling accuracy by 20% in temperature predictions, marking the first training of a vision transformer model to convergence with a sequence length of 188K tokens, using full self-attention.

Tsaris, Aristeidis (aris) [ORNL] (ORCID:0000000277↗

A component decomposition model for evaluating atmospheric effects in remote sensing

A radiance value of a target pixel recorded by a remote sensor can be decomposed into three components: (1) attenuated target signature, (2) pure atmospheric radiation, and (3) the contribution made by the ground through the atmospheric scattering process. Given the meteorological and optical parameters of a layer-structured atmosphere, its transmittance and radiance distribution can be accurately calculated with a plane-parallel radiative transfer model. For a uniform surface, the ground contribution can be obtained by comparing radiances for an atmosphere over a black but nonemitting surface and the same atmosphere with an underlying ground of given albedo or temperature. For an inhomogeneous surface, the first two components remain the same as long as the surface is a plane. The third may be estimated using the locally averaged top-of-atmosphere radiance. An atmospheric point spread function is calculated by a Monte Carlo approach and is used for retrieving the ground signature through a deconvolution procedure.

Li, S.↗

2nd Generation QUATARA Flight Computer Project

Single core flight computer boards have been designed, developed, and tested (DD&T) to be flown in small satellites for the last few years. In this project, a prototype flight computer will be designed as a distributed multi-core system containing four microprocessors running code in parallel. This flight computer will be capable of performing multiple computationally intensive tasks such as processing digital and/or analog data, controlling actuator systems, managing cameras, operating robotic manipulators and transmitting/receiving from/to a ground station. In addition, this flight computer will be designed to be fault tolerant by creating both a robust physical hardware connection and by using a software voting scheme to determine the processor's performance. This voting scheme will leverage on the work done for the Space Launch System (SLS) flight software. The prototype flight computer will be constructed with Commercial Off-The-Shelf (COTS) components which are estimated to survive for two years in a low-Earth orbit.

Falker, Jay↗

Partitioning of unstructured problems for parallel processing

Many large-scale computational problems are based on unstructured computational domains. Primary examples are unstructured grid calculations based on finite volume methods in computational fluid dynamics, or structural analysis problems based on finite element approximations. The question of how to distribute such unstructured computational domains over a large number of processors in a MIMD machine with distributed memory is addressed. A graph theoretical framework for these problems is established. Based on this framework three decomposition algorithms are introduced. In particular a new decomposition algorithm is discussed, which is based on the computation of an eigenvector of the Laplacian matrix associated with the graph. Numerical comparisons on large-scale two- and three-dimensional problems demonstrate the superiority of the new spectral bisection algorithm.

Simon, H. D.↗

Current Practices in Distribution Utility Resilience Planning for Winter Storms

This report is part of a series of hazard-focused case studies examining common practices in electric utility resilience planning. We use standard terminology defining resilience as the ability to anticipate, withstand, absorb, and recover from hazards that cause long duration outages. We distinguish between reliability and resilience using Institute of Electrical and Electronics Engineers (IEEE) 1366-2022, which defines major events as an event that exceeds reasonable design and/or operational limits of the electric power system. Resilience planning is focused on major event days and reliability planning is focused on nonmajor event days. Utility resilience plans are assessed according to common resilience components identified in existing resilience frameworks. The focus of this report is on winter storms in which the primary hazards are heavy snowfall, freezing rain, ice, extreme cold, severe wind, and flooding. These hazards can also contribute to generation shortages, resulting in bulk power system impacts that have consequences for the distribution system, such as load shedding. Stand-alone reports focusing on wildfires and nonwinter storms have been published in parallel with this report. This report can be used as a starting point for understanding potential investment prioritization processes and investment options. This report is intended to improve utility resilience planning by supporting constructive dialogue among utilities, regulators, and other stakeholders.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Flexible User-Defined Domain Decomposition in Kilometer-Scale E3SM Land Model Simulation

The Energy Exascale Earth System Model (E3SM) Land Model (ELM) has been extended to kilometer-scale (km-ELM) resolutions, enabling high-fidelity simulations of terrestrial processes at 1 km x 1 km grid spacing. In ELM, domain decomposition partitions the computational domain across processors, ensuring efficient parallel execution. Currently, round-robin decomposition is applied, providing a straightforward way to distribute computational workload. As ELM continues evolving at the kilometer-scale (km-scale), particularly with integrating lateral flow modeling, decomposition strategies must also account for the increased workload and data movement. This paper introduces a flexible user-defined domain decomposition framework, allowing users to customize domain partitioning based on application requirements. The impact of different decomposition strategies is evaluated across various applications concerning computation, communication, and I/O. Results demonstrate that while 1D partitioning yields superior I/O performance, k-nearest neighbors (KNN) clustering effectively reduces inter-process communication overhead. This study lays the groundwork for scalable partitioning in large-scale land surface simulations, enhancing next-generation Earth system modeling.

Wang, Dali [ORNL] (ORCID:0000000168065108)↗

Adding GPU Support to the Markov Chain Monte Carlo Code Catmip

In geophysics, we are confronted with many under-determined inverse problems. For example, all of our observations of earthquakes are made at the Earth’s surface. So, when we try to infer how slip during an earthquake evolves in space and time, we find that there are many potential slip histories that are consistent with our limited observations and our understanding of earthquake physics. One way to approach these problems is with Bayesian analysis which allows us to infer the ensemble of all potential slip models that satisfy the observations and our prior knowledge of earthquake physics. In Bayesian analysis, our prior knowledge is known as the prior probability density function or prior PDF, the fit to the data is known as the data likelihood, and the target PDF that satisfies both the prior PDF and data likelihood is known as the posterior PDF. However, simulating the posterior PDF typically requires using Markov Chain Monte Carlo (MCMC) to draw tens of billions of random realizations of earthquake slip models, which may not be computationally feasible. To make this and similar geophysical inversions computationally tractable, we developed the Cascading Adaptive Transitional Metropolis In Parallel (CATMIP) algorithm. CATMIP is an efficient parallel Markov Chain Monte Carlo (MCMC) sampler that is used for model fitting and uncertainty quantification in geophysics. Example use cases are earthquake rupture modeling, determining mineral composition on Mars, reconstructing the history of ocean salinity, and historical earthquake relocation. CATMIP employs many parallel instances of the Metropolis algorithm for sampling in a transitioning framework. Transitioning is a process in which a set of random samples at equilibrium with a known probability density function (PDF) are used as seeds for the Markov chains to sample successive target PDFs that incrementally move the distribution from the starting seeds to the final desired PDF that describes the relative plausibility of potential values for the model parameters. The algorithm is implemented as a Master-Worker model employing MPI for communication. The worker processes are loosely coupled with global parameters periodically optimized by the master process. This provides a very high amount of parallelism with little communication between updates. During the presentation we will discuss the history of the algorithm and elaborate the earthquake rupture modeling use case for the CATMIP package. Our first step toward GPU optimization was to optimize the code for the CPU. CPU profiling revealed that most of the compute time is spent in calls to level 2 BLAS routines and calls to GSL random number generators. We revised the algorithm to employ level 3 BLAS routines instead. In our presentation we will describe how this was accomplished. Adding GPU support to CATMIP consisted mostly of replacing the calls to GSL with calls to GPU vendor-provided library routines. A small number of loops were directly implemented in CUDA. In the presentation will provide implementation details. Finally, we will discuss methods for profiling and opportunities for further optimizing GPU execution. By creating a code with the flexibility to run on either a CPU or GPU architecture, CATMIP can be used on systems ranging from large CPU-based HPC environments to single servers with GPU acceleration and everything in between.

HECC↗

Density and Magnetic Field Asymmetric Kelvin‐Helmholtz Instability

Abstract The Kelvin‐Helmholtz (KH) instability can transport mass, momentum, magnetic flux, and energy between the magnetosheath and magnetosphere, which plays an important role in the solar‐wind‐magnetosphere coupling process for different planets. Meanwhile, strong density and magnetic field asymmetry are often present between the magnetosheath (MSH) and magnetosphere (MSP), which could affect the transport processes driven by the KH instability. Our magnetohydrodynamics simulation shows that the KH growth rate is insensitive to the density ratio between the MSP and the MSH in the compressible regime, which is different than the prediction from linear incompressible theory. When the interplanetary magnetic field (IMF) is parallel to the planet's magnetic field, the nonlinear KH instability can drive a double mid‐latitude reconnection (DMLR) process. The total double reconnected flux depends on the KH wavelength and the strength of the lower magnetic field. When the IMF is anti‐parallel to the planet's magnetic field, the nonlinear interaction between magnetic reconnection and the KH instability leads to fast reconnection (i.e., close to Petschek reconnection even without including kinetic physics). However, the peak value of the reconnection rate still follows the asymmetric reconnection scaling laws. We also demonstrate that the DMLR process driven by the KH instability mixes the plasma from different regions and consequently generates different types of velocity distribution functions. We show that the counter‐streaming beams can be simply generated via the change of the flux tube connection and do not require parallel electric fields.

Astronomy & Astrophysics↗

Laser-induced slip casting as an additive manufacturing approach for silicon carbide

Here, this work presents processing silicon carbide (SiC) with the laser-induced slip casting (LIS) additive manufacturing (AM). SiC was stabilized in water with polyethyleneimine (PEI) dispersant, and SiC slurries were made with rheology for LIS printing. High-density ceramic parts were printed, followed by single-step binder burnout and sintering. The printed parts achieved 93–95 % of theoretical density. X-ray computed tomography (XCT) revealed a small distribution of flaws exceeding 100 microns. The mechanical properties were measured in both parallel and perpendicular to the printing layers, and the orientation with layers perpendicular to the bending moment resulted in higher strength compared to the parallel direction. Porosity resulting from processing and large inclusions of boron carbide (B4C) were the root cause of failure in the measured samples. Despite these defects through this effort, this new approach demonstrates promise for green forming of SiC with densities greater than 95 % theoretical and tensile strengths above 250 MPa.

Additive Manufacturing↗

Distributed Macroscopic Traffic Simulation with Open Traffic Models

This paper presents OTM-MPI, an extension of the Open Traffic Models platform (OTM) for running macroscopic traffic simulations in high-performance computing environments. OTM-MPI represents the first open-source, distributed-memory, macroscopic simulation model developed for modern high performance parallel machines and large networks. Macroscopic simulations are appropriate for studying regional traffic scenarios when aggregate trends are of interest, rather than individual vehicle traces. They are also appropriate for studying the routing behavior of classes of vehicles, such as app-informed vehicles. The network partitioning was performed with METIS. Inter-process communication was done with MPI (message-passing interface). Results are provided for two networks: one realistic network which was obtained from Open Street Maps for Chattanooga, TN, and another larger synthetic grid network. The software recorded a speedups of 198x using 256 cores for Chattanooga, and 475x with 1,024 cores for the synthetic network.

macro-scopic traffic simulation↗

Record acceleration of the two-dimensional Ising model using a high-performance wafer-scale engine

The versatility and wide-ranging applicability of the Ising model, originally introduced to study phase transitions in magnetic materials, have made it a cornerstone in statistical physics and a valuable tool for evaluating the performance of emerging computer hardware. Here, we present a novel implementation of the two-dimensional Ising model on Cerebras Wafer-Scale Engine (WSE) – a revolutionary processor that is opening new frontiers in computing. In our deployment of the checkerboard algorithm, we optimized the Ising model to take advantage of the unique WSE architecture. Specifically, we employed a compressed bit representation storing 16 spins on each int16 word, and efficiently distributed the spins over the processing units enabling seamless weak scaling and limiting communications to only immediate neighboring units. Our implementation can handle up to 754 simulations in parallel, achieving an aggregate of over 61.8 trillion flip attempts per second for Ising models with up to 200 million spins. This represents a gain of up to 148 times over previously reported single-devices with a highly optimized implementation on NVIDIA V100 and up to 88 times in productivity compared to NVIDIA H100. Our findings highlight the significant potential of the WSE in scientific computing, particularly in the field of materials modeling.

Ising model↗

Adapting high-level language programs for parallel processing using data flow

EASY-FLOW, a very high-level data flow language, is introduced for the purpose of adapting programs written in a conventional high-level language to a parallel environment. The level of parallelism provided is of the large-grained variety in which parallel activities take place between subprograms or processes. A program written in EASY-FLOW is a set of subprogram calls as units, structured by iteration, branching, and distribution constructs. A data flow graph may be deduced from an EASY-FLOW program.

Standley, Hilda M.↗

Distributed Finite Element Analysis Using a Transputer Network

The principal objective of this research effort was to demonstrate the extraordinarily cost effective acceleration of finite element structural analysis problems using a transputer-based parallel processing network. This objective was accomplished in the form of a commercially viable parallel processing workstation. The workstation is a desktop size, low-maintenance computing unit capable of supercomputer performance yet costs two orders of magnitude less. To achieve the principal research objective, a transputer based structural analysis workstation termed XPFEM was implemented with linear static structural analysis capabilities resembling commercially available NASTRAN. Finite element model files, generated using the on-line preprocessing module or external preprocessing packages, are downloaded to a network of 32 transputers for accelerated solution. The system currently executes at about one third Cray X-MP24 speed but additional acceleration appears likely. For the NASA selected demonstration problem of a Space Shuttle main engine turbine blade model with about 1500 nodes and 4500 independent degrees of freedom, the Cray X-MP24 required 23.9 seconds to obtain a solution while the transputer network, operated from an IBM PC-AT compatible host computer, required 71.7 seconds. Consequently, the $80,000 transputer network demonstrated a cost-performance ratio about 60 times better than the $15,000,000 Cray X-MP24 system.

Watson, James↗

Heating and acceleration of ions with Kappa distribution functions by low‐frequency Alfvén wave

Abstract Heating and acceleration of ions with Kappa distribution functions (with parameter ) in low‐beta plasmas, by a low‐frequency Alfvén wave, is investigated using test‐particle simulations, yielding interesting new results. As long as the Alfvén wave amplitude is sufficiently large, the computed net heating energy of ions becomes independent of the wave frequency and amplitude, always approaching the same value of . The eventual energy of ions is dictated only by the initial ion energy and the ratio of the magnetic field energy density to the plasma density. The heating effect of the Kappa ions increases with . During the heating process, the ions are picked up by the Alfvén wave and pitch angle scattered, forming a quasi‐isotropic spherical shell velocity distribution. The Kappa ions are accelerated in the parallel direction, reaching a bulk flow speed roughly equal to the local Alfvén speed. Higher ‐value in the initial Kappa distribution leads to faster saturation. The above results may explain certain features of the ion heating and acceleration in the solar wind and corona.

Li, Kehua↗

A Robust Parallel Distributed State Estimation for Large Scale Distribution Systems

The growing need and interest in real-time monitoring of large distribution networks motivated by the rapid population of renewable sources, EVs and etc. demand a computationally efficient state estimation framework. Furthermore, this paper presents an improved computational framework for implementing a robust state estimator using a multi-core processor. The main contribution of the paper is the proposed computational framework along with two partitioning strategies which enable fast and robust state estimation for large scale radial and/or meshed distribution systems. Formulation of the proposed method and its implementation are described in detail. Performance of the estimator is tested by simulations first using a small 84-bus radial distribution system. Then the method’s scalability is demonstrated by simulations on two very large scale distribution networks one configured radially and the other meshed each containing over 12,500 buses.

42 ENGINEERING↗