Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

Multiple scattering in reflection nebulae. I - A Monte Carlo approach. II - Uniform plane-parallel nebulae with foreground stars. III - Nebulae with embedded illuminating stars

A method for calculating the surface brightness distribution on a plane-parallel reflection nebula of uniform density illuminated by a star located either in front of, behind, or arbitrarily inside the scattering medium is proposed. The Monte Carlo technique is used to find solutions to the radiative transfer problem. The scattering properties of the nebular particles are parameterized by the albedo for single scattering and a three-parameter analytic phase function. Calculations are then presented for the surface brightness distribution across the face of such nebulae with (1) a foreground star and (2) immersed stars. The calculations include the full effects of multiple scattering, are independent of a particular assumed grain material or size distribution, and are applicable to any wavelength region for which observations can be obtained.

Witt, A. N.↗

The cleft ion fountain - A two-dimensional kinetic model

The transport of ionospheric ions from a source in the polar cleft ionosphere through the polar magnetosphere is investigated using a two-dimensional, kinetic, trajectory-based code. The transport model includes the effects of gravitation, longitudinal magnetic gradient force, convection electric fields, and parallel electric fields. Individual ion trajectories as well as distribution functions and resulting bulk parameters of density, parallel average energy, and parallel flux for a presumed cleft ionosphere source distribution are presented for various conditions to illustrate parametrically the dependences on source energies, convection electric field strengths, ion masses, and parallel electric field strengths. The essential features of the model are consistent with the concept of a cleft-based ion fountain supplying ionospheric ions to the polar magnetosphere, and the resulting plasma distributions and parameters are in general agreement with recent low-energy ion measurements from the DE 1 satellite.

Horwitz, J. L.↗

Distributed Training for High Resolution Images: A Domain and Spatial Decomposition Approach

In this work we developed two Pytorch libraries using the PyTorch RPC interface for distributed deep learning approaches on high resolution images. The spatial decomposition library allows for distributedtraining on very large images, which otherwise won’t be possible on a single GPU. The domain parallelism library allows for distributed training across multiple domain unlabeled data, by leveraging the domain separation architecture. Both of those libraries were tested on the Summit supercomputer at a moderate scale, and we are releasing the code for both of them.

97 MATHEMATICS AND COMPUTING↗

Reducing Interprocessor Dependence in Recoverable Distributed Shared Memory

Checkpointing techniques in parallel systems use dependency tracking and/or message logging to ensure that a system rolls back to a consistent state. Traditional dependency tracking in distributed shared memory (DSM) systems is expensive because of high communication frequency. In this paper we show that, if designed correctly, a DSM system only needs to consider dependencies due to the transfer of blocks of data, resulting in reduced dependency tracking overhead and reduced potential for rollback propagation. We develop an ownership timestamp scheme to tolerate the loss of block state information and develop a passive server model of execution where interactions between processors are considered atomic. With our scheme, dependencies are significantly reduced compared to the traditional message-passing model.

Janssens, Bob↗

Scaling neural simulations in STACS

Abstract As modern neuroscience tools acquire more details about the brain, the need to move towards biological-scale neural simulations continues to grow. However, effective simulations at scale remain a challenge. Beyond just the tooling required to enable parallel execution, there is also the unique structure of the synaptic interconnectivity, which is globally sparse but has relatively high connection density and non-local interactions per neuron. There are also various practicalities to consider in high performance computing applications, such as the need for serializing neural networks to support potentially long-running simulations that require checkpoint-restart. Although acceleration on neuromorphic hardware is also a possibility, development in this space can be difficult as hardware support tends to vary between platforms and software support for larger scale models also tends to be limited. In this paper, we focus our attention on Simulation Tool for Asynchronous Cortical Streams (STACS), a spiking neural network simulator that leverages the Charm++ parallel programming framework, with the goal of supporting biological-scale simulations as well as interoperability between platforms. Central to these goals is the implementation of scalable data structures suitable for efficiently distributing a network across parallel partitions. Here, we discuss a straightforward extension of a parallel data format with a history of use in graph partitioners, which also serves as a portable intermediate representation for different neuromorphic backends. We perform scaling studies on the Summit supercomputer, examining the capabilities of STACS in terms of network build and storage, partitioning, and execution. We highlight how a suitably partitioned, spatially dependent synaptic structure introduces a communication workload well-suited to the multicast communication supported by Charm++. We evaluate the strong and weak scaling behavior for networks on the order of millions of neurons and billions of synapses, and show that STACS achieves competitive levels of parallel efficiency.

59 BASIC BIOLOGICAL SCIENCES↗

Generating local addresses and communication sets for data-parallel programs

Generating local addresses and communication sets is an important issue in distributed-memory implementations of data-parallel languages such as High Performance FORTRAN. We show that, for an array A affinely aligned to a template that is distributed across p processors with a cyclic(k) distribution and a computation involving the regular section A(l:h:s), the local memory access sequence for any processor is characterized by a finite state machine of at most k states. We present fast algorithms for computing the essential information about these state machines, and extend the framework to handle multidimensional arrays. We also show how to generate communication sets using the state machine approach. Performance results show that this solution requires very little run-time overhead and acceptable preprocessing time.

Chatterjee, Siddhartha↗

Generating local addresses and communication sets for data-parallel programs

Generating local addresses and communication sets is an important issue in distributed-memory implementations of data-parallel languages such as High Performance Fortran. We show that for an array A affinely aligned to a template that is distributed across p processors with a cyclic(k) distribution, and a computation involving the regular section A, the local memory access sequence for any processor is characterized by a finite state machine of at most k states. We present fast algorithms for computing the essential information about these state machines, and extend the framework to handle multidimensional arrays. We also show how to generate communication sets using the state machine approach. Performance results show that this solution requires very little runtime overhead and acceptable preprocessing time.

Chatterjee, Siddhartha↗

The ParaScope parallel programming environment

The ParaScope parallel programming environment, developed to support scientific programming of shared-memory multiprocessors, includes a collection of tools that use global program analysis to help users develop and debug parallel programs. This paper focuses on ParaScope's compilation system, its parallel program editor, and its parallel debugging system. The compilation system extends the traditional single-procedure compiler by providing a mechanism for managing the compilation of complete programs. Thus, ParaScope can support both traditional single-procedure optimization and optimization across procedure boundaries. The ParaScope editor brings both compiler analysis and user expertise to bear on program parallelization. It assists the knowledgeable user by displaying and managing analysis and by providing a variety of interactive program transformations that are effective in exposing parallelism. The debugging system detects and reports timing-dependent errors, called data races, in execution of parallel programs. The system combines static analysis, program instrumentation, and run-time reporting to provide a mechanical system for isolating errors in parallel program executions. Finally, we describe a new project to extend ParaScope to support programming in FORTRAN D, a machine-independent parallel programming language intended for use with both distributed-memory and shared-memory parallel computers.

Cooper, Keith D.↗

Efficient computation of aerodynamic influence coefficients for aeroelastic analysis on a transputer network

Aeroelastic analysis is multi-disciplinary and computationally expensive. Hence, it can greatly benefit from parallel processing. As part of an effort to develop an aeroelastic capability on a distributed memory transputer network, a parallel algorithm for the computation of aerodynamic influence coefficients is implemented on a network of 32 transputers. The aerodynamic influence coefficients are calculated using a 3-D unsteady aerodynamic model and a parallel discretization. Efficiencies up to 85 percent were demonstrated using 32 processors. The effect of subtask ordering, problem size, and network topology are presented. A comparison to results on a shared memory computer indicates that higher speedup is achieved on the distributed memory system.

Janetzke, David C.↗

Parametric analysis of hollow conductor parallel and coaxial transmission lines for high frequency space power distribution

A parametric analysis was performed of transmission cables for transmitting electrical power at high voltage (up to 1000 V) and high frequency (10 to 30 kHz) for high power (100 kW or more) space missions. Large diameter (5 to 30 mm) hollow conductors were considered in closely spaced coaxial configurations and in parallel lines. Formulas were derived to calculate inductance and resistance for these conductors. Curves of cable conductance, mass, inductance, capacitance, resistance, power loss, and temperature were plotted for various conductor diameters, conductor thickness, and alternating current frequencies. An example 5 mm diameter coaxial cable with 0.5 mm conductor thickness was calculated to transmit 100 kW at 1000 Vac, 50 m with a power loss of 1900 W, an inductance of 1.45 micron and a capacitance of 0.07 micron-F. The computer programs written for this analysis are listed in the appendix.

Jeffries, K. S.↗

Analysis of Wave and Particle Signatures Observed in Plasma Escape at Venus

Atmospheric gases escape from Venus as neutral and ionized atoms and molecules. Ion escape, considered here, occurs through ion pickup or collective plasma processes. The latter can arise from upward flow of nightside ionospheric plasma into the ionotail, day to night ionospheric flow into the ionotail, and scavenging of ionospheric plasma by ionosphere-magnetosheath instabilities at the ionopause. These plasma processes produce differing signatures in ion velocity and energy distributions and in ULF waves in the magnetic field. Using plasma ion spectra measured by the Pioneer Venus Orbiter (PVO) Orbiter Plasma Analyzer (OPA) and magnetic field fluctuations observed by the PVO Orbiter Magnetometer (OMAG) along with the expected particle and field signatures, various ion escape processes occurring along Pioneer Venus orbits are identified. In particular, OPA ion energy distributions are used in parallel with magnetic field power spectra and wave phase angles derived from OMAG measurements to study the characteristics of escaping ions. The principle ions observed escaping the influence of Venus are H+, He+ and 0'. In the ion energy distributions of the OPA, pickup ions appear hot relative to the much cooler ions flowing away from Venus in the ionotail and in the plasma clouds detached from the ionopause. This energy contrast is particularly evident downstream when PVO crosses the ionotail boundary from the hot solar wind plasma to the much cooler plasma within the tail. Magnetic field signatures accompanying the escaping ions appear as peaks in the power spectra at the corresponding ion cyclotron frequencies. Also, coherent wave trains at the same frequencies are observed in the phase angle plots of magnetic field fluctuations about the mean field.

Hartle, R. E.↗

Parallelizing autotuning for HPC applications: Unveiling the potential of the speculation strategy in Bayesian optimization

In the exascale computing era, tuning High-Performance Computing (HPC) applications has become a significant computational challenge. Although Bayesian optimization (BO) has emerged as a promising tool for HPC performance tuning, the BO workflow is inherently sequential (i.e., one function evaluation at a time) and cannot leverage the huge amount of parallel resources present in modern supercomputers, resulting in a considerable underutilization of their computational capabilities. This paper explores the trade-off between search quality and parallelism in BO, investigating a diverse set of methods. Building upon both previous approaches from the literature and novel methodologies introduced in this work, our study provides a deep analysis to accelerate BO performance tuning. By examining a set of synthetic functions and practical HPC applications, our exploration analyzes the interaction among various BO methods for parallelization, the quantity of parallel resources, the runtime distribution of target HPC applications, and the costs associated with different search orchestration mechanisms that have been overlooked in previous studies. Compared to sequential BO, our novel methodology achieves comparable quality while demonstrating robust scalability in search time as the amount of parallel resources increases; it also outperforms a state-of-the-art tuner, which supports parallelization, achieving up to 3.67x faster search time. We provide high-value insights for practitioners seeking to leverage the power of parallel computing for efficient HPC application tuning. Additionally, to further assist researchers in accelerating the performance tuning of their HPC applications, we provide an extension of an existing open-source tuning framework that incorporates our methods.

Bayesian optimization↗

Energy distribution functions of kilovolt ions in a modified Penning discharge.

The distribution function of ion energy parallel to the magnetic field of a modified Penning discharge has been measured with a retarding potential energy analyzer. These ions escaped through one of the throats of the magnetic mirror geometry. Simultaneous measurements of the ion energy distribution function perpendicular to the magnetic field have been made with a charge-exchange neutral detector. The ion energy distribution functions are approximately Maxwellian, and the parallel and perpendicular kinetic temperatures are equal within experimental error. These results suggest that turbulent processes previously observed in this discharge Maxwellianize the velocity distribution along a radius in velocity space, and result in an isotropic energy distribution.

Roth, J. R.↗

Energy distribution functions of kilovolt ions in a modified Penning discharge.

The distribution function of ion energy parallel to the magnetic field of a modified Penning discharge has been measured with a retarding potential energy analyzer. These ions escaped through one of the throats of the magnetic mirror geometry. Simultaneous measurements of the ion energy distribution function perpendicular to the magnetic field have been made with a charge-exchange neutral detector. The ion energy distribution functions are approximately Maxwellian, and the parallel and perpendicular kinetic temperatures are equal within experimental error. These results suggest that turbulent processes previously observed in this discharge Maxwellianize the velocity distribution along a radius in velocity space, and result in an isotropic energy distribution.

Roth, J. R.↗

Distributed Antenna-Coupled TES for FIR Detectors Arrays

We describe a new architecture for a superconducting detector for the submillimeter and far-infrared. This detector uses a distributed hot-electron transition edge sensor (TES) to collect the power from a focal-plane-filling slot antenna array. The sensors lay directly across the slots of the antenna and match the antenna impedance of about 30 ohms. Each pixel contains many sensors that are wired in parallel as a single distributed TES, which results in a low impedance that readily matches to a multiplexed SQUID readout These detectors are inherently polarization sensitive, with very low cross-polarization response, but can also be configured to sum both polarizations. The dual-polarization design can have a bandwidth of 50The use of electron-phonon decoupling eliminates the need for micro-machining, making the focal plane much easier to fabricate than with absorber-coupled, mechanically isolated pixels. We discuss applications of these detectors and a hybridization scheme compatible with arrays of tens of thousands of pixels.

antenna↗

Performance and Feature Improvements in Parareal-based Power System Dynamic Simulation

In recent years, a novel Parareal-based approach has been developed for fast transient simulations of large power system interconnections. Parareal belongs to the class of Parallel-in-time algorithms for solution of systems of differential-algebraic equations in parallel over an interval of time. The selection of a reasonably fast and accurate coarse solution is crucial to improve the performance of Parareal algorithm. Semi-analytical solution methods are one promising approach to achieve this goal. They have been investigated, and some preliminary results are presented here. In addition, Parareal-based simulator has been expanded to enable co-simulation with OpenDSS, a widely used open-source distribution system simulator. Preserving the parallel nature of the Parareal approach and taking advantage of the parallel capabilities of the latest versions of OpenDSS, each distribution system can be solved in their entirety on different processors in parallel within the main Parareal simulator. This paper also presents the structure of the transmission and distribution co-simulation and some results with different dynamic models of inverter-based resources in the distribution systems.

Park, Byungkwon↗