Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed and parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Evaluation of the Intel iWarp parallel processor for space flight applications

The potential of a DARPA-sponsored advanced processor, the Intel iWarp, for use in future SSF Data Management Systems (DMS) upgrades is evaluated through integration into the Ames DMS testbed and applications testing. The iWarp is a distributed, parallel computing system well suited for high performance computing applications such as matrix operations and image processing. The system architecture is modular, supports systolic and message-based computation, and is capable of providing massive computational power in a low-cost, low-power package. As a consequence, the iWarp offers significant potential for advanced space-based computing. This research seeks to determine the iWarp's suitability as a processing device for space missions. In particular, the project focuses on evaluating the ease of integrating the iWarp into the SSF DMS baseline architecture and the iWarp's ability to support computationally stressing applications representative of SSF tasks.

Hine, Butler P., III↗

On bottleneck partitioning k-ary n-cubes

Graph partitioning is a topic of extensive interest, with applications to parallel processing. In this context graph nodes typically represent computation, and edges represent communication. One seeks to distribute the workload by partitioning the graph so that every processor has approximately the same workload, and the communication cost (measured as a function of edges exposed by the partition) is minimized. Measures of partition quality vary; in this paper we consider a processor's cost to be the sum of its computation and communication costs, and consider the cost of a partition to be the bottleneck, or maximal processor cost induced by the partition. For a general graph the problem of finding an optimal partitioning is intractable. In this paper we restrict our attention to the class of k-art n-cube graphs with uniformly weighted nodes. Given mild restrictions on the node weight and number of processors, we identify partitions yielding the smallest bottleneck. We also demonstrate by example that some restrictions are necessary for the partitions we identify to be optimal. In particular, there exist cases where partitions that evenly partition nodes need not be optimal.

Nicol, David M.↗

Dermal Aged and Fetal Fibroblasts Realign in Response to Mechanical Strain

Integrins specifically recognize and bind extracellular matrix components, providing physical anchor points and functional setpoints. Focal adhesion complexes, containing integrin and cytoskeletal proteins, are potential mechanoreceptors, poised to distribute applied forces through the cytoskeleton. Pursuing the hypothesis that cells both perceive and respond to external force, we applied a stretch/relaxation regimen to normal human fetal and aged dermal fibroblast monolayers cultured on flexible membranes. The frequency and magnitude of the applied force is precisely controlled by the Flexercell Unit(Trademark). A protocol of stretch (20% elongation of the monolayer) at a frequency of 6 cycles/min caused a progressive change from a randomly distributed pattern of cells to a symmetric, radial distribution with cells aligned parallel to the applied force. We have coined the term 'orienteering' as the process of active alignment of cells in response to applied force. Cytochalasin D was added in graded doses to investigate the role of the actin cytoskeleton in force perception and transmission. A clear dose response was found; at high concentrations orienteering was abolished; and the drug's impact was reversible. The two cell strains used were similar in their alignment behavior and in their responses to cytochalasin D. Orienteering was influenced by cell density, and the cell strains studied differed in this respect. Fetal cells, unlike their aged counterparts, failed to orient at high cell density. In both cell strains, mid-density cultures aligned rapidly and sparse cultures lagged. These results indicate that both cell-cell adhesion and cytoskeleton integrity are critical in mediating the orienteering response. Differences between these two cell strains may relate to their expression of extracellular matrix molecules (fibronectin, collagen type 1) integrins and their relative binding affinities.

Sawyer, Christine↗

Analysis of Wave and Particle Signatures Observed in Plasma Escape at Venus

Atmospheric gases escape from Venus as neutral and ionized atoms and molecules. Ion escape, considered here, occurs through ion pickup or collective plasma processes. The latter can arise from upward flow of nightside ionospheric plasma into the ionotail, day to night ionospheric flow into the ionotail, and scavenging of ionospheric plasma by ionosphere-magnetosheath instabilities at the ionopause. These plasma processes produce differing signatures in ion velocity and energy distributions and in ULF waves in the magnetic field. Using plasma ion spectra measured by the Pioneer Venus Orbiter (PVO) Orbiter Plasma Analyzer (OPA) and magnetic field fluctuations observed by the PVO Orbiter Magnetometer (OMAG) along with the expected particle and field signatures, various ion escape processes occurring along Pioneer Venus orbits are identified. In particular, OPA ion energy distributions are used in parallel with magnetic field power spectra and wave phase angles derived from OMAG measurements to study the characteristics of escaping ions. The principle ions observed escaping the influence of Venus are H+, He+ and 0'. In the ion energy distributions of the OPA, pickup ions appear hot relative to the much cooler ions flowing away from Venus in the ionotail and in the plasma clouds detached from the ionopause. This energy contrast is particularly evident downstream when PVO crosses the ionotail boundary from the hot solar wind plasma to the much cooler plasma within the tail. Magnetic field signatures accompanying the escaping ions appear as peaks in the power spectra at the corresponding ion cyclotron frequencies. Also, coherent wave trains at the same frequencies are observed in the phase angle plots of magnetic field fluctuations about the mean field.

Hartle, R. E.↗

LDCM Grid Prototype (LGP)

The LGP successfully demonstrated that grid technology could be used to create a collaboration among research scientists, their science development machines, and distributed data to create a science production system in a nationally distributed environment. Grid technology provides a low cost and effective method of enabling production of science products by the science community. To demonstrate this, the LGP partnered with NASA GSFC scientists and used their existing science algorithms to generate virtual Landsat-like data products using distributed data resources. LGP created 48 output composite scenes with 4 input scenes each for a total of 192 scienes processed in parallel. The demonstration took 12 hours, which beat the requirement by almost 50 percent, well within the LDCM requirement to process 250 scenes per day. The LGP project also showed the successful use of workflow tools to automate the processing. Investing in this technology has led to funding for a ROSES ACCESS proposal. The proposal intends to enable an expert science user to produce products from a number of similar distributed instrument data sets using the Land Cover Change Community-based Processing and Analysis System (LC-ComPS) Toolbox. The LC-ComPS Toolbox is a collection of science algorithms that enable the generation of data with ground resolution on the order of Landsat-class instruments.

Weinstein, Beth↗

Distributed parameter modeling of the structural dynamics of the Solar Array Flight Experiment

A distributed-parameter model of the structural dynamics of the space-shuttle-deployed Solar Array Flight Experiment is developed and used to produce estimates of the modal frequencies and mode shapes. A lumped parameter version of the distributed model is used to estimate model characteristics by analyzing the measured responses of 32 targets. To make the modeling more tenable, a distributed parameter system is used to reduce the number of unknown parameters, a modified Newton-Raphson technique is used for rapid convergence, and a parallel processing supercomputer is used for more efficient computation. The performances of computers with a high-speed serial processor and with a high-speed parallel processor are compared. The best results are obtained with the modeling approach in which maximum likelihood estimation is applied to distributed parameter models.

Taylor, L. W., Jr.↗

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (distributed parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

graph algorithms, high performance comptuing↗

A comparison of queueing, cluster and distributed computing systems

Using workstation clusters for distributed computing has become popular with the proliferation of inexpensive, powerful workstations. Workstation clusters offer both a cost effective alternative to batch processing and an easy entry into parallel computing. However, a number of workstations on a network does not constitute a cluster. Cluster management software is necessary to harness the collective computing power. A variety of cluster management and queuing systems are compared: Distributed Queueing Systems (DQS), Condor, Load Leveler, Load Balancer, Load Sharing Facility (LSF - formerly Utopia), Distributed Job Manager (DJM), Computing in Distributed Networked Environments (CODINE), and NQS/Exec. The systems differ in their design philosophy and implementation. Based on published reports on the different systems and conversations with the system's developers and vendors, a comparison of the systems are made on the integral issues of clustered computing.

Kaplan, Joseph A.↗

Comparison of observed and calculated implanted ion distributions outside Comet Halley's bow shock

This paper compares calculated and measured energy spectra of implanted H(+) and O(+) ions on the assumption that the pickup geometry is quasi-parallel and about 1 percent of the waves generated by the cometary pickup process propagate backward (toward the comet). The model provides a good description of the implanted O(+) and H(+) energy distribution near the pickup energies. The thickness of the implanted ion velocity distribution shells was nearly constant between 2.5 and 1.2 million km (just outside the shock) along the inbound Giotto trajectory. The explanation is that the velocity diffusion coefficient and characteristic diffusion time vary approximately as 1/r and r, respectively, and therefore their product (which determines the velocity shell thickness) remains nearly constant.

Gombosi, T. I.↗

Technology and future ground processing systems

Land-observing satellites with multiple thematic mappers will produce data at rates of 100 to 300 Mbps. When coupled with a high daily scene production rate, these rates will require new approaches to ground processing. Consideration is given here to future downlink rates and data volumes, and requirements peculiar to the future user community are discussed. The advanced technologies required to attain an operational system in the years 1985-1990 are considered, together with advances foreseen in communications, mass storage, bulk memories, and data processing. Using advanced devices, a centralized data processing system capable of handling the 100 Mbps data rate is described. New approaches, among them a parallel pipelined calibration front-end, real-time browse image production, a high bandwidth optical disk archive, regional image broadcast and massively parallel product production, are considered. A distributed system capable of handling the 300 Mbps data rate is then described. Designs for a hub system and a regional processing center are presented.

Wood, B. J.↗

Massive parallelism in the future of science

Massive parallelism appears in three domains of action of concern to scientists, where it produces collective action that is not possible from any individual agent's behavior. In the domain of data parallelism, computers comprising very large numbers of processing agents, one for each data item in the result will be designed. These agents collectively can solve problems thousands of times faster than current supercomputers. In the domain of distributed parallelism, computations comprising large numbers of resource attached to the world network will be designed. The network will support computations far beyond the power of any one machine. In the domain of people parallelism collaborations among large groups of scientists around the world who participate in projects that endure well past the sojourns of individuals within them will be designed. Computing and telecommunications technology will support the large, long projects that will characterize big science by the turn of the century. Scientists must become masters in these three domains during the coming decade.

Denning, Peter J.↗

Distributed memory, GPU accelerated Fock construction for hybrid, Gaussian basis density functional theory

With the growing reliance of modern supercomputers on accelerator-based architecture such a graphics processing units (GPUs), the development and optimization of electronic structure methods to exploit these massively parallel resources has become a recent priority. While significant strides have been made in the development GPU accelerated, distributed memory algorithms for many modern electronic structure methods, the primary focus of GPU development for Gaussian basis atomic orbital methods has been for shared memory systems with only a handful of examples pursing massive parallelism. Here in this work, we present a set of distributed memory algorithms for the evaluation of the Coulomb and exact exchange matrices for hybrid Kohn–Sham DFT with Gaussian basis sets via direct density-fitted (DF-J-Engine) and seminumerical (sn-K) methods, respectively. The absolute performance and strong scalability of the developed methods are demonstrated on systems ranging from a few hundred to over one thousand atoms using up to 128 NVIDIA A100 GPUs on the Perlmutter supercomputer.

97 MATHEMATICS AND COMPUTING↗

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (Distributed Parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve an optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively, the performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

Sattar, Naw Safrin↗

What Multilevel Parallel Programs do when you are not Watching: A Performance Analysis Case Study Comparing MPI/OpenMP, MLP, and Nested OpenMP

With the current trend in parallel computer architectures towards clusters of shared memory symmetric multi-processors, parallel programming techniques have evolved that support parallelism beyond a single level. When comparing the performance of applications based on different programming paradigms, it is important to differentiate between the influence of the programming model itself and other factors, such as implementation specific behavior of the operating system (OS) or architectural issues. Rewriting-a large scientific application in order to employ a new programming paradigms is usually a time consuming and error prone task. Before embarking on such an endeavor it is important to determine that there is really a gain that would not be possible with the current implementation. A detailed performance analysis is crucial to clarify these issues. The multilevel programming paradigms considered in this study are hybrid MPI/OpenMP, MLP, and nested OpenMP. The hybrid MPI/OpenMP approach is based on using MPI [7] for the coarse grained parallelization and OpenMP [9] for fine grained loop level parallelism. The MPI programming paradigm assumes a private address space for each process. Data is transferred by explicitly exchanging messages via calls to the MPI library. This model was originally designed for distributed memory architectures but is also suitable for shared memory systems. The second paradigm under consideration is MLP which was developed by Taft. The approach is similar to MPi/OpenMP, using a mix of coarse grain process level parallelization and loop level OpenMP parallelization. As it is the case with MPI, a private address space is assumed for each process. The MLP approach was developed for ccNUMA architectures and explicitly takes advantage of the availability of shared memory. A shared memory arena which is accessible by all processes is required. Communication is done by reading from and writing to the shared memory.

Jost, Gabriele↗

Optimal mapping of irregular finite element domains to parallel processors

Mapping the solution domain of n-finite elements into N-subdomains that may be processed in parallel by N-processors is an optimal one if the subdomain decomposition results in a well-balanced workload distribution among the processors. The problem is discussed in the context of irregular finite element domains as an important aspect of the efficient utilization of the capabilities of emerging multiprocessor computers. Finding the optimal mapping is an intractable combinatorial optimization problem, for which a satisfactory approximate solution is obtained here by analogy to a method used in statistical mechanics for simulating the annealing process in solids. The simulated annealing analogy and algorithm are described, and numerical results are given for mapping an irregular two-dimensional finite element domain containing a singularity onto the Hypercube computer.

Flower, J.↗

Thermal stresses in chemically hardening elastic media with application to the molding process

A method has been formulated for the determination of thermal stresses in materials which harden in the presence of an exothermic chemical reaction. Hardening is described by the transformation of the material from an inviscid liquid-like state into an elastic solid, where intermediate states consist of a mixture of the two, in a ratio which is determined by the degree of chemical reaction. The method is illustrated in terms of an infinite slab cast between two rigid mold surfaces. It is found that the stress component normal to the slab surfaces vanishes in the residual state, so that removal of the slab from the mold leaves the remaining residual stress unchanged. On the other hand, the residual stress component parallel to the slab surfaces does not vanish. Its distribution is described as a function of the parameters of the hardening process.

Levitsky, M.↗

Electron Bulk Acceleration and Thermalization at Earth's Quasiperpendicular Bow Shock

Electron heating at Earth's quasiperpendicular bow shock has been surmised to be due to the combined effects of a quasistatic electric potential and scattering through wave-particle interaction. Here we report the observation of electron distribution functions indicating a new electron heating process occurring at the leading edge of the shock front. Incident solar wind electrons are accelerated parallel to the magnetic field toward downstream, reaching an electron-ion relative drift speed exceeding the electron thermal speed. The bulk acceleration is associated with an electric field pulse embedded in a whistler-mode wave. The high electron-ion relative drift is relaxed primarily through a nonlinear current-driven instability. The relaxed distributions contain a beam traveling toward the shock as a remnant of the accelerated electrons. Similar distribution functions prevail throughout the shock transition layer, suggesting that the observedacceleration and thermalization is essential to the cross-shock electron heating.

Chen, L.-J.↗