Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

I/O Bottleneck Detection and Tuning: Connecting the Dots using Interactive Log Analysis

Using parallel file systems efficiently is a tricky problem due to inter-dependencies among multiple layers of I/O software, including high-level I/O libraries (HDF5, netCDF, etc.), MPI-IO, POSIX, and file systems (GPFS, Lustre, etc.). Profiling tools such as Darshan collect traces to help understand the I/O performance behavior. However, there are significant gaps in analyzing the collected traces and then applying tuning options offered by various layers of I/O software. Seeking to connect the dots between I/O bottleneck detection and tuning, we propose DXT Explorer, an interactive log analysis tool. In this paper, we present a case study using our interactive log analysis tool to identify and apply various I/O optimizations. We report an evaluation of performance improvement achieved for four I/O kernels extracted from science applications.

Bez, Jean Luca↗

Continental rift evolution and drainage reorganization along the Dead Sea rift since the Miocene

The Dead Sea fault is a section of the Arabian-African plate boundary. Widespread field relations indicate that three major drainage systems (stages) occupied the landscape west of the Dead Sea fault since its initiation at ca. 20 Ma. Specifically, (1) an early to middle Miocene drainage system, only minorly reconfigured by the fault (all sediments of this system belong to the Hazeva Formation); (2) a late Miocene to early Pleistocene fault-parallel drainage system named Paran-Neqarot (all sediments of this system belong to the Arava and Zehiha Formations; and (3) the early Pleistocene to present drainage configuration. The temporal and spatial frameworks of drainage stage 1 are generally constrained by radiometric dating of interfingering volcanic units, and the onset and temporal and spatial frameworks of drainage stage 3 are well constrained by cosmogenic 10 Be surface exposure ages. The timing and longevity of the rift-parallel drainage (stage 2) have until now been evasive to direct dating. The overall time gap between stages 1 and 3 is ~12–13 million years. Thus, an early age (within this time gap) of stage 2 would imply an immediate response of drainage reorganization to rift tectonics, while a later age of this drainage system would imply a delayed response. We present 11 10 Be- 26 Al cosmogenic burial ages of alluvial and colluvial units related to the fault-parallel drainage system (stage 2), which collectively constrain the time of deposition of the Arava Formation sediments in the central Negev to ca. 8 Ma. The general lack of stratigraphic order, together with the large dispersion of ages both across and within the sampling sites, attests to significant recycling of sediments from drainage stage 1 into the Arava Formation deposits. The termination of Arava and Zehiha Formation sediment deposition at ca. 1.8 Ma was determined previously using cosmogenic exposure ages of desert pavements that cover the formations. Combining the previously published data with our new data, we established the longevity and character of the Paran-Neqarot drainage system. In conlcusion, this framework highlights the temporal aspect of drainage system build-up and collapse as it responded to transform and extensional plate boundary tectonics during the Neogene.

58 GEOSCIENCES↗

Parallel Architectures and Parallel Algorithms for Integrated Vision Systems

Computer vision is regarded as one of the most complex and computationally intensive problems. An integrated vision system (IVS) is a system that uses vision algorithms from all levels of processing to perform for a high level application (e.g., object recognition). An IVS normally involves algorithms from low level, intermediate level, and high level vision. Designing parallel architectures for vision systems is of tremendous interest to researchers. Several issues are addressed in parallel architectures and parallel algorithms for integrated vision systems.

Choudhary, Alok Nidhi↗

Large Scale Caching and Streaming of Training Data for Online Deep Learning

The training of deep neural network models on large data remains a difficult problem, despite progress towards scalable techniques. In particular, there is a mismatch between the random but predetermined order in which AI flows select training samples and the streaming I/O patterns for which traditional HPC data storage (e.g., parallel file systems) are designed. In addition, as more data are obtained, it is feasible neither simply to train learning models incrementally, due to catastrophic forgetting (i.e., bias towards new samples), nor to train frequently from scratch, due to prohibitive time and/or resource constraints. In this paper, we study data management techniques that combine caching and streaming with rehearsal support in order to enable efficient access to training samples in both offline training and continual learning. We revisit state-of-art streaming approaches based on data pipelines that transparently handle prefetching, caching, shuffling, and data augmentation, and discuss the challenges and opportunities that arise when combining these methods with data-parallel training techniques. We also report on preliminary experiments that evaluate the I/O overheads involved in accessing the training samples from a parallel file system (PFS) under several concurrency scenarios, highlighting the impact of the PFS on the design of the data pipelines.

data pipelines↗

Proceedings of the Peteflops-Systems Operation Working Review (POWR)

This report constitutes the final technical report. Even as Petaflops performance computing is being achieved for a few important applications on the nation's largest massively parallel processing (MPP) systems, the challenges of realizing far greater performance are being investigated by a strong interdisciplinary team of experts from academia, industry, and government. This team is also exploring the extraordinary opportunity such capabilities would provide for critical areas of strategic national importance. Under what has become known informally as the "Petaflops Initiative," key leaders in research across the areas of device technology, parallel systems architecture, applications and algorithms, and systems software have been exploring the implications, requirements, interrelationships, and trade-offs among these research areas through a series of workshops, studies, and projects sponsored by a number of Federal agencies.

Sterling, Thomas↗

Analytical Assessment of Simultaneous Parallel Approach Feasibility from Total System Error

In a simultaneous paired approach to closely-spaced parallel runways, a pair of aircraft flies in close proximity on parallel approach paths. The aircraft pair must maintain a longitudinal separation within a range that avoids wake encounters and, if one of the aircraft blunders, avoids collision. Wake avoidance defines the rear gate of the longitudinal separation. The lead aircraft generates a wake vortex that, with the aid of crosswinds, can travel laterally onto the path of the trail aircraft. As runway separation decreases, the wake has less distance to traverse to reach the path of the trail aircraft. The total system error of each aircraft further reduces this distance. The total system error is often modeled as a probability distribution function. Therefore, Monte-Carlo simulations are a favored tool for assessing a "safe" rear-gate. However, safety for paired approaches typically requires that a catastrophic wake encounter be a rare one-in-a-billion event during normal operation. Using a Monte-Carlo simulation to assert this event rarity with confidence requires a massive number of runs. Such large runs do not lend themselves to rapid turn-around during the early stages of investigation when the goal is to eliminate the infeasible regions of the solution space and to perform trades among the independent variables in the operational concept. One can employ statistical analysis using simplified models more efficiently to narrow the solution space and identify promising trades for more in-depth investigation using Monte-Carlo simulations. These simple, analytical models not only have to address the uncertainty of the total system error but also the uncertainty in navigation sources used to alert an abort of the procedure. This paper presents a method for integrating total system error, procedure abort rates, avionics failures, and surveillance errors into a statistical analysis that identifies the likely feasible runway separations for simultaneous paired approaches.

Madden, Michael M.↗

SIAM Conference on Parallel Processing for Scientific Computing, 4th, Chicago, IL, Dec. 11-13, 1989, Proceedings

Attention is given to such topics as an evaluation of block algorithm variants in LAPACK and presents a large-grain parallel sparse system solver, a multiprocessor method for the solution of the generalized Eigenvalue problem on an interval, and a parallel QR algorithm for iterative subspace methods on the CM2. A discussion of numerical methods includes the topics of asynchronous numerical solutions of PDEs on parallel computers, parallel homotopy curve tracking on a hypercube, and solving Navier-Stokes equations on the Cedar Multi-Cluster system. A section on differential equations includes a discussion of a six-color procedure for the parallel solution of elliptic systems using the finite quadtree structure, data parallel algorithms for the finite element method, and domain decomposition methods in aerodynamics. Topics dealing with massively parallel computing include hypercube vs. 2-dimensional meshes and massively parallel computation of conservation laws. Performance and tools are also discussed.

Dongarra, Jack↗

Containers for Massive Ensemble of I/O Bound Hierarchical Coupled Simulations

We present our experience using containers to scale up a massive ensemble of coupled I/O bound workloads on the NERSC Cori supercomputer. We describe the design of a hierarchical simulation structure using the Integrated Plasma Simulator (IPS) that enables the flexible execution of coupled simulations at the system, node, and core level using the same coupling abstraction and API. The hierarchical design allows for the node-level execution to be efficiently executed using containers while not impacting the structure of the simulation at the system level. We demonstrate the viability of the approach by presenting experimental results from applications in coupled fusion plasma simulations that illustrate the performance impact of using containers to deploy the node-level workloads, in conjunction with the user mountable XFS file systems to ameliorate the load on the Lustre parallel file system. We also present results from production runs showing the ability of the ensemble simulations to scale to hundreds of Cori Haswell nodes, with little or no overhead.

Elwasif, Wael↗

The relationship of extensional and compressional tectonics to a Precambrian fracture system in the eastern overthrust belt, USA

The central and southern Appalachians have a long history of interrelated extensional and compressional tectonics. It is proposed that each episode was controlled by a reactivation of a fracture system in the Precambrian basement. Proprietary seismic-reflection profiles show a system of down-to-the-east Precambrian extensional faults. When under renewed extension, these faults produce features such as the western border faults of Mesozoic basins, and when under compression, probably produce tectonic ramps in the overlying sedimentary cover rocks as well as the spatially and genetically related Alleghenian folds. This system, which parallels the Appalachian trend, is cut by a system of cross-strike hinge or scissors faults that have probable strike-slip movements. Reactivation of this cross-strike system appears to have produced lateral ramps that connect decollements at different stratigraphic levels and caused abrupt changes in fold wavelength along strikes. Continued reactivation of this cross-strike system is suggested by east-west border faults and Precambrian highs between Mesozoic basins. The present activity of this system is suggested by the fact that more than 35% of recent earthquakes are coincident with cross-strike faults and lateral ramps.

Pohn, H.↗

Phase space simulation of collisionless stellar systems on the massively parallel processor

A numerical technique for solving the collisionless Boltzmann equation describing the time evolution of a self gravitating fluid in phase space was implemented on the Massively Parallel Processor (MPP). The code performs calculations for a two dimensional phase space grid (with one space and one velocity dimension). Some results from calculations are presented. The execution speed of the code is comparable to the speed of a single processor of a Cray-XMP. Advantages and disadvantages of the MPP architecture for this type of problem are discussed. The nearest neighbor connectivity of the MPP array does not pose a significant obstacle. Future MPP-like machines should have much more local memory and easier access to staging memory and disks in order to be effective for this type of problem.

White, Richard L.↗

Parallels between control PDE's (Partial Differential Equations) and systems of ODE's (Ordinary Differential Equations)

System theorists understand that the same mathematical objects which determine controllability for nonlinear control systems of ordinary differential equations (ODEs) also determine hypoellipticity for linear partial differentail equations (PDEs). Moreover, almost any study of ODE systems begins with linear systems. It is remarkable that Hormander's paper on hypoellipticity of second order linear p.d.e.'s starts with equations due to Kolmogorov, which are shown to be analogous to the linear PDEs. Eigenvalue placement by state feedback for a controllable linear system can be paralleled for a Kolmogorov equation if an appropriate type of feedback is introduced. Results concerning transformations of nonlinear systems to linear systems are similar to results for transforming a linear PDE to a Kolmogorov equation.

Hunt, L. R.↗

Parallel solution of pentadiagonal systems using generalized odd-even elimination

A method for the solution of pentadiagonal systems of linear equations is presented. The method is a generalization of ordinary odd-even elimination used for tridiagonal systems. Using n processors, an n x n pentadiagonal system can be solved using the new method (generalized odd-even elimination) in time proportional to log(2) n.

Levit, Creon↗

Merlin - Massively parallel heterogeneous computing

Hardware and software for Merlin, a new kind of massively parallel computing system, are described. Eight computers are linked as a 300-MIPS prototype to develop system software for a larger Merlin network with 16 to 64 nodes, totaling 600 to 3000 MIPS. These working prototypes help refine a mapped reflective memory technique that offers a new, very general way of linking many types of computer to form supercomputers. Processors share data selectively and rapidly on a word-by-word basis. Fast firmware virtual circuits are reconfigured to match topological needs of individual application programs. Merlin's low-latency memory-sharing interfaces solve many problems in the design of high-performance computing systems. The Merlin prototypes are intended to run parallel programs for scientific applications and to determine hardware and software needs for a future Teraflops Merlin network.

Wittie, Larry↗

Airbreathing Propulsion System Analysis Using Multithreaded Parallel Processing

In this paper, parallel processing is used to analyze the mixing, and combustion behavior of hypersonic flow. Preliminary work for a sonic transverse hydrogen jet injected from a slot into a Mach 4 airstream in a two-dimensional duct combustor has been completed [Moon and Chung, 1996]. Our aim is to extend this work to three-dimensional domain using multithreaded domain decomposition parallel processing based on the flowfield-dependent variation theory. Numerical simulations of chemically reacting flows are difficult because of the strong interactions between the turbulent hydrodynamic and chemical processes. The algorithm must provide an accurate representation of the flowfield, since unphysical flowfield calculations will lead to the faulty loss or creation of species mass fraction, or even premature ignition, which in turn alters the flowfield information. Another difficulty arises from the disparity in time scales between the flowfield and chemical reactions, which may require the use of finite rate chemistry. The situations are more complex when there is a disparity in length scales involved in turbulence. In order to cope with these complicated physical phenomena, it is our plan to utilize the flowfield-dependent variation theory mentioned above, facilitated by large eddy simulation. Undoubtedly, the proposed computation requires the most sophisticated computational strategies. The multithreaded domain decomposition parallel processing will be necessary in order to reduce both computational time and storage. Without special treatments involved in computer engineering, our attempt to analyze the airbreathing combustion appears to be difficult, if not impossible.

Schunk, Richard Gregory↗