Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data movement”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Dynamic Load Balancing for Grid Partitioning on a SP-2 Multiprocessor: A Framework

Computational requirements of full scale computational fluid dynamics change as computation progresses on a parallel machine. The change in computational intensity causes workload imbalance of processors, which in turn requires a large amount of data movement at runtime. If parallel CFD is to be successful on a parallel or massively parallel machine, balancing of the runtime load is indispensable. Here a framework is presented for dynamic load balancing for CFD applications, called Jove. One processor is designated as a decision maker Jove while others are assigned to computational fluid dynamics. Processors running CFD send flags to Jove in a predetermined number of iterations to initiate load balancing. Jove starts working on load balancing while other processors continue working with the current data and load distribution. Jove goes through several steps to decide if the new data should be taken, including preliminary evaluate, partition, processor reassignment, cost evaluation, and decision. Jove running on a single EBM SP2 node has been completely implemented. Preliminary experimental results show that the Jove approach to dynamic load balancing can be effective for full scale grid partitioning on the target machine IBM SP2.

Sohn, Andrew↗

Dynamic Load Balancing for Finite Element Calculations on Parallel Computers

Computational requirements of full scale computational fluid dynamics change as computation progresses on a parallel machine. The change in computational intensity causes workload imbalance of processors, which in turn requires a large amount of data movement at runtime. If parallel CFD is to be successful on a parallel or massively parallel machine, balancing of the runtime load is indispensable. Here a frame work is presented for dynamic load balancing for CFD applications, called Jove. One processor is designated as a decision maker Jove while others are assigned to computational fluid dynamics. Processors running CFD send flags to Jove in a predetermined number of iterations to initiate load balancing. Jove starts working on load balancing while other processors continue working with the current data and load distribution. Jove goes through several steps to decide if the new data should be taken, including preliminary evaluate, partition, processor reassignment, cost evaluation, and decision. Jove running on a single SP2 node has been completely implemented. Preliminary experimental results show that the Jove approach to dynamic load balancing can be effective for full scale grid partitioning on the target machine SP2.

Pramono, Eddy↗

Simulations of Recovery of Time-Varying Gravity from DECIGO Pathfinder

We simulated time-varying Earth's gravity field recovered from DPF to evaluate an impact of DPF and future satellite gradiometry mission on earth science. From hydrological water movement data and orbit information, gravity gradients to be measured at altitude about ~500km were generated. Errors caused by atmospheric and oceanic variations and instrumental noise were added. Monthly gravity fields were estimated solving normal equations between spherical harmonic coefficients and simulated gravity gradient data. Simulation results show that DPF likely provides monthly hydrological water storage change with spatial scale between 400 and 1000km. Sensitivities to large scale estimates depends on long-term stability of gravity gradient measurement, and errors in short scale estimates are caused by instrumental noise and imperfections in atmospheric and ocean model. With acceleration noise level is lower than ~5 x 10(exp -14) [m/s2/sqrtHz] at frequency higher than 3mHz, water storage changes at limited small basins will be provided by DPF. To monitor continental scale hydrological water movement, noise level must be lower than ~5 x 10(exp -14) [m/s2/sqrtHz] at frequency higher than 1mHz.

Hasegawa, Takashi↗

Identification of Fixations in Noisy Eye Movements via Recursive Subdivision

When solving problems, multi-person airline crews can choose whether to work together, or to address different aspects of a situation with a divide and conquer strategy. Knowing which of these strategies is most effective may help airlines develop better procedures and training. This paper concentrates on joint attention as a measure of crew coordination. We report results obtained by applying cross recurrence analysis to eye movement data from two-person crews, collected in a flight simulator experiment. The analysis shows that crews exhibit coordinated gaze roughly one sixth of the time, with a tendency for the captain to lead the first officers visual attention. The degree to which crews coordinate their gaze is not significantly correlated with performance ratings assigned by instructors; further research questions and approaches are discussed.

signal processing↗

A stochastic model for eye movements during fixation on a stationary target.

A stochastic model describing small eye movements occurring during steady fixation on a stationary target is presented. Based on eye movement data for steady gaze, the model has a hierarchical structure; the principal level represents the random motion of the image point within a local area of fixation, while the higher level mimics the jump processes involved in transitions from one local area to another. Target image motion within a local area is described by a Langevin-like stochastic differential equation taking into consideration the microsaccadic jumps pictured as being due to point processes and the high frequency muscle tremor, represented as a white noise. The transform of the probability density function for local area motion is obtained, leading to explicit expressions for their means and moments. Evaluation of these moments based on the model is comparable with experimental results.

Vasudevan, R.↗

The application of space technology to practical problems such as those currently facing the mountain sections of the State of Colorado

Rapid growth in small Colorado mountain communities and dangers posed by development in areas that are potentially dangerous to life and property due to natural processes are studied. Special attention was given to snow avalanche, mudflow, rockfall, landslide and flood, as well as the slow continuous and frequently imperceptible form of soil creep and associated mass movement. Data are also given on the relative reliability of ERTS and Skylab imagery and conventional photography in identifying avalanche paths and run out zones.

Ives, J. D.↗

Lineaments in basement terrane of the Peninsular Ranges, Southern California

The author has identified the following significant results. ERTS and Skylab images reveal a number of prominent lineaments in the basement terrane of the Peninsular Ranges, Southern California. The major, well-known, active, northwest trending, right-slip faults are well displayed; northeast and west to west-northwest trending lineaments are also present. Study of large-scale airphotos followed by field investigations have shown that several of these lineaments represent previously unmapped faults. Pitches of striations on shear surfaces of the northeast and west trending faults indicate oblique slip movement; data are insufficient to determine the net-slip. These faults are restricted to the pre-tertiary basement terrane and are truncated by the major northwest trending faults. They may have been formed in response to an earlier stress system. All lineaments observed in the space photography are not due to faulting, and additional detailed geologic investigations are required to determine the nature of the unstudied lineaments, and the history and net-slip of fault-controlled lineaments.

Merifield, P. M.↗

Fault tectonics and earthquake hazards in the Peninsular Ranges, Southern California

The author has identified the following significant results. ERTS and Skylab images reveal a number of prominent lineaments in the basement terrane of the Peninsular Ranges, Southern California. The major, well-known, active, northwest trending, right-slip faults are well displayed, but northeast and west to west-northwest trending lineaments are also present. Study of large-scale airphotos followed by field investigations have shown that several of these lineaments represent previously unmapped faults. Pitches of striations on shear surfaces of the northeast and west trending faults indicate oblique-slip movement; data are insufficient to determine the net-slip. These faults are restricted to the pre-Tertiary basement terrane and are truncated by the major northwest trending faults; therefore, they may have formed in response to an earlier stress system. Future work should be directed toward determining whether the northeast and west trending faults are related to the presently active stress system or to an older inactive system, because this question relates to the earthquake risk in the vicinity of these faults.

Merifield, P. M.↗

On the impact of communication complexity in the design of parallel numerical algorithms

This paper describes two models of the cost of data movement in parallel numerical algorithms. One model is a generalization of an approach due to Hockney, and is suitable for shared memory multiprocessors where each processor has vector capabilities. The other model is applicable to highly parallel nonshared memory MIMD systems. In the second model, algorithm performance is characterized in terms of the communication network design. Techniques used in VLSI complexity theory are also brought in, and algorithm independent upper bounds on system performance are derived for several problems that are important to scientific computation.

Gannon, D.↗

On the impact of communication complexity on the design of parallel numerical algorithms

This paper describes two models of the cost of data movement in parallel numerical alorithms. One model is a generalization of an approach due to Hockney, and is suitable for shared memory multiprocessors where each processor has vector capabilities. The other model is applicable to highly parallel nonshared memory MIMD systems. In this second model, algorithm performance is characterized in terms of the communication network design. Techniques used in VLSI complexity theory are also brought in, and algorithm-independent upper bounds on system performance are derived for several problems that are important to scientific computation.

Gannon, D. B.↗

Statistical dependency in visual scanning

A method to identify statistical dependencies in the positions of eye fixations is developed and applied to eye movement data from subjects who viewed dynamic displays of air traffic and judged future relative position of aircraft. Analysis of approximately 23,000 fixations on points of interest on the display identified statistical dependencies in scanning that were independent of the physical placement of the points of interest. Identification of these dependencies is inconsistent with random-sampling-based theories used to model visual search and information seeking.

Ellis, Stephen R.↗

Compiling global name-space programs for distributed execution

Distributed memory machines do not provide hardware support for a global address space. Thus programmers are forced to partition the data across the memories of the architecture and use explicit message passing to communicate data between processors. The compiler support required to allow programmers to express their algorithms using a global name-space is examined. A general method is presented for analysis of a high level source program and along with its translation to a set of independently executing tasks communicating via messages. If the compiler has enough information, this translation can be carried out at compile-time. Otherwise run-time code is generated to implement the required data movement. The analysis required in both situations is described and the performance of the generated code on the Intel iPSC/2 is presented.

Koelbel, Charles↗

The NAS parallel benchmarks

A new set of benchmarks has been developed for the performance evaluation of highly parallel supercomputers in the framework of the NASA Ames Numerical Aerodynamic Simulation (NAS) Program. These consist of five 'parallel kernel' benchmarks and three 'simulated application' benchmarks. Together they mimic the computation and data movement characteristics of large-scale computational fluid dynamics applications. The principal distinguishing feature of these benchmarks is their 'pencil and paper' specification-all details of these benchmarks are specified only algorithmically. In this way many of the difficulties associated with conventional benchmarking approaches on highly parallel systems are avoided.

Bailey, D. H.↗

Compiling global name-space parallel loops for distributed execution

Distributed memory machines do not provide hardware support for a global address space. Thus programmers are forced to partition the data across the memories of the architecture and use explicit message passing to communicate data between processors. The compiler support required to allow programmers to express their algorithms using a global name-space is examined. A general method is presented for analysis of a high level source program and its translation into a set of independently executing tasks communicating via messages. If the compiler has enough information, this translation can be carried out at compile time. Otherwise, run-time code is generated to implement the required data movement. The analysis required in both situations is described and the performance of the generated code on the Intel iPSC/2 is presented.

Koelbel, Charles↗

Scalable parallel communications

Coarse-grain parallelism in networking (that is, the use of multiple protocol processors running replicated software sending over several physical channels) can be used to provide gigabit communications for a single application. Since parallel network performance is highly dependent on real issues such as hardware properties (e.g., memory speeds and cache hit rates), operating system overhead (e.g., interrupt handling), and protocol performance (e.g., effect of timeouts), we have performed detailed simulations studies of both a bus-based multiprocessor workstation node (based on the Sun Galaxy MP multiprocessor) and a distributed-memory parallel computer node (based on the Touchstone DELTA) to evaluate the behavior of coarse-grain parallelism. Our results indicate: (1) coarse-grain parallelism can deliver multiple 100 Mbps with currently available hardware platforms and existing networking protocols (such as Transmission Control Protocol/Internet Protocol (TCP/IP) and parallel Fiber Distributed Data Interface (FDDI) rings); (2) scale-up is near linear in n, the number of protocol processors, and channels (for small n and up to a few hundred Mbps); and (3) since these results are based on existing hardware without specialized devices (except perhaps for some simple modifications of the FDDI boards), this is a low cost solution to providing multiple 100 Mbps on current machines. In addition, from both the performance analysis and the properties of these architectures, we conclude: (1) multiple processors providing identical services and the use of space division multiplexing for the physical channels can provide better reliability than monolithic approaches (it also provides graceful degradation and low-cost load balancing); (2) coarse-grain parallelism supports running several transport protocols in parallel to provide different types of service (for example, one TCP handles small messages for many users, other TCP's running in parallel provide high bandwidth service to a single application); and (3) coarse grain parallelism will be able to incorporate many future improvements from related work (e.g., reduced data movement, fast TCP, fine-grain parallelism) also with near linear speed-ups.

Maly, K.↗

The NAS parallel benchmarks

A new set of benchmarks was developed for the performance evaluation of highly parallel supercomputers. These benchmarks consist of a set of kernels, the 'Parallel Kernels,' and a simulated application benchmark. Together they mimic the computation and data movement characteristics of large scale computational fluid dynamics (CFD) applications. The principal distinguishing feature of these benchmarks is their 'pencil and paper' specification - all details of these benchmarks are specified only algorithmically. In this way many of the difficulties associated with conventional benchmarking approaches on highly parallel systems are avoided.

Bailey, David↗

Impact of Load Balancing on Unstructured Adaptive Grid Computations for Distributed-Memory Multiprocessors

The computational requirements for an adaptive solution of unsteady problems change as the simulation progresses. This causes workload imbalance among processors on a parallel machine which, in turn, requires significant data movement at runtime. We present a new dynamic load-balancing framework, called JOVE, that balances the workload across all processors with a global view. Whenever the computational mesh is adapted, JOVE is activated to eliminate the load imbalance. JOVE has been implemented on an IBM SP2 distributed-memory machine in MPI for portability. Experimental results for two model meshes demonstrate that mesh adaption with load balancing gives more than a sixfold improvement over one without load balancing. We also show that JOVE gives a 24-fold speedup on 64 processors compared to sequential execution.

Biswas, Rupak↗