Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data movement”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Impact of Load Balancing on Unstructured Adaptive Grid Computations for Distributed-Memory Multiprocessors

The computational requirements for an adaptive solution of unsteady problems change as the simulation progresses. This causes workload imbalance among processors on a parallel machine which, in turn, requires significant data movement at runtime. We present a new dynamic load-balancing framework, called JOVE, that balances the workload across all processors with a global view. Whenever the computational mesh is adapted, JOVE is activated to eliminate the load imbalance. JOVE has been implemented on an IBM SP2 distributed-memory machine in MPI for portability. Experimental results for two model meshes demonstrate that mesh adaption with load balancing gives more than a sixfold improvement over one without load balancing. We also show that JOVE gives a 24-fold speedup on 64 processors compared to sequential execution.

Sohn, Andrew↗

Testing New Programming Paradigms with NAS Parallel Benchmarks

Over the past decade, high performance computing has evolved rapidly, not only in hardware architectures but also with increasing complexity of real applications. Technologies have been developing to aim at scaling up to thousands of processors on both distributed and shared memory systems. Development of parallel programs on these computers is always a challenging task. Today, writing parallel programs with message passing (e.g. MPI) is the most popular way of achieving scalability and high performance. However, writing message passing programs is difficult and error prone. Recent years new effort has been made in defining new parallel programming paradigms. The best examples are: HPF (based on data parallelism) and OpenMP (based on shared memory parallelism). Both provide simple and clear extensions to sequential programs, thus greatly simplify the tedious tasks encountered in writing message passing programs. HPF is independent of memory hierarchy, however, due to the immaturity of compiler technology its performance is still questionable. Although use of parallel compiler directives is not new, OpenMP offers a portable solution in the shared-memory domain. Another important development involves the tremendous progress in the internet and its associated technology. Although still in its infancy, Java promisses portability in a heterogeneous environment and offers possibility to "compile once and run anywhere." In light of testing these new technologies, we implemented new parallel versions of the NAS Parallel Benchmarks (NPBs) with HPF and OpenMP directives, and extended the work with Java and Java-threads. The purpose of this study is to examine the effectiveness of alternative programming paradigms. NPBs consist of five kernels and three simulated applications that mimic the computation and data movement of large scale computational fluid dynamics (CFD) applications. We started with the serial version included in NPB2.3. Optimization of memory and cache usage was applied to several benchmarks, noticeably BT and SP, resulting in better sequential performance. In order to overcome the lack of an HPF performance model and guide the development of the HPF codes, we employed an empirical performance model for several primitives found in the benchmarks. We encountered a few limitations of HPF, such as lack of supporting the "REDISTRIBUTION" directive and no easy way to handle irregular computation. The parallelization with OpenMP directives was done at the outer-most loop level to achieve the largest granularity. The performance of six HPF and OpenMP benchmarks is compared with their MPI counterparts for the Class-A problem size in the figure in next page. These results were obtained on an SGI Origin2000 (195MHz) with MIPSpro-f77 compiler 7.2.1 for OpenMP and MPI codes and PGI pghpf-2.4.3 compiler with MPI interface for HPF programs.

Jin, H.↗

A Latency-Tolerant Partitioner for Distributed Computing on the Information Power Grid

NASA's Information Power Grid (IPG) is an infrastructure designed to harness the power of graphically distributed computers, databases, and human expertise, in order to solve large-scale realistic computational problems. This type of a meta-computing environment is necessary to present a unified virtual machine to application developers that hides the intricacies of a highly heterogeneous environment and yet maintains adequate security. In this paper, we present a novel partitioning scheme. called MinEX, that dynamically balances processor workloads while minimizing data movement and runtime communication, for applications that are executed in a parallel distributed fashion on the IPG. We also analyze the conditions that are required for the IPG to be an effective tool for such distributed computations. Our results show that MinEX is a viable load balancer provided the nodes of the IPG are connected by a high-speed asynchronous interconnection network.

Das, Sajal K.↗

Latency Hiding in Dynamic Partitioning and Load Balancing of Grid Computing Applications

The Information Power Grid (IPG) concept developed by NASA is aimed to provide a metacomputing platform for large-scale distributed computations, by hiding the intricacies of highly heterogeneous environment and yet maintaining adequate security. In this paper, we propose a latency-tolerant partitioning scheme that dynamically balances processor workloads on the.IPG, and minimizes data movement and runtime communication. By simulating an unsteady adaptive mesh application on a wide area network, we study the performance of our load balancer under the Globus environment. The number of IPG nodes, the number of processors per node, and the interconnected speeds are parameterized to derive conditions under which the IPG would be suitable for parallel distributed processing of such applications. Experimental results demonstrate that effective solution are achieved when the IPG nodes are connected by a high-speed asynchronous interconnection network.

Das, Sajal K.↗

Flexible Wing Base Micro Aerial Vehicles: Assessment of Controllability of Micro Air Vehicles

In the last several years, we have developed unique types of micro air vehicles that utilize flexible structures and extensible covering materials. These MAVs can be operated with maximum dimensions as small as 6 inches and carry reasonable payloads, such as video cameras and transmitters. We recently demonstrated the potential of these vehicles by winning the Fourth International Micro Air Vehicle Competition, held at Ft. Huachucha, Arizona in May 2000. The pilots report that these vehicles have unusually smooth flying characteristics and are relatively easy to fly, both in the standard RC mode and "through the camera" when at greater distances. In comparison, they find that similar sized vehicles with more conventional rigid construction require much more input from the pilot just to maintain control. To make these subjective observations more quantitative, we have devised a system that can conveniently record a complete history of all the RC transmitter stick movements during a flight. Post-flight processing of the stick movement data allows for direct comparisons between different types of MAVs when flown by the same pilot, and also comparisons between pilots. Eventually, practical micro air vehicles will be autonomously controlled, but we feel that the smoothest flying and easiest to fly embodiments will also be the most successful in the long run. Comparisons between several types of micro air vehicles will be presented, along with interpretations of the data.

Jenkins, David A.↗

PLUM: Parallel Load Balancing for Adaptive Unstructured Meshes

Mesh adaption is a powerful tool for efficient unstructured-grid computations but causes load imbalance among processors on a parallel machine. We present a novel method called PLUM to dynamically balance the processor workloads with a global view. This paper presents the implementation and integration of all major components within our dynamic load balancing strategy for adaptive grid calculations. Mesh adaption, repartitioning, processor assignment, and remapping are critical components of the framework that must be accomplished rapidly and efficiently so as not to cause a significant overhead to the numerical simulation. A data redistribution model is also presented that predicts the remapping cost on the SP2. This model is required to determine whether the gain from a balanced workload distribution offsets the cost of data movement. Results presented in this paper demonstrate that PLUM is an effective dynamic load balancing strategy which remains viable on a large number of processors.

Oliker, Leonid↗

Parallel Load Balancing for Adaptive Unstructured Meshes

Mesh adaption is a powerful tool for efficient unstructured-grid computations but causes load imbalance among processors on a parallel machine. We describe a novel method to dynamically balance the processor workloads with a global view. Mesh question, repartitioning, processor assignment, and remapping are critical components of the framework that must be accomplished rapidly and efficiently so as not to cause a significant overhead to the numerical simulation. A data redistribution model will also be presented that predicts the remapping cost. This model is required to determine whether the gain from a balanced workload distribution offsets the cost of data movement. Results presented will demonstrate that this is an effective dynamic load balancing strategy which remains viable on a large number of processors.

Biswas, Rupak↗

Functional Testing and Evaluation of Actiwatch Spectrum Devices for Launch on STS-133/ULF5

The Actiwatch Spectrum (AWS) is a wrist-worn device that may be used for obtaining ground or on-orbit light exposure patterns and movement data. The objective of this project was to prepare AWS devices for launch on STS-133/ULF5 by a means of implementing functional tests and engineering evaluations. The data obtained from these tests and evaluations served as a means for detecting any plausible issues that the AWS may encounter while on-orbit. Subsequent steps after detecting anomalies with AWS devices encompassed identifying their root causes and taking the steps needed to mitigate them. As a result of this study, the overall success of sleep/wake research studies for STS-133/ULF5 and future missions will be enhanced.

Rollins, Selisa F.↗

High Data Rate Architecture (HiDRA)

One of the greatest challenges in developing new space technology is in navigating the transition from ground based laboratory demonstration at Technology Readiness Level 6 (TRL-6) to conducting a prototype demonstration in space (TRL-7). This challenge is com- pounded by the relatively low availability of new spacecraft missions when compared with aeronautical craft to bridge this gap, leading to the general adoption of a low-risk stance by mission management to accept new, unproven technologies into the system. Also in consideration of risk, the limited selection and availability of proven space-grade components imparts a severe limitation on achieving high performance systems by current terrestrial technology standards. Finally from a space communications point of view the long duration characteristic of most missions imparts a major constraint on the entire space and ground network architecture, since any new technologies introduced into the system would have to be compliant with the duration of the currently deployed operational technologies, and in some cases may be limited by surrounding legacy capabilities. Beyond ensuring that the new technology is verified to function correctly and validated to meet the needs of the end users the formidable challenge then grows to additionally include: carefully timing the maturity path of the new technology to coincide with a feasible and accepting future mission so it flies before its relevancy has passed, utilizing a limited catalog of available components to their maximum potential to create meaningful and unprecedented new capabilities, designing and ensuring interoperability with aging space and ground infrastructures while simultaneously providing a growth path to the future. The International Space Station (ISS) is approaching 20 years of age. To keep the ISS relevant, technology upgrades are continuously taking place. Regarding communications, the state-of-the-art communication system upgrades underway include high-rate laser terminals. These must interface with the existing, aging data infrastructure. The High Data Rate Architecture (HiDRA) project is designed to provide networked store, carry, and forward capability to optimize data flow through both the existing radio frequency (RF) and new laser communications terminal. The networking capability is realized through the Delay Tolerant Networking (DTN) protocol, and is used for scheduling data movement as well as optimizing the performance of existing RF channels. HiDRA is realized as a distributed FPGA memory and interface controller that is itself controlled by a local computer running DTN software. Thus HiDRA is applicable to other arenas seeking to employ next-generation communications technologies, e.g. deep space. In this paper, we describe HiDRA and its far-reaching research implications.

DTN↗

On the Development and Application of High Data Rate Architecture (HiDRA) in Future Space Networks

Historically, space missions have been severely constrained by their ability to downlink the data they have collected. These constraints are a result of relatively low link rates on the spacecraft as well as limitations on the time during which data can be sent. As part of a coherent strategy to address existing limitations and get more data to the ground more quickly, the Space Communications and Navigation (SCaN) program has been developing an architecture for a future solar system Internet. The High Data Rate Architecture (HiDRA) project is designed to fit into such a future SCaN network. HiDRA's goal is to describe a general packet-based networking capability which can be used to provide assets with efficient networking capabilities while simultaneously reducing the capital costs and operational costs of developing and flying future space systems.Along these lines, this paper begins by reviewing various characteristics of modern satellite design as well as relevant characteristics of emerging technologies (such as free-space optical links capable of working at 100+ Gbps). Next, the paper describes HiDRA's design, and how the system is able to both integrate and support the operation of not only today's high-rate systems, but also the high-rate systems likely to be found in the future. This section also explores both existing and future networking technologies, such as Delay Tolerant Networking (DTN) protocol (RFC4838 citeRFC:1, RFC5050citeRFC:2), and explains how HiDRA supports them. Additionally, this section explores how HiDRA is used for scheduling data movement through both proactive and reactive link management. After this, the paper moves on to explore a reference implementation of HiDRA. This implementation is currently being realized based on a Field Programmable Gate Array (FPGA) memory and interface controller that is itself controlled by a local computer running DTN software. Next, this paper explores HiDRA's natural evolution, which includes an integration path for software-defined networking (SDN) switches. This section also describes considerations for both near-Earth and deep-space instantiations of HiDRA, describing how differences in latencies between the environments will necessarily influence how the system is configured and the networks operate. Finally, this paper describes future work. This section includes a description of a potential ISS implementation which will allow rapid advancement through the technology readiness levels (TRL). This section also explores work being done to support HiDRA's successful implementation and operation in a heterogeneous network: such a network could include communications equipment spanning many vintages and capabilities, and one significant aspect of HiDRA's future development involves balancing compatibility with capability.

DTN↗

Cross Recurrence Analysis as a Measure of Pilots' Coordination Strategy

When solving problems, multi-person airline crews can choose whether to work together, or to address different aspects of a situation with a divide and conquer strategy. Knowing which of these strategies is most effective may help airlines develop better procedures and training. This paper concentrates on joint attention as a measure of crew coordination. We report results obtained by applying cross recurrence analysis to eye movement data from two-person crews, collected in a flight simulator experiment. The analysis shows that crews exhibit coordinated gaze roughly 1/6th of the time, with a tendency for the captain to lead the first officer’s visual attention. The degree to which crews coordinate their gaze is not significantly correlated with performance ratings assigned by instructors; further research questions and approaches are discussed.

cross recurrence analysis↗

Cross Recurrence Analysis as a Measure of Pilots' Coordination Strategy

When solving problems, multi-person airline crews can choose whether to work together, or to address different aspects of a situation with a divide and conquer strategy. Knowing which of these strategies is most effective may help airlines develop better procedures and training. This paper concentrates on joint attention as a measure of crew coordination. We report results obtained by applying cross recurrence analysis to eye movement data from two-person crews, collected in a flight simulator experiment. The analysis shows that crews exhibit coordinated gaze roughly 16th of the time, with a tendency for the captain to lead the first officers visual attention. The degree to which crews coordinate their gaze is not significantly correlated with performance ratings assigned by instructors; further research questions and approaches are discussed.

non-linear models↗

Enabling Execution of a Legacy CFD Mini Application on Accelerators Using OpenMP

We describe the process and outcome of our efforts to port a legacy Fortran benchmark code to heterogeneous GPU-accelerated computing architectures using OpenMP. The benchmark code is one of the multi-zone NAS Parallel Benchmarks (NPB-MZ) called SP-MZ. This “mini-app” mimics the computation and data movement that is found in popular legacy and modern implicit computational fluid dynamics (CFD)solvers. Our objective was to examine how efficiently legacy Fortran codes can be ported to accelerators by leveraging OpenMP directives. We describe the development and optimization process and demonstrate the performance impact of various code modifications. We show select profiling results from the Nvidia nvvp profiler to help others diagnose and overcome performance issues in their own applications. We present results for two compute systems endowed with Nvidia V100 accelerators.

Ioannis Nompelis↗

Towards Gbps Downlinks from Low-Cost Active Phased Arrays

Spacecraft performing science missions use high-directivity antennas to quickly downlink large amounts of collected data. Movement required to keep gimbaled reflector antennas trained on target receivers can impart blockages or vibrations which perturb precision science instruments. Recent manufacturing advances have made Ka-band electronically-steered antenna arrays a cost-effective option to mitigate these challenges, but only if these devices can be proven to deliver comparable quality of service. In this work we present a hardware and software architecture for high-rate links using an active phased array-based terminal in a 1U (10cm cube) form factor. Over-the-air tests in an antenna range characterize two key features of the architecture: high-order (32APSK and above) modulations and dual carriers. We calculate an optimal modulation-dependent input power backoff from array saturation and perform orbital simulations of the link at this operating point. We conclude that reliable links with rates in excess of 1Gbps are possible with current-generation hardware.

Adam Gannon↗

Comparison of Active and Passive Head Impulse Testing of the Horizontal Vestibulo-Ocular Reflex: Exploring the Feasibility of Different Approaches for Spaceflight

INTRODUCTION: Astronauts experience a wide variety of sensorimotor disturbances primarily due to microgravity-induced vestibular adaptations during spaceflight. Head Impulse Testing (HIT) will be conducted during the Complement of Integrated Protocols for Human Exploration Research program(CIPHER) Vestibular Health study to examine changes in the horizontal vestibulo-ocular reflex in response to high-velocity head movements to detect changes in peripheral vestibular function [1,2]. However, correct interpretation will require consideration of potential artifacts and the constraints of conducting this test across different phases of the mission. Therefore, the aims of our study were to (1) examine reliability across different test operators, and (2) compare results of active (aHIT) versus passive (pHIT) approaches to evaluate the feasibility of self-administered versus operator-assisted approaches. A comparison of responses with visual viewing of a wall target (default condition) versus vision occluded evaluated the influence of other oculomotor control influence across conditions, and comparison with computer-generated rotator head impulse tests (rHIT) examined the variability associated with the ocular responses independent of the variability in executing the head movements themselves. METHODS: Seventeen non-astronaut volunteers (male n=12, age=22.9 ± 3.0; female n=5, age=27.0 ± 6.0, mean ± std) completed HIT testing using video-oculography (VOG) goggles and a high-torque rotator system. The test order was counterbalanced across conditions: (1) passive head-on-torso (pHIT, default condition for flight study) using two operators (Op1 and Op2), (2) active head-on-torso (aHIT, subject initiated), and (3) passive head and torso using rotary chair (rHIT). Eye and head movement data were processed to obtain gain in each direction. The main outcome measures were average gain and asymmetry [1], as well as the percentage of acceptable trials not excluded due to insufficient head amplitude or recording artifacts. RESULTS: While the pHIT gains were similar with vision (1.032 ± 0.043) and vision occluded (1.029 ± 0.038) conditions, the percentage of acceptable trials was greater with vision (Op1 = 93.0%, Op2 = 93.2%) versus occluded (Op1 = 70.7%, Op2 = 67.6%). The reliability between operators was greater for pHIT gain (Intraclass Correlation, ICC = 0.66, p = 0.001) than for pHIT asymmetry (ICC = 0.57, p = 0.01). The percentage of acceptable trials reduced by ~20% during the self-administered aHITs for both visual conditions and tended to be lower than computer-generated rHITs (76% versus 83%). While the aHIT gains were not significantly different than pHIT gains for vision or vision occluded conditions, these measures were poorly correlated. Both aHIT and pHIT gains were significantly greater than the rHIT gains, presumably due to the reduced velocities from the rotator. The asymmetry measures were poorly correlated across pHIT, aHIT, and rHIT conditions, although none of the subjects had asymmetries greater than 16%. DISCUSSION: Our findings support the feasibility of using the standard pHIT methodology for spaceflight. Similarities between visual conditions reflect that these responses are mediated by the peripheral lateral canals rather than other oculomotor mechanisms. The inter-tester reliability was acceptable despite differences in the tester training. In addition to concerns of non-vestibular mechanisms introduced during self-administered aHITs [2], operator-assisted pHITs resulted in higher percentage of acceptable trials and should result in more efficient and reliable measures during spaceflight.

M R Ehrenburg↗

Eye/Brain/Task Testbed And Software

Eye/brain/task (EBT) testbed records electroencephalograms, movements of eyes, and structures of tasks to provide comprehensive data on neurophysiological experiments. Intended to serve continuing effort to develop means for interactions between human brain waves and computers. Software library associated with testbed provides capabilities to recall collected data, to process data on movements of eyes, to correlate eye-movement data with electroencephalographic data, and to present data graphically. Cognitive processes investigated in ways not previously possible.

Janiszewski, Thomas↗

Manipulator system performance measurement

Free-flying teleoperator technology for satellite servicing is reported. The rationale for a series of manipulator system tests is presented. Data are reported on movement time in a fine positioning task using two different manipulator systems. The movement time data showed reliable effects of movement direction and index of difficulty. These data were considered to be baseline performance measures for the manipulator system used and modifications to the control and visual systems were suggested for future testing.

Kirkpatrick, M., III↗

Application of Optical Measurement Techniques During Stages of Pregnancy: Use of Phantom High Speed Cameras for Digital Image Correlation (D.I.C.) During Baby Kicking and Abdomen Movements

Paired images were collected using a projected pattern instead of standard painting of the speckle pattern on her abdomen. High Speed cameras were post triggered after movements felt. Data was collected at 120 fps -limited due to 60hz frequency of projector. To ensure that kicks and movement data was real a background test was conducted with no baby movement (to correct for breathing and body motion).

Gradl, Paul↗