Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Micro Channel/Multibus-II Interface Circuit

Micro Channel/Multibus-II interface circuit provides electrical interconnections enabling communications between Micro Channels of IBM Personal System/2 computers and IEEE 1296 standard Multibus-II parallel system bus (iPSB). Made mostly of commercially available parts, interface enables independent Micro Channels to communicate over iPSB without modification.

D'Ambrose, John J.↗

Shared versus distributed memory multiprocessors

The question of whether multiprocessors should have shared or distributed memory has attracted a great deal of attention. Some researchers argue strongly for building distributed memory machines, while others argue just as strongly for programming shared memory multiprocessors. A great deal of research is underway on both types of parallel systems. Special emphasis is placed on systems with a very large number of processors for computation intensive tasks and considers research and implementation trends. It appears that the two types of systems will likely converge to a common form for large scale multiprocessors.

Jordan, Harry F.↗

Compile-time estimation of communication costs in multicomputers

An important problem facing numerous research projects on parallelizing compilers for distributed memory machines is that of automatically determining a suitable data partitioning scheme for a program. Any strategy for automatic data partitioning needs a mechanism for estimating the performance of a program under a given partitioning scheme, the most crucial part of which involves determining the communication costs incurred by the program. A methodology is described for estimating the communication costs at compile-time as functions of the numbers of processors over which various arrays are distributed. A strategy is described along with its theoretical basis, for making program transformations that expose opportunities for combining of messages, leading to considerable savings in the communication costs. For certain loops with regular dependences, the compiler can detect the possibility of pipelining, and thus estimate communication costs more accurately than it could otherwise. These results are of great significance to any parallelization system supporting numeric applications on multicomputers. In particular, they lay down a framework for effective synthesis of communication on multicomputers from sequential program references.

Gupta, Manish↗

Integrating performance data collection, analysis, and visualization

An integrated data collection, analysis, and data visualization environment is described for a specific parallel system - the Intel iPSC/2 hypercube. The data collection components of the environment encompass software event tracing at the operating system with a program level and a hardware-based performance monitoring system used to capture software events. A visualization system based on the X-window environment permits dynamic display and reduction of performance data. A performance data collection, analysis, and visualization environment makes it possible to access the effects of architectural and system software variations.

Malony, Allen D.↗

The NAS parallel benchmarks

A new set of benchmarks has been developed for the performance evaluation of highly parallel supercomputers in the framework of the NASA Ames Numerical Aerodynamic Simulation (NAS) Program. These consist of five 'parallel kernel' benchmarks and three 'simulated application' benchmarks. Together they mimic the computation and data movement characteristics of large-scale computational fluid dynamics applications. The principal distinguishing feature of these benchmarks is their 'pencil and paper' specification-all details of these benchmarks are specified only algorithmically. In this way many of the difficulties associated with conventional benchmarking approaches on highly parallel systems are avoided.

Bailey, D. H.↗

The NAS parallel benchmarks

A new set of benchmarks was developed for the performance evaluation of highly parallel supercomputers. These benchmarks consist of a set of kernels, the 'Parallel Kernels,' and a simulated application benchmark. Together they mimic the computation and data movement characteristics of large scale computational fluid dynamics (CFD) applications. The principal distinguishing feature of these benchmarks is their 'pencil and paper' specification - all details of these benchmarks are specified only algorithmically. In this way many of the difficulties associated with conventional benchmarking approaches on highly parallel systems are avoided.

Bailey, David↗

Parallel computational fluid dynamics '91; Conference Proceedings, Stuttgart, Germany, Jun. 10-12, 1991

A conference was held on parallel computational fluid dynamics and produced related papers. Topics discussed in these papers include: parallel implicit and explicit solvers for compressible flow, parallel computational techniques for Euler and Navier-Stokes equations, grid generation techniques for parallel computers, and aerodynamic simulation om massively parallel systems.

Reinsch, K. G.↗

Exploring the use of I/O nodes for computation in a MIMD multiprocessor

As parallel systems move into the production scientific-computing world, the emphasis will be on cost-effective solutions that provide high throughput for a mix of applications. Cost effective solutions demand that a system make effective use of all of its resources. Many MIMD multiprocessors today, however, distinguish between 'compute' and 'I/O' nodes, the latter having attached disks and being dedicated to running the file-system server. This static division of responsibilities simplifies system management but does not necessarily lead to the best performance in workloads that need a different balance of computation and I/O. Of course, computational processes sharing a node with a file-system service may receive less CPU time, network bandwidth, and memory bandwidth than they would on a computation-only node. In this paper we begin to examine this issue experimentally. We found that high performance I/O does not necessarily require substantial CPU time, leaving plenty of time for application computation. There were some complex file-system requests, however, which left little CPU time available to the application. (The impact on network and memory bandwidth still needs to be determined.) For applications (or users) that cannot tolerate an occasional interruption, we recommend that they continue to use only compute nodes. For tolerant applications needing more cycles than those provided by the compute nodes, we recommend that they take full advantage of both compute and I/O nodes for computation, and that operating systems should make this possible.

Kotz, David↗

NAS Experiences of Porting CM Fortran Codes to HPF on IBM SP2 and SGI Power Challenge

Current Connection Machine (CM) Fortran codes developed for the CM-2 and the CM-5 represent an important class of parallel applications. Several users have employed CM Fortran codes in production mode on the CM-2 and the CM-5 for the last five to six years, constituting a heavy investment in terms of cost and time. With Thinking Machines Corporation's decision to withdraw from the hardware business and with the decommissioning of many CM-2 and CM-5 machines, the best way to protect the substantial investment in CM Fortran codes is to port the codes to High Performance Fortran (HPF) on highly parallel systems. HPF is very similar to CM Fortran and thus represents a natural transition. Conversion issues involved in porting CM Fortran codes on the CM-5 to HPF are presented. In particular, the differences between data distribution directives and the CM Fortran Utility Routines Library, as well as the equivalent functionality in the HPF Library are discussed. Several CM Fortran codes (Cannon algorithm for matrix-matrix multiplication, Linear solver Ax=b, 1-D convolution for 2-D datasets, Laplace's Equation solver, and Direct Simulation Monte Carlo (DSMC) codes have been ported to Subset HPF on the IBM SP2 and the SGI Power Challenge. Speedup ratios versus number of processors for the Linear solver and DSMC code are presented.

Saini, Subhash↗

Reducing Interprocessor Dependence in Recoverable Distributed Shared Memory

Checkpointing techniques in parallel systems use dependency tracking and/or message logging to ensure that a system rolls back to a consistent state. Traditional dependency tracking in distributed shared memory (DSM) systems is expensive because of high communication frequency. In this paper we show that, if designed correctly, a DSM system only needs to consider dependencies due to the transfer of blocks of data, resulting in reduced dependency tracking overhead and reduced potential for rollback propagation. We develop an ownership timestamp scheme to tolerate the loss of block state information and develop a passive server model of execution where interactions between processors are considered atomic. With our scheme, dependencies are significantly reduced compared to the traditional message-passing model.

Janssens, Bob↗

Performance of the Wavelet Decomposition on Massively Parallel Architectures

Traditionally, Fourier Transforms have been utilized for performing signal analysis and representation. But although it is straightforward to reconstruct a signal from its Fourier transform, no local description of the signal is included in its Fourier representation. To alleviate this problem, Windowed Fourier transforms and then wavelet transforms have been introduced, and it has been proven that wavelets give a better localization than traditional Fourier transforms, as well as a better division of the time- or space-frequency plane than Windowed Fourier transforms. Because of these properties and after the development of several fast algorithms for computing the wavelet representation of any signal, in particular the Multi-Resolution Analysis (MRA) developed by Mallat, wavelet transforms have increasingly been applied to signal analysis problems, especially real-life problems, in which speed is critical. In this paper we present and compare efficient wavelet decomposition algorithms on different parallel architectures. We report and analyze experimental measurements, using NASA remotely sensed images. Results show that our algorithms achieve significant performance gains on current high performance parallel systems, and meet scientific applications and multimedia requirements. The extensive performance measurements collected over a number of high-performance computer systems have revealed important architectural characteristics of these systems, in relation to the processing demands of the wavelet decomposition of digital images.

El-Ghazawi, Tarek A.↗

High-Performance Data Analysis Tools for Sun-Earth Connection Missions

The data analysis tool of choice for many Sun-Earth Connection missions is the Interactive Data Language (IDL) by ITT VIS. The increasing amount of data produced by these missions and the increasing complexity of image processing algorithms requires access to higher computing power. Parallel computing is a cost-effective way to increase the speed of computation, but algorithms oftentimes have to be modified to take advantage of parallel systems. Enhancing IDL to work on clusters gives scientists access to increased performance in a familiar programming environment. The goal of this project was to enable IDL applications to benefit from both computing clusters as well as graphics processing units (GPUs) for accelerating data analysis tasks. The tool suite developed in this project enables scientists now to solve demanding data analysis problems in IDL that previously required specialized software, and it allows them to be solved orders of magnitude faster than on conventional PCs. The tool suite consists of three components: (1) TaskDL, a software tool that simplifies the creation and management of task farms, collections of tasks that can be processed independently and require only small amounts of data communication; (2) mpiDL, a tool that allows IDL developers to use the Message Passing Interface (MPI) inside IDL for problems that require large amounts of data to be exchanged among multiple processors; and (3) GPULib, a tool that simplifies the use of GPUs as mathematical coprocessors from within IDL. mpiDL is unique in its support for the full MPI standard and its support of a broad range of MPI implementations. GPULib is unique in enabling users to take advantage of an inexpensive piece of hardware, possibly already installed in their computer, and achieve orders of magnitude faster execution time for numerically complex algorithms. TaskDL enables the simple setup and management of task farms on compute clusters. The products developed in this project have the potential to interact, so one can build a cluster of PCs, each equipped with a GPU, and use mpiDL to communicate between the nodes and GPULib to accelerate the computations on each node.

Messmer, Peter↗

Apparatus and Process for Controlled Nanomanufacturing Using Catalyst Retaining Structures

An apparatus and method for the controlled fabrication of nanostructures using catalyst retaining structures is disclosed. The apparatus includes one or more modified force microscopes having a nanotube attached to the tip portion of the microscopes. An electric current is passed from the nanotube to a catalyst layer of a substrate, thereby causing a localized chemical reaction to occur in a resist layer adjacent the catalyst layer. The region of the resist layer where the chemical reaction occurred is etched, thereby exposing a catalyst particle or particles in the catalyst layer surrounded by a wall of unetched resist material. Subsequent chemical vapor deposition causes growth of a nanostructure to occur upward through the wall of unetched resist material having controlled characteristics of height and diameter and, for parallel systems, number density.

Nguyen, Cattien↗

BioSentinel: Biosensors for Deep-Space Radiation Study

The BioSentinel mission will be deployed on NASA's Exploration Mission 1 (EM-1) in 2018. We will use the budding yeast, Saccharomyces cerevisiae, as a biosensor to study the effect of deep-space radiation on living cells. The BioSentinel mission will be the first investigation of a biological response to space radiation outside Low Earth Orbit (LEO) in over 40 years. Radiation can cause damage such as double stand breaks (DSBs) on DNA. The yeast cell was chosen for this mission because it is genetically controllable, shares homology with human cells in its DNA repair pathways, and can be stored in a desiccated state for long durations. Three yeast strains will be stored dry in multiple microfluidic cards: a wild type control strain, a mutant defective strain that cannot repair DSBs, and a biosensor strain that can only grow if it gets DSB-and-repair events occurring near a specific gene. Growth and metabolic activity of each strain will be measured by a 3-color LED optical detection system. Parallel experiments will be done on the International Space Station and on Earth so that we can compare the results to that of deep space. One of our main objectives is to characterize the microfluidic card activation sequence before the mission. To increase the sensitivity of yeast cells as biosensors, desiccated yeast in each card will be resuspended in a rehydration buffer. After several weeks, the rehydration buffer will be exchanged with a growth medium in order to measure yeast growth and metabolic activity. We are currently working on a time-course experiment to better understand the effects of the rehydration buffer on the response to ionizing radiation. We will resuspend the dried yeast in our rehydration medium over a period of time; then each week, we will measure the viability and ionizing radiation sensitivity of different yeast strains taken from this rehydration buffer. The data obtained in this study will be useful in finalizing the card activation sequence for this mission.

Yeast↗

Robotically Assembled Aerospace Structures: Digital Material Assembly using a Gantry-Type Assembler

This paper evaluates the development of automated assembly techniques for discrete lattice structures using a multi-axis gantry type CNC machine. These lattices are made of discrete components called digital materials. We present the development of a specialized end effector that works in conjunction with the CNC machine to assemble these lattices. With this configuration we are able to place voxels at a rate of 1.5 per minute. The scalability of digital material structures due to the incremental modular assembly is one of its key traits and an important metric of interest. We investigate the build times of a 5x5 beam structure on the scale of 1 meter (325 parts), 10 meters (3,250 parts), and 30 meters (9,750 parts). Utilizing the current configuration with a single end effector, performing serial assembly with a globally fixed feed station at the edge of the build volume, the build time increases according to a scaling law of n4, where n is the build scale. Build times can be reduced significantly by integrating feed systems into the gantry itself, resulting in a scaling law of n3. A completely serial assembly process will encounter time limitations as build scale increases. Automated assembly for digital materials can assemble high performance structures from discrete parts, and techniques such as built in feed systems, parallelization, and optimization of the fastening process will yield much higher throughput.

Mechanical Structure↗

Robotically Assembled Aerospace Structures: Digital Material Assembly using a Gantry-Type Assembler

This paper evaluates the development of automated assembly techniques for discrete lattice structures using a multi-axis gantry type CNC machine. These lattices are made of discrete components called "digital materials." We present the development of a specialized end effector that works in conjunction with the CNC machine to assemble these lattices. With this configuration we are able to place voxels at a rate of 1.5 per minute. The scalability of digital material structures due to the incremental modular assembly is one of its key traits and an important metric of interest. We investigate the build times of a 5x5 beam structure on the scale of 1 meter (325 parts), 10 meters (3,250 parts), and 30 meters (9,750 parts). Utilizing the current configuration with a single end effector, performing serial assembly with a globally fixed feed station at the edge of the build volume, the build time increases according to a scaling law of n4, where n is the build scale. Build times can be reduced significantly by integrating feed systems into the gantry itself, resulting in a scaling law of n3. A completely serial assembly process will encounter time limitations as build scale increases. Automated assembly for digital materials can assemble high performance structures from discrete parts, and techniques such as built in feed systems, parallelization, and optimization of the fastening process will yield much higher throughput.

Manufacturing↗

Oculometric Analysis of Saccadic Compensation for Visual Motion Processing Impairment due to Alcohol and Sleep Disruption

The Visuomotor Control Laboratory at Ames Research Center has developed a 5-minute ocular tracking test that computes 21 largely independent metrics of visuomotor performance, reflecting neural signal processing along a number of distinct pathways through cortex, brainstem, and cerebellum. Human sensorimotor performance is resilient to the challenges and stressors of many operational environments, in part, because overall performance is achieved through multiple parallel systems. Our multidimensional oculometrics allow us to examine impacts on these sub-components separately. To illustrate this, we contrasted the effects of two mild neural stressors, acute sleep-deprivation and low-dose alcohol. We have previously shown that, in both cases, oculometric analysis is a highly sensitive indicator of impairment. Here we quantified not only the observed impact on the performance of one sub-system, smooth pursuit, which uses high-level cortical processing of visual motion to track a moving object, but also the observed (partial) compensation by an evolutionarily older mid-brain and brainstem subsystem, saccades, which generates jumps in eye position to catch up with the target when smooth pursuit is inadequate. Specifically, we examined the dose-response (effect size vs. dose size) of the ground lost (pursuit deficit) and the ground recouped (saccadic compensation) across three separate studies – acute low-dose alcohol administration (16 subjects), acute sleep loss (12 subjects), and acute sleep loss with caffeine intervention (9 subjects). We computed the dose-response slopes using linear regression. The figure below shows that, in the case of acute sleep deprivation, the resulting slopes for ground lost and ground recouped (mean ± SE across subjects) were significantly different (paired t-test, t(11) = 5.17, p < 0.001), indicating poor saccadic compensation. However, when sleep loss was coupled with caffeine ingestion, ground lost was decreased and ground recouped increased such that the slopes were no longer different (t(8) = -0.05, p = 0.965). With alcohol, the two slopes were large albeit not significantly different (t(15) = 0.96, p = 0.351), indicating significant pursuit impairment but effective saccadic compensation. Our findings show that sleep deprivation and alcohol affect oculomotor performance differently. Low-dose alcohol effects appear predominantly cortical, with effective brainstem compensation. Sleep loss and circadian disruption however appears to affect both cortical and brainstem pathways with caffeine providing an effective countermeasure to both effects. Beyond the mere detection of impairment, our oculometric assessment allows us to characterize the nature of the deficit, to provide insight into the neural substrate, and to assess the effectiveness of countermeasures.

pursuit↗

A Study of Parallels Between Antarctica South Pole Traverse Equipment and Lunar/Mars Surface Systems

The parallels between an actual Antarctica South Pole re-supply traverse conducted by the National Science Foundation (NSF) Office of Polar Programs in 2009 have been studied with respect to the latest mission architecture concepts being generated by the United States National Aeronautics and Space Administration (NASA) for lunar and Mars surface systems scenarios. The challenges faced by both endeavors are similar since they must both deliver equipment and supplies to support operations in an extreme environment with little margin for error in order to be successful. By carefully and closely monitoring the manifesting and operational support equipment lists which will enable this South Pole traverse, functional areas have been identified. The equipment required to support these functions will be listed with relevant properties such as mass, volume, spare parts and maintenance schedules. This equipment will be compared to space systems currently in use and projected to be required to support equivalent and parallel functions in Lunar and Mars missions in order to provide a level of realistic benchmarking. Space operations have historically required significant amounts of support equipment and tools to operate and maintain the space systems that are the primary focus of the mission. By gaining insight and expertise in Antarctic South Pole traverses, space missions can use the experience gained over the last half century of Antarctic operations in order to design for operations, maintenance, dual use, robustness and safety which will result in a more cost effective, user friendly, and lower risk surface system on the Moon and Mars. It is anticipated that the U.S Antarctic Program (USAP) will also realize benefits for this interaction with NASA in at least two areas: an understanding of how NASA plans and carries out its missions and possible improved efficiency through factors such as weight savings, alternative technologies, or modifications in training and operations.

Mueller, Robert P.↗