Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

The NAS parallel benchmarks

A new set of benchmarks was developed for the performance evaluation of highly parallel supercomputers. These benchmarks consist of a set of kernels, the 'Parallel Kernels,' and a simulated application benchmark. Together they mimic the computation and data movement characteristics of large scale computational fluid dynamics (CFD) applications. The principal distinguishing feature of these benchmarks is their 'pencil and paper' specification - all details of these benchmarks are specified only algorithmically. In this way many of the difficulties associated with conventional benchmarking approaches on highly parallel systems are avoided.

Bailey, David↗

Parallel computational fluid dynamics '91; Conference Proceedings, Stuttgart, Germany, Jun. 10-12, 1991

A conference was held on parallel computational fluid dynamics and produced related papers. Topics discussed in these papers include: parallel implicit and explicit solvers for compressible flow, parallel computational techniques for Euler and Navier-Stokes equations, grid generation techniques for parallel computers, and aerodynamic simulation om massively parallel systems.

Reinsch, K. G.↗

Exploring the use of I/O nodes for computation in a MIMD multiprocessor

As parallel systems move into the production scientific-computing world, the emphasis will be on cost-effective solutions that provide high throughput for a mix of applications. Cost effective solutions demand that a system make effective use of all of its resources. Many MIMD multiprocessors today, however, distinguish between 'compute' and 'I/O' nodes, the latter having attached disks and being dedicated to running the file-system server. This static division of responsibilities simplifies system management but does not necessarily lead to the best performance in workloads that need a different balance of computation and I/O. Of course, computational processes sharing a node with a file-system service may receive less CPU time, network bandwidth, and memory bandwidth than they would on a computation-only node. In this paper we begin to examine this issue experimentally. We found that high performance I/O does not necessarily require substantial CPU time, leaving plenty of time for application computation. There were some complex file-system requests, however, which left little CPU time available to the application. (The impact on network and memory bandwidth still needs to be determined.) For applications (or users) that cannot tolerate an occasional interruption, we recommend that they continue to use only compute nodes. For tolerant applications needing more cycles than those provided by the compute nodes, we recommend that they take full advantage of both compute and I/O nodes for computation, and that operating systems should make this possible.

Kotz, David↗

NAS Experiences of Porting CM Fortran Codes to HPF on IBM SP2 and SGI Power Challenge

Current Connection Machine (CM) Fortran codes developed for the CM-2 and the CM-5 represent an important class of parallel applications. Several users have employed CM Fortran codes in production mode on the CM-2 and the CM-5 for the last five to six years, constituting a heavy investment in terms of cost and time. With Thinking Machines Corporation's decision to withdraw from the hardware business and with the decommissioning of many CM-2 and CM-5 machines, the best way to protect the substantial investment in CM Fortran codes is to port the codes to High Performance Fortran (HPF) on highly parallel systems. HPF is very similar to CM Fortran and thus represents a natural transition. Conversion issues involved in porting CM Fortran codes on the CM-5 to HPF are presented. In particular, the differences between data distribution directives and the CM Fortran Utility Routines Library, as well as the equivalent functionality in the HPF Library are discussed. Several CM Fortran codes (Cannon algorithm for matrix-matrix multiplication, Linear solver Ax=b, 1-D convolution for 2-D datasets, Laplace's Equation solver, and Direct Simulation Monte Carlo (DSMC) codes have been ported to Subset HPF on the IBM SP2 and the SGI Power Challenge. Speedup ratios versus number of processors for the Linear solver and DSMC code are presented.

Saini, Subhash↗

Reducing Interprocessor Dependence in Recoverable Distributed Shared Memory

Checkpointing techniques in parallel systems use dependency tracking and/or message logging to ensure that a system rolls back to a consistent state. Traditional dependency tracking in distributed shared memory (DSM) systems is expensive because of high communication frequency. In this paper we show that, if designed correctly, a DSM system only needs to consider dependencies due to the transfer of blocks of data, resulting in reduced dependency tracking overhead and reduced potential for rollback propagation. We develop an ownership timestamp scheme to tolerate the loss of block state information and develop a passive server model of execution where interactions between processors are considered atomic. With our scheme, dependencies are significantly reduced compared to the traditional message-passing model.

Janssens, Bob↗

Performance of the Wavelet Decomposition on Massively Parallel Architectures

Traditionally, Fourier Transforms have been utilized for performing signal analysis and representation. But although it is straightforward to reconstruct a signal from its Fourier transform, no local description of the signal is included in its Fourier representation. To alleviate this problem, Windowed Fourier transforms and then wavelet transforms have been introduced, and it has been proven that wavelets give a better localization than traditional Fourier transforms, as well as a better division of the time- or space-frequency plane than Windowed Fourier transforms. Because of these properties and after the development of several fast algorithms for computing the wavelet representation of any signal, in particular the Multi-Resolution Analysis (MRA) developed by Mallat, wavelet transforms have increasingly been applied to signal analysis problems, especially real-life problems, in which speed is critical. In this paper we present and compare efficient wavelet decomposition algorithms on different parallel architectures. We report and analyze experimental measurements, using NASA remotely sensed images. Results show that our algorithms achieve significant performance gains on current high performance parallel systems, and meet scientific applications and multimedia requirements. The extensive performance measurements collected over a number of high-performance computer systems have revealed important architectural characteristics of these systems, in relation to the processing demands of the wavelet decomposition of digital images.

El-Ghazawi, Tarek A.↗

High-Performance Data Analysis Tools for Sun-Earth Connection Missions

The data analysis tool of choice for many Sun-Earth Connection missions is the Interactive Data Language (IDL) by ITT VIS. The increasing amount of data produced by these missions and the increasing complexity of image processing algorithms requires access to higher computing power. Parallel computing is a cost-effective way to increase the speed of computation, but algorithms oftentimes have to be modified to take advantage of parallel systems. Enhancing IDL to work on clusters gives scientists access to increased performance in a familiar programming environment. The goal of this project was to enable IDL applications to benefit from both computing clusters as well as graphics processing units (GPUs) for accelerating data analysis tasks. The tool suite developed in this project enables scientists now to solve demanding data analysis problems in IDL that previously required specialized software, and it allows them to be solved orders of magnitude faster than on conventional PCs. The tool suite consists of three components: (1) TaskDL, a software tool that simplifies the creation and management of task farms, collections of tasks that can be processed independently and require only small amounts of data communication; (2) mpiDL, a tool that allows IDL developers to use the Message Passing Interface (MPI) inside IDL for problems that require large amounts of data to be exchanged among multiple processors; and (3) GPULib, a tool that simplifies the use of GPUs as mathematical coprocessors from within IDL. mpiDL is unique in its support for the full MPI standard and its support of a broad range of MPI implementations. GPULib is unique in enabling users to take advantage of an inexpensive piece of hardware, possibly already installed in their computer, and achieve orders of magnitude faster execution time for numerically complex algorithms. TaskDL enables the simple setup and management of task farms on compute clusters. The products developed in this project have the potential to interact, so one can build a cluster of PCs, each equipped with a GPU, and use mpiDL to communicate between the nodes and GPULib to accelerate the computations on each node.

Messmer, Peter↗

Apparatus and Process for Controlled Nanomanufacturing Using Catalyst Retaining Structures

An apparatus and method for the controlled fabrication of nanostructures using catalyst retaining structures is disclosed. The apparatus includes one or more modified force microscopes having a nanotube attached to the tip portion of the microscopes. An electric current is passed from the nanotube to a catalyst layer of a substrate, thereby causing a localized chemical reaction to occur in a resist layer adjacent the catalyst layer. The region of the resist layer where the chemical reaction occurred is etched, thereby exposing a catalyst particle or particles in the catalyst layer surrounded by a wall of unetched resist material. Subsequent chemical vapor deposition causes growth of a nanostructure to occur upward through the wall of unetched resist material having controlled characteristics of height and diameter and, for parallel systems, number density.

Nguyen, Cattien↗

BioSentinel: Biosensors for Deep-Space Radiation Study

The BioSentinel mission will be deployed on NASA's Exploration Mission 1 (EM-1) in 2018. We will use the budding yeast, Saccharomyces cerevisiae, as a biosensor to study the effect of deep-space radiation on living cells. The BioSentinel mission will be the first investigation of a biological response to space radiation outside Low Earth Orbit (LEO) in over 40 years. Radiation can cause damage such as double stand breaks (DSBs) on DNA. The yeast cell was chosen for this mission because it is genetically controllable, shares homology with human cells in its DNA repair pathways, and can be stored in a desiccated state for long durations. Three yeast strains will be stored dry in multiple microfluidic cards: a wild type control strain, a mutant defective strain that cannot repair DSBs, and a biosensor strain that can only grow if it gets DSB-and-repair events occurring near a specific gene. Growth and metabolic activity of each strain will be measured by a 3-color LED optical detection system. Parallel experiments will be done on the International Space Station and on Earth so that we can compare the results to that of deep space. One of our main objectives is to characterize the microfluidic card activation sequence before the mission. To increase the sensitivity of yeast cells as biosensors, desiccated yeast in each card will be resuspended in a rehydration buffer. After several weeks, the rehydration buffer will be exchanged with a growth medium in order to measure yeast growth and metabolic activity. We are currently working on a time-course experiment to better understand the effects of the rehydration buffer on the response to ionizing radiation. We will resuspend the dried yeast in our rehydration medium over a period of time; then each week, we will measure the viability and ionizing radiation sensitivity of different yeast strains taken from this rehydration buffer. The data obtained in this study will be useful in finalizing the card activation sequence for this mission.

Yeast↗

Robotically Assembled Aerospace Structures: Digital Material Assembly using a Gantry-Type Assembler

This paper evaluates the development of automated assembly techniques for discrete lattice structures using a multi-axis gantry type CNC machine. These lattices are made of discrete components called digital materials. We present the development of a specialized end effector that works in conjunction with the CNC machine to assemble these lattices. With this configuration we are able to place voxels at a rate of 1.5 per minute. The scalability of digital material structures due to the incremental modular assembly is one of its key traits and an important metric of interest. We investigate the build times of a 5x5 beam structure on the scale of 1 meter (325 parts), 10 meters (3,250 parts), and 30 meters (9,750 parts). Utilizing the current configuration with a single end effector, performing serial assembly with a globally fixed feed station at the edge of the build volume, the build time increases according to a scaling law of n4, where n is the build scale. Build times can be reduced significantly by integrating feed systems into the gantry itself, resulting in a scaling law of n3. A completely serial assembly process will encounter time limitations as build scale increases. Automated assembly for digital materials can assemble high performance structures from discrete parts, and techniques such as built in feed systems, parallelization, and optimization of the fastening process will yield much higher throughput.

Mechanical Structure↗

Robotically Assembled Aerospace Structures: Digital Material Assembly using a Gantry-Type Assembler

This paper evaluates the development of automated assembly techniques for discrete lattice structures using a multi-axis gantry type CNC machine. These lattices are made of discrete components called "digital materials." We present the development of a specialized end effector that works in conjunction with the CNC machine to assemble these lattices. With this configuration we are able to place voxels at a rate of 1.5 per minute. The scalability of digital material structures due to the incremental modular assembly is one of its key traits and an important metric of interest. We investigate the build times of a 5x5 beam structure on the scale of 1 meter (325 parts), 10 meters (3,250 parts), and 30 meters (9,750 parts). Utilizing the current configuration with a single end effector, performing serial assembly with a globally fixed feed station at the edge of the build volume, the build time increases according to a scaling law of n4, where n is the build scale. Build times can be reduced significantly by integrating feed systems into the gantry itself, resulting in a scaling law of n3. A completely serial assembly process will encounter time limitations as build scale increases. Automated assembly for digital materials can assemble high performance structures from discrete parts, and techniques such as built in feed systems, parallelization, and optimization of the fastening process will yield much higher throughput.

Manufacturing↗

Oculometric Analysis of Saccadic Compensation for Visual Motion Processing Impairment due to Alcohol and Sleep Disruption

The Visuomotor Control Laboratory at Ames Research Center has developed a 5-minute ocular tracking test that computes 21 largely independent metrics of visuomotor performance, reflecting neural signal processing along a number of distinct pathways through cortex, brainstem, and cerebellum. Human sensorimotor performance is resilient to the challenges and stressors of many operational environments, in part, because overall performance is achieved through multiple parallel systems. Our multidimensional oculometrics allow us to examine impacts on these sub-components separately. To illustrate this, we contrasted the effects of two mild neural stressors, acute sleep-deprivation and low-dose alcohol. We have previously shown that, in both cases, oculometric analysis is a highly sensitive indicator of impairment. Here we quantified not only the observed impact on the performance of one sub-system, smooth pursuit, which uses high-level cortical processing of visual motion to track a moving object, but also the observed (partial) compensation by an evolutionarily older mid-brain and brainstem subsystem, saccades, which generates jumps in eye position to catch up with the target when smooth pursuit is inadequate. Specifically, we examined the dose-response (effect size vs. dose size) of the ground lost (pursuit deficit) and the ground recouped (saccadic compensation) across three separate studies – acute low-dose alcohol administration (16 subjects), acute sleep loss (12 subjects), and acute sleep loss with caffeine intervention (9 subjects). We computed the dose-response slopes using linear regression. The figure below shows that, in the case of acute sleep deprivation, the resulting slopes for ground lost and ground recouped (mean ± SE across subjects) were significantly different (paired t-test, t(11) = 5.17, p < 0.001), indicating poor saccadic compensation. However, when sleep loss was coupled with caffeine ingestion, ground lost was decreased and ground recouped increased such that the slopes were no longer different (t(8) = -0.05, p = 0.965). With alcohol, the two slopes were large albeit not significantly different (t(15) = 0.96, p = 0.351), indicating significant pursuit impairment but effective saccadic compensation. Our findings show that sleep deprivation and alcohol affect oculomotor performance differently. Low-dose alcohol effects appear predominantly cortical, with effective brainstem compensation. Sleep loss and circadian disruption however appears to affect both cortical and brainstem pathways with caffeine providing an effective countermeasure to both effects. Beyond the mere detection of impairment, our oculometric assessment allows us to characterize the nature of the deficit, to provide insight into the neural substrate, and to assess the effectiveness of countermeasures.

pursuit↗

Orchestrating Fault Prediction with Live Migration and Checkpointing

Checkpoint/Restart (C/R) is widely used to provide fault tolerance on High-Performance Computing (HPC) systems. However, Parallel File System (PFS) overhead and failure uncertainty cause significant application overhead. This paper develops an adaptive multi-level C/R model that incorporates a failure prediction and analysis model, which orchestrates failure prediction, checkpointing, checkpoint frequency, and proactive live migration along with the additional benefit of Burst Buffers (BB). It effectively reduces the overheads due to failures, checkpointing, and recovery. Simulation results for the Summit supercomputer yield a reduction of ~20%-86% in application overhead due to BBs, orchestrated failure prediction, and migration. We also observe a ~29% decrease in checkpoint writes to BBs, which can increase the longevity of the BB storage devices.

Behera, Subhendu↗

A Computationally Improved Heuristic Algorithm for Transmission Switching Using Line Flow Thresholds for Load Shed Reduction

We present a computationally improved heuristic algorithm for transmission switching (TS) to recover load shed. Research from the past showed that changing power system topology may control power flows and remove line congestion. Hence, TS may reduce the required load shed. One of the main challenges is to find a potential TS candidate in a suitable time. Here, we propose a novel heuristic method that is capable of finding the potential TS candidate faster than existing algorithms in literature. The proposed method is compatible with both the AC and DC optimal power flows (OPF). Three metrics are used to compare the proposed algorithm with the state-of-the-art from literature to show the speedup and accuracy achieved. The proposed method is implemented on the IEEE 30-bus system, PEGASE 89-bus system, IEEE 118-bus system, and Polish 2383- bus system. The results on the large-scale Polish 2383-bus system shows that the proposed algorithm is scalable to large real-world systems. Parallel computing is implemented to further improve the computational performance of the proposed algorithm.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Bringing OpenCL to Commodity RISC-V CPUs

The importance of open-source hardware has been increasing in recent years with the introduction of the RISC-V Open ISA. This has also accelerated the push for support of the open-source software stack from compiler tools to full-blown operating systems. Parallel computing with today’s Application Programming Interfaces such as OpenCL has proven to be effective at leveraging the parallelism in commodity multi-core processors and programmable parallel accelerators. However, to the best of our knowledge, there is currently no publicly available implementation of OpenCL targeting commodity RISC-V processors that is accessible to the open-source community. Besides opening RISC-V to the existing rich variety of scientific parallel applications, OpenCL also provides access to a unique genre of benchmarks useful in computer architecture research. In this work, we extended an Open-source implementation of OpenCL to target RISC-V CPUs. Our work not only cover commodity multi-core RISC-V processors, but also plethora of low- profile embedded RISC-V CPUs that often do not support atomic instructions or multi-threading.

Tine, Blaise↗

MOOSE: A Modular Platform for Fission and Fusion Multiphysics

The Multiphysics Object-Oriented Simulation Environment (MOOSE) Framework, as well as MOOSE-based simulation tools, have accelerated the development of fission energy and advanced reactor technologies through the United States Department of Energy, Office of Nuclear Science, Nuclear Energy Advanced Modeling & Simulation (NEAMS) Program. MOOSE contains a complete platform of multiphysics simulation capabilities, capable of running on massively parallel systems, and is developed in an open-source manner with great attention paid to high-quality software quality assurance practices. This overall approach could greatly benefit the fusion energy community, which requires rapid design iteration and improvement in order to facilitate the successful development of fusion as an alternative energy source to fossil fuels. In the first half of this talk, applications of MOOSE and MOOSE-based tools for advanced reactor designs will be showcased, as well as MOOSE ecosystem infrastructure (such as the NEAMS Virtual Test Bed) that enables and accelerates fission reactor design. In the second half, a discussion of how the MOOSE approach to modeling and simulation is currently being applied internationally in fusion energy research and development at the United Kingdom Atomic Energy Authority will be discussed, and ongoing/future domestic research efforts will be highlighted.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A Study of Parallels Between Antarctica South Pole Traverse Equipment and Lunar/Mars Surface Systems

The parallels between an actual Antarctica South Pole re-supply traverse conducted by the National Science Foundation (NSF) Office of Polar Programs in 2009 have been studied with respect to the latest mission architecture concepts being generated by the United States National Aeronautics and Space Administration (NASA) for lunar and Mars surface systems scenarios. The challenges faced by both endeavors are similar since they must both deliver equipment and supplies to support operations in an extreme environment with little margin for error in order to be successful. By carefully and closely monitoring the manifesting and operational support equipment lists which will enable this South Pole traverse, functional areas have been identified. The equipment required to support these functions will be listed with relevant properties such as mass, volume, spare parts and maintenance schedules. This equipment will be compared to space systems currently in use and projected to be required to support equivalent and parallel functions in Lunar and Mars missions in order to provide a level of realistic benchmarking. Space operations have historically required significant amounts of support equipment and tools to operate and maintain the space systems that are the primary focus of the mission. By gaining insight and expertise in Antarctic South Pole traverses, space missions can use the experience gained over the last half century of Antarctic operations in order to design for operations, maintenance, dual use, robustness and safety which will result in a more cost effective, user friendly, and lower risk surface system on the Moon and Mars. It is anticipated that the U.S Antarctic Program (USAP) will also realize benefits for this interaction with NASA in at least two areas: an understanding of how NASA plans and carries out its missions and possible improved efficiency through factors such as weight savings, alternative technologies, or modifications in training and operations.

Mueller, Robert P.↗

Proposal for massively parallel data storage system

An architecture for integrating large numbers of data storage units (drives) to form a distributed mass storage system is proposed. The network of interconnected units consists of nodes and links. At each node there resides a controller board, a data storage unit and, possibly, a local/remote user-terminal. The links (twisted-pair wires, coax cables, or fiber-optic channels) provide the communications backbone of the network. There is no central controller for the system as a whole; all decisions regarding allocation of resources, routing of messages and data-blocks, creation and distribution of redundant data-blocks throughout the system (for protection against possible failures), frequency of backup operations, etc., are made locally at individual nodes. The system can handle as many user-terminals as there are nodes in the network. Various users compete for resources by sending their requests to the local controller-board and receiving allocations of time and storage space. In principle, each user can have access to the entire system, and all drives can be running in parallel to service the requests for one or more users. The system is expandable up to a maximum number of nodes, determined by the number of routing-buffers built into the controller boards. Additional drives, controller-boards, user-terminals, and links can be simply plugged into an existing system in order to expand its capacity.

Mansuripur, M.↗