Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel and high performance computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

The Portals 4.3 Network Programming Interface

This report presents a specification for the Portals 4 network programming interface. Portals 4 is intended to allow scalable, high-performance network communication between nodes of a parallel computing system. Portals 4 is well suited to massively parallel processing and embedded systems. Portals 4 represents an adaption of the data movement layer developed for massively parallel processing platforms, such as the 4500-node Intel TeraFLOPS machine. Sandia's Cplant cluster project motivated the development of Version 3.0, which was later extended to Version 3.3 as part of the Cray Red Storm machine and XT line. Version 4 is targeted to the next generation of machines employing advanced network interface architectures that support enhanced offload capabilities.

97 MATHEMATICS AND COMPUTING↗

Parallel-Vector Algorithm For Rapid Structural Anlysis

New algorithm developed to overcome deficiency of skyline storage scheme by use of variable-band storage scheme. Exploits both parallel and vector capabilities of modern high-performance computers. Gives engineers and designers opportunity to include more design variables and constraints during optimization of structures. Enables use of more refined finite-element meshes to obtain improved understanding of complex behaviors of aerospace structures leading to better, safer designs. Not only attractive for current supercomputers but also for next generation of shared-memory supercomputers.

Agarwal, Tarun R.↗

The role of HiPPI switches in mass storage systems: A five year prospective

New standards are evolving which provide the foundation for novel multi-gigabit per second data communication structures. The lowest layer protocols are so generalized that they encourage a wide range of application. Specifically, the ANSI High Performance Parallel Interface (HiPPI) is being applied to computer peripheral attachment as well as general data communication networks. This paper introduces the HiPPI standards suite and technology products which incorporate the standards. The use of simple HiPPI crosspoint switches to build potentially complex extended 'fabrics' is discussed in detail. Several near term applications of the HiPPI technology are briefly described with additional attention to storage systems. Finally, some related standards are mentioned which may further expand the concepts above.

Gilbert, T. A.↗

The role of HiPPI switches in mass storage systems: A five year prospective

New standards are evolving which provide the foundation for multi-gigabit per second data communication structures. The lowest layer protocols are so generalized that they encourage a wide range of application. Specifically, the ANSI High Performance Parallel Interface (HiPPI) is being applied to computer peripheral attachment as well as general data communication networks. The HiPPI Standards suite and technology products which incorporate the standards are introduced. The use of simple HiPPI crosspoint switches to build potentially complex extended 'fabrics' is discussed in detail. Several near term applications of the HiPPI technology are briefly described with additional attention to storage systems. Finally, some related standards are mentioned which may further expand the concepts above.

Gilbert, T. A.↗

USRA/RIACS

The Research Institute for Advanced Computer Science (RIACS) was established by the Universities Space Research Association (USRA) at the NASA Ames Research Center (ARC) on 6 June 1983. RIACS is privately operated by USRA, a consortium of universities with research programs in the aerospace sciences, under a cooperative agreement with NASA. The primary mission of RIACS is to provide research and expertise in computer science and scientific computing to support the scientific missions of NASA ARC. The research carried out at RIACS must change its emphasis from year to year in response to NASA ARC's changing needs and technological opportunities. A flexible scientific staff is provided through a university faculty visitor program, a post doctoral program, and a student visitor program. Not only does this provide appropriate expertise but it also introduces scientists outside of NASA to NASA problems. A small group of core RIACS staff provides continuity and interacts with an ARC technical monitor and scientific advisory group to determine the RIACS mission. RIACS activities are reviewed and monitored by a USRA advisory council and ARC technical monitor. Research at RIACS is currently being done in the following areas: Parallel Computing; Advanced Methods for Scientific Computing; Learning Systems; High Performance Networks and Technology; Graphics, Visualization, and Virtual Environments.

Oliger, Joseph↗

Activities of the Research Institute for Advanced Computer Science

The Research Institute for Advanced Computer Science (RIACS) was established by the Universities Space Research Association (USRA) at the NASA Ames Research Center (ARC) on June 6, 1983. RIACS is privately operated by USRA, a consortium of universities with research programs in the aerospace sciences, under contract with NASA. The primary mission of RIACS is to provide research and expertise in computer science and scientific computing to support the scientific missions of NASA ARC. The research carried out at RIACS must change its emphasis from year to year in response to NASA ARC's changing needs and technological opportunities. Research at RIACS is currently being done in the following areas: (1) parallel computing; (2) advanced methods for scientific computing; (3) high performance networks; and (4) learning systems. RIACS technical reports are usually preprints of manuscripts that have been submitted to research journals or conference proceedings. A list of these reports for the period January 1, 1994 through December 31, 1994 is in the Reports and Abstracts section of this report.

Oliger, Joseph↗

Highly-Parallel, Highly-Compact Computing Structures Implemented in Nanotechnology

In this paper, we describe work in which we are evaluating how the evolving properties of nano-electronic devices could best be utilized in highly parallel computing structures. Because of their combination of high performance, low power, and extreme compactness, such structures would have obvious applications in spaceborne environments, both for general mission control and for on-board data analysis. However, the anticipated properties of nano-devices mean that the optimum architecture for such systems is by no means certain. Candidates include single instruction multiple datastream (SIMD) arrays, neural networks, and multiple instruction multiple datastream (MIMD) assemblies.

Crawley, D. G.↗

Turbomachinery Flows Modeled

Last year, researchers at the NASA Lewis Research Center used the average passage code APNASA to complete the largest three-dimensional simulation of a multistage axial flow compressor to date. Consisting of 29 blade rows, the configuration is typical of those found in aeroengines today. The simulation, which was executed on the High Performance Computing and Communications (HPCC) Program IBM SP2 parallel computer located at the NASA Ames Research Center, took nearly 90 hr to complete. Since the completion of this activity, a fine-grain, parallel version of APNASA has been written by a team of researchers from General Electric, NASA Lewis, and NYMA. Timing studies performed on the SP2 have shown that, with eight processors assigned to each blade row, the simulation time is reduced by a factor of six. For this configuration, the simulation time would be 15 hr. The reduction in computing time indicates that an overnight turnaround of a multistage configuration simulation is feasible. In addition, average passage forms of two-equation turbulence models were formulated. These models are currently being incorporated into APNASA.

Adamczyk, John J.↗

Onboard Autonomous Trajectory Planning for Mars Power Descent

In recent years, there has been an increasing interest in space-qualified processors such as multi-core central processing units and graphics processing units that can withstand the adverse effects of space radiation. These processors can allow parallel programming to perform tasks that typically demand high computational power. One can study guidance schemes that can take advantage of these currently developing processors and provide more robust guidance. Software for Multi-model Autonomous Real-time Trajectories (SMART) guidance can identify robust trajectories by running an onboard Monte Carlo analysis. SMART guidance can take advantage of knowledge updates obtained from the onboard sensors, allowing it to consider the off-nominal cases that it would not typically encounter during the offline trajectory analysis. This work uses the SMART guidance for the powered divert at Mars simulation in Program to Optimize and Simulated Trajectories- II.

Pardha Sai Chadalavada↗

Onboard Autonomous Trajectory Planning for Mars Power Descent

In recent years, there has been an increasing interest in space-qualified processors such as multi-core central processing units and graphics processing units that can withstand the adverse effects of space radiation. These processors can allow parallel programming to perform tasks that typically demand high computational power. One can study guidance schemes that can take advantage of these currently developing processors and provide more robust guidance. Software for Multi-model Autonomous Real-time Trajectories (SMART) guidance can identify robust trajectories by running an onboard Monte Carlo analysis. SMART guidance can take advantage of knowledge updates obtained from the onboard sensors, allowing it to consider the off-nominal cases that it would not typically encounter during the offline trajectory analysis. This work uses the SMART guidance for the powered divert at Mars simulation in Program to Optimize and Simulated Trajectories- II.

Autonomous Planning↗

PandAna: A Python Analysis Framework for Scalable High Performance Computing in High Energy Physics

Modern experiments in high energy physics analyze millions of events recorded in particle detectors to select the events of interest and make measurements of physics parameters. These data can often be stored as tabular data in files with detector information and reconstructed quantities. Current techniques for event selection in these files lack the scalability needed for high performance computing environments. We describe our work to develop a high energy physics analysis framework suitable for high performance computing. This new framework utilizes modern tools for reading files and implicit data parallelism. Framework users analyze tabular data using standard, easy-to-use data analysis techniques in Python while the framework handles the file manipulations and parallelism without the user needing advanced experience in parallel programming. In future versions, we hope to provide a framework that can be utilized on a personal computer or a high performance computing cluster with little change to the user code.

Groh, Micah↗

Parallel-vector solution of large-scale structural analysis problems on supercomputers

A direct linear equation solution method based on the Choleski factorization procedure is presented which exploits both parallel and vector features of supercomputers. The new equation solver is described, and its performance is evaluated by solving structural analysis problems on three high-performance computers. The method has been implemented using Force, a generic parallel FORTRAN language.

Storaasli, Olaf O.↗

Code Optimization and Parallelization on the Origins: Looking from Users' Perspective

Parallel machines are becoming the main compute engines for high performance computing. Despite their increasing popularity, it is still a challenge for most users to learn the basic techniques to optimize/parallelize their codes on such platforms. In this paper, we present some experiences on learning these techniques for the Origin systems at the NASA Advanced Supercomputing Division. Emphasis of this paper will be on a few essential issues (with examples) that general users should master when they work with the Origins as well as other parallel systems.

Chang, Yan-Tyng Sherry↗

Traveler: Navigating Task Parallel Traces for Performance Analysis

Understanding the behavior of software in execution is a key step in identifying and fixing performance issues. This is especially important in high performance computing contexts where even minor performance tweaks can translate into large savings in terms of computational resource use. To aid performance analysis, developers may collect an execution trace —a chronological log of program activity during execution. As traces represent the full history, developers can discover a wide array of possibly previously unknown performance issues, making them an important artifact for exploratory performance analysis. However, interactive trace visualization is difficult due to issues of data size and complexity of meaning. Traces represent nanosecond-level events across many parallel processes, meaning the collected data is often large and difficult to explore. The rise of asynchronous task parallel programming paradigms complicates the relation between events and their probable cause. Here, to address these challenges, we conduct a continuing design study in collaboration with high performance computing researchers. We develop diverse and hierarchical ways to navigate and represent execution trace data in support of their trace analysis tasks. Through an iterative design process, we developed Traveler , an integrated visualization platform for task parallel traces. Traveler provides multiple linked interfaces to help navigate trace data from multiple contexts. We evaluate the utility of Traveler through feedback from users and a case study, finding that integrating multiple modes of navigation in our design supported performance analysis tasks and led to the discovery of previously unknown behavior in a distributed array library.

97 MATHEMATICS AND COMPUTING↗

Scalability study of parallel spatial direct numerical simulation code on IBM SP1 parallel supercomputer

The implementation and the performance of a parallel spatial direct numerical simulation (PSDNS) code are reported for the IBM SP1 supercomputer. The spatially evolving disturbances that are associated with laminar-to-turbulent in three-dimensional boundary-layer flows are computed with the PS-DNS code. By remapping the distributed data structure during the course of the calculation, optimized serial library routines can be utilized that substantially increase the computational performance. Although the remapping incurs a high communication penalty, the parallel efficiency of the code remains above 40% for all performed calculations. By using appropriate compile options and optimized library routines, the serial code achieves 52-56 Mflops on a single node of the SP1 (45% of theoretical peak performance). The actual performance of the PSDNS code on the SP1 is evaluated with a 'real world' simulation that consists of 1.7 million grid points. One time step of this simulation is calculated on eight nodes of the SP1 in the same time as required by a Cray Y/MP for the same simulation. The scalability information provides estimated computational costs that match the actual costs relative to changes in the number of grid points.

Hanebutte, Ulf R.↗

Automatic data partitioning on distributed memory multicomputers

Distributed-memory parallel computers are increasingly being used to provide high levels of performance for scientific applications. Unfortunately, such machines are not very easy to program. A number of research efforts seek to alleviate this problem by developing compilers that take over the task of generating communication. The communication overheads and the extent of parallelism exploited in the resulting target program are determined largely by the manner in which data is partitioned across different processors of the machine. Most of the compilers provide no assistance to the programmer in the crucial task of determining a good data partitioning scheme. A novel approach is presented, the constraints-based approach, to the problem of automatic data partitioning for numeric programs. In this approach, the compiler identifies some desirable requirements on the distribution of various arrays being referenced in each statement, based on performance considerations. These desirable requirements are referred to as constraints. For each constraint, the compiler determines a quality measure that captures its importance with respect to the performance of the program. The quality measure is obtained through static performance estimation, without actually generating the target data-parallel program with explicit communication. Each data distribution decision is taken by combining all the relevant constraints. The compiler attempts to resolve any conflicts between constraints such that the overall execution time of the parallel program is minimized. This approach has been implemented as part of a compiler called Paradigm, that accepts Fortran 77 programs, and specifies the partitioning scheme to be used for each array in the program. We have obtained results on some programs taken from the Linpack and Eispack libraries, and the Perfect Benchmarks. These results are quite promising, and demonstrate the feasibility of automatic data partitioning for a significant class of scientific application programs with regular computations.

Gupta, Manish↗

An Integrated Architecture for Onboard Spacecraft

As increasingly complex scientific and environmental observation spacecraft are deployed, the burden on the downlink assets, and ground-based systems complexity and cost is becoming a major problem. Already, the limitations of communications bandwidth and processing throughput limit the science data gathering, both in volume and in rate. This poses a dilemma to the scientist experimenter forcing choices between data collection and bandwidth/processing/archiving. Advances in ground based processing and space-to-Earth links have fallen behind the requirements for observation data, at increasing rates, over the last few decades. As NASA achieves its 40th anniversary, the ability to observe and capture phenomena of theoretical and practical interest to life on Earth far outstrips the ability to transfer, process, or store these data. NASA recognizes the need to invest on technological advancements that will enable both the space and ground systems to address the limitations. Spacecraft onboard computing power is a clear one. The capability of creating data products onboard the spacecraft adds a new level of flexibility to address the more demanding observation needs. Current spacecraft computing power is limited and incapable of addressing the needs of the new generation of observation satellites because extensive onboard data processing is required. Traditional spacecraft architectures only collect, package, and transmit to Earth the data acquired by multiple instruments. Conversely, the experience on developing ground data systems shows the need for high performance computing systems to process and create information from the instrumentation data. The expectation is that supercomputing technology is required to enable spacecraft to create information onboard. Moving supercomputing capability onboard spacecraft requires an approach that considers an integrated data architecture. Otherwise, it may simply convert a compute-bound problem into a communications bound problem, as has been shown numerous times in the context of massively parallel architectures. What is left to determine are the technologies that will enable spacecraft high performance computing.

Figueiredo, Marco A.↗

EBS Task Force: Task 9/FEBEX Modeling Final Report: Thermo-Hydrological Modeling with PFLOTRAN

This report outlines Sandia National Laboratories modeling studies applied to Stage 1 and Stage 2 of the Full-scale Engineered Barriers Experiment in Crystalline Host Rock (FEBEX) in situ test for the SKB EBS Task Force Task 9. The FEBEX test was a full-scale test conducted over ~18 years at the Grimsel, Switzerland Underground Research Laboratory (URL) managed by NAGRA. It involved emplacing simulated waste packages, in the form of welded cylindrical heaters, inside a tunnel in crystalline granitic rock and surrounded by a bentonite barrier and cement plug. Sensors emplaced within the bentonite monitored the wetting-up, heating, and drying out of the bentonite barrier, and the large resulting data set provides an excellent opportunity for validation of multiphysics Thermal-Hydrological (TH), Thermal-Hydrologic-Chemical (THC), and Thermal-Hydrological-Mechanical (THM) modeling approaches for underground nuclear waste storage and the performance of engineered bentonite barriers. The present status of the EBS Task Force is finalizing Task 9, which follows years of modeling studies of the FEBEX test, by many notable modeling teams (Gens et al., 2009; Sanchez et al. 2010; 2012; Samper et al., 2018). These modeling studies generally use two-dimensional axisymmetric meshes, ignoring threedimensional effects, gravity and asymmetric wetting and dry out of the bentonite engineered barrier. This study investigates these effects with use of the PFLOTRAN THC code with massively parallel computational methods in modeling FEBEX Stage 1 and Stage 2 results. The PFLOTRAN numerical code is an open source, state-of-the-art, massively parallel subsurface flow and reactive transport code operating in a high-performance computing environment (Hammond et al., 2014). Section 2 describes the applied partial differential equations describing mass, momentum and energy balance used in this study, considerations derived by assuming phase equilibrium between gas and liquid phases, constitutive equations for granite, cement plug, and bentonite domains, and specific approaches for use inthe PFLOTRAN code. Section 3 describes the geometry, meshing, and model set-up. Section 4 describes modeling results, Section 5 compares modeling results to field testing data, and Section 6 gives conclusions. The Appendix provides detailed information required by the EBSTask Force for final reporting.

42 ENGINEERING↗