Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “concurrent computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

PythonFOAM: In-situ data analyses with OpenFOAM and Python

Here, we outline the development of a general-purpose Python-based data analysis tool for OpenFOAM. Our implementation relies on the construction of OpenFOAM applications that have bindings to data analysis libraries in Python. Double precision data in OpenFOAM is cast to a NumPy array using the NumPy C-API and Python modules may then be used for arbitrary data analysis and manipulation on flow-field information. We highlight how the proposed wrapper may be used for an in-situ online singular value decomposition (SVD) implemented in Python and accessed from the OpenFOAM solver PimpleFOAM. Here, 'in-situ' refers to a programming paradigm that allows for a concurrent computation of the data analysis on the same computational resources utilized for the partial differential equation solver. In addition, to demonstrate parallel deployments, we deploy a distributed SVD, which collects snapshot data across the ranks of a distributed simulation to compute the global left singular vectors. Crucially, both OpenFOAM and Python share the same message passing interface (MPI) communicator for this deployment which allows Python objects and functions to exchange NumPy arrays across ranks. Subsequently, we provide scaling assessments of this distributed SVD on multiple nodes of Intel Broadwell and KNL architectures for canonical test cases such as the large eddy simulations of a backward facing step and a channel flow at friction Reynolds number of 395. Finally, we demonstrate the deployment of a deep neural network for compressing the flow-field information using an autoencoder to demonstrate an ability to use state-of-the-art machine learning tools in the Python ecosystem.

97 MATHEMATICS AND COMPUTING↗

NekRS, a GPU-accelerated spectral element Navier–Stokes solver

The development of NekRS, a GPU-oriented thermal-fluids simulation code based on the spectral element method (SEM) is described. For performance portability, the code is based on the open concurrent compute abstraction and leverages scalable developments in the SEM code Nek5000 and in libParanumal, which is a library of high-performance kernels for high-order discretizations and PDE-based miniapps. Critical performance sections of the Navier–Stokes time advancement are addressed. Performance results on several platforms are presented here, including scaling to 27,648 V100s on OLCF Summit, for calculations of up to 60B gridpoints.

97 MATHEMATICS AND COMPUTING↗

Structural Simluation Toolkit (SST) v.12.0

The Structural Simulation Toolkit (SST) was developed to explore innovations in highly concurrent computing systems where the instruction set architecture (ISA), micro-architecture, and memory interact with the programming model and communications system. The package provides a fully modular design for extensive exploration of an individual system parameter as well as a parallel simulation environment based on message passing interface (MPI) which enable a high level of performance as well as the ability to look at large systems. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Rodrigues, ArunF.↗

A solar wind turbulence event during the Voyager 1978 solar conjunction profiled via new DSN radio science

A radio science data capability within the DSN Tracking System is described. This capability consists of routine provision of phase fluctuation data concurrently computed over several different time scales. This capability was used to observe phase fluctuation spectral characteristics during a rapid increase in solar wind turbulence that occurred during a July 23, 1978 track of the Voyager 1 spacecraft by Deep Space Station 11. It is suggested that the capability will prove useful in studies of variations of solar wind phase fluctuation spectral characteristics with, for instance, parameters such as the solar cycle and radial distance.

Berman, A. L.↗

Design of a verifiable subset for HAL/S

An attempt to evaluate the applicability of program verification techniques to the existing programming language, HAL/S is discussed. HAL/S is a general purpose high level language designed to accommodate the software needs of the NASA Space Shuttle project. A diversity of features for scientific computing, concurrent and real-time programming, and error handling are discussed. The criteria by which features were evaluated for inclusion into the verifiable subset are described. Individual features of HAL/S with respect to these criteria are examined and justification for the omission of various features from the subset is provided. Conclusions drawn from the research are presented along with recommendations made for the use of HAL/S with respect to the area of program verification.

Browne, J. C.↗

Feasibility study for convertible engine torque converter

The feasibility study has shown that a dump/fill type torque converter has excellent potential for the convertible fan/shaft engine. The torque converter space requirement permits internal housing within the normal flow path of a turbofan engine at acceptable engine weight. The unit permits operating the engine in the turboshaft mode by decoupling the fan. To convert to turbofan mode, the torque converter overdrive capability bring the fan speed up to the power turbine speed to permit engagement of a mechanical lockup device when the shaft speed are synchronized. The conversion to turbofan mode can be made without drop of power turbine speed in less than 10 sec. Total thrust delivered to the aircraft by the proprotor, fan, and engine during tansient can be controlled to prevent loss of air speed or altitude. Heat rejection to the oil is low, and additional oil cooling capacity is not required. The turbofan engine aerodynamic design is basically uncompromised by convertibility and allows proper fan design for quiet and efficient cruise operation. Although the results of the feasibility study are exceedingly encouraging, it must be noted that they are based on extrapolation of limited existing data on torque converters. A component test program with three trial torque converter designs and concurrent computer modeling for fluid flow, stress, and dynamics, updated with test results from each unit, is recommended.

Source record↗

Comparing barrier algorithms

A barrier is a method for synchronizing a large number of concurrent computer processes. After considering some basic synchronization mechanisms, a collection of barrier algorithms with either linear or logarithmic depth are presented. A graphical model is described that profiles the execution of the barriers and other parallel programming constructs. This model shows how the interaction between the barrier algorithms and the work that they synchronize can impact their performance. One result is that logarithmic tree structured barriers show good performance when synchronizing fixed length work, while linear self-scheduled barriers show better performance when synchronizing fixed length work with an imbedded critical section. The linear barriers are better able to exploit the process skew associated with critical sections. Timing experiments, performed on an eighteen processor Flex/32 shared memory multiprocessor, that support these conclusions are detailed.

Arenstorf, Norbert S.↗

Concurrent Image Processing Executive (CIPE). Volume 1: Design overview

The design and implementation of a Concurrent Image Processing Executive (CIPE), which is intended to become the support system software for a prototype high performance science analysis workstation are described. The target machine for this software is a JPL/Caltech Mark 3fp Hypercube hosted by either a MASSCOMP 5600 or a Sun-3, Sun-4 workstation; however, the design will accommodate other concurrent machines of similar architecture, i.e., local memory, multiple-instruction-multiple-data (MIMD) machines. The CIPE system provides both a multimode user interface and an applications programmer interface, and has been designed around four loosely coupled modules: user interface, host-resident executive, hypercube-resident executive, and application functions. The loose coupling between modules allows modification of a particular module without significantly affecting the other modules in the system. In order to enhance hypercube memory utilization and to allow expansion of image processing capabilities, a specialized program management method, incremental loading, was devised. To minimize data transfer between host and hypercube, a data management method which distributes, redistributes, and tracks data set information was implemented. The data management also allows data sharing among application programs. The CIPE software architecture provides a flexible environment for scientific analysis of complex remote sensing image data, such as planetary data and imaging spectrometry, utilizing state-of-the-art concurrent computation capabilities.

Lee, Meemong↗

Proceedings of the NASA Conference on Space Telerobotics, volume 2

These proceedings contain papers presented at the NASA Conference on Space Telerobotics held in Pasadena, January 31 to February 2, 1989. The theme of the Conference was man-machine collaboration in space. The Conference provided a forum for researchers and engineers to exchange ideas on the research and development required for application of telerobotics technology to the space systems planned for the 1990s and beyond. The Conference: (1) provided a view of current NASA telerobotic research and development; (2) stimulated technical exchange on man-machine systems, manipulator control, machine sensing, machine intelligence, concurrent computation, and system architectures; and (3) identified important unsolved problems of current interest which can be dealt with by future research.

Rodriguez, Guillermo↗

Proceedings of the NASA Conference on Space Telerobotics, volume 3

The theme of the Conference was man-machine collaboration in space. The Conference provided a forum for researchers and engineers to exchange ideas on the research and development required for application of telerobotics technology to the space systems planned for the 1990s and beyond. The Conference: (1) provided a view of current NASA telerobotic research and development; (2) stimulated technical exchange on man-machine systems, manipulator control, machine sensing, machine intelligence, concurrent computation, and system architectures; and (3) identified important unsolved problems of current interest which can be dealt with by future research.

Rodriguez, Guillermo↗

Comparing barrier algorithms

A barrier is a method for synchronizing a large number of concurrent computer processes. After considering some basic synchronization mechanisms, a collection of barrier algorithms with either linear or logarithmic depth are presented. A graphical model is described that profiles the execution of the barriers and other parallel programming constructs. This model shows how the interaction between the barrier algorithms and the work that they synchronize can impact their performance. One result is that logarithmic tree structured barriers show good performance when synchronizing fixed length work, while linear self-scheduled barriers show better performance when synchronizing fixed length work with an imbedded critical section. The linear barriers are better able to exploit the process skew associated with critical sections. Timing experiments, performed on an eighteen processor Flex/32 shared memory multiprocessor that support these conclusions, are detailed.

Arenstorf, Norbert S.↗

Efficiency of group implicit concurrent algorithms for transient finite element analysis

The performance of group implicit algorithms is assessed on actual concurrent computers. It is shown that, as the number of subdomains is increased, performance enhancements are derived from two sources: the increased parallelism in the computations; and a reduction in equation solving effort. Moreover, these two performance enhancements are synergistic, in the sense that the corresponding speed-ups are multiplied, rather than merely added. Simulations on a 32-node hypercube are presented for which the interprocessor communications efficiencies obtained are consistently in excess of 90 percent.

Ortiz, M.↗

Numerical studies of electron dynamics in oblique quasi-perpendicular collisionless shock waves

Linear and nonlinear electron damping of the whistler precursor wave train to low Mach number quasi-perpendicular oblique shocks is studied using a one-dimensional electromagnetic plasma simulation code with particle electrons and ions. In some parameter regimes, electrons are observed to trap along the magnetic field lines in the potential of the whistler precursor wave train. This trapping can lead to significant electron heating in front of the shock for low beta(e). Use of a 64-processor hypercube concurrent computer has enabled long runs using realistic mass ratios in the full particle in-cell code and thus simulate shock parameter regimes and phenomena not previously studied numerically.

Liewer, P. C.↗

Task Description Language

Task Description Language (TDL) is an extension of the C++ programming language that enables programmers to quickly and easily write complex, concurrent computer programs for controlling real-time autonomous systems, including robots and spacecraft. TDL is based on earlier work (circa 1984 through 1989) on the Task Control Architecture (TCA). TDL provides syntactic support for hierarchical task-level control functions, including task decomposition, synchronization, execution monitoring, and exception handling. A Java-language-based compiler transforms TDL programs into pure C++ code that includes calls to a platform-independent task-control-management (TCM) library. TDL has been used to control and coordinate multiple heterogeneous robots in projects sponsored by NASA and the Defense Advanced Research Projects Agency (DARPA). It has also been used in Brazil to control an autonomous airship and in Canada to control a robotic manipulator.

Simmons, Reid↗

Instabilities in the Wake of Roughness on a Flat Plate in a Quiet Supersonic Tunnel

Roughness-induced transition is an unavoidable reality in practical high-speed vehicles. Typical prediction of transition due to roughness include algebraic correlations and, more recently, semi-empirical methods. In the NASA Langley Research Center Supersonic Low Disturbance Tunnel, a Mach 3.5 quiet tunnel, several transition experiments have been performed in the past decade to better understand the mechanisms by which small roughness causes transition in a supersonic boundary layer. The study started first with isolated roughness elements of different planforms and shapes and progressed to increasingly more complicated geometries before arriving at a pseudorandom roughness, defined by an analytic function. Concurrent computational efforts progressed with these studies as well, starting with the use of linear stability theory and progressing to the use of harmonic linearized Navier-Stokes to predict the growth of boundary layer stabilities.

Amanda Chou↗

Characterization of concurrent processing

Computer architectures designed for concurrent processing are characterized by the number of processing elements, ensemble speed, random access memory, input/output routes, and modes of operation. The important attributes of processing tasks are then identified, and some processing stratagems are examined. It is shown that the greater the complexity of a given task, the wider the range of possible stratagems which can accomplish the task. For relatively simple tasks, the optimum stratagem can be found by analytical reasoning. For more complex tasks, however, optimum scheduling techniques may have to be employed for the assignment of segments of the task to the available processing elements.

Utku, S.↗

Report on the feasibility of hypercube concurrent processing systems in computational fluid dynamics

The feasibility of using hypercube-connected concurrent processor systems for problems in computational fluid dynamics is studied. Both explicit and implicit numerical methods are considered and several alternative implementations of these methods are evaluated on concurrent processor systems. A Lax-Wendroff explicit method was designed and implemented for the Navier-Stokes equations. The code runs on the Intel iPSC concurrent processor system. Tests of this code show that it is reasonably efficient. The Beam and Warming implicit factored method was designed and implemented for Berger's equation. Preliminary tests show that the efficiency of code is poor.

Bruno, J.↗

Reliability models for dataflow computer systems

The demands for concurrent operation within a computer system and the representation of parallelism in programming languages have yielded a new form of program representation known as data flow (DENN 74, DENN 75, TREL 82a). A new model based on data flow principles for parallel computations and parallel computer systems is presented. Necessary conditions for liveness and deadlock freeness in data flow graphs are derived. The data flow graph is used as a model to represent asynchronous concurrent computer architectures including data flow computers.

Kavi, K. M.↗