Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Asynchronous iterations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

A model of asynchronous iterative algorithms for solving large, sparse, linear systems

Solving large, sparse, linear systems of equations is one of the fundamental problems in large scale scientific and engineering computation. A model of a general class of asynchronous, iterative solution methods for linear systems is developed. In the model, the system is solved by creating several cooperating tasks that each compute a portion of the solution vector. This model is then analyzed to determine the expected intertask data transfer and task computational complexity as functions of the number of tasks. Based on the analysis, recommendations for task partitioning are made. These recommendations are a function of the sparseness of the linear system, its structure (i.e., randomly sparse or banded), and dimension.

Reed, D. A.↗

Parallel, iterative solution of sparse linear systems: Models and architectures

A model of a general class of asynchronous, iterative solution methods for linear systems is developed. In the model, the system is solved by creating several cooperating tasks that each compute a portion of the solution vector. A data transfer model predicting both the probability that data must be transferred between two tasks and the amount of data to be transferred is presented. This model is used to derive an execution time model for predicting parallel execution time and an optimal number of tasks given the dimension and sparsity of the coefficient matrix and the costs of computation, synchronization, and communication. The suitability of different parallel architectures for solving randomly sparse linear systems is discussed. Based on the complexity of task scheduling, one parallel architecture, based on a broadcast bus, is presented and analyzed.

Reed, D. A.↗

Parallel, iterative solution of sparse linear systems - Models and architectures

Solving large, sparse, linear systems of equations is a fundamental problem in large scale scientific and engineering computation. A model of a general class of asynchronous, iterative solution methods for linear systems is developed. In the model, the system is solved by creating several cooperating tasks that each compute a portion of the solution vector. A data transfer model predicting both the probability that data must be transferred between two tasks and the amount of data to be transferred is presented. This model is used to derive an execution time model for predicting parallel execution time and an optimal number of tasks given the dimension and sparsity of the coefficient matrix and the costs of computation, synchronization, and communication. The suitability of different parallel architectures for solving randomly sparse linear systems is discussed. Based on the complexity of task scheduling, one parallel architecture, based on a broadcast bus, is presented and analyzed.

Reed, D. A.↗

Multiple grid problems on concurrent-processing computers

Three computer codes were studied which make use of concurrent processing computer architectures in computational fluid dynamics (CFD). The three parallel codes were tested on a two processor multiple-instruction/multiple-data (MIMD) facility at NASA Ames Research Center, and are suggested for efficient parallel computations. The first code is a well-known program which makes use of the Beam and Warming, implicit, approximate factored algorithm. This study demonstrates the parallelism found in a well-known scheme and it achieved speedups exceeding 1.9 on the two processor MIMD test facility. The second code studied made use of an embedded grid scheme which is used to solve problems having complex geometries. The particular application for this study considered an airfoil/flap geometry in an incompressible flow. The scheme eliminates some of the inherent difficulties found in adapting approximate factorization techniques onto MIMD machines and allows the use of chaotic relaxation and asynchronous iteration techniques. The third code studied is an application of overset grids to a supersonic blunt body problem. The code addresses the difficulties encountered when using embedded grids on a compressible, and therefore nonlinear, problem. The complex numerical boundary system associated with overset grids is discussed and several boundary schemes are suggested. A boundary scheme based on the method of characteristics achieved the best results.

Eberhardt, D. S.↗

Asynchronous sequential circuit design using pass transistor iterative logic arrays

The iterative logic array (ILA) is introduced as a new architecture for asynchronous sequential circuits. This is the first ILA architecture for sequential circuits reported in the literature. The ILA architecture produces a very regular circuit structure. Moreover, it is immune to both 1-1 and 0-0 crossovers and is free of hazards. This paper also presents a new critical race free STT state assignment which produces a simple form of design equations that greatly simplifies the ILA realizations.

Liu, M. N.↗

Quasi-perfect FIFO: Synchronous or asynchronous with application in controller design for the UNICON laser memory

The first-in-first-out memory buffer (FIFO), is an elastic digital memory whose main application is in data buffering between devices operating at different rates. Data written into the top is moved autonomously down toward the bottom of the FIFO to the lowest unoccupied location, and data read from the bottom of the FIFO will cause data from the top to move autonomously down toward the bottom. The FIFO is available in MOS LSI asynchronous form with data rate in the 1 MHz region. The FIFO described yields a simple high-speed iterative implementation, either synchronous of asynchronous. Because of this simple iterative structure, the FIFO is expandable in both number of words and bits per word, and it is attractive from the viewpoint of integrated-circuit production. For the synchronous FIFO, a model was built and successfully used in the controller for the UNICON laser memory. For the asynchronous FIFO, a model was built and also successfully used in a high-performance magnetic tape controller.

Lim, R. S.↗

Enabling Communication Between Astronauts and Ground Teams for Space Exploration Missions

Over the last four years, Playbook’s Mission Log has evolved to become an enabling capability for analog missions that simulate deep space, exploration missions with communication transmission latency. Playbook is a planning and execution web-application for mission operations, aggregating multiple sources of information for astronauts to execute the mission in one place: timeline, procedures, chat interface. Playbook’s Mission Log provides a multimedia chat software interface with unique features and functionalities that support asynchronous communication between analog astronauts and ground support teams. This paper describes the iterative design the Mission Log has undergone based on user observations and solicited feedback. Key features include indicators that help users cope with asynchronous communication as well as aids that assist teams coordinate work. Future work and capabilities are outlined, which build upon the increased use of the Mission Log as a communication and coordination tool for space exploration.

communication↗

Saturation: An efficient iteration strategy for symbolic state-space generation

This paper presents a novel algorithm for generating state spaces of asynchronous systems using Multi-valued Decision Diagrams. In contrast to related work, the next-state function of a system is not encoded as a single Boolean function, but as cross-products of integer functions. This permits the application of various iteration strategies to build a system's state space. In particular, this paper introduces a new elegant strategy, called saturation, and implements it in the tool SMART. On top of usually performing several orders of magnitude faster than existing BDD-based state-space generators, the algorithm's required peak memory is often close to the nal memory needed for storing the overall state spaces.

Ciardo, Gianfranco↗

Exploiting parallel computing with limited program changes using a network of microcomputers

Network computing and multiprocessor computers are two discernible trends in parallel processing. The computational behavior of an iterative distributed process in which some subtasks are completed later than others because of an imbalance in computational requirements is of significant interest. The effects of asynchronus processing was studied. A small existing program was converted to perform finite element analysis by distributing substructure analysis over a network of four Apple IIe microcomputers connected to a shared disk, simulating a parallel computer. The substructure analysis uses an iterative, fully stressed, structural resizing procedure. A framework of beams divided into three substructures is used as the finite element model. The effects of asynchronous processing on the convergence of the design variables are determined by not resizing particular substructures on various iterations.

Rogers, J. L., Jr.↗

SIAM Conference on Parallel Processing for Scientific Computing, 4th, Chicago, IL, Dec. 11-13, 1989, Proceedings

Attention is given to such topics as an evaluation of block algorithm variants in LAPACK and presents a large-grain parallel sparse system solver, a multiprocessor method for the solution of the generalized Eigenvalue problem on an interval, and a parallel QR algorithm for iterative subspace methods on the CM2. A discussion of numerical methods includes the topics of asynchronous numerical solutions of PDEs on parallel computers, parallel homotopy curve tracking on a hypercube, and solving Navier-Stokes equations on the Cedar Multi-Cluster system. A section on differential equations includes a discussion of a six-color procedure for the parallel solution of elliptic systems using the finite quadtree structure, data parallel algorithms for the finite element method, and domain decomposition methods in aerodynamics. Topics dealing with massively parallel computing include hypercube vs. 2-dimensional meshes and massively parallel computation of conservation laws. Performance and tools are also discussed.

Dongarra, Jack↗

Performance Optimization Methods for a Memory-Bound, Unstructured-Grid CFD Application on Massively Parallel GPU Platforms

Computational performance of the FUN3D unstructured-grid computational fluid dynamics (CFD) application on massively parallel GPU environments is memory-bound and highly dependent upon efficient reads from and atomic updates to the irregular cell-, edge-, and node-based data structures. In this talk, we present recent efforts into optimizing select performance-critical kernels on NVIDIA Tesla V100 and A100 GPUs and AMD CDNA MI100 GPUs. A novel use of L2 cache residency controls and asynchronous loads into on-chip shared memory are explored on the A100 GPU for the sparse iterative solver, which is dominated by mixed-precision, sparse matrix vector multiplication. Demonstrations show that these methods improve global memory bandwidth utilization by 13.5% on the A100 GPU. Several techniques are also presented that use registers and/or shared memory to facilitate array transposition and aggregation which combine to reduce the frequency and increase the cache efficiency of floating-point atomic updates to the irregular data structures. These methods are demonstrated to improve the kernel throughput by nearly 500% on select kernels on the AMD MI100 over atomic updates directly to global memory. Overall, both V100 and A100 GPUs outperformed the MI100 GPU on kernels dominated by double-precision atomic updates; however, the techniques demonstrated here reduced the performance gap and improved the MI100 performance.

GPU CPU unstructured CFD memory↗

Project Integration Architecture: Distributed Lock Management, Deadlock Detection, and Set Iteration

The migration of the Project Integration Architecture (PIA) to the distributed object environment of the Common Object Request Broker Architecture (CORBA) brings with it the nearly unavoidable requirements of multiaccessor, asynchronous operations. In order to maintain the integrity of data structures in such an environment, it is necessary to provide a locking mechanism capable of protecting the complex operations typical of the PIA architecture. This paper reports on the implementation of a locking mechanism to treat that need. Additionally, the ancillary features necessary to make the distributed lock mechanism work are discussed.

Jones, William Henry↗

A Parallel Particle Swarm Optimization Algorithm Accelerated by Asynchronous Evaluations

A parallel Particle Swarm Optimization (PSO) algorithm is presented. Particle swarm optimization is a fairly recent addition to the family of non-gradient based, probabilistic search algorithms that is based on a simplified social model and is closely tied to swarming theory. Although PSO algorithms present several attractive properties to the designer, they are plagued by high computational cost as measured by elapsed time. One approach to reduce the elapsed time is to make use of coarse-grained parallelization to evaluate the design points. Previous parallel PSO algorithms were mostly implemented in a synchronous manner, where all design points within a design iteration are evaluated before the next iteration is started. This approach leads to poor parallel speedup in cases where a heterogeneous parallel environment is used and/or where the analysis time depends on the design point being analyzed. This paper introduces an asynchronous parallel PSO algorithm that greatly improves the parallel e ciency. The asynchronous algorithm is benchmarked on a cluster assembled of Apple Macintosh G5 desktop computers, using the multi-disciplinary optimization of a typical transport aircraft wing as an example.

Venter, Gerhard↗

Real-Time Science Decisioning During High Tempo-High Intensity Mission Operations and the Role of Analogs

Introduction: NASA’s VIPER mission presents a unique operational paradigm within the history of robotic spaceflight. The proximity of the Moon to the Earth and the terrain elements (surface characteristics, light/shadow dynamics, communication links) of the lunar South Polar landing site create unprecedented operational conditions between these two planetary bodies. Apollo era lunar science and exploration included humans in situ to operate instruments and assimilate observational inputs in real-time. Previous lunar orbital missions have worked to operational timescales, e.g., decisional timelines and communication exchanges, that were weeks in length. Mars rover missions have worked to operational timescales, e.g., decisional timelines and communication exchanges between Mars and Earth, that were hours, days, and weeks in length. In the case of the VIPER mission, our operational decisioning for rover driving and instrument commanding will be compressed to minute-scale timeframes. These operational conditions directly impact the manner and speed with which the VIPER Science Team (VST) is required to synthesize and analyze data and produce timely science-driven decisions throughout surface mission operations. The VST shall provide mission enhancing scientific input to guide rover traverse planning and drill site confirmation and selection throughout surface operations. Further, the VST input will be of vital importance to the mission’s ability to maximize science return and to meet broader NASA objectives for future lunar in-situ resource utilization (ISRU)and exploration activities. The VST co-located in the Mission Science Center (MSC) will be responsive to the tactical operational cadence of the Mission Operations Center (MOC) and will provide further strategic and Long-Term Planning (LTP) guidance to the mission. The VIPER Science Operations & Integration(SO&I)team has developed an architecture that is focused on the infusion of science-decisioning into the operational framework and execution cadence of VIPER. NASA analog research has played a significant role in the construction of the VIPER science operations systems. As an example, the SO&I team has led analog missions that have focused on bringing together expertise in the sciences (natural, applied and social) and in operations in service of learning how to build and hold together interdisciplinary work environments and what tools are needed to support high tempo, high intensity integrated decisioning. These experiences have provided an essential foundation of knowledge to the VIPER team. Those analogs that specifically influenced the VIPER science operations construct were identified through a process of comparative analysis to prioritize those that offered relevance in whole or in part, and those that did not. The analog research output that provided extensibility to the VIPER science operations architecture included remote teams of humans and robots in cooperation (synchronous and asynchronous) with simulated earthbound systems, engineering and science teams, and the integrated assembly of tools that supported scientific analysis and data synthesis and provided infrastructure for the remote testing framework. Analogs which included real-time data monitoring, synthesis, visualization and access in a democratized and operationalized manner were of particular interest to the development of the VIPER MSC toolset both in terms of the technology and the processes used to develop the supporting infrastructure. We anticipate that each subsequent mission to the lunar south pole, whether with robots or humans, will be able to optimize science and exploration return by evolving strategies to infuse real-time collaborative science-decisioning. Furthermore, these efforts will result in a foundation for science operations development in support of human-robotic exploration of deep space and Mars. NASA analogs can continue to provide the opportunity to prepare, test and iterate on the operational concepts and tools that will support these ever-expanding space exploration efforts. Our presentation will include an overview of the VIPER Science Operations & Integration development process and specifics on what aspects of analog research have had a significant impact on our work systems.

D S S Lim↗