Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44

Numerical eigen-spectrum slicing, accurate orthogonal eigen-basis, and mixed-precision eigenvalue refinement using OpenMP data-dependent tasks and accelerator offload

Performing a variety of numerical computations efficiently and, at the same time, in a portable fashion requires both an overarching design followed by a number of implementation strategies. All of these are exemplified below as we present transitioning the PLASMA numerical library from relying on dependence-driven large tasks to achieving utilization of fine grain tasking and offload to hardware accelerators while keeping its core dependence sets: OpenMP source code pragmas and runtime for most system-level functionality and basic low-level numerical kernels provided directly by hardware vendors or open source projects with vendor contributions. We also present new algorithmic methods and their efficient parallel implementations including fine grained tasking for eigen-spectrum slicing and offload for mixed-precision eigenvalue refinement. We provide performance, scaling, and numerical results showing sizable gains over the available solutions from either the open source and vendor-provided packages.

Luszczek, Piotr↗

Efficient quadrature rules for finite element discretizations of nonlocal equations

In this paper we design efficient quadrature rules for finite element discretizations of nonlocal diffusion problems with compactly supported kernel functions. Two of the main challenges in nonlocal modeling and simulations are the prohibitive computational cost and the nontrivial implementation of discretization schemes, especially in three-dimensional settings. In this work we circumvent both challenges by introducing a parametrized mollifying function that improves the regularity of the integrand, utilizing an adaptive integration technique, and exploiting parallelization. We first showthat the “mollified” solution converges to the exact one as the mollifying parameter vanishes, then we illustrate the consistency and accuracy of the proposed method on several two- and three-dimensional test cases. Furthermore, we demonstrate the good scaling properties of the parallel implementation of the adaptive algorithm and we compare the proposed method with recently developed techniques for efficient finite element assembly.

97 MATHEMATICS AND COMPUTING↗

Cold Plasma Measurements

We have continued the simulation campaign in support of our ongoing magnetospheric cold plasma research project. This project aims to develop the next-generation particle instruments to measure the properties of the cold particle populations in the Earth’s magnetosphere. For this purpose, simulations have been performed with a Particle-In-Cell (PIC) code called the Curvilinear PIC (CPIC). The code is formulated in curvilinear geometry and couples the standard PIC algorithm with algorithms for the generation and adaptation of the underlaying computational mesh. It conforms to complex objects like spacecraft and it can place more grid points in regions where higher resolution is needed. The code also features a scalable solver based on the multigrid algorithm and it is fully parallelized via domain decomposition and MPI.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Multiplexed photoinjector optimization for high-repetition-rate free-electron lasers

The multiplexing capabilities of superconducting X-ray free-electron lasers (FELs) have gained much attention in recent years. The demanding requirements for photon properties from multiple undulator lines necessitate more flexible beam manipulation techniques to achieve the goal of “beam on demand”. In this paper, we investigate a multiplexed configuration for the photoinjector of high-repetition-rate FELs that aims to simultaneously provide low-emittance electron beams of different charges. A parallel, multi-objective genetic algorithm is implemented for the photoinjector parameter optimization. The proposed configuration could drastically enhance the flexibility of beam manipulation to improve multiplexing capabilities and realize the full potential of the facility.

47 OTHER INSTRUMENTATION↗

A design study of a signal detection system

A system is described which can aid in the search for radio signals from extraterrestrial sources, or in other applications characterized by low signal-to-noise ratios and very high data rates. The system follows a multichannel (16 million bin) spectrum analyzer, and has critical processing, system control, and memory fuctions. The design includes a moderately rich set of algorithms to be used in parallel to detect signals of unknown form. A multi-threshold approach is used to obtain high and low signal sensitivities. Relatively compact and transportable memory systems are specified.

Healy, T. J.↗

Geometric registration and rectification of spaceborne SAR imagery

This paper describes the development of automated location and geometric rectification techniques for digitally processed synthetic aperture radar (SAR) imagery. A software package has been developed that is capable of determining the absolute location of an image pixel to within 60 m using only the spacecraft ephemeris data and the characteristics of the SAR data collection and processing system. Based on this location capability algorithms have been developed that geometrically rectify the imagery, register it to a common coordinate system and mosaic multiple frames to form extended digital SAR maps. These algorithms have been optimized using parallel processing techniques to minimize the operating time. Test results are given using Seasat SAR data.

Curlander, J. C.↗

ICASE Computer Science Program

The Institute for Computer Applications in Science and Engineering computer science program is discussed in outline form. Information is given on such topics as problem decomposition, algorithm development, programming languages, and parallel architectures.

Source record↗

A class Hierarchical, object-oriented approach to virtual memory management

The Choices family of operating systems exploits class hierarchies and object-oriented programming to facilitate the construction of customized operating systems for shared memory and networked multiprocessors. The software is being used in the Tapestry laboratory to study the performance of algorithms, mechanisms, and policies for parallel systems. Described here are the architectural design and class hierarchy of the Choices virtual memory management system. The software and hardware mechanisms and policies of a virtual memory system implement a memory hierarchy that exploits the trade-off between response times and storage capacities. In Choices, the notion of a memory hierarchy is captured by abstract classes. Concrete subclasses of those abstractions implement a virtual address space, segmentation, paging, physical memory management, secondary storage, and remote (that is, networked) storage. Captured in the notion of a memory hierarchy are classes that represent memory objects. These classes provide a storage mechanism that contains encapsulated data and have methods to read or write the memory object. Each of these classes provides specializations to represent the memory hierarchy.

Russo, Vincent F.↗

Distributed neural control of a hexapod walking vehicle

There has been a long standing interest in the design of controllers for multilegged vehicles. The approach is to apply distributed control to this problem, rather than using parallel computing of a centralized algorithm. Researchers describe a distributed neural network controller for hexapod locomotion which is based on the neural control of locomotion in insects. The model considers the simplified kinematics with two degrees of freedom per leg, but the model includes the static stability constraint. Through simulation, it is demonstrated that this controller can generate a continuous range of statically stable gaits at different speeds by varying a single control parameter. In addition, the controller is extremely robust, and can continue the function even after several of its elements have been disabled. Researchers are building a small hexapod robot whose locomotion will be controlled by this network. Researchers intend to extend their model to the dynamic control of legs with more than two degrees of freedom by using data on the control of multisegmented insect legs. Another immediate application of this neural control approach is also exhibited in biology: the escape reflex. Advanced robots are being equipped with tactile sensing and machine vision so that the sensory inputs to the robot controller are vast and complex. Neural networks are ideal for a lower level safety reflex controller because of their extremely fast response time. The combination of robotics, computer modeling, and neurobiology has been remarkably fruitful, and is likely to lead to deeper insights into the problems of real time sensorimotor control.

Beer, R. D.↗

A discrete decentralized variable structure robotic controller

A decentralized trajectory controller for robotic manipulators is designed and tested using a multiprocessor architecture and a PUMA 560 robot arm. The controller is made up of a nominal model-based component and a correction component based on a variable structure suction control approach. The second control component is designed using bounds on the difference between the used and actual values of the model parameters. Since the continuous manipulator system is digitally controlled along a trajectory, a discretized equivalent model of the manipulator is used to derive the controller. The motivation for decentralized control is that the derived algorithms can be executed in parallel using a distributed, relatively inexpensive, architecture where each joint is assigned a microprocessor. Nonlinear interaction and coupling between joints is treated as a disturbance torque that is estimated and compensated for.

Tumeh, Zuheir S.↗

A natural partitioning scheme for parallel simulation of multibody systems

A parallel partitioning scheme based on physical-coordinate variables is presented to systematically eliminate system constraint forces and yield the equations of motion of multibody dynamics systems in terms of their independent coordinates. Key features of the present scheme include an explicit determination of the independent coordinates, a parallel construction of the null space matrix of the constraint Jacobian matrix, an easy incorporation of the previously developed two-stage staggered solution procedure, and Schur complement based parallel preconditioned conjugate gradient numerical algorithm.

Chiou, J. C.↗

Electrostatic Particle-In-Cell Code For Hypercube Computer

Code simulates two-dimensional motions of plasma particles in self-consistent electrostatic field and externally applied magnetic field developed for execution on Mark IIIfp hypercube computer. Based on generalization of one-dimensional particle-in-cell algorithm applicable to many different parallel computing architectures. Intermediate product of continuing effort to speed particle-in-cell computations by taking advantage of distributed-memory parallel computers like those of hypercube class.

Ferraro, Robert D.↗

Optical computing at NASA Ames Research Center

Optical computing research at NASA Ames Research Center seeks to utilize the capability of analog optical processing, involving free-space propagation between components, to produce natural implementations of algorithms requiring large degrees of parallel computation. Potential applications being investigated include robotic vision, planetary lander guidance, aircraft engine exhaust analysis, analysis of remote sensing satellite multispectral images, control of space structures, and autonomous aircraft inspection.

Reid, Max B.↗

A natural partitioning scheme for parallel simulation of multibody systems

A parallel partitioning scheme based on physical-co-ordinate variables is presented to systematically eliminate system constraint forces and yield the equations of motion of multibody dynamics systems in terms of their independent coordinates. Key features of the present scheme include an explicit determination of the independent coordinates, a parallel construction of the null space matrix of the constraint Jacobian matrix, an easy incorporation of the previously developed two-stage staggered solution procedure and a Schur complement based parallel preconditioned conjugate gradient numerical algorithm.

Chiou, J. C.↗

Real-time dynamics simulation of the Cassini spacecraft using DARTS. Part 1: Functional capabilities and the spatial algebra algorithm

This paper describes the Dynamics Algorithms for Real-Time Simulation (DARTS) real-time hardware-in-the-loop dynamics simulator for the National Aeronautics and Space Administration's Cassini spacecraft. The spacecraft model consists of a central flexible body with a number of articulated rigid-body appendages. The demanding performance requirements from the spacecraft control system require the use of a high fidelity simulator for control system design and testing. The DARTS algorithm provides a new algorithmic and hardware approach to the solution of this hardware-in-the-loop simulation problem. It is based upon the efficient spatial algebra dynamics for flexible multibody systems. A parallel and vectorized version of this algorithm is implemented on a low-cost, multiprocessor computer to meet the simulation timing requirements.

Jain, A.↗

Proceedings of the Peteflops-Systems Operation Working Review (POWR)

This report constitutes the final technical report. Even as Petaflops performance computing is being achieved for a few important applications on the nation's largest massively parallel processing (MPP) systems, the challenges of realizing far greater performance are being investigated by a strong interdisciplinary team of experts from academia, industry, and government. This team is also exploring the extraordinary opportunity such capabilities would provide for critical areas of strategic national importance. Under what has become known informally as the "Petaflops Initiative," key leaders in research across the areas of device technology, parallel systems architecture, applications and algorithms, and systems software have been exploring the implications, requirements, interrelationships, and trade-offs among these research areas through a series of workshops, studies, and projects sponsored by a number of Federal agencies.

Sterling, Thomas↗

The IRAS Galaxy Atlas (IGA)

In 1993 we proposed a project to NASA having the goal of producing a new infrared map of our Galaxy. In particular, we proposed to reprocess the IRAS data taken in the early 1980's using modern image processing algorithms and the large Intel parallel computers of the Center for Advanced Computing Research, (at that time called the Caltech Concurrent Supercomputing Facilities - CCSF). The rationale was simple: what took approximately 100 days on a typical workstation would take less than a day on the multi-processor parallel computers, thus making a high-resolution infrared atlas of the Galaxy feasible.

Prince, Thomas A.↗

Supercomputing 2002: NAS Demo Abstracts

The hyperwall is a new concept in visual supercomputing, conceived and developed by the NAS Exploratory Computing Group. The hyperwall will allow simultaneous and coordinated visualization and interaction of an array of processes, such as a the computations of a parameter study or the parallel evolutions of a genetic algorithm population. Making over 65 million pixels available to the user, the hyperwall will enable and elicit qualitatively new ways of leveraging computers to accomplish science. It is currently still unclear whether we will be able to transport the hyperwall to SC02. The crucial display frame still has not been completed by the metal fabrication shop, although they promised an August delivery. Also, we are still working the fragile node issue, which may require transplantation of the compute nodes from the present 2U cases into 3U cases. This modification will increase the present 3-rack configuration to 5 racks.

Parks, John↗