Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “program processors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Programable Pipelined-Image Processor

Computer serves as pipelined processor for imagery or other two-dimensional digital data. Processor does feature extraction, smoothing, edge detection, texture measurement, and stereoscoptic area correlation. Also plans routes for obstacle avoidance by robots and solves two-dimensional partial differential equations. Image processor consists of modular units: each includes set of computing elements of types particularly useful in pipelined-image processing. Flexible interconnection scheme used to route data to subsequent stages of pipeline.

Gennery, D. B.↗

Managing Complexity in Next Generation Robotic Spacecraft: From a Software Perspective

This presentation highlights the challenges in the design of software to support robotic spacecraft. Robotic spacecraft offer a higher degree of autonomy, however currently more capabilities are required, primarily in the software, while providing the same or higher degree of reliability. The complexity of designing such an autonomous system is great, particularly while attempting to address the needs for increased capabilities and high reliability without increased needs for time or money. The efforts to develop programming models for the new hardware and the integration of software architecture are highlighted.

Multi - threaded programming↗

Transient Finite Element Computations on a Variable Transputer System

A parallel program to analyze transient finite element problems was written and implemented on a system of transputer processors. The program uses the explicit time integration algorithm which eliminates the need for equation solving, making it more suitable for parallel computations. An interprocessor communication scheme was developed for arbitrary two dimensional grid processor configurations. Several 3-D problems were analyzed on a system with a small number of processors.

Smolinski, Patrick J.↗

Measuring the capabilities of quantum computers

Quantum computers can now run interesting programs, but each processor’s capability—the set of programs that it can run successfully—is limited by hardware errors. These errors can be complicated, making it difficult to accurately predict a processor’s capability. Benchmarks can be used to measure capability directly, but current benchmarks have limited flexibility and scale poorly to many-qubit processors. We show how to construct scalable, efficiently verifiable benchmarks based on any program by using a technique that we call circuit mirroring. With it, we construct two flexible, scalable volumetric benchmarks based on randomized and periodically ordered programs. We use these benchmarks to map out the capabilities of twelve publicly available processors, and to measure the impact of program structure on each one. We find that standard error metrics are poor predictors of whether a program will run successfully on today’s hardware, and that current processors vary widely in their sensitivity to program structure.

97 MATHEMATICS AND COMPUTING↗

A programmable computer interface for CAMAC

An interface has been developed for CAMAC instrumentation systems that implements data transfers controlled either by the computer CPU or by an autonomous (data-channel) processor in the interface unit. The data channel processor executes programs stored in the computer memory. These programs consist of standard CAMAC module commands plus special control characters and commands for the processor itself. The interface was built for the PDP-15 computer, which has an 18-bit word structure, but both 18- and 24-bit data transfers can be made. A software system has been written that exploits the many features of the processor.

Bercaw, R. W.↗

A scan processor as an aid to program documentation

Program documentation was separated into two catagories: documentation for program use and documentation for program analysis, modification, or extension. The symbol scanner, flow chart production subsystem, decision table subsystem, and scan processor are also described.

Oliver, P.↗

SPAR reference manual

The functions and operating rules of the SPAR system, which is a group of computer programs used primarily to perform stress, buckling, and vibrational analyses of linear finite element systems, were given. The following subject areas were discussed: basic information, structure definition, format system matrix processors, utility programs, static solutions, stresses, sparse matrix eigensolver, dynamic response, graphics, and substructure processors.

Whetstone, W. D.↗

System support software for the Space Ultrareliable Modular Computer (SUMC)

The highly transportable programming system designed and implemented to support the development of software for the Space Ultrareliable Modular Computer (SUMC) is described. The SUMC system support software consists of program modules called processors. The initial set of processors consists of the supervisor, the general purpose assembler for SUMC instruction and microcode input, linkage editors, an instruction level simulator, a microcode grid print processor, and user oriented utility programs. A FORTRAN 4 compiler is undergoing development. The design facilitates the addition of new processors with a minimum effort and provides the user quasi host independence on the ground based operational software development computer. Additional capability is provided to accommodate variations in the SUMC architecture without consequent major modifications in the initial processors.

Hill, T. E.↗

Techniques for recovering from errors when executing software applications on parallel processors

In various embodiments, a software program uses hardware features of a parallel processor to checkpoint a context associated with an execution of a software application on the parallel processor. The software program uses a preemption feature of the parallel processor to cause the parallel processor to stop executing instructions in accordance with the context. The software program then causes the parallel processor to collect state data associated with the context. After generating a checkpoint based on the state data, the software program causes the parallel processor to resume executing instructions in accordance with the context.

Hukerikar, Saurabh↗

Modula-2*: An extension of Modula-2 for highly parallel programs

Parallel programs should be machine-independent, i.e., independent of properties that are likely to differ from one parallel computer to the next. Extensions are described of Modula-2 for writing highly parallel, portable programs meeting these requirements. The extensions are: synchronous and asynchronous forms of forall statement; and control of the allocation of data to processors. Sample programs written with the extensions demonstrate the clarity of parallel programs when machine-dependent details are omitted. The principles of efficiently implementing the extensions on SIMD, MIMD, and MSIMD machines are discussed. The extensions are small enough to be integrated easily into other imperative languages.

Tichy, Walter F.↗

What Multilevel Parallel Programs do when you are not Watching: A Performance Analysis Case Study Comparing MPI/OpenMP, MLP, and Nested OpenMP

With the current trend in parallel computer architectures towards clusters of shared memory symmetric multi-processors, parallel programming techniques have evolved that support parallelism beyond a single level. When comparing the performance of applications based on different programming paradigms, it is important to differentiate between the influence of the programming model itself and other factors, such as implementation specific behavior of the operating system (OS) or architectural issues. Rewriting-a large scientific application in order to employ a new programming paradigms is usually a time consuming and error prone task. Before embarking on such an endeavor it is important to determine that there is really a gain that would not be possible with the current implementation. A detailed performance analysis is crucial to clarify these issues. The multilevel programming paradigms considered in this study are hybrid MPI/OpenMP, MLP, and nested OpenMP. The hybrid MPI/OpenMP approach is based on using MPI [7] for the coarse grained parallelization and OpenMP [9] for fine grained loop level parallelism. The MPI programming paradigm assumes a private address space for each process. Data is transferred by explicitly exchanging messages via calls to the MPI library. This model was originally designed for distributed memory architectures but is also suitable for shared memory systems. The second paradigm under consideration is MLP which was developed by Taft. The approach is similar to MPi/OpenMP, using a mix of coarse grain process level parallelization and loop level OpenMP parallelization. As it is the case with MPI, a private address space is assumed for each process. The MLP approach was developed for ccNUMA architectures and explicitly takes advantage of the availability of shared memory. A shared memory arena which is accessible by all processes is required. Communication is done by reading from and writing to the shared memory.

Jost, Gabriele↗

A package for 3-D unstructured grid generation, finite-element flow solution and flow field visualization

A set of computer programs for 3-D unstructured grid generation, fluid flow calculations, and flow field visualization was developed. The grid generation program, called VGRID3D, generates grids over complex configurations using the advancing front method. In this method, the point and element generation is accomplished simultaneously, VPLOT3D is an interactive, menudriven pre- and post-processor graphics program for interpolation and display of unstructured grid data. The flow solver, VFLOW3D, is an Euler equation solver based on an explicit, two-step, Taylor-Galerkin algorithm which uses the Flux Corrected Transport (FCT) concept for a wriggle-free solution. Using these programs, increasingly complex 3-D configurations of interest to aerospace community were gridded including a complete Space Transportation System comprised of the space-shuttle orbitor, the solid-rocket boosters, and the external tank. Flow solutions were obtained on various configurations in subsonic, transonic, and supersonic flow regimes.

Parikh, Paresh↗

Real-time SAR image processing onboard a Venus orbiting spacecraft

The potential use of real-time SAR processing to produce 200-meter resolution imagery onboard a 1983 Venus Orbiter Imaging Radar (VOIR) spacecraft is discussed. The current NASA SAR processor development program and its relationship to the VOIR application are described. VOIR SAR processing requirements are defined in terms of a nominal baseline design evolving from a 1977 VOIR mission study by JPL. A candidate onboard SAR processor architecture compatible with the VOIR requirements is described. Detailed implementation characteristics, based on currently available integrated circuits, are estimated in terms of chip count, weight, and power.

Arens, W. E.↗

Characterization of Quantum Frequency Processors

Frequency-bin qubits possess unique synergies with wavelength-multiplexed lightwave communications, suggesting valuable opportunities for quantum networking with the existing fiber-optic infrastructure. Although the coherent manipulation of frequency-bin states requires highly controllable multi-spectral-mode interference, the quantum frequency processor (QFP) provides a scalable path for gate synthesis leveraging standard telecom components. Here, we summarize the state of the art in experimental QFP characterization. Distinguishing between physically motivated “open box” approaches that treat the QFP as a multiport interferometer, and “black box” approaches that view the QFP as a general quantum operation, we highlight the assumptions and results of multiple techniques, including quantum process tomography of a tunable beamsplitter—to our knowledge the first full process tomography of any frequency-bin operation. Our findings should inform future characterization efforts as the QFP increasingly moves beyond proof-of-principle tabletop demonstrations toward integrated devices and deployed quantum networking experiments.

42 ENGINEERING↗

Automatic Generation of High-Performance Convolution Kernels on ARM CPUs for Deep Learning

In this work, we present FastConv, a template-based code auto-generation open source library that can automatically generate high-performance deep learning convolution kernels of arbitrary matrices/tensors shapes. FastConv is based on the Winograd algorithm, which is reportedly the highest performing algorithm for the time-consuming convolution layers of convolutional neural networks. ARM CPUs cover a wide range designs and specifications, from embedded devices to HPC-grade CPUs. The leads to the dilemma of how to consistently optimize Winograd-based convolution solvers for convolution layers of different shapes. FastConv addresses this problem by using templates to auto-generate multiple shapes of tuned kernels variants suitable for skinny tall matrices. As a performance portable library, FastConv transparently searches for the best combination of kernel shapes, cache tiles, scheduling of loop orders, packing strategies, access patterns, and online/offline computations. Auto-tuning is used to search the parameter configuration space for the best performance for a given target architecture and problem size. The experiments with layer-wise evaluation on the VGG--16 model confirms a 1.25x performance gains is got by tuning the Winograd library. Integrated comparison results shows 1.02x to 1.40x, 1.14x to 2.17x, and 1.22x and 2.48x speedup is achieved over NNPACK, Arm NN, and FeatherCNN on the Kunpeng 920 beside few cases. Furthermore, problem size performance portability experiments with various convolution shapes shows that FastConv achieves 1.2x to 1.7x speedup and 2x to 22x speedup over NNPACK and ARM NN inference engine using Winograd on Kunpeng 920 . CPU performance portability evaluation on the VGG--16 show an average speedup over NNPACK of 1.42x, 1.21x, 1.26x, 1.37x, 2.26x, and 11.02x is observed on Kunpeng 920, Snapdragon 835, 855, 888, Apple M1, and AWS Graviton2, respectively.

97 MATHEMATICS AND COMPUTING↗

Formulation of consumables management models: Test plan for the mission planning processor working model

The test plan and test procedures to be used in the verification and validation of the software being implemented in the mission planning processor working model program are documented. The mission planning processor is a user oriented tool for consumables management and is part of the total consumables subsystem management concept. An overview of the working model is presented. Execution of the test plan will comprehensively exercise the working model software. An overview of the test plan, including a testing schedule, is presented along with the test plan for the unit, module, and system levels. The criteria used to validate the working model results for each consumables subsystem is discussed.

Connelly, L. C.↗