Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “processors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Peripheral processors for high-speed simulation

This paper describes some of the results of a study directed to the specification and procurement of a new cockpit simulator for an advanced class of helicopters. A part of the study was the definition of a challenging benchmark problem, and detailed analyses of it were made to assess the suitability of a variety of simulation techniques. The analyses showed that a particularly cost-effective approach to the attainment of adequate speed for this extremely demanding application is to employ a large minicomputer acting as host and controller for a special-purpose digital peripheral processor. Various realizations of such peripheral processors, all employing state-of-the-art electronic circuitry and a high degree of parallelism and pipelining, are available or under development. The types of peripheral processors array processors, simulation-oriented processors, and arrays of processing elements - are analyzed and compared. They are particularly promising approaches which should be suitable for high-speed simulations of all kinds, the cockpit simulator being a case in point.

Karplus, W. J.↗

On nonlinear finite element analysis in single-, multi- and parallel-processors

Numerical solution of nonlinear equilibrium problems of structures by means of Newton-Raphson type iterations is reviewed. Each step of the iteration is shown to correspond to the solution of a linear problem, therefore the feasibility of the finite element method for nonlinear analysis is established. Organization and flow of data for various types of digital computers, such as single-processor/single-level memory, single-processor/two-level-memory, vector-processor/two-level-memory, and parallel-processors, with and without sub-structuring (i.e. partitioning) are given. The effect of the relative costs of computation, memory and data transfer on substructuring is shown. The idea of assigning comparable size substructures to parallel processors is exploited. Under Cholesky type factorization schemes, the efficiency of parallel processing is shown to decrease due to the occasional shared data, just as that due to the shared facilities.

Utku, S.↗

Acoustooptic linear algebra processors - Architectures, algorithms, and applications

Architectures, algorithms, and applications for systolic processors are described with attention to the realization of parallel algorithms on various optical systolic array processors. Systolic processors for matrices with special structure and matrices of general structure, and the realization of matrix-vector, matrix-matrix, and triple-matrix products and such architectures are described. Parallel algorithms for direct and indirect solutions to systems of linear algebraic equations and their implementation on optical systolic processors are detailed with attention to the pipelining and flow of data and operations. Parallel algorithms and their optical realization for LU and QR matrix decomposition are specifically detailed. These represent the fundamental operations necessary in the implementation of least squares, eigenvalue, and SVD solutions. Specific applications (e.g., the solution of partial differential equations, adaptive noise cancellation, and optimal control) are described to typify the use of matrix processors in modern advanced signal processing.

Casasent, D.↗

The CSM testbed matrix processors internal logic and dataflow descriptions

This report constitutes the final report for subtask 1 of Task 5 of NASA Contract NAS1-18444, Computational Structural Mechanics (CSM) Research. This report contains a detailed description of the coded workings of selected CSM Testbed matrix processors (i.e., TOPO, K, INV, SSOL) and of the arithmetic utility processor AUS. These processors and the current sparse matrix data structures are studied and documented. Items examined include: details of the data structures, interdependence of data structures, data-blocking logic in the data structures, processor data flow and architecture, and processor algorithmic logic flow.

Regelbrugge, Marc E.↗

The computational structural mechanics testbed generic structural-element processor manual

The usage and development of structural finite element processors based on the CSM Testbed's Generic Element Processor (GEP) template is documented. By convention, such processors have names of the form ESi, where i is an integer. This manual is therefore intended for both Testbed users who wish to invoke ES processors during the course of a structural analysis, and Testbed developers who wish to construct new element processors (or modify existing ones).

Stanley, Gary M.↗

Fault tolerant, radiation hard, high performance digital signal processor

An architecture has been developed for a high-performance VLSI digital signal processor that is highly reliable, fault-tolerant, and radiation-hard. The signal processor, part of a spacecraft receiver designed to support uplink radio science experiments at the outer planets, organizes the connections between redundant arithmetic resources, register files, and memory through a shuffle exchange communication network. The configuration of the network and the state of the processor resources are all under microprogram control, which both maps the resources according to algorithmic needs and reconfigures the processing should a failure occur. In addition, the microprogram is reloadable through the uplink to accommodate changes in the science objectives throughout the course of the mission. The processor will be implemented with silicon compiler tools, and its design will be verified through silicon compilation simulation at all levels from the resources to full functionality. By blending reconfiguration with redundancy the processor implementation is fault-tolerant and reliable, and possesses the long expected lifetime needed for a spacecraft mission to the outer planets.

Holmann, Edgar↗

DMS processor evolution study

Increasing the processor performance and capability has not only become a wish list item for the Space Station Freedom (SSF) Data Management System (DMS), but also a necessity. There are many commercially available processors which have superior performance compared to the 386, but we cannot only consider performance when selecting a processor for the DMS. The processor is the 'foundation' of the DMS. It will affect the local bus, system bus, interface unit, global network, etc., that have been selected for the DMS. Besides, once the processor instruction set architecture (ISA) is selected, all of the important system software will be implemented based on this ISA. Selecting an ISA that has a strong commercial support and can be upgraded in the 30-year life cycle of the Space Station Freedom is one of the most important items for the DMS.

Liu, Yuan-Kwei↗

Towards the formal specification of the requirements and design of a processor interface unit

Work to formally specify the requirements and design of a Processor Interface Unit (PIU), a single-chip subsystem providing memory interface, bus interface, and additional support services for a commercial microprocessor within a fault-tolerant computer system, is described. This system, the Fault-Tolerant Embedded Processor (FTEP), is targeted towards applications in avionics and space requiring extremely high levels of mission reliability, extended maintenance free operation, or both. The approaches that were developed for modeling the PIU requirements and for composition of the PIU subcomponents at high levels of abstraction are described. These approaches were used to specify and verify a nontrivial subset of the PIU behavior. The PIU specification in Higher Order Logic (HOL) is documented in a companion NASA contractor report entitled 'Towards the Formal Specification of the Requirements and Design of a Processor Interfacs Unit - HOL Listings.' The subsequent verification approach and HOL listings are documented in NASA contractor report entitled 'Towards the Formal Verification of the Requirements and Design of a Processor Interface Unit' and NASA contractor report entitled 'Towards the Formal Verification of the Requirements and Design of a Processor Interface Unit - HOL Listings.'

Fura, David A.↗

Modeling heterogeneous processor scheduling for real time systems

A new model is presented to describe dataflow algorithms implemented in a multiprocessing system. Called the resource/data flow graph (RDFG), the model explicitly represents cyclo-static processor schedules as circuits of processor arcs which reflect the order that processors execute graph nodes. The model also allows the guarantee of meeting hard real-time deadlines. When unfolded, the model identifies statically the processor schedule. The model therefore is useful for determining the throughput and latency of systems with heterogeneous processors. The applicability of the model is demonstrated using a space surveillance algorithm.

Leathrum, J. F.↗

ELIPS: Toward a Sensor Fusion Processor on a Chip

The paper presents the concept and initial tests from the hardware implementation of a low-power, high-speed reconfigurable sensor fusion processor. The Extended Logic Intelligent Processing System (ELIPS) processor is developed to seamlessly combine rule-based systems, fuzzy logic, and neural networks to achieve parallel fusion of sensor in compact low power VLSI. The first demonstration of the ELIPS concept targets interceptor functionality; other applications, mainly in robotics and autonomous systems are considered for the future. The main assumption behind ELIPS is that fuzzy, rule-based and neural forms of computation can serve as the main primitives of an "intelligent" processor. Thus, in the same way classic processors are designed to optimize the hardware implementation of a set of fundamental operations, ELIPS is developed as an efficient implementation of computational intelligence primitives, and relies on a set of fuzzy set, fuzzy inference and neural modules, built in programmable analog hardware. The hardware programmability allows the processor to reconfigure into different machines, taking the most efficient hardware implementation during each phase of information processing. Following software demonstrations on several interceptor data, three important ELIPS building blocks (a fuzzy set preprocessor, a rule-based fuzzy system and a neural network) have been fabricated in analog VLSI hardware and demonstrated microsecond-processing times.

Daud, Taher↗

Multi-Core Processor Memory Contention Benchmark Analysis Case Study

Multi-core processors dominate current mainframe, server, and high performance computing (HPC) systems. This paper provides synthetic kernel and natural benchmark results from an HPC system at the NASA Goddard Space Flight Center that illustrate the performance impacts of multi-core (dual- and quad-core) vs. single core processor systems. Analysis of processor design, application source code, and synthetic and natural test results all indicate that multi-core processors can suffer from significant memory subsystem contention compared to similar single-core processors.

Simon, Tyler↗

Multiple Embedded Processors for Fault-Tolerant Computing

A fault-tolerant computer architecture has been conceived in an effort to reduce vulnerability to single-event upsets (spurious bit flips caused by impingement of energetic ionizing particles or photons). As in some prior fault-tolerant architectures, the redundancy needed for fault tolerance is obtained by use of multiple processors in one computer. Unlike prior architectures, the multiple processors are embedded in a single field-programmable gate array (FPGA). What makes this new approach practical is the recent commercial availability of FPGAs that are capable of having multiple embedded processors. A working prototype (see figure) consists of two embedded IBM PowerPC 405 processor cores and a comparator built on a Xilinx Virtex-II Pro FPGA. This relatively simple instantiation of the architecture implements an error-detection scheme. A planned future version, incorporating four processors and two comparators, would correct some errors in addition to detecting them.

Bolotin, Gary↗

Atmospheric Correction Inter-comparison eXercise, ACIX-II Land: An Assessment of Amospheric Correction Processors for Landsat 8 and Sentinel-2 Over Land

The correction of the atmospheric effects on optical satellite images is essential for quantitative and multi-temporal remote sensing applications. In order to study the performance of the state-of-the-art methods in an integrated way, a voluntary and open-access benchmark Atmospheric Correction Inter-comparison eXercise (ACIX) was initiated in 2016 in the frame of Committee on Earth Observation Satellites (CEOS) Working Group on Calibration & Validation (WGCV). The first exercise was extended in a second edition wherein twelve atmospheric correction (AC) processors, a substantially larger testing dataset and additional validation metrics were involved. The sites for the inter-comparison analysis were defined by investigating the full catalogue of the Aerosol Robotic Network (AERONET) sites for coincident measurements with satellites' overpass. Although there were more than one hundred sites for Copernicus Sentinel-2 and Landsat 8 acquisitions, the analysis presented in this paper concerns only the common matchups amongst all processors, reducing the number to 79 and 62 sites respectively. Aerosol Optical Depth (AOD) and Water Vapour (WV) retrievals were consequently validated based on the available AERONET observations. The processors mostly succeeded in retrieving AOD for relatively light to medium aerosol loading (AOD < 0.2) with uncertainties <0.08, while the overall uncertainty values were typically 0.23 ± 0.15. Better performances were observed for WV retrievals with >90% of the results falling within the suggested empirical specifications and with the Root Mean Square Error (RMSE) being mostly <0.25 g/cm2. Regarding Surface Reflectance (SR) validation two main approaches were followed. For the first one, a simulated SR reference dataset was computed over all of the test sites by using the 6SV (Second Simulation of the Satellite Signal in the Solar Spectrum vector code) full radiative transfer modelling (RTM) and AERONET measurements for the required aerosol variables and water vapour content. The performance assessment demonstrated that the retrievals were not biased for most of the bands. The uncertainties ranged from approximately 0.003 to 0.01 (excluding B01) for the best performing processors in both sensors' analyses. For the second one, measurements from the radiometric calibration network RadCalNet over La Crau (France) and Gobabeb (Namibia) were involved in the validation. The performance of the processors was in general consistent across all bands for both sensors and with low standard deviations (<0.04) between on-site and estimated surface reflectance. Overall, our study provides a good insight of AC algorithms' performance to developers and users, pointing out similarities and differences for AOD, WV and SR retrievals. Such validation though still lacks of ground-based measurements of known uncertainty to better assess and characterize the uncertainties in SR retrievals.

Atmospheric correction↗

An Efficient Solution Method for Multibody Systems with Loops Using Multiple Processors

This paper describes a multibody dynamics algorithm formulated for parallel implementation on multiprocessor computing platforms using the divide-and-conquer approach. The system of interest is a general topology of rigid and elastic articulated bodies with or without loops. The algorithm divides the multibody system into a number of smaller sets of bodies in chain or tree structures, called "branches" at convenient joints called "connection points", and uses an Order-N (O (N)) approach to formulate the dynamics of each branch in terms of the unknown spatial connection forces. The equations of motion for the branches, leaving the connection forces as unknowns, are implemented in separate processors in parallel for computational efficiency, and the equations for all the unknown connection forces are synthesized and solved in one or several processors. The performances of two implementations of this divide-and-conquer algorithm in multiple processors are compared with an existing method implemented on a single processor.

Multibody dynamics↗

Mitigating cosmic-ray-like correlated events with a modular quantum processor

Quantum processors based on superconducting qubits are being scaled to larger qubit numbers, enabling the implementation of small-scale quantum error-correction codes. However, catastrophic chip-scale correlated errors have been observed in these processors, attributed to, e.g., cosmic ray impacts, which challenge conventional error-correction codes such as the surface code. These events are characterized by a temporary but pronounced suppression of the qubit-energy relaxation times. Here, in this study, we explore the potential for modular quantum computing architectures to mitigate such correlated energy decay events. We measure cosmic-ray-like events in a quantum processor comprising a motherboard and two flip-chip bonded daughterboard modules, each module containing two superconducting qubits. We monitor the appearance of correlated qubit decay events within a single module and across the physically separated modules. We find that while decay events within one module are strongly correlated (over 85%), events in separate modules only display approximately 2% correlations. We also report coincident decay events in the motherboard and in either of the two daughterboard modules, providing further insight into the nature of these decay events. These results suggest that modular architectures, combined with bespoke errorcorrection codes, offer a promising approach for protecting future quantum processors from chip-scale correlated errors.

Wu, Xuntao [Univ. of Chicago, IL (United States)] ↗

Efficient frequency allocation for superconducting quantum processors using improved optimization techniques

Building on previous research on frequency allocation optimization for superconducting circuit quantum processors, this work incorporates several techniques to improve overall solution quality. Here, we introduce constraints and imposed edgewise differences help to improve the optimization results. We also introduce optimization variables for the orientation of each edge, defined as the direction from the control qubit to the target qubit, to be chosen during optimization. To scale up to larger processors, multimodule designs are employed with various boundary conditions, thereby enhancing the collective yield. These enhancements allow for greater flexibility in processor design by eliminating the need for handpicked orientations. We support the efficient assembly of large processors with dense connectivity by choosing the best boundary conditions. Examples demonstrate that, at low computational cost, this optimization approach finds a frequency configuration for a square chip with over 1000 qubits and over 10% yield at much larger dispersion levels than required by previous approaches.

Zhang, Zewen [Argonne National Laboratory (ANL), A↗

Quantum Computer-Aided Design: Digital Quantum Simulation of Quantum Processors

With the increasing size of quantum processors, submodules that constitute the processor hardware will become too large to accurately simulate on a classical computer. Therefore, one would soon have to fabricate and test each new design primitive and parameter choice in time-consuming coordination between design, fabrication, and experimental validation. Here we show how one can design and test the performance of next-generation quantum hardware—by using existing quantum computers. Focusing on superconducting transmon processors as a prominent hardware platform, we compute the static and dynamic properties of individual and coupled transmons. We show how the energy spectra of transmons can be obtained by variational hybrid quantum-classical algorithms that are well suited for near-term noisy quantum computers. In addition, single- and two-qubit gate simulations are demonstrated via Suzuki-Trotter decomposition. Our methods pave a promising way towards designing candidate quantum processors when the demands of calculating submodule properties exceed the capabilities of classical computing resources.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Mitigation of Cosmic Rays-Induced Errors in Superconducting Quantum Processors

Environmental radioactivity and cosmic-rays have recently been identified as a source of decoherence in super-conducting quantum bits (qubits). In particular, the absorption of cosmic-ray muons and gamma rays emitted by naturally occurring radioactive isotopes in the qubit substrate leads to correlated errors in superconducting quantum processors, posing significant challenges to quantum error correction. To enable quantum computing to scale, it is therefore necessary the devel-opment of mitigation strategies to prevent, or keep under control, error bursts due to particle impacts in the chip. While most environmental radioactive sources can be effectively suppressed using dedicated shielding, cosmic-ray muons, with their high penetration capability, can only be mitigated by moving the entire facility in a deep underground laboratory. This work explores the potential for developing a novel class of quantum processors equipped with an active veto system to protect superconducting-based quantum computers from the detrimental effects of atmospheric muons. Such a device would enable the identification of an atmospheric muon interaction within the processor and veto all operations performed during the occurrence of such an interaction. By demonstrating high detection efficiency and negligible dead time, we aim to establish that the future of quantum processors can be envisioned in above-around facilities.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗