Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Fault Tolerant Algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Synthesis of Single Qutrit Circuits from Clifford + R Gates

The Clifford + R gate-set is a promising basis for fault-tolerant synthesis of qutrit unitaries. We present an algorithm for approximating an arbitrary single-qutrit unitary with a circuit over the Clifford + R gates. Moreover, we analyze its complexity and obtain the non-Clifford gates cost.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Distributed asynchronous microprocessor architectures in fault tolerant integrated flight systems

The paper discusses the implementation of fault tolerant digital flight control and navigation systems for rotorcraft application. It is shown that in implementing fault tolerance at the systems level using advanced LSI/VLSI technology, aircraft physical layout and flight systems requirements tend to define a system architecture of distributed, asynchronous microprocessors in which fault tolerance can be achieved locally through hardware redundancy and/or globally through application of analytical redundancy. The effects of asynchronism on the execution of dynamic flight software is discussed. It is shown that if the asynchronous microprocessors have knowledge of time, these errors can be significantly reduced through appropiate modifications of the flight software. Finally, the papear extends previous work to show that through the combined use of time referencing and stable flight algorithms, individual microprocessors can be configured to autonomously tolerate intermittent faults.

Dunn, W. R.↗

Evolvable Hardware for Space Applications

This article surveys the research of the Evolvable Systems Group at NASA Ames Research Center. Over the past few years, our group has developed the ability to use evolutionary algorithms in a variety of NASA applications ranging from spacecraft antenna design, fault tolerance for programmable logic chips, atomic force field parameter fitting, analog circuit design, and earth observing satellite scheduling. In some of these applications, evolutionary algorithms match or improve on human performance.

Lohn, Jason↗

Evolvable Hardware for Space Applications

This article surveys the research of the Evolvable Systems Group at NASA Ames Research Center. Over the past few years, our group has developed the ability to use evolutionary algorithms in a variety of NASA applications ranging from spacecraft antenna design, fault tolerance for programmable logic chips, atomic force field parameter fitting, analog circuit design, and earth observing satellite scheduling. In some of these applications, evolutionary algorithms match or improve on human performance.

Lohn, Jason↗

Reliable and Efficient Parallel Processing Algorithms and Architectures for Modern Signal Processing

Least-squares (LS) estimations and spectral decomposition algorithms constitute the heart of modern signal processing and communication problems. Implementations of recursive LS and spectral decomposition algorithms onto parallel processing architectures such as systolic arrays with efficient fault-tolerant schemes are the major concerns of this dissertation. There are four major results in this dissertation. First, we propose the systolic block Householder transformation with application to the recursive least-squares minimization. It is successfully implemented on a systolic array with a two-level pipelined implementation at the vector level as well as at the word level. Second, a real-time algorithm-based concurrent error detection scheme based on the residual method is proposed for the QRD RLS systolic array. The fault diagnosis, order degraded reconfiguration, and performance analysis are also considered. Third, the dynamic range, stability, error detection capability under finite-precision implementation, order degraded performance, and residual estimation under faulty situations for the QRD RLS systolic array are studied in details. Finally, we propose the use of multi-phase systolic algorithms for spectral decomposition based on the QR algorithm. Two systolic architectures, one based on triangular array and another based on rectangular array, are presented for the multiphase operations with fault-tolerant considerations. Eigenvectors and singular vectors can be easily obtained by using the multi-pase operations. Performance issues are also considered.

Liu, Kuojuey Ray↗

Simultaneous estimation of multiple eigenvalues with short-depth quantum circuit on early fault-tolerant quantum computers

We introduce a multi-modal, multi-level quantum complex exponential least squares (MM-QCELS) method to simultaneously estimate multiple eigenvalues of a quantum Hamiltonian on early fault-tolerant quantum computers. Our theoretical analysis demonstrates that the algorithm exhibits Heisenberg-limited scaling in terms of circuit depth and total cost. Notably, the proposed quantum circuit utilizes just one ancilla qubit, and with appropriate initial state conditions, it achieves significantly shorter circuit depths compared to circuits based on quantum phase estimation (QPE). Numerical results suggest that compared to QPE, the circuit depth can be reduced by around two orders of magnitude under several settings for estimating ground-state and excited-state energies of certain quantum systems.

97 MATHEMATICS AND COMPUTING↗

Microprocessor arrays for large scale computation

An important new direction in computer architecture centers around the achievement of very high computational power (capacity, speed and reliability) through the use of tens of thousands of microprocessors, micromemories, and switch modules, all interconnected into a large homogeneous network using one of certain advanced connection schemes. When surrounded and supported by conventional computers and memories, such a machine holds potential for out-performing both conventional and array-based computers of the mid-1980's by one to two orders of magnitude, at least for particular classes of applications amenable to high parallelism, such as aerodynamic simulation. The homogeneous feature of this machine concept also implies size extendability, fault tolerance, and improved flexibility to handle a variety of algorithms of interest. Current work is addressing the design of technologically efficient interconnection configurations and the development of new computation algorithms that are especially efficient for highly parallel computation.

Kautz, W. H.↗

Design and Implementation of Replicated Object Layer

One of the widely used techniques for construction of fault tolerant applications is the replication of resources so that if one copy fails sufficient copies may still remain operational to allow the application to continue to function. This thesis involves the design and implementation of an object oriented framework for replicating data on multiple sites and across different platforms. Our approach, called the Replicated Object Layer (ROL) provides a mechanism for consistent replication of data over dynamic networks. ROL uses the Reliable Multicast Protocol (RMP) as a communication protocol that provides for reliable delivery, serialization and fault tolerance. Besides providing type registration, this layer facilitates distributed atomic transactions on replicated data. A novel algorithm called the RMP Commit Protocol, which commits transactions efficiently in reliable multicast environment is presented. ROL provides recovery procedures to ensure that site and communication failures do not corrupt persistent data, and male the system fault tolerant to network partitions. ROL will facilitate building distributed fault tolerant applications by performing the burdensome details of replica consistency operations, and making it completely transparent to the application.Replicated databases are a major class of applications which could be built on top of ROL.

Koka, Sudhir↗

Generalized Cycle Benchmarking Algorithm for Characterizing Midcircuit Measurements

Midcircuit measurements (MCMs) are crucial ingredients in the development of fault-tolerant quantum computation. While there have been rapid experimental progresses in realizing MCMs, a systematic method for characterizing noisy MCMs is still under exploration. In this work, we develop a cycle benchmarking (CB)-type algorithm to characterize noisy MCMs. The key idea is to use a joint Fourier transform on the classical and quantum registers and then estimate parameters in the Fourier space, analogous to Pauli fidelities used in CB-type algorithms for characterizing the Pauli-noise channel of Clifford gates. Furthermore, we develop a theory of the noise learnability of MCMs, which determines what information can be learned about the noise model (in the presence of state preparation and terminating measurement noise) and what cannot, which shows that all learnable information can be learned using our algorithm. As an application, we show how to use the learned information to test the independence between measurement noise and state-preparation noise in an MCM. Finally, we conduct numerical simulations to illustrate the practical applicability of the algorithm. Similar to other CB-type algorithms, we expect the algorithm to provide a useful toolkit that is of experimental interest. Published by the American Physical Society 2025

Zhang, Zhihan (ORCID:0009000862907691)↗

Analysis of fault-tolerant neurocontrol architectures

The fault-tolerance of analog parallel distributed implementations of a multivariable aircraft neurocontroller is analyzed by simulating weight and neuron failures in a simplified scheme of analog processing based on the functional architecture of the ETANN chip (Electrically Trainable Artificial Neural Network). The neural information processing is found to be only partially distributed throughout the set of weights of the neurocontroller synthesized with the backpropagation algorithm. Although the degree of distribution of the neural processing, and consequently the fault-tolerance of the neurocontroller, could be enhanced using Locally Distributed Weight and Neuron Approaches, a satisfactory level of fault-tolerance could only be obtained by retraining the degrated VLSI neurocontroller. The possibility of maintaining neurocontrol performance and stability in the presence of single weight of neuron failures was demonstrated through an automated retraining procedure of the neurocontroller based on a pre-programmed choice and sequence of the training parameters.

Troudet, T.↗

Nearly optimal state preparation for quantum simulations of lattice gauge theories

Here, we present several improvements to the recently developed ground-state preparation algorithm based on the quantum eigenvalue transformation for unitary matrices (QETU), apply this algorithm to a lattice formulation of U(1) gauge theory in (2+1) dimensions, as well as propose an alternative application of QETU, a highly efficient preparation of Gaussian distributions. The QETU technique was originally proposed as an algorithm for nearly optimal ground-state preparation and ground-state energy estimation on early fault-tolerant devices. It uses the time-evolution input model, which can potentially overcome the large overall prefactor in the asymptotic gate cost arising in similar algorithms based on the Hamiltonian input model. We present modifications to the original QETU algorithm that significantly reduce the cost for the cases of both exact and Trotterized implementation of the time evolution circuit. We use QETU to prepare the ground state of a U(1) lattice gauge theory in two spatial dimensions, explore the dependence of computational resources on the desired precision and system parameters, and discuss the applicability of our results to general lattice gauge theories. We also demonstrate how the QETU technique can be utilized for preparing Gaussian distributions and wave packets in a way which outperforms existing algorithms for as little as n q ≳ 2–5 qubits.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Verification of the FtCayuga fault-tolerant microprocessor system. Volume 1: A case study in theorem prover-based verification

The design and formal verification of a hardware system for a task that is an important component of a fault tolerant computer architecture for flight control systems is presented. The hardware system implements an algorithm for obtaining interactive consistancy (byzantine agreement) among four microprocessors as a special instruction on the processors. The property verified insures that an execution of the special instruction by the processors correctly accomplishes interactive consistency, provided certain preconditions hold. An assumption is made that the processors execute synchronously. For verification, the authors used a computer aided design hardware design verification tool, Spectool, and the theorem prover, Clio. A major contribution of the work is the demonstration of a significant fault tolerant hardware design that is mechanically verified by a theorem prover.

Srivas, Mandayam↗

Development of an interface for an ultrareliable fault-tolerant control system and an electronic servo-control unit

The NASA Ames Research Center sponsors a research program for the investigation of Intelligent Flight Control Actuation systems. The use of artificial intelligence techniques in conjunction with algorithmic techniques for autonomous, decentralized fault management of flight-control actuation systems is explored under this program. The design, development, and operation of the interface for laboratory investigation of this program is documented. The interface, architecturally based on the Intel 8751 microcontroller, is an interrupt-driven system designed to receive a digital message from an ultrareliable fault-tolerant control system (UFTCS). The interface links the UFTCS to an electronic servo-control unit, which controls a set of hydraulic actuators. It was necessary to build a UFTCS emulator (also based on the Intel 8751) to provide signal sources for testing the equipment.

Shaver, Charles↗

Ion Coulomb Crystals in Storage Rings for Quantum Information Science

Quantum information science is a growing field that promises to take computing into a new age of higher performance and larger scale computing as well as being capable of solving problems classical computers are incapable of solving. The outstanding issue in practical quantum computing today is scaling up the system while maintaining interconnectivity of the qubits and low error rates in qubit operations to be able to implement error correction and fault-tolerant operations. Trapped ion qubits offer long coherence times that allow error correction. However, error correction algorithms require large numbers of qubits to work properly. We can potentially create many thousands (or more) of qubits with long coherence states in a storage ring. For example, a circular radio-frequency quadrupole, which acts as a large circular ion trap and could enable larger scale quantum computing. Such a Storage Ring Quantum Computer (SRQC) would be a scalable and fault tolerant quantum information system, composed of qubits with very long coherence lifetimes. With computing demands potentially outpacing the supply of high-performance systems, quantum computing could bring innovation and scientific advances to particle physics and other DOE supported programs. Increased support of R$\&$D in large scale ion trap quantum computers would allow the timely exploration of this exciting new scalable quantum computer. The R$\&$D program could start immediately at existing facilities and would include the design and construction of a prototype SRQC. We invite feedback from and collaboration with the particle physics and quantum information science communities.

43 PARTICLE ACCELERATORS↗

Blockchain for Fault-Tolerant Grid Operations

Radial topology and vast geographic coverage make distribution systems prone to widespread power outages upon the failure of a single (or multiple) upstream component. Fault-handling algorithms depend heavily on correct estimations of the system’s state to effectively isolate the affected area and reduce the number of affected customers while maintaining operational safety. The work described here leverages the core features of distributed, consensus-based decision-making processes and the immutability of blockchain, and demonstrates their value in improving fault-tolerant grid operations. In this work, blockchain was used to create a trusted data-sharing platform that enables independent actors to reconstruct the system state; this enables distributed resources to make intelligent decisions with limited knowledge. Although the process requires data sharing, its algorithms have been designed to limit the amount of private information that is exchanged, which helps preserve business-sensitive data and maintain customer privacy. In addition, by reducing the information that must be shared, the communication requirements are also reduced; (however, an in-depth analysis of the communication requirements is beyond the scope of this project). The proposed use cases are intended to represent a foundational basis for third parties to develop functional solutions that can eventually be deployed in the field. To further provide guidance, the envisioned use cases have incorporated design requirements that consider the blockchain characteristics and a need to limit information from surrounding resources, which preserve the assumption and the possibility that such resources could belong to different entities. This report presents a detailed design of the three use cases with the tools needed to enable the analysis being tested. The implemented gross error detection method can detect mismatches when the error exceeds 3.8 times the sensor’s rated accuracy. Detection of the circuit breaker state successfully identified the correct states across all simulation tests. A distribution-system power-flow solution in the simulator OpenDSS generally possesses a convergency tolerance of 0.01% on the voltage magnitude. The evaluation of possible reconnection using voltage magnitude—preserving the data ownership—has a voltage magnitude difference smaller than 0.001% from the OpenDSS result. The results preserving data ownership have a difference within the expected power flow tolerance with full knowledge of the system, which surpasses expectations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Benchmarking quantum logic operations relative to thresholds for fault tolerance

Contemporary methods for benchmarking noisy quantum processors typically measure average error rates or process infidelities. However, thresholds for fault-tolerant quantum error correction are given in terms of worst-case error rates—defined via the diamond norm—which can differ from average error rates by orders of magnitude. One method for resolving this discrepancy is to randomize the physical implementation of quantum gates, using techniques like randomized compiling (RC). In this work, we use gate set tomography to perform precision characterization of a set of two-qubit logic gates to study RC on a superconducting quantum processor. We find that, under RC, gate errors are accurately described by a stochastic Pauli noise model without coherent errors, and that spatially correlated coherent errors and non-Markovian errors are strongly suppressed. We further show that the average and worst-case error rates are equal for randomly compiled gates, and measure a maximum worst-case error of 0.0197(3) for our gate set. Our results show that randomized benchmarks are a viable route to both verifying that a quantum processor’s error rates are below a fault-tolerance threshold, and to bounding the failure rates of near-term algorithms, if—and only if—gates are implemented via randomization methods which tailor noise.

97 MATHEMATICS AND COMPUTING↗

Fault-tolerant wait-free shared objects

A concurrent system consists of processes and shared objects. Previous research focused on the problem of tolerating process failure. We study the complementary problem of tolerating failures. We divide object failures into two broad classes: responsive and non-responsive. With responsive failures, a faulty object responds to every invocation, but responses may be incorrect. With non-responsive failures, a faulty object may also 'hang' without responding. For each class, we consider crash, and arbitrary types of failures. For each type of failure, we are seeking a universal implementation for fault-tolerant wait-free shared objects. We present (deterministic) implementations for all types of responsive failures, including arbitrary failures. In contrast, we show that even the most benign type of non-responsive failures requires the use of randomization. Of special interest is the problem of implementing fault-tolerant objects using only objects of the same type. We present such fault-tolerant self-implementations for many common object types. Graceful degradation is a desirable property of fault-tolerant implementations: the implemented object never fails more severely than the base objects it is derived from, even if all the base objects fail. For several failure models, we show whether this property can be achieved, and, if so, how. In addition to the above possibility/impossibility results, we also consider the resources complexity of fault-tolerant implementations. In many cases, we present lower bounds and give matching algorithms.

Jayanti, Prasad↗