Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Fault Tolerant Algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Fault-tolerant grid frequency measurement algorithm during transients

Many critical electric grid operations rely on accurate grid frequency measurements. Unfortunately, the measurement accuracy can be easily undermined by power system transient faults. During a power system transient fault, the power grid voltages and currents are usually highly distorted by high-frequency components. What is worse, the power grid signals could have discontinuity during some system transient faults such as phase angle jump, and the discontinuity could result in large measurement errors to state-of-the-art grid measurement algorithms. In this study, a fault-tolerant grid frequency measurement algorithm during transients is proposed. The new algorithm consists of two stages. The first stage is a transient detector, and it can detect the occurrence of system transient faults instantaneously. The second stage is the intelligent frequency estimator, and it will adapt its measurements according to the transient detector. The performance of the algorithm is evaluated under different steady-state and transient conditions. Both dependability and security of the fault-tolerant algorithm are assessed by using PSCAD simulation data and IEEE Standard test data.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Low-overhead transversal fault tolerance for universal quantum computation

Fast, reliable logical operations are essential for realizing useful quantum computers. By redundantly encoding logical qubits into many physical qubits and using syndrome measurements to detect and correct errors, we can achieve low logical error rates. However, for many practical quantum error correction codes such as the surface code, owing to syndrome measurement errors, standard constructions require multiple extraction rounds—of the order of the code distance d—for fault-tolerant computation, particularly considering fault-tolerant state preparation. Here we show that logical operations can be performed fault-tolerantly with only a constant number of extraction rounds for a broad class of quantum error correction codes, including the surface code with magic state inputs and feedforward, to achieve ‘transversal algorithmic fault tolerance’. Through the combination of transversal operations7 and new strategies for correlated decoding, despite only having access to partial syndrome information, we prove that the deviation from the ideal logical measurement distribution can be made exponentially small in the distance, even if the instantaneous quantum state cannot be made close to a logical codeword because of measurement errors. We supplement this proof with circuit-level simulations in a range of relevant settings, demonstrating the fault tolerance and competitive performance of our approach. Furthermore, our work sheds new light on the theory of quantum fault tolerance and has the potential to reduce the space–time cost of practical fault-tolerant quantum computation by over an order of magnitude.

Zhou, Hengyun [QuEra Computing, Boston, MA (United↗

Adapting Secure MultiParty Computation to Support Machine Learning in Radio Frequency Sensor Networks

In this project we developed and validated algorithms for privacy-preserving linear regression using a new variant of Secure Multiparty Computation (MPC) we call "Hybrid MPC" (hMPC). Our variant is intended to support low-power, unreliable networks of sensors with low-communication, fault-tolerant algorithms. In hMPC we do not share training data, even via secret sharing. Thus, agents are responsible for protecting their own local data. Only the machine learning (ML) model is protected with information-theoretic security guarantees against honest-but-curious agents. There are three primary advantages to this approach: (1) after setup, hMPC supports a communication-efficient matrix multiplication primitive, (2) organizations prevented by policy or technology from sharing any of their data can participate as agents in hMPC, and (3) large numbers of low-power agents can participate in hMPC. We have also created an open-source software library named "Cicada" to support hMPC applications with fault-tolerance. The fault-tolerance is important in our applications because the agents are vulnerable to failure or capture. We have demonstrated this capability at Sandia's Autonomy New Mexico laboratory through a simple machine-learning exercise with Raspberry Pi devices capturing and classifying images while flying on four drones.

42 ENGINEERING↗

Analysis and design of algorithm-based fault-tolerant systems

An important consideration in the design of high performance multiprocessor systems is to ensure the correctness of the results computed in the presence of transient and intermittent failures. Concurrent error detection and correction have been applied to such systems in order to achieve reliability. Algorithm Based Fault Tolerance (ABFT) was suggested as a cost-effective concurrent error detection scheme. The research was motivated by the complexity involved in the analysis and design of ABFT systems. To that end, a matrix-based model was developed and, based on that, algorithms for both the design and analysis of ABFT systems are formulated. These algorithms are less complex than the existing ones. In order to reduce the complexity further, a hierarchical approach is developed for the analysis of large systems.

Nair, V. S. Sukumaran↗

What FM can offer DFCS design

The results of aircrafts and spacecrafts flight tests are reported. It is shown that the problems of Digital Flight Control Systems (DFCS) are the problems of systems whose complexity has exceeded the reach of the intellectual tools employed. It is also shown that intuition, experience, and techniques derived from mechanical and analog systems are insufficient for complex, integrated, digital systems. Formal Methods (FM) of computer science can offer DFCS systematic techniques for the construction of trustworthy software, including: techniques for the precise specification of requirements and the development of designs; systematic approaches to the design and structuring of distributed and concurrent systems; fault tolerance algorithms; and systematic methods of testing and analytic methods of verification.

Rushby, John↗

A Methodology for Evaluating Artifacts Produced by a Formal Verification Process

The goal of this study is to produce a methodology for evaluating the claims and arguments employed in, and the evidence produced by formal verification activities. To illustrate the process, we conduct a full assessment of a representative case study for the Enabling Technology Development and Demonstration (ETDD) program. We assess the model checking and satisfiabilty solving techniques as applied to a suite of abstract models of fault tolerant algorithms which were selected to be deployed in Orion, namely the TTEthernet startup services specified and verified in the Symbolic Analysis Laboratory (SAL) by TTTech. To this end, we introduce the Modeling and Verification Evaluation Score (MVES), a metric that is intended to estimate the amount of trust that can be placed on the evidence that is obtained. The results of the evaluation process and the MVES can then be used by non-experts and evaluators in assessing the credibility of the verification results.

Siminiceanu, Radu I.↗

Fault isolation and fault-tolerant control for nonlinear stochastic distribution control systems with multiplicative faults

Here, in this paper, a fault isolation, diagnosis and fault tolerant control algorithm is proposed for nonlinear multiple multiplicative faults stochastic distribution control systems employing Takagi–Sugeno fuzzy system. To obtain the detailed fault information, a fault detection algorithm is introduced to discover the fault occurrence time. Then a fault isolation observer is built to produce the residual, and the error system is separated to subsystems affected only by disturbance and multiplicative faults. Moreover, a fault estimation scheme is presented to obtain the fault magnitude information. When faults occur, the system output probability density function will deviate from the desired distribution. So the model predictive control fault tolerant control scheme is needed to minimize the impact of faults as much as possible to make sure that the post fault output probability density function track the desired probability density function. The validity of the designed algorithm is demonstrated through a simulation example, where the fault tolerant control algorithm ensures that the system output probability density function still track the given output probability density function despite the complex case of multiple multiplicative faults occurring simultaneously.

42 ENGINEERING↗

Formal verification of a fault tolerant clock synchronization algorithm

A formal specification and mechanically assisted verification of the interactive convergence clock synchronization algorithm of Lamport and Melliar-Smith is described. Several technical flaws in the analysis given by Lamport and Melliar-Smith were discovered, even though their presentation is unusally precise and detailed. It seems that these flaws were not detected by informal peer scrutiny. The flaws are discussed and a revised presentation of the analysis is given that not only corrects the flaws but is also more precise and easier to follow. Some of the corrections to the flaws require slight modifications to the original assumptions underlying the algorithm and to the constraints on its parameters, and thus change the external specifications of the algorithm. The formal analysis of the interactive convergence clock synchronization algorithm was performed using the Enhanced Hierarchical Development Methodology (EHDM) formal specification and verification environment. This application of EHDM provides a demonstration of some of the capabilities of the system.

Rushby, John↗

3D Coded SUMMA: Communication-Efficient and Robust Parallel Matrix Multiplication

In this paper, we propose a novel fault-tolerant parallel matrix multiplication algorithm called 3D Coded SUMMA that achieves higher failure-tolerance than replication-based schemes for the same amount of redundancy. This work bridges the gap between recent developments in coded computing and fault-tolerance in high-performance computing (HPC). The core idea of coded computing is the same as algorithm-based fault-tolerance (ABFT), which is weaving redundancy in the computation using error-correcting codes. In particular, we show that MatDot codes, an innovative code construction for parallel matrix multiplications, can be integrated into three-dimensional SUMMA (Scalable Universal Matrix Multiplication Algorithm [30]) in a communication-avoiding manner. To tolerate any two node failures, the proposed 3D Coded SUMMA requires ~50% less redundancy than replication, while the overhead in execution time is only about 5–10%.

97 MATHEMATICS AND COMPUTING↗

Modifying the Asynchronous Jacobi Method for Data Corruption Resilience

Moving scientific computation from high-performance computing (HPC) and cloud computing (CC) environments to devices on the edge, i.e., physically near instruments of interest, has received tremendous interest in recent years. Such edge computing environments can operate on data in situ, offering enticing benefits over data aggregation to HPC and CC facilities that include avoiding costs of transmission, increased data privacy, and real-time data analysis. Because of the inherent unreliability of edge computing environments, new fault-tolerant approaches must be developed before the benefits of edge computing can be realized. Motivated by algorithm-based fault tolerance, a variant of the asynchronous Jacobi (ASJ) method is developed that achieves resilience to data corruption by rejecting solution approximations from neighbor devices according to a bound derived from convergence theory. Numerical results on a two-dimensional Poisson problem show that the new rejection criterion, along with a novel approximation to the shortest path length on which the criterion depends, restores convergence for the ASJ variant in the presence of certain types data corruption. Numerical results are obtained for when the singular values in the analytic bound are approximated. Additional linear systems are also explored, one with a more dense sparsity pattern and one that includes advection. All results indicate that successful resilience to data corruption depends on whether the bound tightens fast enough to reject corrupted data before the iteration evolution deviates significantly from that predicted by the convergence theory defining the bound. This observation generalizes to future work on algorithm-based fault tolerance for other asynchronous algorithms, including upcoming approaches that leverage Krylov subspaces.

97 MATHEMATICS AND COMPUTING↗

Survivable algorithms and redundancy management in NASA's distributed computing systems

The design of survivable algorithms requires a solid foundation for executing them. While hardware techniques for fault-tolerant computing are relatively well understood, fault-tolerant operating systems, as well as fault-tolerant applications (survivable algorithms), are, by contrast, little understood, and much more work in this field is required. We outline some of our work that contributes to the foundation of ultrareliable operating systems and fault-tolerant algorithm design. We introduce our consensus-based framework for fault-tolerant system design. This is followed by a description of a hierarchical partitioning method for efficient consensus. A scheduler for redundancy management is introduced, and application-specific fault tolerance is described. We give an overview of our hybrid algorithm technique, which is an alternative to the formal approach given.

Malek, Miroslaw↗

Simulating and Detecting Radiation-Induced Errors for Onboard Machine Learning

Spacecraft processors and memory are subjected to high radiation doses and therefore employ radiation-hardened components. However, these components are orders of magnitude more expensive than typical desktop components, and they lag years behind in terms of speed and size. We have integrated algorithm-based fault tolerance (ABFT) methods into onboard data analysis algorithms to detect radiation-induced errors, which ultimately may permit the use of spacecraft memory that need not be fully hardened, reducing cost and increasing capability at the same time. We have also developed a lightweight software radiation simulator, BITFLIPS, that permits evaluation of error detection strategies in a controlled fashion, including the specification of the radiation rate and selective exposure of individual data structures. Using BITFLIPS, we evaluated our error detection methods when using a support vector machine to analyze data collected by the Mars Odyssey spacecraft. We found ABFT error detection for matrix multiplication is very successful, while error detection for Gaussian kernel computation still has room for improvement.

data analysis↗

Fault-tolerant grid frequency measurement algorithm during transients

A system determines the frequency of grid signals corresponding to an electrical grid in real time. The system includes a transient detector that monitors a grid signal from a voltage meter or a current meter connected to the electrical grid. The system produces, in real time and at a sampling rate, a deviation signal indicative of a periodicity of the monitored grid signal. The system determines, over one or more cycles of the monitored grid signal, a measurement signal corresponding to the deviation signal. The system determines a frequency signal that corresponds a frequency estimation of the monitored signal by applying a frequency estimation when values of the measurement signal are less than a deviation threshold and maintaining the frequency signal at a constant value when values of the measured signal equal or exceeds the deviation threshold.

Zhan, Lingwei↗

Advanced Symbolic Analysis Tools for Fault-Tolerant Integrated Distributed Systems

The project aims to develop advanced model-checking algorithms and tools to automate the verification of fault-tolerant distributed systems for avionics. We present a new method called Property-Directed K-Induction (PD-KIND) for synthesizing K-inductive invariants of state-transition systems. PD-KIND builds upon Satifiability Modulo Theories (SMT) to generalize Bradley's IC3 method and its variants. This method is implemented in a new tool called SALLY. Case studies show that PD-KIND can automatically verify fault-tolerant algorithms under a variety of fault models and that SALLY is competitive with other SMT-based model checkers.

Dutertre, Bruno↗

A probabilistic model for the evaluation of fault-tolerant multiprocessor systems using concurrent error detection

A probabilistic model to evaluate fault-tolerant multiprocessor systems has been developed. The matrix-based model and the analysis algorihtms based on it are described. Various probabilities associated with an algorithm-based fault tolerance system are discussed and the fault coverage of a given check is derived analytically and illustrated with examples. The probability matrices that are formed by introducing the spatial probabilities into the matrix model are considered. Based on these matrices, a technique is developed to determine the combined coverage of multiple numbers of checks. Examples of the analysis of systems using the model are given.

Nair, V. S. S.↗