Establishing Fault Tolerance for a Class of Systems by Experiment
A long-standing problem in system verification is establishing fault tolerance at the ultra-high level by experiment. It is considered impossible because of system complexity and the enormous number of trials needed. This paper considers the problem for a class of digital systems that use redundancy to achieve reliability. The class is the systems that operate for a period of time without maintenance followed by a maintenance check that replaces components identified as faulty. The paper considers simulating a natural life test where a natural life test observes a number of operating periods. If the system does not fail during the test, it can be said to have a certain reliability at a certain confidence level. The approach in this paper is to make the simulated life test more efficient while maintaining realism by integrating structural arguments, information on fault occurrence, and fault injection in the lab. The major result of this paper is constructing a global fault model using the failure rate of the components and proving theorems about the model that tell how many, what kind, when, and where to inject faults. A simple example illustrates applying the theorems.