Engineering PapersSearch

Engineering topics

Allan L White

Publications and source records attributed to Allan L White.

Establishing Fault Tolerance for a Class of Systems by Experiment

A long-standing problem in system verification is establishing fault tolerance at the ultra-high level by experiment. It is considered impossible because of system complexity and the enormous number of trials needed. This paper considers the problem for a class of digital systems that use redundancy to achieve reliability. The class is the systems that operate for a period of time without maintenance followed by a maintenance check that replaces components identified as faulty. The paper considers simulating a natural life test where a natural life test observes a number of operating periods. If the system does not fail during the test, it can be said to have a certain reliability at a certain confidence level. The approach in this paper is to make the simulated life test more efficient while maintaining realism by integrating structural arguments, information on fault occurrence, and fault injection in the lab. The major result of this paper is constructing a global fault model using the failure rate of the components and proving theorems about the model that tell how many, what kind, when, and where to inject faults. A simple example illustrates applying the theorems.

design of experiments

State reduction for semi-Markov reliability models

Semi-Markov processes have proved to be an effective and convenient tool to construct models of systems that achieve reliability by redundancy and reconfiguration. These models are able to depict complex system architectures and to capture the dynamics of fault arrival and system recovery. A disadvantage of this approach is that the models can be extremely large, which poses both a model construction and a computational problem. Techniques are needed to reduce the model size. Because these systems are used in critical applications where failure can be expensive, there must be an analytically derived bound for the error produced by the model reduction technique. Automatic model generation programs have been written to help the reliability analyst produce models of complex systems. Because of the importance of these programs, the model reduction technique needs to be precise and easily implemented. This paper presents a model reduction technique called trimming that can be applied to a popular class of systems. An error bound for the trimming procedure is derived that uses readily available system parameters. The trimming procedure is precisely described and appears easy to implement in a model generation program.

Reduced order systems

Verifying Autonomous Air Traffic Algorithms in the Presence of Sensor Error and Flight Perturbations

We offer a method to verify algorithms for air traffic control in the presence of sensor errors and flight perturbations. Since errors and perturbations are described by stochastic processes, verification by analytical methods is difficult. We turn to Monte Carlo simulation, but there are two major obstacles. Since air traffic algorithms are safety-critical, the algorithm must be established at a very high confidence level which requires an enormous number of trials. Another is that a simulation attempts to show that an algorithm correctly handles a hazard that might appear, but there is a lack of information about the frequency of occurrence of hazards in the airspace. This paper offers a solution to these and other obstacles. It recognizes there may be a number of algorithms that address different hazards, and the results of the different simulations must be combined to satisfy a global reliability requirement at a high confidence level. To demonstrate the feasibility of the approach, we apply it to an example, an initial effort in collision avoidance.

verification