Engineering PapersSearch

Engineering topics

Lee, Y. H.

Publications and source records attributed to Lee, Y. H..

33 records · Page 2

Integrated analysis of error detection and recovery

An integrated modeling and analysis of error detection and recovery is presented. When fault latency and/or error latency exist, the system may suffer from multiple faults or error propagations which seriously deteriorate the fault-tolerant capability. Several detection models that enable analysis of the effect of detection mechanisms on the subsequent error handling operations and the overall system reliability were developed. Following detection of the faulty unit and reconfiguration of the system, the contaminated processes or tasks have to be recovered. The strategies of error recovery employed depend on the detection mechanisms and the available redundancy. Several recovery methods including the rollback recovery are considered. The recovery overhead is evaluated as an index of the capabilities of the detection and reconfiguration mechanisms.

Shin, K. G.

A comparison study of spectral energetics analysis using varius FGGE 3b data

The atmospheric spectral energetics studies performed before the First GARP Global Experiment (FGGE) were reviewed, and the deficiencies of these studies were pointed out. The data generated by the FGGE IIIb analyses of the European Center for Medium Range Weather Forecasts, the Geophysical Fluid Dynamics Laboratory, and the Goddard Laboratory for Atmospheric Sciences over the entire globe during the FGGE summer and winter are used to analyze the atmospheric spectral energetics. Comparisons were made between the results using three FGGE IIIb analyses, and between previous and current studies with the spectral energetics of the Southern Hemisphere.

Chen, T. C.

Modeling and measurement of fault-tolerant multiprocessors

The workload effects on computer performance are addressed first for a highly reliable unibus multiprocessor used in real-time control. As an approach to studing these effects, a modified Stochastic Petri Net (SPN) is used to describe the synchronous operation of the multiprocessor system. From this model the vital components affecting performance can be determined. However, because of the complexity in solving the modified SPN, a simpler model, i.e., a closed priority queuing network, is constructed that represents the same critical aspects. The use of this model for a specific application requires the partitioning of the workload into job classes. It is shown that the steady state solution of the queuing model directly produces useful results. The use of this model in evaluating an existing system, the Fault Tolerant Multiprocessor (FTMP) at the NASA AIRLAB, is outlined with some experimental results. Also addressed is the technique of measuring fault latency, an important microscopic system parameter. Most related works have assumed no or a negligible fault latency and then performed approximate analyses. To eliminate this deficiency, a new methodology for indirectly measuring fault latency is presented.

Shin, K. G.

Analysis of the impact of error detection on computer performance

Conventionally, reliability analyses either assume that a fault/error is detected immediately following its occurrence, or neglect damages caused by latent errors. Though unrealistic, this assumption was imposed in order to avoid the difficulty of determining the respective probabilities that a fault induces an error and the error is then detected in a random amount of time after its occurrence. As a remedy for this problem a model is proposed to analyze the impact of error detection on computer performance under moderate assumptions. Error latency, the time interval between occurrence and the moment of detection, is used to measure the effectiveness of a detection mechanism. This model is used to: (1) predict the probability of producing an unreliable result, and (2) estimate the loss of computation due to fault and/or error.

Shin, K. C.

Optimal design and use of retry in fault tolerant real-time computer systems

A new method to determin an optimal retry policy and for use in retry of fault characterization is presented. An optimal retry policy for a given fault characteristic, which determines the maximum allowable retry durations to minimize the total task completion time was derived. The combined fault characterization and retry decision, in which the characteristics of fault are estimated simultaneously with the determination of the optimal retry policy were carried out. Two solution approaches were developed, one based on the point estimation and the other on the Bayes sequential decision. The maximum likelihood estimators are used for the first approach, and the backward induction for testing hypotheses in the second approach. Numerical examples in which all the durations associated with faults have monotone hazard functions, e.g., exponential, Weibull and gamma distributions are presented. These are standard distributions commonly used for modeling analysis and faults.

Lee, Y. H.

Design and evaluation of a fault-tolerant multiprocessor using hardware recovery blocks

A fault-tolerant multiprocessor with a rollback recovery mechanism is discussed. The rollback mechanism is based on the hardware recovery block which is a hardware equivalent to the software recovery block. The hardware recovery block is constructed by consecutive state-save operations and several state-save units in every processor and memory module. When a fault is detected, the multiprocessor reconfigures itself to replace the faulty component and then the process originally assigned to the faulty component retreats to one of the previously saved states in order to resume fault-free execution. A mathematical model is proposed to calculate both the coverage of multi-step rollback recovery and the risk of restart. A performance evaluation in terms of task execution time is also presented.

Lee, Y. H.

Analysis of backward error recovery for concurrent processes with recovery blocks

Three different methods of implementing recovery blocks (RB's). These are the asynchronous, synchronous, and the pseudo recovery point implementations. Pseudo recovery points so that unbounded rollback may be avoided while maintaining process autonomy are proposed. Probabilistic models for analyzing these three methods under standard assumptions in computer performance analysis, i.e., exponential distributions for related random variables were developed. The interval between two successive recovery lines for asynchronous RB's mean loss in computation power for the synchronized method, and additional overhead and rollback distance in case PRP's are used were estimated.

Shin, K. G.

A unified method for evaluating real-time computer controllers: A case study

A real time control system consists of a synergistic pair, that is, a controlled process and a controller computer. Performance measures for real time controller computers are defined on the basis of the nature of this synergistic pair. A case study of a typical critical controlled process is presented in the context of new performance measures that express the performance of both controlled processes and real time controllers (taken as a unit) on the basis of a single variable: controller response time. Controller response time is a function of current system state, system failure rate, electrical and/or magnetic interference, etc., and is therefore a random variable. Control overhead is expressed as a monotonically nondecreasing function of the response time and the system suffers catastrophic failure, or dynamic failure, if the response time for a control task exceeds the corresponding system hard deadline, if any. A rigorous probabilistic approach is used to estimate the performance measures. The controlled process chosen for study is an aircraft in the final stages of descent, just prior to landing. First, the performance measures for the controller are presented. Secondly, control algorithms for solving the landing problem are discussed and finally the impact of the performance measures on the problem is analyzed.

Shin, K. G.

High-energy electron-induced damage production at room temperature in aluminum-doped silicon

DLTS and EPR measurements are reported on aluminum-doped silicon that was irradiated at room temperature with high-energy electrons. Comparisons are made to comparable experiments on boron-doped silicon. Many of the same defects observed in boron-doped silicon are also observed in aluminum-doped silicon, but several others were not observed, including the aluminum interstitial and aluminum-associated defects. Damage production modeling, including the dependence on aluminum concentration, is presented.

Corbett, J. W.

Defect distribution near the surface of electron-irradiated silicon

The surface-defect distributions of electron-irradiated n-type silicon have been investigated using a transient capacitance technique. Schottky, p-n junction, and MOS structures were used in profiling the defect distributions. Surface depletions of defects observed were attributed to the vacancy distribution, but not that of oxygen, and other capture centers' distributions. The vacancy diffusion length at 300 K was estimated to be about 3-6 microns.

Wang, K. L.

EPR and transient capacitance studies on electron-irradiated silicon solar cells

One and two ohm-cm solar cells irradiated with 1 MeV electrons at 30 C were studied using both EPR and transient capacitance techniques. In 2 ohm-cm cells, Si-G6 and Si-G15 EPR spectra and majority carrier trapping levels at (E sub V + 0.23) eV and (E sub V + 0.38) eV were observed, each of which corresponded to the divacancy and the carbon-oxygen-vacancy complex, respectively. In addition, a boron-associated defect with a minority carrier trapping level at (E sub C -0.27) eV was observed. In 1 ohm-cm cells, the G15 spectrum and majority carrier trap at (E sub V + 0.38) eV were absent and an isotropic EPR line appeared at g = 1.9988 (+ or - 0.0003); additionally, a majority carrier trapping center at (E sub V + 0.32) eV, was found which could be associated with impurity lithium. The formation mechanisms of these defects are discussed according to isochronal annealing data in electron-irradiated p-type silicon.

Lee, Y. H.

An application of cluster detection to scene analysis

Certain arrangements of local features in a scene tend to group together and to be seen as units. It is suggested that in some instances, this phenomenon might be interpretable as a process of cluster detection in a graph-structured space derived from the scene. This idea is illustrated using a class of scenes that contain only horizontal and vertical line segments.

Rosenfeld, A. H.