Engineering PapersSearch

SEARCH · Engineering Papers

Results for “TOLERANCE”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Design and validation of fault-tolerant flight systems

Flight systems must be validated to show that they are consistent with the requirements of their intended applications. While high reliability is difficult to validate, the additional complexity of fault tolerance further compounds the validation problem. The objective of NASA’s research is to develop a methodology for designing validatable fault-tolerant systems. Under the design-for-validation philosophy, emphasis is placed on developing validation methods that can be incorporated into the design process right from the start and design methods and guidance which, while incorporating fault tolerance, can assure validatability. This paper examines the statistical issues of validating highly reliable, fault tolerant system. There are many problems associated with traditional methods of designing and validating these potentially complex hardware and software systems. Useful design-for-validation methods, which include structured specification and design methodologies, mathematical proof techniques, analytical modeling, simulation and emulation, and physical testing, are discussed. Important design issues associated with fault tolerance are presented along with the related validation concerns which must be addressed. Experience has shown that synchronization and Byzantine resilience must accompany fault tolerance. Other design attributes associated with fault tolerance may be used by a designer on the basis of cost, weight, performance, and validation considerations.

Computer systems

Fault-tolerant wait-free shared objects

A concurrent system consists of processes and shared objects. Previous research focused on the problem of tolerating process failure. We study the complementary problem of tolerating failures. We divide object failures into two broad classes: responsive and non-responsive. With responsive failures, a faulty object responds to every invocation, but responses may be incorrect. With non-responsive failures, a faulty object may also 'hang' without responding. For each class, we consider crash, and arbitrary types of failures. For each type of failure, we are seeking a universal implementation for fault-tolerant wait-free shared objects. We present (deterministic) implementations for all types of responsive failures, including arbitrary failures. In contrast, we show that even the most benign type of non-responsive failures requires the use of randomization. Of special interest is the problem of implementing fault-tolerant objects using only objects of the same type. We present such fault-tolerant self-implementations for many common object types. Graceful degradation is a desirable property of fault-tolerant implementations: the implemented object never fails more severely than the base objects it is derived from, even if all the base objects fail. For several failure models, we show whether this property can be achieved, and, if so, how. In addition to the above possibility/impossibility results, we also consider the resources complexity of fault-tolerant implementations. In many cases, we present lower bounds and give matching algorithms.

Jayanti, Prasad

Parallel fault-tolerant robot control

A shared memory multiprocessor architecture is used to develop a parallel fault-tolerant robot controller. Several versions of the robot controller are developed and compared. A robot simulation is also developed for control observation. Comparison of a serial version of the controller and a parallel version without fault tolerance showed the speedup possible with the coarse-grained parallelism currently employed. The performance degradation due to the addition of processor fault tolerance was demonstrated by comparison of these controllers with their fault-tolerant versions. Comparison of the more fault-tolerant controller with the lower-level fault-tolerant controller showed how varying the amount of redundant data affects performance. The results demonstrate the trade-off between speed performance and processor fault tolerance.

Hamilton, D. L.

Software fault tolerance in computer operating systems

This chapter provides data and analysis of the dependability and fault tolerance for three operating systems: the Tandem/GUARDIAN fault-tolerant system, the VAX/VMS distributed system, and the IBM/MVS system. Based on measurements from these systems, basic software error characteristics are investigated. Fault tolerance in operating systems resulting from the use of process pairs and recovery routines is evaluated. Two levels of models are developed to analyze error and recovery processes inside an operating system and interactions among multiple instances of an operating system running in a distributed environment. The measurements show that the use of process pairs in Tandem systems, which was originally intended for tolerating hardware faults, allows the system to tolerate about 70% of defects in system software that result in processor failures. The loose coupling between processors which results in the backup execution (the processor state and the sequence of events occurring) being different from the original execution is a major reason for the measured software fault tolerance. The IBM/MVS system fault tolerance almost doubles when recovery routines are provided, in comparison to the case in which no recovery routines are available. However, even when recovery routines are provided, there is almost a 50% chance of system failure when critical system jobs are involved.

Iyer, Ravishankar K.

Damage Tolerance Issues as Related to Metallic Rotorcraft Dynamic Components

In this paper issues related to the use of damage tolerance in life managing rotorcraft dynamic components are reviewed. In the past, rotorcraft fatigue design has combined constant amplitude tests of full-scale parts with flight loads and usage data in a conservative manner to provide "safe life" component replacement times. In contrast to the safe life approach over the past twenty years the United States Air Force and several other NATO nations have used damage tolerance design philosophies for fixed wing aircraft to improve safety and reliability. The reliability of the safe life approach being used in rotorcraft started to be questioned shortly after presentations at an American Helicopter Society's specialist meeting in 1980 showed predicted fatigue lives for a hypothetical pitch-link problem to vary from a low of 9 hours to a high in excess of 2594 hours. This presented serious cost, weight, and reliability implications. Somewhat after the U.S. Army introduced its six nines reliability on fatigue life, attention shifted towards using a possible damage tolerance approach to the life management of rotorcraft dynamic components. The use of damage tolerance in life management of dynamic rotorcraft parts will be the subject of this paper. This review will start with past studies on using damage tolerance life management with existing helicopter parts that were safe life designed. Also covered will be a successful attempt at certifying a tail rotor pitch rod using damage tolerance, which was designed using the safe life approach. The FAA review of rotorcraft fatigue design and their recommendations along with some on-going U.S. industry research in damage tolerance on rotorcraft will be reviewed.

Everett, R. A., Jr.

Damage Tolerance Issues as Related to Metallic Rotorcraft Dynamic Components

In this paper issues related to the use of damage tolerance in life managing rotorcraft dynamic components are reviewed. In the past, rotorcraft fatigue design has combined constant amplitude tests of full-scale parts with flight loads and usage data in a conservative manner to provide "safe life" component replacement times. In contrast to the safe life approach over the past twenty years the United States Air Force and several other NATO nations have used damage tolerance design philosophies for fixed wing aircraft to improve safety and reliability. The reliability of the safe life approach being used in rotorcraft started to be questioned shortly after presentations at an American Helicopter Society's specialist meeting in 1980 showed predicted fatigue lives for a hypothetical pitch-link problem to vary from a low of 9 hours to a high in excess of 2594 hours. This presented serious cost, weight, and reliability implications. Somewhat after the U.S. Army introduced its six nines reliability on fatigue life, attention shifted towards using a possible damage tolerance approach to the life management of rotorcraft dynamic components. The use of damage tolerance in life management of dynamic rotorcraft parts will be the subject of this paper. This review will start with past studies on using damage tolerance life management with existing helicopter parts that were safe life designed. Also covered will be a successful attempt at certifying a tail rotor pitch rod using damage tolerance, which was designed using the safe life approach. The FAA review of rotorcraft fatigue design and their recommendations along with some on-going U.S. industry research in damage tolerance on rotorcraft will be reviewed. Finally, possible problems and future needs for research will be highlighted.

Everett, R. A., Jr.

IRON-TOLERANT CYANOBACTERIA: IMPLICATIONS FOR ASTROBIOLOGY

The review is dedicated to the new group of extremophiles - iron tolerant cyanobacteria. The authors have analyzed earlier published articles about the ecology of iron tolerant cyanobacteria and their diversity. It was concluded that contemporary iron depositing hot springs might be considered as relative analogs of Precambrian environment. The authors have concluded that the diversity of iron-tolerant cyanobacteria is understudied. The authors also analyzed published data about the physiological peculiarities of iron tolerant cyanobacteria. They made the conclusion that iron tolerant cyanobacteria may oxidize reduced iron through the photosystem of cyanobacteria. The involvement of both Reaction Centers 1 and 2 is also discussed. The conclusion that iron tolerant protocyanobacteria could be involved in banded iron formations generation is also proposed. The possible mechanism of the transition from an oxygenic photosynthesis to an oxygenic one is also discussed. In the final part of the review the authors consider the possible implications of iron tolerant cyanobacteria for astrobiology.

Brown, Igor I.

Statistical Tolerance and Clearance Analysis for Assembly

Tolerance is inevitable because manufacturing exactly equal parts is known to be impossible. Furthermore, the specification of tolerances is an integral part of product design since tolerances directly affect the assemblability, functionality, manufacturability, and cost effectiveness of a product. In this paper, we present statistical tolerance and clearance analysis for the assembly. Our proposed work is expected to make the following contributions: (i) to help the designers to evaluate products for assemblability, (ii) to provide a new perspective to tolerance problems, and (iii) to provide a tolerance analysis tool which can be incorporated into a CAD or solid modeling system.

statistical tolerance clearance analysis Monte-Car

+Gz tolerance after 14 days bed rest and the effects of rehydration.

Measurement of the reduction in centrifugation tolerance after two weeks of bedrest with moderate daily exercise, with an attempt to determine if rehydration improves +Gz tolerance. There were significant reductions in +Gz tolerance during bedrest periods at three acceleration levels. Rehydration resulted in a significant increase in tolerance at 2.1G but it did not restore tolerance to control levels. Rehydration did not affect tolerance at 3.2G and 3.8G.

Greenleaf, J. E.

Systems approach to software fault tolerance

Computing systems are employed for aerospace applications with high reliability requirements. In order to provide the needed reliability, it was necessary to make use of computing systems with fault-tolerance characteristics. Traditionally, fault tolerance is achieved through the use of hardware redundance. However, fault-tolerant techniques based on suitable software design considerations have also been developed. The present paper is concerned with the major issues arising in the context of an application of fault-tolerant software techniques to dynamic systems. Attention is given to fault-tolerant flight software, software component stability, system stability with fault-tolerant software, the preservation of functional performance, N-version vs. recovery blocks in flight software, systems-based software, static and dynamic models, static and dynamic consistency tests, and recovery block initialization.

Caglayan, A. K.

Intelligent failure-tolerant control

An overview of failure-tolerant control is presented, beginning with robust control, progressing through parallel and analytical redundancy, and ending with rule-based systems and artificial neural networks. By design or implementation, failure-tolerant control systems are 'intelligent' systems. All failure-tolerant systems require some degrees of robustness to protect against catastrophic failure; failure tolerance often can be improved by adaptivity in decision-making and control, as well as by redundancy in measurement and actuation. Reliability, maintainability, and survivability can be enhanced by failure tolerance, although each objective poses different goals for control system design. Artificial intelligence concepts are helpful for integrating and codifying failure-tolerant control systems, not as alternatives but as adjuncts to conventional design methods.

Stengel, Robert F.

A fault-tolerant intelligent robotic control system

This paper describes the concept, design, and features of a fault-tolerant intelligent robotic control system being developed for space and commercial applications that require high dependability. The comprehensive strategy integrates system level hardware/software fault tolerance with task level handling of uncertainties and unexpected events for robotic control. The underlying architecture for system level fault tolerance is the distributed recovery block which protects against application software, system software, hardware, and network failures. Task level fault tolerance provisions are implemented in a knowledge-based system which utilizes advanced automation techniques such as rule-based and model-based reasoning to monitor, diagnose, and recover from unexpected events. The two level design provides tolerance of two or more faults occurring serially at any level of command, control, sensing, or actuation. The potential benefits of such a fault tolerant robotic control system include: (1) a minimized potential for damage to humans, the work site, and the robot itself; (2) continuous operation with a minimum of uncommanded motion in the presence of failures; and (3) more reliable autonomous operation providing increased efficiency in the execution of robotic tasks and decreased demand on human operators for controlling and monitoring the robotic servicing routines.

Marzwell, Neville I.

Avoiding and tolerating latency in large-scale next-generation shared-memory multiprocessors

A scalable solution to the memory-latency problem is necessary to prevent the large latencies of synchronization and memory operations inherent in large-scale shared-memory multiprocessors from reducing high performance. We distinguish latency avoidance and latency tolerance. Latency is avoided when data is brought to nearby locales for future reference. Latency is tolerated when references are overlapped with other computation. Latency-avoiding locales include: processor registers, data caches used temporally, and nearby memory modules. Tolerating communication latency requires parallelism, allowing the overlap of communication and computation. Latency-tolerating techniques include: vector pipelining, data caches used spatially, prefetching in various forms, and multithreading in various forms. Relaxing the consistency model permits increased use of avoidance and tolerance techniques. Each model is a mapping from the program text to sets of partial orders on program operations; it is a convention about which temporal precedences among program operations are necessary. Information about temporal locality and parallelism constrains the use of avoidance and tolerance techniques. Suitable architectural primitives and compiler technology are required to exploit the increased freedom to reorder and overlap operations in relaxed models.

Probst, David K.

A minimum cost tolerance allocation method for rocket engines and robust rocket engine design

Rocket engine design follows three phases: systems design, parameter design, and tolerance design. Systems design and parameter design are most effectively conducted in a concurrent engineering (CE) environment that utilize methods such as Quality Function Deployment and Taguchi methods. However, tolerance allocation remains an art driven by experience, handbooks, and rules of thumb. It was desirable to develop and optimization approach to tolerancing. The case study engine was the STME gas generator cycle. The design of the major components had been completed and the functional relationship between the component tolerances and system performance had been computed using the Generic Power Balance model. The system performance nominals (thrust, MR, and Isp) and tolerances were already specified, as were an initial set of component tolerances. However, the question was whether there existed an optimal combination of tolerances that would result in the minimum cost without any degradation in system performance.

Gerth, Richard J.

Dynamic Leg Exercise Improves Tolerance to Lower Body Negative Pressure

These results clearly demonstrate that dynamic leg exercise against the footward force produced by LBNP substantially improves tolerance to LBNP, and that even cyclic ankle flexion without load bearing also increases tolerance. This exercise-induced increase of tolerance was actually an underestimate, because subjects who completed the tolerance test while exercising could have continued for longer periods. Exercise probably increases LBNP tolerance by multiple mechanisms. Tolerance was increased in part by skeletal muscle pumping venous blood from the legs. Rosenhamer and Linnarsson and Rosenhamer also deduced this for subjects cycling during centrifugation, although no measurements of leg volume were made in those studies: they found that male subjects cycling at 98 W could endure 3 Gz centrifugation longer than when they remained relaxed during centrifugation. Skeletal muscle pumping helps maintain cardiac filling pressure by opposing gravity-, centrifugation-, or LBNP-induced accumulation of blood and extravascular fluid in the legs.

Watenpaugh, D. E.

The Design of a Fault-Tolerant COTS-Based Bus Architecture

In this paper, we report our experiences and findings on the design of a fault-tolerant bus architecture comprised of two COTS buses, the IEEE 1394 and the 12C. This fault-tolerant bus is the backbone system bus for the avionics architecture of the X2000 program at the Jet Propulsion Laboratory. COTS buses are attractive because of the availability of low cost commercial products. However, they are not specifically designed for highly reliable applications such as long-life deep-space missions. The X2000 design team has devised a multi-level fault tolerance approach to compensate for this shortcoming of COTS buses. First, the approach enhances the fault tolerance capabilities of the IEEE 1394 and 12 C buses by adding a layer of fault handling hardware and software. Second, algorithms are developed to enable the IEEE 1394 and the 12 C buses assist each other to isolate and recovery from faults. Third, the set of IEEE 1394 and 12 C buses is duplicated to further enhance system reliability. The X2000 design team has paid special attention to guarantee that all fault tolerance provisions will not cause the bus design to deviate from the commercial standard specifications. Otherwise, the economic attractiveness of using COTS will be diminished. The hardware and software design of the X2000 fault-tolerant bus are being implemented and flight hardware will be delivered to the ST4 and Europa Orbiter missions.

Chau, Savio N.

Fault-Tolerant Heat Exchanger

A compact, lightweight heat exchanger has been designed to be fault-tolerant in the sense that a single-point leak would not cause mixing of heat-transfer fluids. This particular heat exchanger is intended to be part of the temperature-regulation system for habitable modules of the International Space Station and to function with water and ammonia as the heat-transfer fluids. The basic fault-tolerant design is adaptable to other heat-transfer fluids and heat exchangers for applications in which mixing of heat-transfer fluids would pose toxic, explosive, or other hazards: Examples could include fuel/air heat exchangers for thermal management on aircraft, process heat exchangers in the cryogenic industry, and heat exchangers used in chemical processing. The reason this heat exchanger can tolerate a single-point leak is that the heat-transfer fluids are everywhere separated by a vented volume and at least two seals. The combination of fault tolerance, compactness, and light weight is implemented in a unique heat-exchanger core configuration: Each fluid passage is entirely surrounded by a vented region bridged by solid structures through which heat is conducted between the fluids. Precise, proprietary fabrication techniques make it possible to manufacture the vented regions and heat-conducting structures with very small dimensions to obtain a very large coefficient of heat transfer between the two fluids. A large heat-transfer coefficient favors compact design by making it possible to use a relatively small core for a given heat-transfer rate. Calculations and experiments have shown that in most respects, the fault-tolerant heat exchanger can be expected to equal or exceed the performance of the non-fault-tolerant heat exchanger that it is intended to supplant (see table). The only significant disadvantages are a slight weight penalty and a small decrease in the mass-specific heat transfer.

Izenson, Michael G.

Fault Tolerance Middleware for a Multi-Core System

Fault Tolerance Middleware (FTM) provides a framework to run on a dedicated core of a multi-core system and handles detection of single-event upsets (SEUs), and the responses to those SEUs, occurring in an application running on multiple cores of the processor. This software was written expressly for a multi-core system and can support different kinds of fault strategies, such as introspection, algorithm-based fault tolerance (ABFT), and triple modular redundancy (TMR). It focuses on providing fault tolerance for the application code, and represents the first step in a plan to eventually include fault tolerance in message passing and the FTM itself. In the multi-core system, the FTM resides on a single, dedicated core, separate from the cores used by the application. This is done in order to isolate the FTM from application faults and to allow it to swap out any application core for a substitute. The structure of the FTM consists of an interface to a fault tolerant strategy module, a responder module, a fault manager module, an error factory, and an error mapper that determines the severity of the error. In the present reference implementation, the only fault tolerant strategy implemented is introspection. The introspection code waits for an application node to send an error notification to it. It then uses the error factory to create an error object, and at this time, a severity level is assigned to the error. The introspection code uses its built-in knowledge base to generate a recommended response to the error. Responses might include ignoring the error, logging it, rolling back the application to a previously saved checkpoint, swapping in a new node to replace a bad one, or restarting the application. The original error and recommended response are passed to the top-level fault manager module, which invokes the response. The responder module also notifies the introspection module of the generated response. This provides additional information to the introspection module that it can use in generating its next response. For example, if the responder triggers an application rollback and errors are still occurring, the introspection module may decide to recommend an application restart.

Some, Raphael R.