Engineering Papers⌕ Search

NASA NTRS · 20210014179

Going beyond reliability to achieve robustness

Abstract

Reliability is the ability to perform well and consistently. More formally, reliability is defined as the mathematical probability that a system does not fail during a specified time period under its specified operating conditions. The specified operating conditions often go beyond the nominal environment to include variations and challenges encountered in operational use. The difficulty is that systems are often operated outside of their specified operating conditions and, if they fail, the designers are in theory blameless. Unanticipated damaging events include internal failures, external disruptions in supporting systems, accidents, and repurposing. The most common explanation of a system failure is human error, which is usually the first assumption of the system designers. Robustness is the capability to perform without failure under a wide range of conditions that go beyond the specified operating conditions. The first step towards improving robustness would be to expand the system’s specified operating conditions to include a wider range of anticipated challenges, especially human error. Beyond this, there is a need for general approach to reduce the impact of unanticipated future events, the unknown unknowns, by improving the system’s general ability to cope. Robustness can be improved by providing additional processing capacity, larger flow control buffers, increased backup storage, more online redundancy, and more capable supervisory monitoring and control.

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Harry W Jones. Going beyond reliability to achieve robustness. https://ntrs.nasa.gov/citations/20210014179

Cite the original work for its findings. Save a collection to share your selection of sources.