Engineering PapersSearch

Engineering topics

Harry W Jones

Publications and source records attributed to Harry W Jones.

At least 19 records

High Reliability at Minimum Cost

This paper investigates the minimum cost of improving the reliability of complex technical systems. The two major methods to improve reliability are redesigning the system for higher reliability or providing redundant components to replace failed elements. The costs of redesign for reliability or adding redundancy are estimated. The most cost-effective combination for high reliability can be identified. The cost of increasing the intrinsic reliability of a system can be modeled as cost proportional to 1/(system failure rate) a , where the exponent “a” measures the difficulty of increasing reliability. The “a” exponent can vary from 0.25 to about 2.5. Operational reliability can also be increased by using redundant systems. The failure rate for N parallel redundant units is (system failure rate) N . The cost of redundancy is N times the system cost. The total redundant system cost is proportional to N/(system failure rate) a . The cost of redundancy increases as N gets larger, but larger N allows a higher system failure rate, which reduces the system design cost. There is a certain N, a certain level of redundancy, that has the minimum cost to achieve the required overall redundant system failure rate. The minimum cost for the redundant system is achieved at the optimum level of redundancy. The N for minimum cost is equal to -a ln (redundant system failure rate). The minimum cost of the N redundant systems is proportional to N * (original system failure rate) a . The optimum redesigned individual system failure rate is proportional to exp (-1/a), so the greater the difficulty, the higher the optimum individual system failure rate. Increasing the intrinsic reliability of a system encounters diminishing returns and at some point it becomes more cost-effective to add redundancy. The difficulty of increasing intrinsic system reliability determines the optimum design for high reliability at minimum cost.

reliability

Redundancy: How Many Unreliable Spares are Needed for High Reliability and Confidence?

This paper investigates the number of redundant units needed to achieve high reliability with high confidence. The approach is developed for the case when the system failure rate is too high for a single unit to provide the required reliability over the mission duration. To achieve high reliability, N redundant units can be used, one operating unit and N – 1 spares. If the unit failure rate is f, the mission length is L, and f * L is small (not the case assumed here), the unit failure probability over the mission duration is F1 = f * L << 1. In this case, the probability that all N units will fail is Ffail = F1 N , and the needed redundancy N = LN(F)/LN(F1). For the case of large f * L assumed here, F1 = f * L > 1, and F1 is the expected number of failures during the mission. (When F1 = f * L << 1, F1 is the probability that a unit will fail during the mission. When F1 = f * L > 1, F1 is the expected number of failures during the mission.) The needed redundancy, N, to achieve the required N redundant unit reliability, FN, can be computed using the cumulative Poisson distribution with mean equal to F1. The number of spares, N - 1, is increased until the probability - that the total number of failures will be less than N -1 - is equal to the required reliability. The confidence that this reliability can be achieved can be computed using the cumulative Poisson distribution or the chi-square distribution. Since the measured unit failure rate, f, has some probabilistic uncertainty, the actual failure rate will be randomly higher or lower. This means that the reliability of the N redundant systems will be overestimated about half the time. Adding more redundant units increases the confidence that the required reliability will be achieved. For a fixed number of redundant units, the expected reliability and confidence can be traded off, since lower reliability goals will be achieved with higher confidence. Both the desired reliability and confidence can be specified as initial requirements and the needed number of redundant units estimated using the measured failure rate.

Redundancy

The abcd Reliability Growth Model

This paper presents a modification of the well-known Duane-Crow reliability growth model. In the abcd reliability growth model, the initial period of exponential decline of the failure rate in the Duane-Crow model may be followed by a period of constant failure rate. Data often show that an exponential decline in failures is followed by a constant failure rate. If a growth model including only the initial period of exponential decline is applied to increasingly longer failure rate data sets, the data will include longer periods of constant failure rate, and the estimated reliability growth rate will decline from an initially high value down toward zero. Using the Duane-Crow model without extending it to include a possible period of constant failure rate may create the mistaken impression that the initial reliability growth continues forever, but at an ever decreasing rate.

Reliability growth

The abcd Reliability Growth Model

This paper presents a modification of the well-known Duane-Crow reliability growth model. In the abcd reliability growth model, the initial period of exponential decline of the failure rate in the Duane-Crow model may be followed by a period of constant failure rate. Data often show that an exponential decline in failures is followed by a constant failure rate. If a growth model including only the initial period of exponential decline is applied to increasingly longer failure rate data sets, the data will include longer periods of constant failure rate, and the estimated reliability growth rate will decline from an initially high value down toward zero. Using the Duane-Crow model without extending it to include a possible period of constant failure rate may create the mistaken impression that the initial reliability growth continues forever, but at an ever decreasing rate.

Reliability growth

Lower Level Repair Can Easily Fail Due to High Complexity

The International Space Station (ISS) uses Orbital Replacement Units (ORU’s) to repair failures on orbit. Using ORU’s reduces the crew time required to repair failures, but several copies of each ORU must be stored on ISS to ensure system availability. A typical ORU contains many components and has significant mass, but each ORU can repair only a single component failure. A full set of the ORU internal components could repair many different failures. Lower level assembly or component repair should reduce total spares mass. Successful electronics repair experiments were conducted on ISS. However, implementing component level repair would require a significant effort. The systems must be designed so they can be repaired during a mission, considering component layout and accessibility. The repair procedures must be developed and repair facilities, tools, and diagnostic and test instruments provided. Tracing a fault to a component is much more difficult than isolating it to an ORU. Replacing a component is much more difficult than replacing an ORU. Some problems with lower level repair are discussed. The mass savings of lower level repair will not save as much launch cost as before since launch cost has recently been reduced by an order of magnitude. Most system failures are not component failures that can be fixed by replacing a component but are due to system level problems. Repair and maintenance should be planned as part of an overall maintainability design. The risk that a lower level repair will fail is considerably greater than when using ORUs. With modern high reliability packaged systems, failure diagnosis and repair has become a lost art. However, diagnosis and repair data from the 1960’s show that increasing complexity often causes much longer diagnosis and repair times and may prevent successful repair. Increasing complexity by using lower level repair directly increases system cost, failure rate, crew time for repair, and the risk of an unrepairable system failure.

Harry W Jones

Metrics in Space Life Support Technology Selection

Engineering metrics are useful in space life support technology selection, but they must be carefully used. Metrics are only part of a complete system trade-off. Metrics do harm if they cause neglect of other important technical, organizational, or intuitive decision factors. Two metrics have damaged space life support, closure and Equivalent Systems Mass (ESM). Closure measures the fraction of the required system inputs that are produced by recycling system outputs. Increasing closure produces diminishing returns and becomes increasingly expensive. Increasing closure does not directly contribute to providing better life support. ESM measures the total launch mass required to provide life support. ESM includes the mass of the system hardware and of its power, cooling, pressurized volume, spares, and logistics. ESM predicts launch costs, but recently launch costs have been reduced by a factor of 20 or more. System development cost for space hardware is often much greater than launch cost. The past nearly exclusive use of ESM has led to the neglect of Life Cycle Cost (LCC), reliability, cost, and the other engineering factors. Closure and ESM have misguided space life support technology selection for more than twenty years and have adversely affected the expenditure of 100’s of millions of dollars. Metrics can be effectively used three ways in space life support technology selection: 1. A small set of key engineering metrics for preliminary screening. 2. A full set of engineering to guide technical selection. 3. Combining engineering metrics with organizational, political, and intuitive decision factors to understand technology selection. The past emphasis on closure and ESM served to support recycling life support over resupply and built on the intuitive appeal of a human ecosystem in space.

Harry W Jones

Applying the System Complexity Metric (SCM)

A fundamental cause of difficulty in larger engineering projects is their inherent complexity. An impression of complexity occurs if a system is simply difficult to understand, so that there is no obvious mental model that correctly predicts its behavior. Higher complexity is usually associated with higher cost and higher failure rate. Complexity is indicated by a system having more and diverse components, multiple interactions and feedback loops, transients and dynamic behavior, and often the emergence of unanticipated failure modes. Identifying and removing these signs of complexity should reduce complexity and improve performance. Here we limit complexity measurement to the number of components and their interactions. A System Complexity Metric (SCM) is defined as equal to the sum of the number of parts in a system, N, plus the sum of the one-way interconnections between them, I. SCM = N + I. The SCM is easily determined by direct inspection of system block diagrams. Previous work found that life support system cost was directly proportional to SCM and that failure rate increased faster than SCM squared. SCM can be used to compare systems or to guide their redesign to reduce cost and failure rate. Carbon dioxide removal systems will be analyzed using SCM, cost, and failure rate.

Harry W Jones

The optimum cost-effective test time for redundant systems with specified reliability and confidence

This paper investigates the optimum test time to determine the number of redundant units needed to achieve high reliability with high confidence. Newly designed systems often have high initial failure rates which can be reduced by testing to find failure modes and remove them by redesign. To accurately estimate the required number of redundant units, the test time must be extended to accurately determine the failure rate. If the measured failure rate is used, there is a 50% chance that the actual hardware failure rate is higher. Using the measured failure rate gives only a 50% confidence that the failure rate and number of spares are not too low. After the test, given the measured failure rate and the desired reliability, the number of spares can be determined and the confidence in the reliability computed. Instead of accepting the reliability results of a fixed duration test, it is possible to set the requirements for both the redundant reliability and the confidence level and then compute the test time needed to minimize the total cost required to achieve these requirements. The confidence that the redundant reliability is not too low is increased by using a higher than measured failure rate to increase the number of spares. The higher number of spares increases cost. Longer test time reduces the variance in the failure rate and the increase in the number of spares, so that test cost increases and spares cost decreases. The total cost is the sum of the cost of the test time and the spares. The optimum test time produces the minimum total cost for the system failure rate, mission length, and required reliability and confidence level. Longer testing is justified by reduced cost. Some examples are given.

Harry W Jones

Going beyond reliability to robustness and resilience in space systems

The words reliability, robustness, and resilience, are often used interchangeably to describe tough and dependable systems but the distinctions between them suggest how to design more serviceable space systems. Reliability is simply the quality of consistently performing well. A system that dependably meets its design requirements in the specified environments is reliable. The designers may not consider themselves responsible for failures under unanticipated conditions. Robustness is the capability of performing without failure under a wide range of conditions, which can go beyond the expected range to include possible off-nominal conditions. Resilience is the ability to recover from or adapt to damaging events, such as failures, accidents, external disruptions, and repurposing. Such changes are usually unanticipated. They often invalidate the usual operating assumptions and cause system failure. Reliability, robustness, and resilience describe dependable performance under increasingly difficult conditions, first the specified environment, then a wider possible environment, and finally unanticipated damaging events. These three are increasingly desirable and increasingly difficult to achieve. Engineering for resilience would design systems that can ignore or repair failures, survive accidents, and recover from disruptions. Increasing the resilience of space systems, the ability to perform after unanticipated events, would greatly increase space crew safety. Improving reliability and robustness can be done by dealing with known sources of problems, but improving resilience requires implementing a general approach to reducing the impact of unknown future events. Two contrasting approaches are reducing system complexity and adding supervisory control. The need for resilience has been claimed for decades but little has been accomplished. Systems designers assume that they understand requirements, technologies, designs, architectures, integration, testing, operations, and environments. The potential problems of changes, failures, accidents, unknown environments, and unknown unknowns are ignored. Systems designers are typically overconfident and ignore the need for robustness and resilience.

Harry W Jones

Space Settlement Should Use 1 g Shielded Space Platforms, Not the 1/6 or 1/3 g, Radiation Exposed Moon or Mars

Previous missions have subjected astronauts to confined space, weightlessness, and increased radiation. These impair astronaut comfort, performance, and future health. The exploration and future settlement of space will depend on the long term presence of individual humans. This requires the development of space platforms where humans can work and live in health for many years, perhaps generations. A livable space platform must provide adequate volume, gravity, and radiation shielding. This seems easier to do in deep space or Low Earth Orbit (LEO) than on the surface of the moon or Mars. Permanently habitable deep space platforms will enable scientific and technical research, space tourism, space mission preparation, space industry development, and military surveillance and operations. The first fully habitable space station would probably be located in LEO for convenience and lower cost.

Harry W Jones

The problem of partial gravity on the Moon and Mars

Astronauts who spend many weeks or months in space in zero g suffer serious health problems including muscle atrophy, cardiovascular deconditioning, bone calcium loss, impaired vision, and immune system change. The debilitating effects of weightlessness were first demonstrated on the early Skylab, Salyut, and Mir missions, but it was too optimistically hoped that in-flight exercise and resistance training could prevent these problems. Similar problems are anticipated in the partial gravity of the Moon and Mars. Partial gravity exposure below 0.4 g seems too low to maintain musculoskeletal and cardiopulmonary conditioning in the long term. Some studies show a strong correlation between heart rate, oxygen consumption, net metabolic rate and simulated gravity from 0 to 1 g. Exposure to Moon and Mars gravities probably will cause less severe physiological deconditioning than experienced in 0 g, but the benefit of partial gravity seems likely to be roughly proportional to the 1/6 or 1/3 gravity experienced. As in 0 g, exercise countermeasures seem necessary but insufficient to preserve all physiological systems to a 1 g standard.

Harry W Jones

The Challenger disaster was caused by an Apollo decision

NASA’s view of risk changed between early Apollo and the Space Shuttle. Risk was a known serious problem at the beginning of Apollo and the risk estimates were disturbingly high. To avoid public concern, risk analysis was discontinued. Risk analysis was avoided in Shuttle, leading to an unnecessarily risky design. The immediate cause of the Challenger tragedy was the mistaken decision to launch in cold weather. The fundamental cause was the high risk of the Shuttle design. Before Challenger, management thought and testified that the probability of an accident was 1 in 100,000. After Challenger, Probabilistic Risk Analysis (PRA) found a roughly 1 in 100 chance of a Shuttle failure. The recent Orion design uses the safer Apollo approach, with a hardened capsule, launch abort escape, and the crew placed above the rocket tanks and engines. During Apollo it was estimated that, “assuming all elements from propulsion to rendezvous and life support were done as well or better than ever before, that 30 astronauts would be lost before 3 were returned safely to the Earth.” The chance of astronaut survival was only 10%. After the Apollo 1 tragedy, the awareness of risk led to an intense focus on achieving safety. “The only possible explanation for the astonishing success – no losses in space and on time – was that every participant at every level in every area far exceeded the norm of human capabilities.” During Apollo, a NASA PRA found that the chance of success was “less than 5 percent.” The NASA Administrator felt that “the numbers could do irreparable harm,” and discontinued numerical risk assessment.

Harry W Jones

Supercritical Water Oxidation (SCWO) for the treatment of spacecraft waste.

Supercritical water oxidation (SCWO) is a powerful waste treatment that has been considered as a substitute or complement to other potential spacecraft waste handling methods including jettisoning, compacted sealed storage, drying, and others. Combustible trash and garbage is mixed with water and then heated and pressurized above the critical point of water (705 °F and 218 atmospheres). The waste is nearly all oxidized to produce carbon dioxide, water and nitrogen. NOx is not formed. The combustion energy can be used to heat the input stream. SCWO is relatively complex, expensive, and hazardous compared to other methods. Oxidation of compounds containing chlorine and sulfur produces hydrochloric and sulfuric acid which can cause corrosion in the reactor. Neutralizing these acids produces salts which can precipitate and cause clogging. The major engineering challenges to implementing SCWO are chemical corrosion and fouling due to the deposition of salts. SCWO was first proposed for general waste treatment and for spacecraft in the 1980’s and extensive research has been carried out.

Harry W Jones

Managing spacecraft waste

Waste is a universal problem, on Earth, in spacecraft, and for any closed ecological system. Waste must be processed and is often recycled to recover resources. Many different approaches and technologies are used. Spacecraft waste is derived from the spacecraft logistics supplies, the materials provided for use by the crew. The composition of spacecraft logistics and the resulting waste depend on the mission and its duration and use of recycling. Spacecraft waste is a more serious problem on long duration missions because of the large logistics supplies consumed and the difficulty of storing or disposing of waste. The quantity and composition of waste can vary and may require flexible management. Most missions will produce about two kilograms per crewmember per day of trash, consisting of food waste, plastic, paper, packaging, hygiene wipes and many other supplies used and discarded by the crew. The waste is bulky, messy and difficult to store since the wet waste can decompose and produce odors and an accumulation of pathogenic bacteria. Many different methods have been proposed for managing spacecraft waste.

Harry W Jones

Can water mined on the moon save cost for life support?

The water needed for space life support has been so far provided by direct supply from Earth on most missions or by recycling on the International Space Station (ISS). With the recent reduction in launch cost, Earth resupply has become more attractive than recycling for lunar missions with smaller crews and briefer missions requiring less than 1,000 crew days. The third alternative supply method would use water mined from the lunar regolith. Continuing lunar research suggests that the moon may have more water, more widely distributed than previously expected. The cost of water mined on the moon will depend on its form and concentration, the machinery and operations required, and the overall demand for water. The use of oxygen and hydrogen for propulsion may greatly exceed the need for crew life support. The conditions under which water mined on the moon would be less expensive that water provided by resupply or recycling are described.

Harry W Jones

Going beyond reliability to achieve robustness

Reliability is the ability to perform well and consistently. More formally, reliability is defined as the mathematical probability that a system does not fail during a specified time period under its specified operating conditions. The specified operating conditions often go beyond the nominal environment to include variations and challenges encountered in operational use. The difficulty is that systems are often operated outside of their specified operating conditions and, if they fail, the designers are in theory blameless. Unanticipated damaging events include internal failures, external disruptions in supporting systems, accidents, and repurposing. The most common explanation of a system failure is human error, which is usually the first assumption of the system designers. Robustness is the capability to perform without failure under a wide range of conditions that go beyond the specified operating conditions. The first step towards improving robustness would be to expand the system’s specified operating conditions to include a wider range of anticipated challenges, especially human error. Beyond this, there is a need for general approach to reduce the impact of unanticipated future events, the unknown unknowns, by improving the system’s general ability to cope. Robustness can be improved by providing additional processing capacity, larger flow control buffers, increased backup storage, more online redundancy, and more capable supervisory monitoring and control.

Harry W Jones

A Misperception in Reliability Growth Modelling

The Duane reliability growth model is n(t)/t = k t^-alpha (1) The reliability growth rate is alpha, the downward slope of n(t)/t versus t. It usually varies from 0.2 to 0.6. k is a constant. Crow used a 56-failure data set to illustrate reliability growth.1 A graphical Duane model fit to this data gives n(t)/t = 0.640 t^-0.283 (2) A problem in using the Duane-Crow reliability growth model is that it assumes that reliability growth continues and the failure rate decreases throughout the test period. It is more usual that reliability growth stops when the failure rated is low enough. Growth testing is often followed by testing with a low constant failure rate due to rare or uncorrectable failure modes. As more and more low constant rate acceptable failures accumulate after the period of reliability growth, the reliability growth time exponent alpha decreases toward zero. This occurs if constant rate failures are treated as occurring during the reliability growth period. It is more accurate to model a period of initial reliability growth followed by testing without repair to more accurately determine the final constant failure rate. This is done in the abcd model. n(t)/t = a t^-b + c from t = 0 to td (3) = c + d after td, where d = a td^-b (4) The term a t^-b describes the continuous reliability growth that continues out to time td and c is the constant uncorrected failure rate. The parameter d represents an additional constant failure rate due to correctable but uncorrected failure modes. After the reliability growth process is terminated, the failure rate n(t)/t = c + d.

Harry W Jones