Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Trends in Human Spaceflight: Failure Tolerance, High Reliability and Correlated Failure History

In a half century of human spaceflight, NASA has continuously refined agency safety and reliability requirements in response to mission demands, critical failures, and technology development. Early spacecraft, including Mercury, Gemini and Apollo vehicles, were highly reliant on dissimilar redundancy and demonstrated test margins. Later programs, such as the reusable Space Transportation System (STS) and International Space Station (ISS), introduced probabilistic studies and isolated two-failure tolerance to improve robustness at the expense of added complexity. More recently, the Orion Multi-Program Crew Vehicle (MPCV) program adopted universal single-failure tolerance with two categorical exceptions; Zero-Failure Tolerant (0FT) and Design for Minimum Risk (DFMR) hardware. Failure tolerance variances are defined and managed in accordance with agency human-rating requirements, and require concurrence from program Technical Authorities (TA) as well as the MPCV Safety and Mission Assurance Safety and Engineering Review Panel (MSERP). To understand and reaffirm standards applied to Apollo, Space Shuttle and Orion vehicles, Orion and Deep Space Gateway Safety and Mission Assurance (S&MA) representatives conducted accelerated research to compare unique safety and reliability criteria against ground and flight anomalies, based on information contained in post-mission reports and the Problem Reporting and Corrective Action (PRACA) database. In some cases, high-profile failures and narrow escapes have reinforced decisions to maintain or adapt safety requirements. In others, empirical trends have highlighted the need for vigilance and innovative safety guidelines. Given the inability to achieve absolute compliance with evolving safety and reliability requirements, the team conducted a targeted review of DFMR and 0FT propulsion elements within the framework of changing system design, inspection, materials and process developments to formulate conclusions on technological maturity, failure density, and net changes in safety risk. Based on the aggregate performance of high-reliability and failure-tolerant systems, the authors have attempted to establish best practices and guidelines to inform future program decisions. On a somewhat cautionary note, this study is not intended to direct a universal set of requirements for future missions based on prior lessons learned. Spacecraft safety is a multi-variable problem, and attempts to mitigate past failures will not guarantee future success. However, this assessment offers a retrospective review of policy changes, implementation and effectiveness. In the future, NASA, European Space Agency (ESA) and industry partners may benefit from a more robust correlation between requirements and performance, as space-faring nations work toward more challenging, complex and long-duration commercial and deep-space ventures.

Green, Carrie↗

Hypervelocity Impact Testing and MMOD Risk Reduction

Purpose: (1) Provide data to develop, update, and/or verify ballistic limit equations used in the MMOD risk assessment. (2) Provide data used to compare two or more shielding options to reduce MMOD risk. (3) Determine failure modes and failure criteria for hardware: (a) Failure modes: how hardware fails (pressure vessels, pressurized lines, electronic hardware, power cables). (b) Failure criteria: quantify damage level that results in hardware failure (for example: depth of penetration into pressure vessel that results in leak or burst).

Lear, Dana↗

Unstable insulation resistance in ceramic capacitors

The effects of low and varying values of insulation resistance (IR) and their significance in the cause of failure in ceramic capacitors were investigated. Two specific instances involving hardware failure involving ceramic capacitors were examined. It is shown that monolithic, multilayer ceramic capacitors may exhibit low and unstable IR at low voltages, or exhibit no voltage. Other significant results are reported.

Holladay, A. M.↗

An experimental evaluation of software redundancy as a strategy for improving reliability

The strategy of using multiple versions of independently developed software as a means to tolerate residual software design faults is suggested by the success of hardware redundancy for tolerating hardware failires. Although, as generally accepted, the independence of hardware failures resulting from physical wearout can lead to substantial increases in reliability for redundant hardware structures, a similar conclusion is not immediate for software. The degree to which design faults are manifested as independent failures determines the effectiveness of redundancy as a method for improving software reliability. Interest in multi-version software centers on whether it provides an adequate measure of increased reliability to warrant its use in critical applications. The effectiveness of multi-version software is studied by comparing estimates of the failure probabilities of these systems with the failure probabilities of single versions. The estimates are obtained under a model of dependent failures and compared with the estimates obtained when failures are assumed to be independent. The experimental results are based on twenty versions of an aerospace application developed and certified by sixty programmers from four universities. Descriptions of the application, development and certifications processes, and operational evaluation are given together with an analysis of the twenty versions.

Eckhardt, Dave E.↗

The optimum cost-effective test time for redundant systems with specified reliability and confidence

This paper investigates the optimum test time to determine the number of redundant units needed to achieve high reliability with high confidence. Newly designed systems often have high initial failure rates which can be reduced by testing to find failure modes and remove them by redesign. To accurately estimate the required number of redundant units, the test time must be extended to accurately determine the failure rate. If the measured failure rate is used, there is a 50% chance that the actual hardware failure rate is higher. Using the measured failure rate gives only a 50% confidence that the failure rate and number of spares are not too low. After the test, given the measured failure rate and the desired reliability, the number of spares can be determined and the confidence in the reliability computed. Instead of accepting the reliability results of a fixed duration test, it is possible to set the requirements for both the redundant reliability and the confidence level and then compute the test time needed to minimize the total cost required to achieve these requirements. The confidence that the redundant reliability is not too low is increased by using a higher than measured failure rate to increase the number of spares. The higher number of spares increases cost. Longer test time reduces the variance in the failure rate and the increase in the number of spares, so that test cost increases and spares cost decreases. The total cost is the sum of the cost of the test time and the spares. The optimum test time produces the minimum total cost for the system failure rate, mission length, and required reliability and confidence level. Longer testing is justified by reduced cost. Some examples are given.

Harry W Jones↗

Extraction-Separation Performance and Dynamic Modeling of Orion Test Vehicles with Adams Simulation: 3rd Edition

NASA's Orion Capsule Parachute Assembly System (CPAS) Project is now in the qualification phase of testing, and the Adams simulation has continued to evolve to model the complex dynamics experienced during the test article extraction and separation phases of flight. The ability to initiate tests near the upper altitude limit of the Orion parachute deployment envelope requires extractions from the aircraft at 35,000 ft-MSL. Engineering development phase testing of the Parachute Test Vehicle (PTV) carried by the Carriage Platform Separation System (CPSS) at altitude resulted in test support equipment hardware failures due to increased energy caused by higher true airspeeds. As a result, hardware modifications became a necessity requiring ground static testing of the textile components to be conducted and a new ground dynamic test of the extraction system to be devised. Force-displacement curves from static tests were incorporated into the Adams simulations, allowing prediction of loads, velocities and margins encountered during both flight and ground dynamic tests. The Adams simulation was then further refined by fine tuning the damping terms to match the peak loads recorded in the ground dynamic tests. The failure observed in flight testing was successfully replicated in ground testing and true safety margins of the textile components were revealed. A multi-loop energy modulator was then incorporated into the system level Adams simulation model and the effect on improving test margins be properly evaluated leading to high confidence ground verification testing of the final design solution.

Varela, Jose G.↗

Integrated Hardware and Software for No-Loss Computing

When an algorithm is distributed across multiple threads executing on many distinct processors, a loss of one of those threads or processors can potentially result in the total loss of all the incremental results up to that point. When implementation is massively hardware distributed, then the probability of a hardware failure during the course of a long execution is potentially high. Traditionally, this problem has been addressed by establishing checkpoints where the current state of some or part of the execution is saved. Then in the event of a failure, this state information can be used to recompute that point in the execution and resume the computation from that point. A serious problem arises when one distributes a problem across multiple threads and physical processors is that one increases the likelihood of the algorithm failing due to no fault of the scientist but as a result of hardware faults coupled with operating system problems. With good reason, scientists expect their computing tools to serve them and not the other way around. What is novel here is a unique combination of hardware and software that reformulates an application into monolithic structure that can be monitored in real-time and dynamically reconfigured in the event of a failure. This unique reformulation of hardware and software will provide advanced aeronautical technologies to meet the challenges of next-generation systems in aviation, for civilian and scientific purposes, in our atmosphere and in atmospheres of other worlds. In particular, with respect to NASA s manned flight to Mars, this technology addresses the critical requirements for improving safety and increasing reliability of manned spacecraft.

James, Mark↗

SSME Propellant Path Leak Detection

The complicated high-pressure cycle of the space shuttle main engine (SSME) propellant path provides many opportunities for external propellant path leaks while the engine is running. This mode of engine failure may be detected and analyzed with sufficient speed to save critical engine test hardware from destruction. The leaks indicate hardware failures which will damage or destroy an engine if undetected; therefore, detection of both cryogenic and hot gas leaks is the objective of this investigation. The primary objective of this phase of the investigation is the experimental validation of techniques for detecting and analyzing propellant path external leaks which have a high probability of occurring on the SSME. The selection of candidate detection methods requires a good analytic model for leak plumes which would develop from external leaks and an understanding of radiation transfer through the leak plume. One advanced propellant path leak detection technique is obtained by using state-of-the-art technology infrared (IR) thermal imaging systems combined with computer, digital image processing, and expert systems for the engine protection. The feasibility of IR leak plume detection is evaluated on subscale simulated laboratory plumes to determine sensitivity, signal to noise, and general suitability for the application.

Roger Crawford↗

First incremental buy for Increment 2 of the Space Transportation System (STS)

Thiokol manufactured and delivered 9 flight motors to KSC on schedule. All test flights were successful. All spent SRMs were recovered. Design, development, manufacture, and delivery of required transportation, handling, and checkout equipment to MSFC and to KSC were completed on schedule. All items of data required by DPD 400 were prepared and delivered as directed. In the system requirements and analysis area, the point of departure from Buy 1 to the operational phase was developed in significant detail with a complete set of transition documentation available. The documentation prepared during the Buy 1 program was maintained and updated where required. The following flight support activities should be continued through other production programs: as-built materials usage tracking on all flight hardware; mass properties reporting for all flight hardware until sample size is large enough to verify that the weight limit requirements were met; ballistic predictions and postflight performance assessments for all production flights; and recovered SRM hardware inspection and anomaly identification. In the safety, reliability, and quality assurance area, activities accomplished were assurance oriented in nature and specifically formulated to prevent problems and hardware failures. The flight program to date has adequately demonstrated the success of this assurance approach. The attention focused on details of design, analysis, manufacture, and inspection to assure the production of high-quality hardware has resulted in the absence of flight failures. The few anomalies which did occur were evaluated, design or manufacturing changes incorporated, and corrective actions taken to preclude recurrence.

Source record↗

Orbiter wheel and tire certification

The orbiter wheel and tire development has required a unique series of certification tests to demonstrate the ability of the hardware to meet severe performance requirements. Early tests of the main landing gear wheel using conventional slow roll testing resulted in hardware failures. This resulted in a need to conduct high velocity tests with crosswind effects for assurance that the hardware was safe for a limited number of flights. Currently, this approach and the conventional slow roll and static tests are used to certify the wheel/tire assembly for operational use.

Campbell, C. C., Jr.↗

ISS Regenerative Life Support: Challenges and Success in the Quest for Long-Term Habitability in Space

This presentation will discuss the International Space Station s (ISS) Regenerative Environmental Control and Life Support System (ECLSS) operations with discussion of the on-orbit lessons learned, specifically regarding the challenges that have been faced as the system has expanded with a growing ISS crew. Over the 10 year history of the ISS, there have been numerous challenges, failures, and triumphs in the quest to keep the crew alive and comfortable. Successful operation of the ECLSS not only requires maintenance of the hardware, but also management of the station resources in case of hardware failure or missed re-supply. This involves effective communication between the primary International Partners (NASA and Roskosmos) and the secondary partners (JAXA and ESA) in order to keep a reserve of the contingency consumables and allow for re-supply of failed hardware. The ISS ECLSS utilizes consumables storage for contingency usage as well as longer-term regenerative systems, which allow for conservation of the expensive resources brought up by re-supply vehicles. This long-term hardware, and the interactions with software, was a challenge for Systems Engineers when they were designed and require multiple operational workarounds in order to function continuously. On a day-to-day basis, the ECLSS provides big challenges to the on console controllers. Main challenges involve the utilization of the resources that have been brought up by the visiting vehicles prior to undocking, balance of contributions between the International Partners for both systems and resources, and maintaining balance between the many interdependent systems, which includes providing the resources they need when they need it. The current biggest challenge for ECLSS is the Regenerative ECLSS system, which continuously recycles urine and condensate water into drinking water and oxygen. These systems were brought to full functionality on STS-126 (ULF-2) mission. Through system failures and recovery, the ECLSS console has learned how to balance the water within the systems, store and use water for contingencies, and continue to work with the International Partners for short-term failures. Through these challenges and the system failures, the most important lesson learned has been the importance of redundancy and operational workarounds. It is only because of the flexibility of the hardware and the software that flight controllers have the opportunity to continue operating the system as a whole for mission success.

Bazley, Jesse A.↗

Space Station Freedom integrated fault model

A demonstration of an integrated fault propagation model for Space Station Freedom is described. The demonstration uses a HyperCard graphical interface to show how failures can propagate from one component to another, both within a system and between systems. It also shows how hardware failures can impact certain defined functions like reboost, atmosphere maintenance or collision avoidance. The demonstration enables the user to view block diagrams for the various space station systems using an overview screen, and interactively choose a component and see what single or dual failure combinations can cause it to fail. It also allows the user to directly view the fault model, which is a collection of drawing and text listings accessible from a guide screen. Fault modeling provides a useful technique for analyzing individual systems and also interactions between systems in the presence of multiple failures so that a complete picture of failure tolerance and component criticality can be achieved.

Becker, Fred J.↗

Independent Orbiter Assessment (IOA): FMEA/CIL assessment

The McDonnell Douglas Astronautics Company (MDAC) was selected to perform an Independent Orbiter Assessment (IOA) of the Failure Modes and Effects Analysis (FMEA) and Critical Items List (CIL). Direction was given by the Orbiter and GFE Projects Office to perform the hardware analysis and assessment using the instructions and ground rules defined in NSTS 22206. The IOA analysis featured a top-down approach to determine hardware failure modes, criticality, and potential critical items. To preserve independence, the analysis was accomplished without reliance upon the results contained within the NASA and Prime Contractor FMEA/CIL documentation. The assessment process compared the independently derived failure modes and criticality assignments to the proposed NASA post 51-L FMEA/CIL documentation. When possible, assessment issues were discussed and resolved with the NASA subsystem managers. Unresolved issues were elevated to the Orbiter and GFE Projects Office manager, Configuration Control Board (CCB), or Program Requirements Control Board (PRCB) for further resolution. The most important Orbiter assessment finding was the previously unknown stuck autopilot push-button criticality 1/1 failure mode. The worst case effect could cause loss of crew/vehicle when the microwave landing system is not active. It is concluded that NASA and Prime Contractor Post 51-L FMEA/CIL documentation assessed by IOA is believed to be technically accurate and complete. All CIL issues were resolved. No FMEA issues remain that have safety implications. Consideration should be given, however, to upgrading NSTS 22206 with definitive ground rules which more clearly spell out the limits of redundancy.

Hinsdale, L. W.↗

Method of Testing and Predicting Failures of Electronic Mechanical Systems

A method employing a knowledge base of human expertise comprising a reliability model analysis implemented for diagnostic routines is disclosed. The reliability analysis comprises digraph models that determine target events created by hardware failures human actions, and other factors affecting the system operation. The reliability analysis contains a wealth of human expertise information that is used to build automatic diagnostic routines and which provides a knowledge base that can be used to solve other artificial intelligence problems.

Iverson, David L.↗

Independent Orbiter Assessment (IOA): FMEA/CIL assessment

The results of the Independent Orbiter Assessment (IOA) of the Failure Modes and Effects Analysis (FMEA) and Critical Items List (CIL) are presented. Direction was given by the Orbiter and GFE Projects Office to perform the hardware analysis and assessment using the instructions and ground rules defined in NSTS 22206. The IOA analysis features a top-down approach to determine hardware failure modes, criticality, and potential critical items. To preserve independence, the anlaysis was accomplished without reliance upon the results contained within the NASA and prime contractor FMEA/CIL documentation. The assessment process compares the independently derived failure modes and criticality assignments to the proposed NASA Post 51-L FMEA/CIL documentation. When possible, assessment issues are discussed and resolved with the NASA subsystem managers. The assessment results for each subsystem are summarized. The most important Orbiter assessment finding was the previously unknown stuck autopilot push-button criticality 1/1 failure mode, having a worst case effect of loss of crew/vehicle when a microwave landing system is not active.

Saiidi, Mo J.↗

Test and evaluation of the generalized gate logic system simulator

The results of the initial testing of the Generalized Gate Level Logic Simulator (GGLOSS) are discussed. The simulator is a special purpose fault simulator designed to assist in the analysis of the effects of random hardware failures on fault tolerant digital computer systems. The testing of the simulator covers two main areas. First, the simulation results are compared with data obtained by monitoring the behavior of hardware. The circuit used for these comparisons is an incomplete microprocessor design based upon the MIL-STD-1750A Instruction Set Architecture. In the second area of testing, current simulation results are compared with experimental data obtained using precursors of the current tool. In each case, a portion of the earlier experiment is confirmed. The new results are then viewed from a different perspective in order to evaluate the usefulness of this simulation strategy.

Miner, Paul S.↗

Earth to Moon Transfer: Direct vs Via Libration Points (L1, L2)

For some three decades, the Apollo-style mission has served as a proven baseline technique for transporting flight crews to the Moon and back with expendable hardware. This approach provides an optimal design for expeditionary missions, emphasizing operational flexibility in terms of safely returning the crew in the event of a hardware failure. However, its application is limited essentially to low-latitude lunar sites, and it leaves much to be desired as a model for exploratory and evolutionary programs that employ reusable space-based hardware. This study compares the performance requirements for a lunar orbit rendezvous mission type with one using the cislunar libration point (L1) as a stopover and staging point for access to arbitrary sites on the lunar surface. For selected constraints and mission objectives, it contrasts the relative uniformity of performance cost when the L1 staging point is used with the wide variation of cost for the Apollo-style lunar orbit rendezvous.

Condon, Gerald L.↗

ISS Internal Active Thermal Control System (IATCS) Coolant Remediation Project

The IATCS coolant has experienced a number of anomalies in the time since the US Lab was first activated on Flight 5A in February 2001. These have included: 1) a decrease in coolant pH, 2) increases in inorganic carbon, 3) a reduction in phosphate buffer concentration, 4) an increase in dissolved nickel and precipitation of nickel salts, and 5) increases in microbial concentration. These anomalies represent some risk to the system, have been implicated in some hardware failures and are suspect in others. The ISS program has conducted extensive investigations of the causes and effects of these anomalies and has developed a comprehensive program to remediate the coolant chemistry of the on-orbit system as well as provide a robust and compatible coolant solution for the hardware yet to be delivered. The remediation steps include changes in the coolant chemistry specification, development of a suite of new antimicrobial additives, and development of devices for the removal of nickel and phosphate ions from the coolant. This paper presents an overview of the anomalies, their known and suspected system effects, their causes, and the actions being taken to remediate the coolant.

Morrison, Russell H.↗