Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “failure recovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Failure Analysis and Recovery of a 50mm Highly Elastic Intermetallic NiTi Ball Bearing for an ISS Application

The ISS Distillation Assembly centrifuge is the pathfinder application for 50mm bore, deep-groove ball bearings made from the highly elastic intermetallic material 60NiTi. Superior corrosion and shock resistance are required to withstand the acidic wastewater exposure and heavy spacecraft launch related loads that challenge conventional steel bearings. During early ground testing one bearing unexpectedly and catastrophically failed after operating for only 200 hours of run time. A second bearing running on the same shaft was completely unaffected. A thorough investigation into the root cause of the failure determined that an excessively tight press-fit of the bearing outer race coupled with NiTis relatively low elastic modulus were key contributing factors. The proposed failure mode was successfully duplicated by experiment. To further corroborate the root cause theory, a successful bearing life test using improved installation practices (selective fitting) was conducted. The results show that NiTi bearings are suitable for space applications provided that care is taken to accommodate their unique material characteristics.

intermetallics↗

Failure Analysis and Recovery of a 50 MM Highly Elastic Intermetallic NiTi Ball Bearing for an ISS Application

The ISS Distillation Assembly centrifuge is the pathfinder application for 50mm bore, deep-groove ball bearings made from the highly elastic intermetallic material 60NiTi. Superior corrosion and shock resistance are required to withstand the acidic wastewater exposure and heavy spacecraft launch related loads that challenge conventional steel bearings. During early ground testing one bearing unexpectedly and catastrophically failed after operating for only 200 hours of run time. A second bearing running on the same shaft was completely unaffected. A thorough investigation into the root cause of the failure determined that an excessively tight press-fit of the bearing outer race coupled with NiTis relatively low elastic modulus were key contributing factors. The proposed failure mode was successfully duplicated by experiment. To further corroborate the root cause theory, a successful bearing life test using improved installation practices (selective fitting) was conducted. The results show that NiTi bearings are suitable for space applications provided that care is taken to accommodate their unique material characteristics.

failure analyses↗

Failure Analysis and Recovery of a 50-mm Highly Elastic Intermetallic NiTi Ball Bearing for an ISS Application

Ball bearings used inside the ISS Distillation Assembly centrifuge require superior corrosion and shock resistance to withstand acidic wastewater exposure and heavy spacecraft launch related loads. These requirements challenge conventional steel bearings and provide an ideal pathfinder application for 50-mm bore, deep-groove ball bearings made from the corrosion immune and highly elastic intermetallic material 60NiTi. During early ground testing in 2014 one 60NiTi bearing unexpectedly and catastrophically failed after operating for only 200 hr. A second bearing running on the same shaft was completely unaffected. An investigation into the root cause of the failure determined that an excessively tight press fit of the bearing outer race coupled with NiTi's relatively low elastic modulus were key contributing factors. The proposed failure mode was successfully replicated by experiment. To further corroborate the root cause theory, a successful bearing life test using improved installation practices (selective fitting) was conducted. The results show that NiTi bearings are suitable for space applications provided that care is taken to accommodate their unique material characteristics.

DellaCorte, Christopher↗

Optimal Aircraft Control Upset Recovery With and Without Component Failures

This paper treats the problem of recovering sustainable nondescending (safe) flight in a transport aircraft after one or more of its control effectors fail. Such recovery can be a challenging goal for many transport aircraft currently in the operational fleet for two reasons. First, they have very little redundancy in their means of generating control forces and moments. These aircraft have, as primary control surfaces, a single rudder and pairwise elevators and aileron/spoiler units that provide yaw, pitch, and roll moments with sufficient bandwidth to be used in stabilizing and maneuvering the airframe. Beyond this, throttling the engines can provide additional moments, but on a much slower time scale. Other aerodynamic surfaces, such as leading and trailing edge flaps, are only intended to be placed in a position and left, and are, hence, very slow-moving. Because of this, loss of a primary control surface strongly degrades the controllability of the vehicle, particularly when the failed effector becomes stuck in a non-neutral position where it exerts a disturbance moment that must be countered by the remaining operating effectors. The second challenge in recovering safe flight is that these vehicles are not agile, nor can they tolerate large accelerations. This is of special importance when, at the outset of the recovery maneuver, the aircraft is flying toward the ground, as is frequently the case when there are major control hardware failures. Recovery of safe flight is examined in this paper in the context of trajectory optimization. For a particular transport aircraft, and a failure scenario inspired by an historical air disaster, recovery scenarios are calculated with and without control surface failures, to bring the aircraft to safe flight from the adverse flight condition that it had assumed, apparently as a result of contact with a vortex from a larger aircraft's wake. An effort has been made to represent relevant airframe dynamics, acceleration limits, and actuator limits faithfully, since these contribute to the lack of agility and control power that plays an important role in defining what can be achieved with the vehicle when it is in extremis.

Sparks, Dean W.↗

Techniques for Improving Pilot Recovery from System Failures

This project examined the application of intelligent cockpit systems to aid air transport pilots at the tasks of reacting to in-flight system failures and of planning and then following a safe four dimensional trajectory to the runway threshold during emergencies. Two studies were conducted. The first examined pilot performance with a prototype awareness/alerting system in reacting to on-board system failures. In a full-motion, high-fidelity simulator, Army helicopter pilots were asked to fly a mission during which, without warning or briefing, 14 different failures were triggered at random times. Results suggest that the amount of information pilots require from such diagnostic systems is strongly dependent on their training; for failures they are commonly trained to react to with a procedural response, they needed only an indication of which failure to follow, while for 'un-trained' failures, they benefited from more intelligent and informative systems. Pilots were also found to over-rely on the system in conditions were it provided false or mis-leading information. In the second study, a proof-of-concept system was designed suitable for helping pilots replan their flights in emergency situations for quick, safe trajectory generation. This system is described in this report, including: the use of embedded fast-time simulation to predict the trajectory defined by a series of discrete actions; the models of aircraft and pilot dynamics required by the system; and the pilot interface. Then, results of a flight simulator evaluation with airline pilots are detailed. In 6 of 72 simulator runs, pilots were not able to establish a stable flight path on localizer and glideslope, suggesting a need for cockpit aids. However, results also suggest that, to be operationally feasible, such an aid must be capable of suggesting safe trajectories to the pilot; an aid that only verified plans entered by the pilot was found to have significantly detrimental effects on performance and pilot workload. Results also highlight that the trajectories suggested by the aid must capture the context of the emergency; for example, in some emergencies pilots were willing to violate flight envelope limits to reduce time in flight - in other emergencies the opposite was found.

Pritchett, Amy R.↗

A resilient network recovery framework against cascading failures with deep graph learning

Because of the increasing importance and dependencies of infrastructure networks and the potential for massive cascading failures in real-world network systems, maintenance optimization to effectively reduce system performance loss caused by diverse disruptions is of significant interest among researchers and practitioners. In this work, a new recovery framework was developed to rapidly identify important system components for maintenance to improve network resilience against cascading failures. Here this work provides distinct advantages to determine an optimal maintenance priority by combining real-time network structure importance with other maintenance prioritization based on customer preference. This approach adopts structural graph embedding and deep reinforcement learning to extract real-time network topology information (such as minimum vertex cover) to update the maintenance priority during the recovery process. Based on the case studies on synthetic networks and a US airport network, the proposed recovery framework with real-time network topology awareness shows better performance than other maintenance prioritization strategies regarding resilience enhancement. This work improves the understanding of how the changing network structure influences maintenance effects. It also provides insights of the practical usefulness of advanced deep learning on helping optimal maintenance prioritization to effectively reduce the intensity and extent of cascading failures.

42 ENGINEERING↗

Performance evaluation of redundant disk array support for transaction recovery

Redundant disk arrays provide a way of achieving rapid recovery from media failures with a relatively low storage cost for large scale data systems requiring high availability. Here, we propose a method for using redundant disk arrays to support rapid recovery from system crashes and transaction aborts in addition to their role in providing media failure recovery. A twin page scheme is used to store the parity information in the array so that the time for transaction commit processing is not degraded. Using an analytical model, we show that the proposed method achieves a significant increase in the throughput of database systems using redundant disk arrays by reducing the number of recovery operations needed to maintain the consistency of the database.

Mourad, Antoine N.↗

Recovery issues in databases using redundant disk arrays

Redundant disk arrays provide a way for achieving rapid recovery from media failures with a relatively low storage cost for large scale database systems requiring high availability. In this paper we propose a method for using redundant disk arrays to support rapid recovery from system crashes and transaction aborts in addition to their role in providing media failure recovery. A twin page scheme is used to store the parity information in the array so that the time for transaction commit processing is not degraded. Using an analytical model, we show that the proposed method achieves a significant increase in the throughput of database systems using redundant disk arrays by reducing the number of recovery operations needed to maintain the consistency of the database.

Mourad, Antoine N.↗

Database recovery using redundant disk arrays

Redundant disk arrays provide a way for achieving rapid recovery from media failures with a relatively low storage cost for large scale database systems requiring high availability. In this paper a method is proposed for using redundant disk arrays to support rapid-recovery from system crashes and transaction aborts in addition to their role in providing media failure recovery. A twin page scheme is used to store the parity information in the array so that the time for transaction commit processing is not degraded. Using an analytical model, it is shown that the proposed method achieves a significant increase in the throughput of database systems using redundant disk arrays by reducing the number of recovery operations needed to maintain the consistency of the database.

Mourad, Antoine N.↗

A simulator investigation of engine failure compensation for powered-lift STOL aircraft

A piloted simulator investigation of various engine failure compensation concepts for powered-lift STOL aircraft was carried out at the Ames Research Center. The purpose of this investigation was to determine the influence of engine failure compensation on recovery from an engine failure during the landing approach and on the precision of the STOL landing. The various concepts include: (1) cockpit warning lights to cue the pilot of an engine failure, (2) programmed thrust and roll trim compensation, (3) thrust command and (4) flight-path stabilization. The aircraft simulated was a 150 passenger four-engine, externally blown flap civil STOL transport having a 90 psf wing loading and a .56 thrust to weight ratio. Results of the simulation indicate that the combination of thrust command and flight-path stabilization offered the best engine-out landing performance in turbulence and did so over the entire range of altitudes for which engine failures occurred.

Nieuwenhuijse, A. W.↗

Virtually-synchronous communication based on a weak failure suspector

Failure detectors (or, more accurately Failure Suspectors (FS)) appear to be a fundamental service upon which to build fault-tolerant, distributed applications. This paper shows that a FS with very weak semantics (i.e., that delivers failure and recovery information in no specific order) suffices to implement virtually-synchronous communication (VSC) in an asynchronous system subject to process crash failures and network partitions. The VSC paradigm is particularly useful in asynchronous systems and greatly simplifies building fault-tolerant applications that mask failures by replicating processes. We suggest a three-component architecture to implement virtually-synchronous communication: (1) at the lowest level, the FS component; (2) on top of it, a component (2a) that defines new views; and (3) a component (2b) that reliably multicasts messages within a view. The issues covered in this paper also lead to a better understanding of the various membership service semantics proposed in recent literature.

Schiper, Andre↗

ISS Regenerative Life Support: Challenges and Success in the Quest for Long-Term Habitability in Space

This presentation will discuss the International Space Station s (ISS) Regenerative Environmental Control and Life Support System (ECLSS) operations with discussion of the on-orbit lessons learned, specifically regarding the challenges that have been faced as the system has expanded with a growing ISS crew. Over the 10 year history of the ISS, there have been numerous challenges, failures, and triumphs in the quest to keep the crew alive and comfortable. Successful operation of the ECLSS not only requires maintenance of the hardware, but also management of the station resources in case of hardware failure or missed re-supply. This involves effective communication between the primary International Partners (NASA and Roskosmos) and the secondary partners (JAXA and ESA) in order to keep a reserve of the contingency consumables and allow for re-supply of failed hardware. The ISS ECLSS utilizes consumables storage for contingency usage as well as longer-term regenerative systems, which allow for conservation of the expensive resources brought up by re-supply vehicles. This long-term hardware, and the interactions with software, was a challenge for Systems Engineers when they were designed and require multiple operational workarounds in order to function continuously. On a day-to-day basis, the ECLSS provides big challenges to the on console controllers. Main challenges involve the utilization of the resources that have been brought up by the visiting vehicles prior to undocking, balance of contributions between the International Partners for both systems and resources, and maintaining balance between the many interdependent systems, which includes providing the resources they need when they need it. The current biggest challenge for ECLSS is the Regenerative ECLSS system, which continuously recycles urine and condensate water into drinking water and oxygen. These systems were brought to full functionality on STS-126 (ULF-2) mission. Through system failures and recovery, the ECLSS console has learned how to balance the water within the systems, store and use water for contingencies, and continue to work with the International Partners for short-term failures. Through these challenges and the system failures, the most important lesson learned has been the importance of redundancy and operational workarounds. It is only because of the flexibility of the hardware and the software that flight controllers have the opportunity to continue operating the system as a whole for mission success.

Bazley, Jesse A.↗

Simulation-Based Recovery Action Analysis Using the EMRALD Dynamic Risk Assessment Tool

A recovery action is defined as the action that prevents deviant conditions from producing unwanted effects. It generally indicates a kind of countermeasure performed in response to a failure of human action. The recovery actions especially play an important role in complex systems like nuclear power plants (NPPs), which consist of highly sophisticated controllers to ensure that desired performance and safety must be achieved and maintained. This is because a combination of human error and its recovery failure may be able to cause a catastrophic effect on a system. Analyzing recovery actions has been a critical part of HRA, which is a technique to evaluate human errors and provide human error probabilities (HEPs) for application in probabilistic safety assessment (PSA). If recovery actions are not adequately analyzed and applied to PSA models, the PSA results may be under-estimated or be not able to reasonably account for the failure of human actions in the context of PSA. For this reason, some regulatory documents such as ASME/ANS RA-Sb-2013 by the American Society for Mechanical Engineers and the American Nuclear Society and NUREG-1792 by U.S. Nuclear Regulatory Commission have emphasized the importance of recovery analysis within the HRA. A couple of existing HRA methods, such as the Technique for Human Error-Rate Prediction (THERP), the Cause-Based Decision Tree (CBDT), and the Korean Standard HRA (K-HRA), have respectively suggested their own approaches to the HRA recovery analysis. However, there are a couple of limitations to treating recovery actions using only the current HRA methods available. The biggest limitation is that the existing recovery analysis does not explicitly consider a variety of recovery action types and recovery sequences as they occur in actual NPPs. To handle the limitations of existing recovery analysis, this study proposes a simulation-based recovery analysis method using the Event Modeling Risk Assessment Using Linked Diagram (EMRALD) software. The EMRALD software is a dynamic simulation tool for PSA. It supports realistic and dynamic modeling of human actions as they would be performed at NPPs. It is also favorable to simultaneously model the specific moment at which an action is performed, the time it takes to perform the action, and the failure probability of that action. In this paper, a detailed methodology for modeling recovery actions in the simulation platform is proposed with a couple of examples. Then, outputs from the simulation are discussed as reviewing if this novel approach can complement the challenges of existing recovery analyses.

99 GENERAL AND MISCELLANEOUS↗

Galileo spacecraft anomaly and safing recovery

A high-level anomaly recovery plan which identifies the steps necessary to recover from a spacecraft 'Safing' incident was developed for the Galileo spacecraft prior to launch. Since launch, a total of four in-flight anomalies have lead to entry into a system fault protection 'Safing' routine which has required the Galileo flight team to refine and execute the recovery plan. These failures have allowed the flight team to develop an efficient recovery process when permanent spacecraft capability degradation is minimal and the cause of the anomaly is quickly diagnosed. With this previous recovery experience and the very focused boundary conditions of a specific potential failure, a Gaspra asteroid recovery plan was designed to be implemented in as quickly as forty hours (desired goal). This paper documents the work performed above, however, the Galileo project remains challenged to develop a generic detailed recovery plan which can be implemented in a relatively short time to configure the spacecraft to a nominal state prior to future high priority mission objectives.

Basilio, Ralph R.↗

Technical Specification Surveillance Interval Extension Using Self-Diagnostics

As part of the Light Water Reactor Sustainability program, an ongoing research effort is being conducted on technical specifications surveillance interval extension of digital equipment in nuclear power plants. The research team is led by Idaho National Laboratory and includes Pacific Northwest National Laboratory, Technology Resources, and Oak Ridge National Laboratory. This research focuses on developing methods for applying the U.S. Nuclear Regulatory Commission (NRC)–approved guidance to implement a licensee-controlled, risk-informed surveillance frequency change program in digital instrumentation and control (I&C) systems that include self-diagnostics and online monitoring (OLM) capabilities. Although approved methods exist for extending technical specifications (TS) surveillance test intervals (STIs) for general equipment, including analog I&C equipment, gaps remain in technology and guidance on crediting newer digital equipment’s internal self-diagnostics and OLM characteristics. Previous research described a general methodology for crediting internal self-diagnostics for extending surveillance test intervals. The methodology used self-diagnostics to detect—and credited recovery from—failure. Self-diagnostics were also applied for performance monitoring during the extended surveillance interval. This report discusses the status of recent activities to evaluate the previously developed methodology using a pilot study. Although both a utility partner for a pilot study and a specific digital asset were identified in FY2020, delays in obtaining proprietary information resulted in a limited ability to fully evaluate the methodology, and further interactions were complicated by the COVID pandemic. Therefore, at that time, the use of public-domain information—along with current processes for surveillance interval extension through a surveillance frequency control program—identified the need to fully assess diagnostic coverage as part of the pilot study. Furthermore, self-diagnostics were also identified as a potential option to replace the drift analyses conducted as part of current STI extension procedures. In FY2022, the project was reconstituted with the industry partner, and information and data were made available by the industry partner to the research team for review. The shared information included failure event descriptions and data for a digital I&C system since its implementation, as well as recent STI extension interval reports developed by the utility partner on that digital I&C system. This report presents an evaluation of this information and data and describes an application of the proposed methodology cited above. The methodology seeks to take advantage of the self-diagnostics and OLM capabilities to reduce risk or reduce the level of qualitative monitoring assessment needed to perform a risk-informed STI extension using existing NRC approved guidance or both. Addressing these issues of STI extension by crediting self-diagnostics is likely to result in benefits for current and future nuclear power plant (NPP) operations, including lowering the barriers to adoption of digital I&C systems and increasing cost savings by deferring or eliminating unneeded preventive maintenance (tasks or checks or activities). Specifically, self-diagnostic and OLM capabilities of newer digital equipment being installed in non-safety and safety applications are designed to detect failures, provide early warning of potential failures, and notify plant operators to take appropriate action to reduce out of service (OOS) time thus protecting safety margins. Moreover, the equipment is expected to provide information that time-related operational degradation is identified early to ensure timely and planned corrective actions instead of a reactive and unplanned approach ahead of an extended-surveillance interval.

42 ENGINEERING↗

The KATE shell: An implementation of model-based control, monitor and diagnosis

The conventional control and monitor software currently used by the Space Center for Space Shuttle processing has many limitations such as high maintenance costs, limited diagnostic capabilities and simulation support. These limitations have caused the development of a knowledge based (or model based) shell to generically control and monitor electro-mechanical systems. The knowledge base describes the system's structure and function and is used by a software shell to do real time constraints checking, low level control of components, diagnosis of detected faults, sensor validation, automatic generation of schematic diagrams and automatic recovery from failures. This approach is more versatile and more powerful than the conventional hard coded approach and offers many advantages over it, although, for systems which require high speed reaction times or aren't well understood, knowledge based control and monitor systems may not be appropriate.

Cornell, Matthew↗

Transforming AdaPT to Ada

This paper describes how the main features of the proposed Ada language extensions intended to support distribution, and offered as possible solutions for Ada9X can be implemented by transformation into standard Ada83. We start by summarizing the features proposed in a paper (Gargaro et al, 1990) which constitutes the definition of the extensions. For convenience we have called the language in its modified form AdaPT which might be interpreted as Ada with partitions. These features were carefully chosen to provide support for the construction of executable modules for execution in nodes of a network of loosely coupled computers, but flexibly configurable for different network architectures and for recovery following failure, or adapting to mode changes. The intention in their design was to provide extensions which would not impact adversely on the normal use of Ada, and would fit well in style and feel with the existing standard. We begin by summarizing the features introduced in AdaPT.

Goldsack, Stephen J.↗