Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “failure recovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Engineering challenges of in-flight spacecraft - Voyager: A case study

Some of the engineering problems encountered during the post-launch phase of interplanetary space missions are described, with emphasis given to the Voyager missions. The major in-flight modifications in Voyager spacecraft's operational capability with respect to communications, payload, and navigation systems are discussed. Attention is given to the instances of 'failure workaround' including: recovery from a failed receiver, recovery from a seized scan platform actuator, and modifications to the Attitude Articulation and Control Subsystem (AACS) software during the Saturn encounter. A detailed line drawing of the Voyager spacecraft is provided.

Jones, C. P.↗

An approximation formula for a class of fault-tolerant computers

An approximation formula is derived for the probability of failure for fault-tolerant process-control computers. These computers use redundancy and reconfiguration to achieve high reliability. Finite-state Markov models capture the dynamic behavior of component failure and system recovery, and the approximation formula permits an estimation of system reliability by an easy examination of the model.

White, A. L.↗

Biconic cargo return vehicle with an advanced recovery system. Volume 1: Conceptual design

The conceptual design of the biconic Cargo Return Vehicle (CRV) is presented. The CRV will be able to meet all of the Space Station Freedom (SSF's) resupply needs. Worth note is the absence of a backup recovery chute in case of Advanced Recovery System (ARS) failure. The high reliability of ram-air parachutes does not warrant the penalty weight that such a system would create on successful missions. The CRV will launch vertically integrated with an Liquid Rocket Booster (LRB) vehicle and meets all NASA restrictions on fuel type for all phases of the mission. Because of the downscaled Orbital Maneuvering Vehicle (OMV) program, the CRV has been designed to be able to transfer cargo by docking directly to the Space Station Freedom as well as with OMV assistance. The CRV will cover enough crossrange to reach its primary landing site, Edwards Airforce Base, and all secondary landing sites with the exception of one orbit. Transportation back to KSC will be via the Boeing Super Guppy. Due to difficulties with man-rating the CRV, it will not be used in a CERV role. A brief summary of the CRV's specifications is given.

Source record↗

Attitude and articulation control for CRAF/Cassini

The Comet Rendezvous/Asteroid Flyby (CRAF) and Cassini planetary missions provide exciting pointing and control challenges. The mission and science objectives, and an attitude and articulation control concept designed to meet these challenges, are described. CRAF/Cassini mission characteristics which drive pointing and control include: close range flybys of asteroids and icy satellites; Huygens probe guidance and communication; Saturn orbit insertion; comet rendezvous and orbit insertion; closed loop target tracking from a comet orbit perturbed by gas and dust pressure; fine spacecraft pointing for Titan radar mapping and Earth communications; requirements for autonomous failure detection; isolation; recovery; and 13.5 year lifetime. The philosophy and approach chosen to meet these challenges and the overall control architecture are addressed, including operational and autonomous safe modes. Critical functions are highlighted, such as charge coupled device imaging of stars and extended bodies which provide references for inertial and target referenced pointing respectively. Tradeoffs and rationale for the selection and location of sensors and actuators are reviewed.

Bell, C. E.↗

Post irradiation effects (PIE) in integrated circuits

Post-irradiation effects (PIE) ranging from normal recovery to catastrophic failure have been observed in integrated circuits during the PIE period. Data presented show failure due to rebound after a 10 krad(Si) dose. In particular, five device types are investigated with varying PIE response. Special attention has been given to the HI1-507A analog multiplexer because its PIE response is extreme. X-ray diffraction has been uniquely employed to measure physical stress in the HI1-507A metallization. An attempt has been made to show a relationship between stress relaxation and radiation effects. All data presented support the current MIL-STD Method 1019.4 but demonstrate the importance of performing PIE measurements, even when mission doses are as low as 10 krad(Si).

Shaw, D. C.↗

A measurement-based performability model for a multiprocessor system

A measurement-based performability model based on real error-data collected on a multiprocessor system is described. Model development from the raw errror-data to the estimation of cumulative reward is described. Both normal and failure behavior of the system are characterized. The measured data show that the holding times in key operational and failure states are not simple exponential and that semi-Markov process is necessary to model the system behavior. A reward function, based on the service rate and the error rate in each state, is then defined in order to estimate the performability of the system and to depict the cost of different failure types and recovery procedures.

Ilsueh, M. C.↗

Timeliner: Automating Procedures on the ISS

Timeliner has been developed as a tool to automate procedural tasks. These tasks may be sequential tasks that would typically be performed by a human operator, or precisely ordered sequencing tasks that allow autonomous execution of a control process. The Timeliner system includes elements for compiling and executing sequences that are defined in the Timeliner language. The Timeliner language was specifically designed to allow easy definition of scripts that provide sequencing and control of complex systems. The execution environment provides real-time monitoring and control based on the commands and conditions defined in the Timeliner language. The Timeliner sequence control may be preprogrammed, compiled from Timeliner "scripts," or it may consist of real-time, interactive inputs from system operators. In general, the Timeliner system lowers the workload for mission or process control operations. In a mission environment, scripts can be used to automate spacecraft operations including autonomous or interactive vehicle control, performance of preflight and post-flight subsystem checkouts, or handling of failure detection and recovery. Timeliner may also be used for mission payload operations, such as stepping through pre-defined procedures of a scientific experiment.

Brown, Robert↗

Cycle Time Reduction in Trapped Mercury Ion Atomic Frequency Standards

The use of the mercury ion isotope (201)Hg(+) was examined for an atomic clock. Taking advantage of the faster optical pumping time in (201)Hg(+) reduces both the state preparation and the state readout times, thereby decreasing the overall cycle time of the clock and reducing the impact of medium-term LO noise on the performance of the frequency standard. The spectral overlap between the plasma discharge lamp used for (201)Hg(+) state preparation and readout is much larger than that of the lamp used for the more conventional (199)Hg(+). There has been little study of (201)Hg(+) for clock applications (in fact, all trapped ion clock work in mercury has been with (199)Hg(+); however, recently the optical pumping time in (201)Hg(+) has been measured and found to be 0.45 second, or about three times faster than in (199)Hg(+) due largely to the better spectral overlap. This can be used to reduce the overall clock cycle time by over 2 seconds, or up to a factor of 2 improvement. The use of the (201)Hg(+) for an atomic clock is totally new. Most attempts to reduce the impact of LO noise have focused on reducing the interrogation time. In the trapped ion frequency standards built so far at JPL, the optical pumping time is already at its minimum so that no enhancement can be had by shortening it. However, by using (201)Hg(+), this is no longer the case. Furthermore, integrity monitoring, the mechanism that determines whether the clock is functioning normally, cannot happen faster than the clock cycle time. Therefore, a shorter cycle time will enable quicker detection of failure modes and recovery from them.

Burt, Eric A.↗

Distributed Prognostics and Health Management with a Wireless Network Architecture

A heterogeneous set of system components monitored by a varied suite of sensors and a particle-filtering (PF) framework, with the power and the flexibility to adapt to the different diagnostic and prognostic needs, has been developed. Both the diagnostic and prognostic tasks are formulated as a particle-filtering problem in order to explicitly represent and manage uncertainties in state estimation and remaining life estimation. Current state-of-the-art prognostic health management (PHM) systems are mostly centralized in nature, where all the processing is reliant on a single processor. This can lead to a loss in functionality in case of a crash of the central processor or monitor. Furthermore, with increases in the volume of sensor data as well as the complexity of algorithms, traditional centralized systems become for a number of reasons somewhat ungainly for successful deployment, and efficient distributed architectures can be more beneficial. The distributed health management architecture is comprised of a network of smart sensor devices. These devices monitor the health of various subsystems or modules. They perform diagnostics operations and trigger prognostics operations based on user-defined thresholds and rules. The sensor devices, called computing elements (CEs), consist of a sensor, or set of sensors, and a communication device (i.e., a wireless transceiver beside an embedded processing element). The CE runs in either a diagnostic or prognostic operating mode. The diagnostic mode is the default mode where a CE monitors a given subsystem or component through a low-weight diagnostic algorithm. If a CE detects a critical condition during monitoring, it raises a flag. Depending on availability of resources, a networked local cluster of CEs is formed that then carries out prognostics and fault mitigation by efficient distribution of the tasks. It should be noted that the CEs are expected not to suspend their previous tasks in the prognostic mode. When the prognostics task is over, and after appropriate actions have been taken, all CEs return to their original default configuration. Wireless technology-based implementation would ensure more flexibility in terms of sensor placement. It would also allow more sensors to be deployed because the overhead related to weights of wired systems is not present. Distributed architectures are furthermore generally robust with regard to recovery from node failures.

Goebel, Kai↗

Capture Latch Assembly for the NASA Docking System

Final Paper and not the abstract is attached. a summary of the Design, Development, and Qualification of the Capture Latch Assembly (CLA) for the NASA Docking System Block 1 (NDSB1). The CLA is an integral part of the Soft Capture System (SCS) of the NDSB1, serving the purpose of connecting the mating SCS Rings of two docking vehicles. The paper will present an overview of the function of the CLA and its basic concept of operations, including a summary of the major components of the CLA. The development, qualification, and production of the CLA will then be described. Particular focus will be provided on two major issues that occurred during production and qualification of the CLA. The first issue was failures of the CLA Motor (CLM) during acceptance testing (AT). The failures of the CLM were ultimately determined to be due to design defects and manufacturing errors in the motor commutation sensor assembly. The second issue was failure of the secondary release mechanism, or Contingency Capture Latch Release (CCLR) mechanism during development and qualification testing. The CCLR failures were found to be a result of excess free play in the release mechanism, resulting in wear leading to galling inside the release mechanism. An overview of each failure will be provided, along with a summary of the failure investigation and recovery process. Finally, Lessons Learned from each of the major issues and the overall development of the Capture Latch will be presented.

Capture Latch↗

Automated Impact Assessment: A New Approach to ISS Payload Operations Anomaly Response

The International Space Station (ISS) Payload Operations and Integration Center (POIC) is undergoing rapid growth as the space station program focuses on science and commercial activities. The ISS is expanding its onboard capabilities to support additional science activities. In parallel, the POIC is expanding the capabilities of our operations tools to support the higher pace of payload activities being executed each week. An effect of these changes is that anomaly resolution has become more challenging. In the event of a real-time system fault, operators are responsible for analyzing telemetry displays, anomaly monitoring tools, documentation, and system models in order to produce failure impacts and recovery strategies. This approach to operations relies on the operator to ingest, process, and analyze information from an array of deterministic sources to provide actionable data on impacted systems and activities. Changing the existing approach of anomaly response is necessary if the ISS community is to succeed in the age of science and commercialization of space. The creation of a tool that captures deterministic technical systems knowledge and integrates existing telemetry, documentation, and planning information will allow the burden of impact assessment to be automated, thereby allowing the operator to focus on non-deterministic tasks, such as recovering failed systems and restoring critical payload operations.

Hall, R. Mason↗

Effects of Communication Delay on Human Spaceflight Missions

Missions onboard the International Space Station rely on the real-time availability of a large ground team of system experts to command the vehicle, solve safety-critical problems, and guide the crew during complex operations. Also, in Low Earth Orbit (LEO), supplies can be sent and crews evacuated quite quickly if needed. Future missions Beyond Low Earth Orbit (BLEO) will not have this 24/7, real-time safety net as communication latency increases, resupply difficulty increases, and evacuation opportunities diminish. There are few, if any, terrestrial analogs for human spaceflight missions BLEO that reflect the conditions—including extreme environments, long mission durations, and small crew sizes – that make these missions so high risk. Studies on specific conditions, such as communication delays and asynchronous interactions, have been performed in NASA Earth-based analog missions and have found that communication delays can disrupt ground-crew interactions and adversely impact team performance. However, there are gaps and limitations in studies conducted to date, notably on human spacecraft system failure response and recovery, the impacts of shorter lunar-relevant communication delays on complex operations, and the effectiveness of countermeasures. The work presented here breaks down real anomalies that occurred on ISS and Apollo missions and creates example scenarios fort Lunar Surface and Mars missions to explore the impact of communication delays of varying length on onboard operations and mission outcomes. Our analyses indicate that short communication delays (e.g., seconds to a minute) adversely impact the ability for ground to provide real-time oversight and guidance and to catch quickly emerging problems in time. Longer communication delays (e.g., up to 40 minutes on Mars missions) call for a shift of responsibility for tactical operations from ground to crew; crew must make time-critical decisions independently and respond to time-critical vehicle anomalies to prevent consequences.

human-systems integration↗

Effects of Communication Delay on Human Spaceflight Missions

Missions onboard the International Space Station rely on the real-time availability of a large ground team of system experts to command the vehicle, solve safety-critical problems, and guide the crew during complex operations. Also, in Low Earth Orbit (LEO), supplies can be sent and crews evacuated quite quickly if needed. Future missions Beyond Low Earth Orbit (BLEO) will not have this 24/7, real-time safety net as communication latency increases, resupply difficulty increases, and evacuation opportunities diminish. There are few, if any, terrestrial analogs for human spaceflight missions BLEO that reflect the conditions—including extreme environments, long mission durations, and small crew sizes – that make these missions so high risk. Studies on specific conditions, such as communication delays and asynchronous interactions, have been performed in NASA Earth-based analog missions and have found that communication delays can disrupt ground-crew interactions and adversely impact team performance. However, there are gaps and limitations in studies conducted to date, notably on human spacecraft system failure response and recovery, the impacts of shorter lunar-relevant communication delays on complex operations, and the effectiveness of countermeasures. The work presented here breaks down real anomalies that occurred on ISS and Apollo missions and creates example scenarios fort Lunar Surface and Mars missions to explore the impact of communication delays of varying length on onboard operations and mission outcomes. Our analyses indicate that short communication delays (e.g., seconds to a minute) adversely impact the ability for ground to provide real-time oversight and guidance and to catch quickly emerging problems in time. Longer communication delays (e.g., up to 40 minutes on Mars missions) call for a shift of responsibility for tactical operations from ground to crew; crew must make time-critical decisions independently and respond to time-critical vehicle anomalies to prevent consequences.

human-systems integration↗

Proof that timing requirements of the FDDI token ring protocol are satisfied

The fiber distributed data interface (FDDI) is an ANSI draft proposed standard for a 100 Mbit/s fiber-optic token ring. The FDDI timed token access protocol provides dynamic adjustment of the load offered to the ring, with the goal of maintaining a specified token rotation time and of providing a guaranteed upper bound on time between successive arrivals of the token at a station. FDDI also provides automatic recovery when errors occur. The bound on time between successive token arrivals is guaranteed only if the token rotates quickly enough to satisfy timer requirements in each station when all ring resources are functioning properly. Otherwise, recovery would be initiated unnecessarily. The purpose of this paper is to prove that FDDI timing requirements are satisfied, i.e., the token rotates quickly enough to prevent initiation of recovery unless there is failure of a physical resource or unless the network management entity within a station initiates the recovery process.

Johnson, Marjory J.↗

Integrating static PRA information with risk informed safety margin characterization (RISMC) simulation methods

The overall objective of the project was to develop a computationally feasible and user-friendly mechanized process to integrate traditional probabilistic risk assessment (PRA) and dynamic PRA (DPRA) results. Starting with the systematic identification of items in an existing PRA that need dynamic augmentation, the project used a generic 4-loop pressurized reactor (PWR) and 3-loop PWR as example plants. Station blackout (SBO) and large break loss of coolant accident SBLOCA) were selected as the example initiating events. Using the traditional event-tree (ET)/fault-tree (FT) methodology augmented by dynamic evet tree approach, the potential consequences of the initiating events were simulated with RELAP-3D and MELCOR/RASCAL codes to cover Level 1 through Level 3 of PRA. RAVEN and ADAPT software were used to generate Level 1 simulations with RELAP-3D and Level 2/3 simulations with MELCOR (Level 2)/RASCAL (Level 3), respectively. Example branching conditions (BCs) for SBO included AC power recovery time, valve repair failure time, reactor coolant pump leak time/break size and emergency power supply duration to a total of 9. Example BCs for LOCA included off-site power recovery time, diesel generator power recovery time, auxiliary feed water system operation time, safety relief valve failure to open upon demand, reactor coolant pump seal break time and size to a total of 21. Each RELAP-3D simulation (9,587 scenarios) was labelled OK or Core Damage based on the maximum allowed peak clad temperature (2,100oF). Each MELCOR simulation (4610 scenarios) was labeled as Bin over 10rem or Bin 0-10rem based on the dose at the site boundary. The scenarios were clustered based on the criteria above using the mean shift methodology. Classical PRA (CPRA) and DPRA results were compared to identify the ET sequences that need DPRA augmentation. Several approaches were proposed for the incorporation of these sequences into CPRA using clustering with the mean shift methodology, restructuring the CPRA ETs by adding new BCs/sequences, and using the concept of a limit surface. Procedures for decision making regarding the possible consequences of an initiating event (e.g. core damage or not, site evacuation or not) were developed using a convolutional neural network (CNN), a recurrent neural network (RNN) and a transformer neural network (TNN). The project has led to two PhD degrees, three archival journal papers and five refereed conference proceedings.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Electro-Thermo-Mechanical (ETM) Study on Submarine Dynamic Power Cables (SDPC) for Offshore Wind Transmission

Understanding the failure mechanisms of submarine dynamic power cables (SDPC) is critical for innovative design to meet the 2035 cost reduction target of U.S. DOE Floating Offshore Wind Shot. This is important because the current design suffers a significant failure rate in the field. This project conducted a systematic electro-thermo-mechanical (ETM) experimental study on the power cores extracted from a 15kV power cable with three cores of copper conductor, ethylene propylene rubber (EPR) insulation, and continuously corrugated welded aluminum armor (CCWA). An ETM testing system was developed by integrating a high voltage (HV) amplifier, two ceramic heaters, and a rod-plate transverse compression setup into a mechanical testing machine. Increasing temperature from room temperature (RT, 22 degree C) to 90 degree C resulted in the 67% decrease in the failure mechanical load as defined by the dielectric breakdown. Under creep mechanical loading, the dielectric breakdown time was decreased by 70% for a given mechanical load when the specimen temperature increased from RT to 90oC. The failure strain in both monotonic and creep mechanical loading modes was related to the maximum mechanical load applied. Although the dielectric breakdown occurred in the ETM test, the post-test measurement revealed an impressive recovery of electric resistance. The failure analysis on cross section of tested specimens indicated a sizable gap across insulation layer near the area between copper conductor and loading rod, which apparently resulted from online dielectric failure and offline elastic recovery of components including conductor and insulation layer.

17 WIND ENERGY↗

Enhanced Component Performance Study: Emergency Diesel Generators 1998-2020

This report presents an enhanced performance evaluation of emergency power system (EPS) and high-pressure core spray (HPCS) emergency diesel generators (EDGs) at U.S. commercial nuclear power plants. This report evaluates component performance over time using (1) Institute of Nuclear Power Operations (INPO) Industry Reporting and Information System (IRIS) data from 1998 through 2020 and (2) maintenance unavailability performance data from Mitigating Systems Performance Index (MSPI) Basis Document data from 2002 through 2020. The objective is to show estimates of current failure probabilities and rates related to EDGs, trend these data on an annual basis, determine if the current data are consistent with the probability distributions currently recommended for use in NRC probabilistic risk assessments, show how the reliability data differ for different EDG manufacturers and for EDGs with different ratings; and summarize the subcomponents, causes, detection methods, and recovery associated with each EDG failure mode. The EDG failure modes considered are fail to start (FTS), fail to load and run (FTLR), and fail to run after one hour of operation (FTR>1H). Engineering analyses were performed with respect to time period and failure mode without regard to the actual number of EDGs at each plant. The factors analyzed are: subcomponent, failure cause, detection method, recovery, manufacturer, and EDG rating. The following increasing trends were identified for EDGs for the most recent 10-year period: • HPCS EDG FTS failure probability • EPS and HPCS EDG frequency of start demands (demands per reactor year) • EPS and HPCS EDG frequency of FTLR demands. The following decreasing trends were identified for EDGs for the most recent 10-year period: • EPS EDG FTS failure probability • EPS EDG FTLR failure probability • EPS EDG unavailability • EPS and HPCS EDG frequency of FTLR events (failures per reactor year).

99 GENERAL AND MISCELLANEOUS↗

Enhanced Component Performance Study: Emergency Diesel Generators 1998–2022

This report presents an enhanced performance evaluation of the emergency power system (EPS) and high-pressure core spray (HPCS) emergency diesel generators (EDGs) at U.S. commercial nuclear power plants. This report evaluates component performance over time using (1) Institute of Nuclear Power Operations (INPO) Industry Reporting and Information System (IRIS) data from 1998 through 2022 and (2) maintenance unavailability performance data from Mitigating Systems Performance Index (MSPI) Basis Document data from 2002 through 2022. The objective is to show estimates of current failure probabilities and rates related to EDGs, trend these data on an annual basis, determine if the current data are consistent with the probability distributions currently recommended for use in Nuclear Regulatory Commission (NRC) probabilistic risk assessments, show how the reliability data differ for different EDG manufacturers and for EDGs with different ratings; and summarize the subcomponents, causes, detection methods, and recovery associated with each EDG failure mode. The EDG failure modes considered are fail to start (FTS), fail to load and run (FTLR), and fail to run after one hour of operation (FTR>1H). Engineering analyses were performed with respect to time-period and failure mode without regard to the actual number of EDGs at each plant. The factors analyzed include subcomponent, failure cause, detection method, recovery, manufacturer, and EDG rating. The following increasing trends were identified for EDGs for the most recent 10-year period: (1) HPCS EDG FTS failure probability, (2) EPS and HPCS EDG frequency of start demands (demands per reactor year), and (3) EPS and HPCS EDG frequency of FTLR demands. The following decreasing trends were identified for EDGs for the most recent 10-year period: (1) EPS EDG FTS failure probability, (2) EPS EDG FTLR failure probability, (3) EPS EDG unavailability, and (4) EPS and HPCS EDG frequency of FTLR events (failures per reactor year).

22 GENERAL STUDIES OF NUCLEAR REACTORS↗