Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “failure recovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Generative Design for Resilience of Interdependent Network Systems

Abstract Interconnected complex systems usually undergo disruptions due to internal uncertainties and external negative impacts such as those caused by harsh operating environments or regional natural disaster events. To maintain the operation of interconnected network systems under both internal and external challenges, design for resilience research has been conducted from both enhancing the reliability of the system through better designs and improving the failure recovery capabilities. As for enhancing the designs, challenges have arisen for designing a robust system due to the increasing scale of modern systems and the complicated underlying physical constraints. To tackle these challenges and design a resilient system efficiently, this study presents a generative design method that utilizes graph learning algorithms. The generative design framework contains a performance estimator and a candidate design generator. The generator can intelligently mine good properties from existing systems and output new designs that meet predefined performance criteria while the estimator can efficiently predict the performance of the generated design for a fast iterative learning process. Case studies results based on synthetic supply chain networks and power systems from the IEEE dataset have illustrated the applicability of the developed method for designing resilient interdependent network systems.

Engineering↗

A resilient network recovery framework against cascading failures with deep graph learning

Because of the increasing importance and dependencies of infrastructure networks and the potential for massive cascading failures in real-world network systems, maintenance optimization to effectively reduce system performance loss caused by diverse disruptions is of significant interest among researchers and practitioners. In this work, a new recovery framework was developed to rapidly identify important system components for maintenance to improve network resilience against cascading failures. Here this work provides distinct advantages to determine an optimal maintenance priority by combining real-time network structure importance with other maintenance prioritization based on customer preference. This approach adopts structural graph embedding and deep reinforcement learning to extract real-time network topology information (such as minimum vertex cover) to update the maintenance priority during the recovery process. Based on the case studies on synthetic networks and a US airport network, the proposed recovery framework with real-time network topology awareness shows better performance than other maintenance prioritization strategies regarding resilience enhancement. This work improves the understanding of how the changing network structure influences maintenance effects. It also provides insights of the practical usefulness of advanced deep learning on helping optimal maintenance prioritization to effectively reduce the intensity and extent of cascading failures.

42 ENGINEERING↗

Simulation-Based Recovery Action Analysis Using the EMRALD Dynamic Risk Assessment Tool

A recovery action is defined as the action that prevents deviant conditions from producing unwanted effects. It generally indicates a kind of countermeasure performed in response to a failure of human action. The recovery actions especially play an important role in complex systems like nuclear power plants (NPPs), which consist of highly sophisticated controllers to ensure that desired performance and safety must be achieved and maintained. This is because a combination of human error and its recovery failure may be able to cause a catastrophic effect on a system. Analyzing recovery actions has been a critical part of HRA, which is a technique to evaluate human errors and provide human error probabilities (HEPs) for application in probabilistic safety assessment (PSA). If recovery actions are not adequately analyzed and applied to PSA models, the PSA results may be under-estimated or be not able to reasonably account for the failure of human actions in the context of PSA. For this reason, some regulatory documents such as ASME/ANS RA-Sb-2013 by the American Society for Mechanical Engineers and the American Nuclear Society and NUREG-1792 by U.S. Nuclear Regulatory Commission have emphasized the importance of recovery analysis within the HRA. A couple of existing HRA methods, such as the Technique for Human Error-Rate Prediction (THERP), the Cause-Based Decision Tree (CBDT), and the Korean Standard HRA (K-HRA), have respectively suggested their own approaches to the HRA recovery analysis. However, there are a couple of limitations to treating recovery actions using only the current HRA methods available. The biggest limitation is that the existing recovery analysis does not explicitly consider a variety of recovery action types and recovery sequences as they occur in actual NPPs. To handle the limitations of existing recovery analysis, this study proposes a simulation-based recovery analysis method using the Event Modeling Risk Assessment Using Linked Diagram (EMRALD) software. The EMRALD software is a dynamic simulation tool for PSA. It supports realistic and dynamic modeling of human actions as they would be performed at NPPs. It is also favorable to simultaneously model the specific moment at which an action is performed, the time it takes to perform the action, and the failure probability of that action. In this paper, a detailed methodology for modeling recovery actions in the simulation platform is proposed with a couple of examples. Then, outputs from the simulation are discussed as reviewing if this novel approach can complement the challenges of existing recovery analyses.

99 GENERAL AND MISCELLANEOUS↗

Efficient Sampling of Complex Interdependent and Multiplex Networks

Efficient sampling of interdependent and multiplex infrastructure networks is critical for effectively applying failure and recovery algorithms in real-world settings, as well as to generate property-preserving reduced-order graph-based ensembles that address topological uncertainties. In this paper, we first explore the performance, i.e. the success in preserving graph properties, of graph sampling algorithms for interdependent and multiplex networks with synthetic and real-world graphs. We simulate sampling algorithms under different parameter settings. These settings include probabilistic graph generators, coupling patterns, and various performance metrics. Our results show that while Random Node and Random Walk sampling algorithms perform best for interdependent networks, Random Edge and Forest Fire sampling algorithms perform best for multiplex networks. Second, we propose and implement a novel similarity-based sampling algorithm for multiplex networks that samples only log(N) number of layers of an N-layer multiplex network while yielding computational savings with performance guarantees. Experimental results show that similarity sampling outperforms complete sampling of all layers while decreasing performance costs from a linear scale to a logarithmic one. Our results also indicate that similarity-based sampling outperforms complete sampling and random selection in nearly all scenarios when tested with real-world data.

Subasi, Omer↗

Technical Specification Surveillance Interval Extension Using Self-Diagnostics

As part of the Light Water Reactor Sustainability program, an ongoing research effort is being conducted on technical specifications surveillance interval extension of digital equipment in nuclear power plants. The research team is led by Idaho National Laboratory and includes Pacific Northwest National Laboratory, Technology Resources, and Oak Ridge National Laboratory. This research focuses on developing methods for applying the U.S. Nuclear Regulatory Commission (NRC)–approved guidance to implement a licensee-controlled, risk-informed surveillance frequency change program in digital instrumentation and control (I&C) systems that include self-diagnostics and online monitoring (OLM) capabilities. Although approved methods exist for extending technical specifications (TS) surveillance test intervals (STIs) for general equipment, including analog I&C equipment, gaps remain in technology and guidance on crediting newer digital equipment’s internal self-diagnostics and OLM characteristics. Previous research described a general methodology for crediting internal self-diagnostics for extending surveillance test intervals. The methodology used self-diagnostics to detect—and credited recovery from—failure. Self-diagnostics were also applied for performance monitoring during the extended surveillance interval. This report discusses the status of recent activities to evaluate the previously developed methodology using a pilot study. Although both a utility partner for a pilot study and a specific digital asset were identified in FY2020, delays in obtaining proprietary information resulted in a limited ability to fully evaluate the methodology, and further interactions were complicated by the COVID pandemic. Therefore, at that time, the use of public-domain information—along with current processes for surveillance interval extension through a surveillance frequency control program—identified the need to fully assess diagnostic coverage as part of the pilot study. Furthermore, self-diagnostics were also identified as a potential option to replace the drift analyses conducted as part of current STI extension procedures. In FY2022, the project was reconstituted with the industry partner, and information and data were made available by the industry partner to the research team for review. The shared information included failure event descriptions and data for a digital I&C system since its implementation, as well as recent STI extension interval reports developed by the utility partner on that digital I&C system. This report presents an evaluation of this information and data and describes an application of the proposed methodology cited above. The methodology seeks to take advantage of the self-diagnostics and OLM capabilities to reduce risk or reduce the level of qualitative monitoring assessment needed to perform a risk-informed STI extension using existing NRC approved guidance or both. Addressing these issues of STI extension by crediting self-diagnostics is likely to result in benefits for current and future nuclear power plant (NPP) operations, including lowering the barriers to adoption of digital I&C systems and increasing cost savings by deferring or eliminating unneeded preventive maintenance (tasks or checks or activities). Specifically, self-diagnostic and OLM capabilities of newer digital equipment being installed in non-safety and safety applications are designed to detect failures, provide early warning of potential failures, and notify plant operators to take appropriate action to reduce out of service (OOS) time thus protecting safety margins. Moreover, the equipment is expected to provide information that time-related operational degradation is identified early to ensure timely and planned corrective actions instead of a reactive and unplanned approach ahead of an extended-surveillance interval.

42 ENGINEERING↗

DMTN-260: Failure Modes and Error Handling for Prompt Processing

The Prompt Processing system will be responsible for processing roughly a thousand visits per night, and distributing the results in near real time, for at least ten years of Rubin Observatory operations. As such, it must be highly robust to algorithmic, network, and infrastructure failures, ranging from momentary glitches to extended downtimes. DMTN-219 introduced the initial design for the Prompt Processing framework; this document expands on the design to address expected failure modes and recovery strategies for each.

79 ASTRONOMY AND ASTROPHYSICS↗

Integrating static PRA information with risk informed safety margin characterization (RISMC) simulation methods

The overall objective of the project was to develop a computationally feasible and user-friendly mechanized process to integrate traditional probabilistic risk assessment (PRA) and dynamic PRA (DPRA) results. Starting with the systematic identification of items in an existing PRA that need dynamic augmentation, the project used a generic 4-loop pressurized reactor (PWR) and 3-loop PWR as example plants. Station blackout (SBO) and large break loss of coolant accident SBLOCA) were selected as the example initiating events. Using the traditional event-tree (ET)/fault-tree (FT) methodology augmented by dynamic evet tree approach, the potential consequences of the initiating events were simulated with RELAP-3D and MELCOR/RASCAL codes to cover Level 1 through Level 3 of PRA. RAVEN and ADAPT software were used to generate Level 1 simulations with RELAP-3D and Level 2/3 simulations with MELCOR (Level 2)/RASCAL (Level 3), respectively. Example branching conditions (BCs) for SBO included AC power recovery time, valve repair failure time, reactor coolant pump leak time/break size and emergency power supply duration to a total of 9. Example BCs for LOCA included off-site power recovery time, diesel generator power recovery time, auxiliary feed water system operation time, safety relief valve failure to open upon demand, reactor coolant pump seal break time and size to a total of 21. Each RELAP-3D simulation (9,587 scenarios) was labelled OK or Core Damage based on the maximum allowed peak clad temperature (2,100oF). Each MELCOR simulation (4610 scenarios) was labeled as Bin over 10rem or Bin 0-10rem based on the dose at the site boundary. The scenarios were clustered based on the criteria above using the mean shift methodology. Classical PRA (CPRA) and DPRA results were compared to identify the ET sequences that need DPRA augmentation. Several approaches were proposed for the incorporation of these sequences into CPRA using clustering with the mean shift methodology, restructuring the CPRA ETs by adding new BCs/sequences, and using the concept of a limit surface. Procedures for decision making regarding the possible consequences of an initiating event (e.g. core damage or not, site evacuation or not) were developed using a convolutional neural network (CNN), a recurrent neural network (RNN) and a transformer neural network (TNN). The project has led to two PhD degrees, three archival journal papers and five refereed conference proceedings.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Electro-Thermo-Mechanical (ETM) Study on Submarine Dynamic Power Cables (SDPC) for Offshore Wind Transmission

Understanding the failure mechanisms of submarine dynamic power cables (SDPC) is critical for innovative design to meet the 2035 cost reduction target of U.S. DOE Floating Offshore Wind Shot. This is important because the current design suffers a significant failure rate in the field. This project conducted a systematic electro-thermo-mechanical (ETM) experimental study on the power cores extracted from a 15kV power cable with three cores of copper conductor, ethylene propylene rubber (EPR) insulation, and continuously corrugated welded aluminum armor (CCWA). An ETM testing system was developed by integrating a high voltage (HV) amplifier, two ceramic heaters, and a rod-plate transverse compression setup into a mechanical testing machine. Increasing temperature from room temperature (RT, 22 degree C) to 90 degree C resulted in the 67% decrease in the failure mechanical load as defined by the dielectric breakdown. Under creep mechanical loading, the dielectric breakdown time was decreased by 70% for a given mechanical load when the specimen temperature increased from RT to 90oC. The failure strain in both monotonic and creep mechanical loading modes was related to the maximum mechanical load applied. Although the dielectric breakdown occurred in the ETM test, the post-test measurement revealed an impressive recovery of electric resistance. The failure analysis on cross section of tested specimens indicated a sizable gap across insulation layer near the area between copper conductor and loading rod, which apparently resulted from online dielectric failure and offline elastic recovery of components including conductor and insulation layer.

17 WIND ENERGY↗

Enhanced Component Performance Study: Emergency Diesel Generators 1998-2020

This report presents an enhanced performance evaluation of emergency power system (EPS) and high-pressure core spray (HPCS) emergency diesel generators (EDGs) at U.S. commercial nuclear power plants. This report evaluates component performance over time using (1) Institute of Nuclear Power Operations (INPO) Industry Reporting and Information System (IRIS) data from 1998 through 2020 and (2) maintenance unavailability performance data from Mitigating Systems Performance Index (MSPI) Basis Document data from 2002 through 2020. The objective is to show estimates of current failure probabilities and rates related to EDGs, trend these data on an annual basis, determine if the current data are consistent with the probability distributions currently recommended for use in NRC probabilistic risk assessments, show how the reliability data differ for different EDG manufacturers and for EDGs with different ratings; and summarize the subcomponents, causes, detection methods, and recovery associated with each EDG failure mode. The EDG failure modes considered are fail to start (FTS), fail to load and run (FTLR), and fail to run after one hour of operation (FTR>1H). Engineering analyses were performed with respect to time period and failure mode without regard to the actual number of EDGs at each plant. The factors analyzed are: subcomponent, failure cause, detection method, recovery, manufacturer, and EDG rating. The following increasing trends were identified for EDGs for the most recent 10-year period: • HPCS EDG FTS failure probability • EPS and HPCS EDG frequency of start demands (demands per reactor year) • EPS and HPCS EDG frequency of FTLR demands. The following decreasing trends were identified for EDGs for the most recent 10-year period: • EPS EDG FTS failure probability • EPS EDG FTLR failure probability • EPS EDG unavailability • EPS and HPCS EDG frequency of FTLR events (failures per reactor year).

99 GENERAL AND MISCELLANEOUS↗

Enhanced Component Performance Study: Emergency Diesel Generators 1998–2022

This report presents an enhanced performance evaluation of the emergency power system (EPS) and high-pressure core spray (HPCS) emergency diesel generators (EDGs) at U.S. commercial nuclear power plants. This report evaluates component performance over time using (1) Institute of Nuclear Power Operations (INPO) Industry Reporting and Information System (IRIS) data from 1998 through 2022 and (2) maintenance unavailability performance data from Mitigating Systems Performance Index (MSPI) Basis Document data from 2002 through 2022. The objective is to show estimates of current failure probabilities and rates related to EDGs, trend these data on an annual basis, determine if the current data are consistent with the probability distributions currently recommended for use in Nuclear Regulatory Commission (NRC) probabilistic risk assessments, show how the reliability data differ for different EDG manufacturers and for EDGs with different ratings; and summarize the subcomponents, causes, detection methods, and recovery associated with each EDG failure mode. The EDG failure modes considered are fail to start (FTS), fail to load and run (FTLR), and fail to run after one hour of operation (FTR>1H). Engineering analyses were performed with respect to time-period and failure mode without regard to the actual number of EDGs at each plant. The factors analyzed include subcomponent, failure cause, detection method, recovery, manufacturer, and EDG rating. The following increasing trends were identified for EDGs for the most recent 10-year period: (1) HPCS EDG FTS failure probability, (2) EPS and HPCS EDG frequency of start demands (demands per reactor year), and (3) EPS and HPCS EDG frequency of FTLR demands. The following decreasing trends were identified for EDGs for the most recent 10-year period: (1) EPS EDG FTS failure probability, (2) EPS EDG FTLR failure probability, (3) EPS EDG unavailability, and (4) EPS and HPCS EDG frequency of FTLR events (failures per reactor year).

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Enhanced Component Performance Study: Emergency Diesel Generators 1998-2024

This report presents an enhanced performance evaluation of the emergency power system (EPS) and high-pressure core spray (HPCS) emergency diesel generators (EDGs) at U.S. commercial nuclear power plants. This report evaluates component performance over time using (1) Institute of Nuclear Power Operations (INPO) Industry Reporting and Information System (IRIS) data from 1998 through 2024 and (2) maintenance unavailability performance data from Mitigating Systems Performance Index (MSPI) Basis Document data from 2002 through 2024. The objective is to show estimates of current failure probabilities and rates related to EDGs, trend these data on an annual basis, determine if the current data are consistent with the probability distributions currently recommended for use in Nuclear Regulatory Commission (NRC) probabilistic risk assessments, show how the reliability data differ for different EDG manufacturers and for EDGs with different ratings; and summarize the subcomponents, causes, detection methods, and recovery associated with each EDG failure mode. The EDG failure modes considered are fail to start (FTS), fail to load and run (FTLR), and fail to run after one hour of operation (FTR>1H). Engineering analyses were performed with respect to time-period and failure mode without regard to the actual number of EDGs at each plant. The factors analyzed include subcomponent, failure cause, detection method, recovery, manufacturer, and EDG rating. The following increasing trends were identified for EDGs for the most recent 10-year period: • EPS and HPCS EDG frequency of start demands (demands per reactor year) • EPS and HPCS EDG frequency of FTLR demands • EPS and HPCS EDG frequency of run>1H hours. The following decreasing trends were identified for EDGs for the most recent 10-year period: • EPS EDG FTR>1H failure rate • EPS EDG unreliability • EPS and HPCS EDG frequency of FTR>1H events (failures per reactor year).

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Probing the evolution of fault properties during the seismic cycle with deep learning

We use seismic waves that pass through the hypocentral region of the 2016 M6.5 Norcia earthquake together with Deep Learning (DL) to distinguish between foreshocks, aftershocks and time-to-failure (TTF). Binary and N-class models defined by TTF correctly identify seismograms in test with > 90% accuracy. We use raw seismic records as input to a 7 layer CNN model to perform the classification. Here we show that DL models successfully distinguish seismic waves pre/post mainshock in accord with lab and theoretical expectations of progressive changes in crack density prior to abrupt change at failure and gradual postseismic recovery. Performance is lower for band-pass filtered seismograms (below 10 Hz) suggesting that DL models learn from the evolution of subtle changes in elastic wave attenuation. Tests to verify that our results indeed provide a proxy for fault properties included DL models trained with the wrong mainshock time and those using seismic waves far from the Norcia mainshock; both show degraded performance. Our results demonstrate that DL models have the potential to track the evolution of fault zone properties during the seismic cycle. If this result is generalizable it could improve earthquake early warning and seismic hazard analysis.

58 GEOSCIENCES↗

EVALUATION OF HRA METHODOLOGIES FOR APPLICATION IN SDP WORK

This study critically evaluates human reliability analysis (HRA) methodologies applicable to regulatory probabilistic safety assessment (PSA) model, with a particular focus on their role in supporting the significance determination process (SDP) in nuclear safety assessment. Firstly, three widely utilized HRA methods – IDHEAS-ECA, SPAR-H, and ASEP/THERP – were qualitatively and quantitatively assessed. Qualitative assessments were conducted using attributes from the NEA/CSNI/R(2015)1 report, while quantitative evaluations employed regression and correlation analyses to compare predicted human error probabilities (HEPs) against empirical data. Results reveal distinct strengths, for example, IDHEAS-ECA’s robust predictive accuracy and K-HRA’s alignment with operational practices. In addition, dependency analysis and recovery analysis were critically evaluated. For dependency analysis, the methods’ handling of inter-task dependencies and their impact on HEPs were examined, while recovery analysis highlighted strategies for mitigating failure events. Furthermore, strategies were proposed to evaluate performance-shaping factors under conditions of reduced human performance, such as stress, fatigue, or cognitive overload, addressing specific challenges faced in SDP evaluations. Human errors from KINS’s operational performance information system event reports were evaluated as a case study. This study identifies gaps and provides actionable insights to ensure their validity and applicability in SDP HRA applications. This paper is a part of research conducted by KINS, and it should be noted that this result does not represent the regulatory position of KINS.

99 - GENERAL AND MISCELLANEOUS↗

Seamless Transition of Critical Infrastructures using Droop Controlled Grid-forming Inverters

Seamless recovery of power to critical infrastructures, after grid failure, is a crucial need arising in scenarios that are increasingly becoming more frequent. Here, this article proposes a seamless transition strategy using a single and unified mode-dependent droop controlled grid-forming inverters. The control strategy achieves the following objectives: 1) regulates the output active and reactive power by the droop-controlled inverters to a desired value while operating in on-grid mode; 2) seamless transition and recovery of power injections into the load after grid failure by inverters that operates in grid-forming mode all the time; 3) requires only a single bit of information on the grid/network status for the mode transition. A framework for assessing the stability of the system and to guide the choice of parameters for controllers is developed using control-oriented modeling. A controller hardware-in-the-loop-based real-time simulation study on a test system based on the realistic electrical network of a commercial-scale medical center is conducted for initial prototyping of the control strategy. A hardware experiment is conducted with two 3 - $\phi$, 480 -V, 125 -kVA grid-forming inverters, a 3 - $\phi$, 480 -V, 270 -kVA grid simulator, a physical grid switch, and a physical load bank. The experimental data establishes the effectiveness of the always grid-forming operation and control of inverters in meeting power delivery objectives when on-grid and off-grid under various kinds of loads and scenarios while minimizing transients during transitions. Furthermore, performance comparison with existing strategies showcases the advantage of the proposed strategy.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Automated, reliable, and efficient continental-scale replication of 7.3 petabytes of computational simulation data: A case study

We report on our experiences replicating 7.3 petabytes (PB) of Earth System Grid Federation (ESGF) computational simulation data from Lawrence Livermore National Laboratory (LLNL) in California to Argonne National Laboratory (ANL) in Illinois and Oak Ridge National Laboratory (ORNL) in Tennessee—a task motivated by a need for increased reliability, capacity, and performance. This task presented significant challenges: the need to move 29 million files twice under time pressure from aging storage hardware; a source file system bottleneck limiting throughput to 1.5 GB/s; frequent site maintenance windows; and the need for complete reliability at scale. We addressed these challenges using a simple replication tool that invoked Globus to transfer large bundles of files while tracking progress in a database, dynamically rerouting transfers to work around maintenance periods and file system limitations. Under the covers, Globus organized transfers to make efficient use of the high-speed Energy Sciences network (ESnet) and the data transfer nodes deployed at participating sites, and also addressed security, integrity checking, and recovery from a variety of transient failures. This success demonstrates the considerable benefits that can accrue from the adoption of performant data replication infrastructure. The replication tool is available at https://github.com/esgf2-us/data-replication-tools.

Globus↗

Performance Losses and Current-Driven Recovery from Cation Contaminants in PEM Water Electrolysis

Water contaminants are a common cause of failure for polymer electrolyte membrane (PEM) electrolyzers in the field as well as a confounding factor in research on cell performance and durability. In this study, we investigated the performance impacts of feed water containing representative tap water cations at concentrations ranging from 0.5–500 μ M, with conductivities spanning from ASTM Type II to tap-water levels. We present multiple diagnostic signatures to help identify the presence of contaminants in PEM electrolysis cells. Through analysis of polarization curves and impedance spectroscopy to understand the origins of performance losses, we found that a switch from the acidic to alkaline hydrogen evolution mechanism is a key factor in contaminated cell behavior. Finally, we demonstrated that this mechanism switching can be harnessed to remove cation contaminants and recover cell performance without the use of an acid wash. We demonstrated near-complete recovery of cells contaminated with sodium and calcium, and partial recovery of a cell contaminated with iron, which was further investigated by post-mortem microscopy. The improved understanding of contaminant impacts from this work can inform development of strategies to mitigate or recover performance losses as well as improve the consistency and rigor of electrolysis research.

30 DIRECT ENERGY CONVERSION↗