Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “SYSTEM FAILURE”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Reliability Analysis of Power Grids Considering Component Failures of Variable Energy Resources

This paper proposes an improved model for the reliability assessment of power systems considering component failures of variable energy resources (VER). The inherent intermittency of VER such as solar photovoltaic (PV) and wind farms, along with their susceptibility to component failures, present significant challenges to reliable system operation. These issues, combined with power grid operation and network constraints, complicate the reliable operation of VER-integrated power systems. Here, to address these concerns, this paper introduces a reliability assessment framework that considers VER input variability, its impact on component availability, and their resulting impact on overall system reliability. Stochastic models based on discrete Markov processes are developed to incorporate variable irradiance, wind speeds, and their effects on PV and wind component failure rates. A next-event and state transition-based approach is then developed to integrate the stochastic models into a mixed-timing sequential Monte Carlo simulation framework for composite reliability assessment. Case studies on the RTS-GMLC system demonstrate the effectiveness of the proposed model in evaluating the reliability of VER-integrated systems.

Pandit, Dilip [Sandia National Laboratories (SNL-N↗

Operational Experience and Development of a Reliability Model for the ATR Demineralizer System (Slides)

The Advanced Test Reactor (ATR) is a light water reactor with aluminum clad driver fuel. Strict limits are implemented on pH, conductivity, and filterable solids to assure the performance of the driver fuel limit corrosion of the primary coolant system (PCS) pressure boundary components. A bypass demineralizer system is used to maintain PCS coolant within allowable bands for each of the aforementioned variables. Unlike many other reactors, the ATR does not have a filtration system. As such, failure of ion exchange resin in the bypass demineralizer can defeat the purpose of the system and result in high filterable solids in the PCS. Power operations with high filterable solids is prohibited by a technical safety requirement. This paper discusses recent failures of the ATR bypass demineralizer system. The testing and analysis done to deduce the cause of the failures is covered, as well as short term mitigative steps that were taken. To prevent the likelihood of future issues, historical data from the ATR and industry is used to develop a reliability model of the ATR bypass demineralizer system. Evidence of radiolytic decomposition of ion exchange resin was found in the cause investigation. As such, special attention is given to ion exchange resin failure due to high radiation environments. The reliability model is used to develop the ATR testing and maintenance program. Specifically, limits on resin service life are discussed.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

From Ensemble Climate to Ensemble Impacts

Many climate-risk tools rely on ensemble mean projections or endpoint climate snapshots to characterize future hazards. Although convenient for communication, these representations remove the statistical, temporal, and physical information that real infrastructure systems respond to. Infrastructure degradation and failure arise from extremes, sequences, cumulative stress, compound hazards, and nonlinear fragility relationships, none of which survive ensemble averaging or temporal compression. Power-system failure statistics and cascading failure models further show that infrastructure risk is dominated by tail events and path-dependent dynamics rather than by mean conditions. This paper demonstrates why ensemble mean or endpoint-only climate representations are mathematically and physically inconsistent with engineering-grade risk analysis. We outline a model-resolved, time-series-based workflow that preserves extremes, variability, and sequencing by propagating each climate-model realization independently through hazard formation, exposure, fragility, and cascading failure mechanisms. Taking the ensemble of impacts—rather than the ensemble of climate—provides a defensible, physically coherent foundation for infrastructure resilience planning, regulatory compliance, and long-term investment decisions.

54 - ENVIRONMENTAL SCIENCES/GLOBAL CLIMATE CHANGE ↗

Shock-induced twinning/detwinning and spall failure in Cu–Ta nanolaminates at atomic scales

Here, this study provides new insights into the role of interfaces on the deformation and failure mechanisms in shock-loaded Cu–Ta–Cu trilayer system. The thickness of the Ta layer, piston velocities, and shock pulse durations were varied to explore the impact of impedance mismatch and loading conditions on spallation behavior and twin formation. It was found that the interfaces play a crucial role in the dynamic response of these multilayered systems since secondary reflection waves generated at the interfaces significantly affected the peak stress and pressure profiles, influencing void nucleation and failure modes. In the trilayer systems, failure predominantly occurred at interfaces and within the Ta layer, with void nucleation sites and twinning behavior being markedly different compared to single-crystal Cu and Ta. Increasing the Ta layer thickness modified the wave interactions, leading to different failure locations. Higher piston velocities were associated with increased spall strength by enhancing wave interactions and void formation, particularly at the interfaces and within the Ta layer, under specific configurations. Additionally, shorter shock pulse durations facilitated earlier initiation of the release fan, reducing twin formation and altering the failure dynamics by accelerating twin annihilation and pressure release.

36 MATERIALS SCIENCE↗

Root Cause Correlation Analysis of Software Failures via Orthogonal Defect Classification and Natural Language Processing

Systems theoretic process analysis (STPA) is becoming an increasingly popular technique to assess how complex digital software systems can fail. Rather than defining failures by their observable failure events, which may be sparse especially for safety rated nuclear digital instrumentation and control systems (DI&C), failures are defined as postulated unsafe actions under specific contextual conditions. This permits a top-down analysis of system hazards and identifies whether imposed constraints and requirements can sufficiently address undesirable hazards. However, STPA is a qualitative approach at identifying inadequacies in the development process and cannot currently be used to quantify unsafe action likelihoods for probabilistic risk assessment. Therefore, in this work, we examine the root causes of software failure and explore whether a consistent correlation can be linked to specific unsafe action classes. We implement Lbl2Vec, an unsupervised document classification and retrieval algorithm, on a database of 4,096 software defect reports acquired from various open-source software systems. By analyzing sentence structure, embedded labels, and word vectors, we show that certain defect types positively correlate to specific unsafe action classes over others. The correlations developed can be used to estimate the failure probability of safety intended DI&C systems which provides a licensing basis for nuclear plant modernization efforts.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Root Cause Correlation Analysis of Software Failures via Orthogonal Defect Classification and Natural Language Processing

Systems theoretic process analysis (STPA) is becoming an increasingly popular technique to assess how complex digital software systems can fail. Rather than defining failures by their observable failure events, which may be sparse especially for safety rated nuclear digital instrumentation and control systems (DI&C), failures are defined as postulated unsafe actions under specific contextual conditions. This permits a top-down analysis of system hazards and identifies whether imposed constraints and requirements can sufficiently address undesirable hazards. However, STPA is a qualitative approach at identifying inadequacies in the development process and cannot currently be used to quantify unsafe action likelihoods for probabilistic risk assessment. Therefore, in this work, we examine the root causes of software failure and explore whether a consistent correlation can be linked to specific unsafe action classes. We implement Lbl2Vec, an unsupervised document classification and retrieval algorithm, on a database of 4,096 software defect reports acquired from various open-source software systems. By analyzing sentence structure, embedded labels, and word vectors, we show that certain defect types positively correlate to specific unsafe action classes over others. The correlations developed can be used to estimate the failure probability of safety intended DI&C systems which provides a licensing basis for nuclear plant modernization efforts.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Ultrasonic Characterization of Lithium Ion Thermal Runaway Conditions for Real-time Ultrasonic Enabled BMS Integration

The demand for energy storage is growing, and lithium-ion batteries are a promising technology to meet this need due to their high power/energy density, high round-trip efficiency, rapid response time, and portability. However, recent catastrophic events caused by thermal runaway have slowed their adoption, highlighting the need for an early warning system for battery failure. In this work, ultrasound is used to detect physical changes in 950 mAh batteries by identifying material property changes independent of voltage and current. [Ultrasound signal features (e.g., time of flight, maximum frequency component) were extracted as the batteries were cycled and subjected to both constant current and constant voltage overcharge and were used to develop two metrics identifying failure: a warning to detect the start of overcharge and an emergency stop (E-stop) to immediately take the battery out of service. The identification method involved locating magnitude differences of several ultrasound features compared to baseline operation considering different currents and temperatures, and the warning/stopping metrics were consistent across all experiments. For an average overcharge time of 140 minutes, the average warning was issued 124 minutes before the failure and the average E-stop was triggered 94 minutes before failure. As a test of using ultrasound for early warning detection, a battery was forced into overcharge and returned to normal cycling conditions based on the previously determined warning metric. Both the voltage profile and the ultrasound measurements returned to their baseline behavior, indicating that ultrasonic detection can not only identify battery failure before a catastrophic event, but can also provide early enough warnings such that overcharges can be detected and corrected quickly enough so the battery does not need to be decommissioned.]

25 ENERGY STORAGE↗

Energy storage planning for enhanced resilience of power systems against wildfires and heatwaves

Extreme weather events pose significant risks to power grid stability due to their severe consequences and potential for widespread failures. Energy storage systems hold great potential for enhancing grid resilience against such events by providing reliable power during peak demand periods. However, accurately quantifying the size, location, and investment costs of new energy storage assets is a complex task, as energy storage planning decisions depend on the investment choices of other generation technologies and the integration of new transmission projects. Here, this paper presents a novel capacity expansion planning framework that simultaneously optimizes investments in energy storage, generation, and transmission, determining their optimal size, location, and type, while incorporating extreme weather events into long-term planning. More specifically, our stress-event-informed planning framework integrates the impact of heatwaves and wildfires into the planning process, identifying least-cost investment solutions that comply with policy goals and enhance grid resilience. The proposed framework employs machine-learning-based modeling to project heatwave-induced loads and performance-based risk assessment to evaluate wildfire-driven transmission line derates. Using industry-standard datasets to accurately represent the transmission topology of the Western Interconnection (WI) system, the proposed framework is applied to the WI 40-zone system, with investment decisions reported for the years 2030, 2035, and 2040. Simulation results reveal that with just a 10% increase in investment costs, resilience against extreme events can be significantly improved, with investment decisions heavily favoring energy storage, particularly 4-hour energy storage systems.

25 ENERGY STORAGE↗

Coincident learning for unsupervised anomaly detection of scientific instruments

Abstract Anomaly detection is an important task for complex scientific experiments and other complex systems (e.g. industrial facilities, manufacturing), where failures in a sub-system can lead to lost data, poor performance, or even damage to components. While scientific facilities generate a wealth of data, labeled anomalies may be rare (or even nonexistent), and expensive to acquire. Unsupervised approaches are therefore common and typically search for anomalies either by distance or density of examples in the input feature space (or some associated low-dimensional representation). This paper presents a novel approach called coincident learning for anomaly detection (CoAD), which is specifically designed for multi-modal tasks and identifies anomalies based on coincident behavior across two different slices of the feature space. We define an unsupervised metric, F ^ β , out of analogy to the supervised classification F β statistic. CoAD uses F ^ β to train an anomaly detection algorithm on unlabeled data , based on the expectation that anomalous behavior in one feature slice is coincident with anomalous behavior in the other. The method is illustrated using a synthetic outlier data set and a MNIST-based image data set, and is compared to prior state-of-the-art on two real-world tasks: a metal milling data set and our motivating task of identifying RF station anomalies in a particle accelerator.

43 PARTICLE ACCELERATORS↗

System Engineers and Decisions: It?s All about Knowledge

In order to guarantee that a system meets adequate levels of reliability and availability, system performances are continuously monitored and analyzed thanks to the technological advancements driving the Industry 4.0 revolution. An Industry 4.0 approach is typically based on advanced statistical, big data mining, machine learning, and internet-of-things methods designed to detect anomalies in the behavior of system, detect the most likely failure modes, and provide indications to system engineers on when maintenance activities should be performed before system performance are deemed unacceptable (which can be generated by diagnostic and prognostic methods). However, these analyses, which are designed to automatize and increase the efficacy of the system maintenance program, require large amount of data which can come in various forms: numeric, textual, images, sounds etc. Such data constitutes the historic knowledge benchmark to track system performances and support system engineer decisions. Here we claim that data is not sufficient to support this kind of analyses when applied to systems characterized by complex architectures and behaviors. Robust system engineer decisions require the ability to understand the system operational context that lies behind the observed data elements. In this respect, system models are in fact necessary to “put data in context” and capture relationships between data elements. Industry 4.0 methods require in fact contextual knowledge as a basis upon which hypotheses can be generated and assumptions tested. In our view, for complex systems, model-based system engineering (MBSE) models can afford this contextual knowledge, as they are typically used to describe systems architecture and dynamic behaviors. System knowledge is here intended as the blending of collected data and system architecture which takes the form of a “knowledge graph”. A knowledge graph is a database which consists of a large set of nodes (in our case an entity can be either a data or an MBSE element) which are linked to each other. The types of nodes and links follow a pre-defined topology, sometimes also refers as an ontology, that is designed to fit the actual decisions that needs to be performed. We show here how a knowledge graph can be defined to support system engineer maintenance decisions and how the same graph can be built based on system MBSE models and pre-processed data from numeric (through anomaly detections and diagnostic methods) and textual elements (through technical language processing TLP).

97 - MATHEMATICS AND COMPUTING↗

Formulation and Performance Evaluation of Epoxy Sealant Systems for Double-Shell Tank Bottom Refurbishment

The performance of epoxy sealants used in the refurbishment of double-shell tank (DST) systems requires balancing processability, thermomechanical stability, and adhesion to cementitious substrates. This study incorporates Heloxy 8 as a reactive diluent into Westlake 862 epoxy to tailor workability and cured-state properties. Rheological time-sweep analysis demonstrates that increasing the diluent content significantly reduces complex viscosity and extends workability, thereby improving pumpability and flow for large-area applications. However, the targeted 2-hour processing window is not fully achieved. Differential scanning calorimetry (DSC) confirms that all formulations cure at room temperature to glass transition temperatures ( T g ) at least 20 °C above the maximum DST operating temperature (27 °C), thereby ensuring service in the glassy regime. Dynamic mechanical analysis (DMA) reveals formulation-dependent reductions in tan delta and increases in storage modulus, indicating increasingly elastic and mechanically stable networks with diluent incorporation. Pull-off adhesion testing shows that modified formulations (70–90% Westlake epoxy) exhibit significantly higher adhesion strengths than the unmodified system. Grout cohesive failure indicates that interfacial bonding exceeds substrate strength. Collectively, these results demonstrate that controlled reactive diluent incorporation enables optimization of processing behavior, interfacial adhesion, and thermomechanical performance, supporting the suitability of the modified epoxy systems as durable sealant layers for cementitious barrier applications in hazardous waste containment infrastructure.

Differential scanning calorimetry↗

Application of Banking Scoring and Rating for Coherent Risk Measures in Electricity Systems ABSCORES

This project developed a framework for asset and system risk management that can be incorporated into current electricity system operations to improve economic efficiency and establish an Electric Assets Risk Bureau. We leveraged scoring and ratings from banking and financial institutions alongside current optimization methods in dispatching power systems to help system operators and electricity markets schedule resources. This approach is based on the observation that there are major discrepancies between the power scheduled by a system operator and the actual power generated/consumed. These discrepancies—exacerbated by unplanned contingencies (e.g., natural disasters)—are caused by multiple factors, including the different financial, environmental and risk preferences of power producers, consumers, and aggregators. We developed a framework that counteracts two failures in electricity system operations: imperfect information and missing markets for products. The technical approach included five tasks. Tasks 1 and 2 supported the development of risk scores at the asset level with historical data collected for this project. Tasks 3, 4, and 5 incorporated scoring into decision-making at the system level. The proposed effort achieved PERFORM's Program Objectives because the proposed outputs and algorithms do not exist in the electricity industry and are an innovative approach to managing risk. Since the acknowledged need to better assess and act upon risk profiles for grid assets has not been met by the industry, this project will also impact ARPA-E's Mission Areas, including improving energy efficiency and giving the U.S. a technological lead in advanced energy technologies.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Enhanced Component Performance Study: Motor Driven Pumps 1998-2024

This report presents an enhanced performance evaluation of motor-driven pumps (MDPs) at U.S. commercial nuclear power plants. The data used in this study are based on the operating experience failure reports from calendar year 1998 through 2024 as reported in the Institute of Nuclear Power Operations (INPO) Industry Reporting and Information System (IRIS). The MDP failure modes considered for standby systems are fail to start (FTS), fail to run (FTR) for one hour of operation (FTR=1H), FTR after one hour of operation (FTR>1H), and for normally running systems FTS and FTR. An eight-hour unreliability estimate is also calculated and trended. The component reliability estimates and the reliability data are trended for the most recent 10-year period while yearly estimates for reliability are provided for the entire study period.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Enhanced Component Performance Study: Turbine-Driven Pumps 1998-2024

This report presents an enhanced performance evaluation of turbine driven pumps (TDPs) at U.S. commercial nuclear power plants. The data used in this study are based on the operating experience failure reports from calendar year 1998 through 2024 as reported in the Institute of Nuclear Power Operations (INPO) Industry Reporting and Information System (IRIS). The TDP failure modes considered for standby systems are fail to start (FTS), fail to run (FTR) for one hour of operation (FTR=1H), FTR after one hour of operation (FTR>1H), and for normally running systems FTS and FTR. An eight hour unreliability estimate is also calculated and trended. The component reliability estimates and the reliability data are trended for the most recent 10 year period while yearly estimates for reliability are provided for the entire study period. No increasing trends were identified for TDPs for the most recent 10 year period.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗