Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “SYSTEM FAILURE”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

First-of-a-Kind Risk-Informed Digital Twin for Operational Decision Making

A digital twin (DT) is a digital model or a collection of models of a physical entity. DTs in the nuclear arena can be used from plant design through decommissioning. Decisions are typically a priori or made offline. Risk-informed decision making is identifying what can go wrong, its frequency, and the consequences of its failure. Ideally risk-informed decision making reflects the current state of the plant and provides a decision in real time. Traditionally, probabilistic risk assessments (PRAs) evaluate the failures of safety systems, the risk of core damage, and the offsite dose as the consequence. However, this DT evaluates the decisions on the control side rather than the protection side. It uses the same risk methods to probabilistically inform the decision-making process but in a different way. Rather than evaluating the risk of core damage, this DT evaluates the likelihood of avoiding a trip set point while maintaining plant safety. Performance-based assessments are identified via its probabilistic evaluation of operational alternatives based on system status. Because the purpose of the control system is to maintain system variables within prescribed operating ranges, upsets or challenges that can exceed a trip set point resulting in a plant transient and a challenge to plant mitigating systems based on actual plant conditions, are evaluated to safely maintain the plant within the operating ranges. The probabilistic portion of the model is autonomously and automatically adjusted, and the metric of interest (i.e. likelihood of avoiding a trip set point) is recalculated. The digital representation of the physical system (i.e. the DT) performs a deterministic performance–based assessment of the probabilistically identified alternatives identified to validate the probabilistic assessment. A decision-making algorithm selects the appropriate option based on the probabilistic and deterministic assessments and transmits a control signal to a component(s) to initiate a corrective action or informs an operator of its decision.

digital twin↗

Unsupervised Process Anomaly Detection and Identification Using the Leave-One-Variable-Out Approach

Automated anomaly detection and identification can signal equipment issues and pinpoint causes in large-scale industrial systems. For systems with limited failure history, unsupervised machine learning methods can be utilized as they do not require past failures. This study introduces the leave-one-variable-out (LOVO) model, which masks one variable at a time to predict the others, learning underlying process correlations. Detection performance was assessed with synthetic and experimental data, while identification performance used only synthetic data due to its ability to generate labeled anomaly types. For detection using synthetic data, the LOVO model generally outperformed comparative models; while using experimental data, the comparative methods outperformed the LOVO model. However, the comparative methods required selecting a latent size, and these conclusions pertain to using the optimal size. In practice, it would not be feasible to always select the optimal value, and incorrect selections impacted performance. In contrast, the LOVO model does not require a latent space. For identification using synthetic data, the LOVO model was slightly outperformed in interpretability and repeatability but still demonstrated impressive results. These outcomes suggest that the LOVO model is an effective model and may be more easily implemented without the challenging tuning process of selecting a latent size.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Impacts of PV Module Connector Failures on Cost and Performance of Utility Scale Photovoltaic Systems

The reliability, cost and performance of electrical connectors are a concern in all types of electrical systems, and demands on connectors used on photovoltaic (PV) systems include that connectors maintain electrical conductivity and physical strength, endure ultraviolet sunlight and high ambient temperature, and resist moisture and chemical intrusion over a very long (>25 year) performance period. Connector failures increase operation and maintenance (O&M) costs and reduce plant production, but connector failure can also cause safety and liability problems, which are of greater concern. This work results from a three-year collaboration between Sandia National Laboratories (SNL), the Electric Power Research Institute (EPRI), and the National Renewable Energy Laboratory (NREL) and funded by the U.S. Department of Energy (DOE) Solar Energy Technology Office (SETO) under Agreements #39035 and #38531 "Connector Reliability Across the US Solar Sector." a multi-pronged investigation of PV connector health across the US (see https://energy.sandia.gov/pvconnectors/). This report presents derivation of a Techno-Economic Analysis (TEA) that models failure modes and frequencies (how often failure occurs), estimates O&M costs and lost production associated with connector failures, and then calculates the effect that PV module connectors can have on Levelized Cost of Energy (LCOE). The model is informed with initial data from quantitative assessment of failure rates, root causes and mechanisms, in-situ diagnostics and data collection, lab-based forensics, and interviews with PV connector manufacturers and plant operators. SNL conducted site inspections at multiple utility-scale sites in different climates and subjected field samples of new, used, and degraded connectors to visual and electrical characterization. EPRI conducted metallurgical analysis of the pin and sleeve conductors to study failure-induced morphological and compositional changes. There is in general a shortage of statistically valid data, but data from PVROM database maintained by SNL was sufficient to ascertain failure rates and lost production as well as provide qualitative insight in its curated maintenance records. This report details the structure of the mathematical model but the sources of data to inform the model will continue to evolve. Analysis of a 100 MW PV plant is provided as an example of the use of the model, with results indicating that connectors are responsible for Annualized O&M Costs of $\$$71,933/year; Annualized Unit O&M Costs of $\$$0.72/kW/year; that a Reserve Account of $\$$187,220 should be available to fund repairs related to connectors; that connectors add $\$$1,494,004 to the Net Present Value of the O&M Costs (project life); and that O&M related to connectors adds about $\$$0.00088/kWh to the Levelized Cost of Energy. The impact of this model is to provide a tool to make the US solar sector more robust by quantifying and monetizing the reliability risks to utility-scale PV systems posed by poorly installed, mismatched and/or poorly designed and manufactured connectors. The TEA provides a model incorporating failure statistics, O&M cost data, and lost production into a single figure of merit, informing decisions and enabling practitioners to optimize cost and performance trade-offs. Stakeholders include connector manufacturers, system designers and equipment specifiers, standards bodies, installers and O&M providers, investors and insurance underwriters. This report supports continued growth of PV predicated on assurances that properly installed and maintained PV system connectors are safe and reliable. The project team is proposing future work including accelerated testing of connectors and expanding the approach taken here to other PV system components, such as TEA for rapid shut-down devices.

14 SOLAR ENERGY↗

Lifetime extension of legacy CEBAF LLRF hardware

A significant portion of the Low-Level Radio Frequency (LLRF) hardware in Jefferson Lab’s CEBAF is from the original construction of the facility using 1980’s CAMAC technology. Of the fifty-three zones in CEBAF, thirty-six of them are legacy hardware. The age of the legacy system has led to difficulties in maintaining the hardware due to parts going obsolete without suitable drop in replacements. Continued operation of the legacy system is required as the installation of LLRF 3.0 systems is costly and cannot be completed in a short period of time with the available resources. The most pressing failure in the legacy system was a failing buffer card, which is responsible for communication between the EPICs network and individual RF control modules. A new buffer card was designed as a transparent, drop in, replacement so that upgrades are simply a matter of swapping the existing legacy hardware. This buffer card upgrades a single point failure component and promises to extend the operable lifetime of CEBAF’s legacy systems.

Accelerator Physics↗

Safety Hazards of Batteries and Hydrogen Storage Systems in Proximity

This report addresses the safety concerns and mitigations for battery failures and their impact on hydrogen storage systems. Through an analysis of failure modes, this report highlights the risks posed by thermal runaway and chemical emissions caused by batteries. Although rare, battery thermal failure events may prompt the opening of the relief valve on the hydrogen tank. Strategies such as battery management systems, thermal management systems, and multiple thermally activated pressure relief devices can mitigate these risks. Potential simulations and experiments to better quantify the unique risks posed by lithium-ion batteries near tanks are suggested. Improving safety standards will enable integration of batteries and hydrogen storage systems in various energy storage technologies.

08 HYDROGEN↗

BuildingQA: A Benchmark for Natural Language Question Answering over Building Knowledge Graphs

Graph-based representations of building metadata using ontologies like Brick are vital for smart building applications, but querying them remains a challenge for practitioners. Knowledge Graph Question Answering (KGQA) systems, meant to retrieve answers from natural language questions, traditionally require large-scale training data, making them ill-suited for the specialized and data-scarce building domain. The advent of Large Language Models (LLMs) offers a paradigm shift, enabling zero-shot natural language querying without building/domain-specific training. Yet, there is no standardized benchmark for building-specific KGQA which can guide and validate research in this area. To address this gap, our work makes three primary contributions. First, we introduce the BuildingQA Benchmark Dataset, constructed through a multi-stage process of collecting practitioner data, augmenting it with LLMs for linguistic diversity, and curating a final set of 188 questions across 4 buildings. Second, we characterize the benchmark's complexity and ambiguity, introducing a novel method to quantify its "lexical gap" and providing a four-stage diagnostic framework for analyzing how systems fail. Third, we benchmark zero-shot LLM-powered KGQA systems to establish baseline performance and analyze their failure modes. Our evaluation reveals that top-performing systems achieve a maximum F1 score of only 0.38. This result does not indicate a failure of these powerful systems, but rather underscores the unique challenges posed by our benchmark. It demonstrates a critical performance gap, showing that current methods successful on general KGs struggle with the specific lexical and structural nuances of the building domain. BuildingQA1 thus provides the benchmark dataset and foundational analysis needed to drive the development of novel, domain-aware methods required to unlock the use of semantic data in buildings.

Mulayim, Ozan Baris↗

Analysis and Mitigation of Cascading Failures Using a Stochastic Interaction Graph with Eigen-analysis

In studies on complex network systems using graph theory, eigen-analysis is typically performed on an undirected graph model of the network. However, when analyzing cascading failures in a power system, the interactions among failures suggest the need for a directed graph beyond the topology of the power system to model directions of failure propagation. To accurately quantify failure interactions for effective mitigation strategies, this paper proposes a stochastic interaction graph model and associated eigen-analysis. Different types of modes on failure propagations are defined and characterized by the eigenvalues of a stochastic interaction matrix, whose absolute values are unity, zero, or in between. Finding and interpreting these modes helps identify the probable patterns of failure propagation, either local or widespread, and the participating components based on eigenvectors. Then, by lowering the failure probabilities of critical components highly participating in a mode of widespread failures, cascading can be mitigated. Here, the validity of the proposed stochastic interaction graph model, eigen-analysis and the resulting mitigation strategies is demonstrated using simulated cascading failure data on an NPCC 140-bus system.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Enhanced Component Performance Study: Motor-Operated Valves 1998-2024

This report presents an enhanced performance evaluation of motor-operated valves (MOVs) at U.S. commercial nuclear power plants. The data used in this study are based on the operating experience failure reports from calendar year 1998 through 2024 as reported in the Institute of Nuclear Power Operations (INPO) Industry Reporting and Information System (IRIS). The MOV failure modes considered are fail to open or close (FTOC), fail to operate or control (FTOP), and spurious operation (SO). The component reliability estimates and the reliability data are trended for the most recent 10-year period while yearly estimates for reliability are provided for the entire study period. The following increasing trend was identified for MOVs for the most recent 10-year period: • Low-demand MOV frequency of FTOC demands (demands per reactor year). The following decreasing trends were identified for MOVs for the most recent 10-year period: • Low-demand MOV FTOC failure probability • High-demand MOV SO failure rate • Low-demand MOV frequency of FTOC events (failures per reactor year) • High-demand MOV frequency of SO events (failures per reactor year).

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

From Failure to Insight: Analyzing Disk Breakdowns in Large-Scale HPC Environments

Disk failure data provides valuable insights for preventing failures, enhancing storage robustness, guiding system design and deployment, and ensuring reliable operations at data centers. This paper introduces two disk failure datasets collected from large-scale HPC production environments over the past five years, comprising over 5,000 failure records from more than 40,000 disks. We analyzed these datasets across multiple dimensions, including temporal, spatial, and relational trends, and performed a comprehensive reliability assessment. Our analysis yielded numerous observations and insights that influence various operational aspects of HPC storage systems. We believe this study offers a holistic understanding of disk failure trends likely to interest the HPC storage community.

George, Anjus↗

Risk-Aware Measurement Synchronization and Recovery for DSSE With Heterogeneous Data Sources

Power distribution systems are increasingly integrating heterogeneous sensors with varying data reporting rates and types, which pose challenges to achieving observability at the desired temporal resolution of distribution system state estimation (DSSE). Multisensor failures caused by extreme events exacerbate these issues, introducing substantial uncertainties into DSSE. This article proposes a novel solution to these challenges by ensuring high-resolution system observability despite heterogeneous data sources and multisensor failures. First, a deep learning architecture combining long short-term memory (LSTM) and graph convolutional network (GCN) is employed to synchronize meters with different reporting rates, aiming to achieve system observability. A random-walk-model-based approach is introduced to generate pseudo-measurements while properly characterizing their uncertainties under multisensor failures. Finally, a disaster-risk-informed observability metric (RiOM) is defined to quantify the uncertainty associated with state estimation results. The proposed framework offers deeper insights into the system observability on the fly compared with conventional analysis. The effectiveness of the framework is demonstrated on an IEEE standard test case and a large-scale real-world distribution feeder in mid-Minnesota in the U.S.

97 MATHEMATICS AND COMPUTING↗

SCEPTRE: A Cyber-Physical Emulation Capability

Cyber-physical systems form a critical but vulnerable backbone to US critical infrastructure. Recent high-profile cyber-attacks have shown the need for increased assessment and hardening of these systems. However, such assessments and investigations into advanced technologies to harden these systems is difficult due to their operational nature. Instead, modeling of these systems is heavily leveraged. Investigation into these complex systems and their potential cascading failures requires comprehensive modeling of both the cyber and physical components of the system. This paper introduces SCEPTRE, an emulation capability to address this need.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Electrical Resistivity Tomography data from 2016 to 2018 at the Lower Montane site in the East River Watershed, Colorado

This dataset contains time-lapse Electrical Resistivity Tomography (ERT) data along a transect located on the northeast-facing hillslope at the lower montane site (Pumphouse site) in the upper East River Watershed. The monitoring dataset covers the period from November 2, 2016, to August 6, 2018. In addition, the archive also contains a baseline dataset from October 9, 2016. The ERT transect consisted of 128 electrodes with an electrode spacing of 1.25 m. The acquisition system was located in the middle of the transect, about 50 m on one side, and included an MPT (Multi-Phase Technologies) ERT system, a mini computer, and batteries with solar panels. Acquisition occurred daily under normal circumstances. The first 16 electrodes (from the upper end of the transect) could not be used after the cable was damaged during the 2017–2018 winter. Also, due to multiple failures in the power system, the temporal resolution of the data is much lower in 2018 compared to 2016 and 2017. The data have been processed and used in Dafflon et al., 2023, and the baseline dataset was used in Falco et al., 2019 (see reference list). This archive contains the measurements (ER.zip containing csv files) for each of the 326 acquisition times and a filtered version where only electrodes 17 to 128 are included (ERT_sm.zip containing csv files). The archive also contains the baseline dataset and two acquisitions with full reciprocals (ERT_RB.zip containing csv files), as well as all the raw MPT files (ERT_raw_MTP.zip). The geometry (electrode position and elevation) is provided in Universal Transverse Mercator (UTM) 13N Geoid2012AB in the file named ERT_Location.csv. The archive contains 1 *.csv data files, four *.zip files, and three metadata *.csv files. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Small-Signal Stability Constrained Optimal Power Flow of Inverter-Dominated Power Systems with Flexible Operation Mode Selection

Given the intermittence and low inertia nature of inverter-based resources (IBRs), modern power systems with high penetration of IBRs challenge the conventional optimal power flow (OPF) analysis and the system may experience unexpected failures if stability constraints are not incorporated. This study proposes a small-signal stability-constrained OPF (SSSC-OPF) with flexible operation mode selection between grid-forming (GFM) and grid-following (GFL) modes for IBRs to address these challenges. The approach aims to maintain system stability with a sufficient stability margin while minimizing operation costs. The effectiveness of the proposed method is validated through extensive case studies on the IEEE 14-bus system. The results demonstrate that the proposed method is able to support system-level power flow analysis, reduce generation costs, and ensure stability under various disturbances.

grid-following↗

Development of mechanistic, microstructure-informed BISON models for fission product-induced failure mechanisms in advanced nuclear fuels

The US continues to prioritize commercial deployment of advanced reactors—particularly those utilizing U-Zr metallic and TRISO particle fuels. Fission product-induced failure mechanisms for these systems include fuel–cladding chemical interaction in metallic fuel rods, which threatens cladding integrity, and Pd penetration in TRISO particles, which degrades SiC layer properties. Empirical models can be applied within the bounds of existing irradiation databases, but their utility is limited when considering new designs or investigating fundamental material behaviors. The Nuclear Energy Advanced Modeling and Simulation Program is therefore developing mechanistic models for these behaviors, which are expected to aid in fuel design, assist in development of failure mitigation strategies, and provide support for qualification and licensing. These efforts leverage multiscale capabilities to capture the microstructural features and processes that govern these behaviors. This talk summarizes past and ongoing efforts to develop and validate these models for the BISON fuel performance code.

FCCI↗

Relief Zones Enhance the Durability of Ultrathin Membranes in Electrochemical Conversion Devices

Premature failures in electrochemical conversion systems often result when membrane electrode assemblies (MEAs) use ultrathin (≤15 μm-thick) polymer electrolyte membranes, susceptible to mechanical degradation from stress concentrations arising from device-level integration. Herein, relief zones were developed to mitigate mechanical degradation by alleviating excess and nonuniform compression across active areas. Relief zones, created through ablation of carbonaceous diffusion media, enable seamless adaptation across MEA dimensions without need for hardware modifications. Demonstrated using fuel cells as a case study, accelerated stress tests revealed a 6-fold lifetime improvement (∼1500 h) compared to conventional edge-protected MEAs, decoupling device-level engineering effects from material limitations.

accelerated stress test↗

NSTX-U liquid metal core-edge facility (LMCE)

NSTX-U/LMCE will provide a unique and world-leading research facility to address the primary challenge to delivering economic and timely magnetic fusion energy, namely the need to develop a power and particle exhaust and first-wall system that can withstand very high edge heat fluxes, maximize energy confinement, and avoid the production of large masses of solid eroded first-wall material. The NSTX-U/LMCE facility will assess the ability of liquid metals (LMs) – especially liquid lithium – to provide a new boundary condition for magnetic fusion systems, to extend the lifetime of the plasma facing components (PFCs) and improve core plasma confinement. Such capability is needed to establish the basis for next-step fusion facilities including fusion pilot plants, and to maintain U.S. world leadership in core-edge integration research. NSTX-U/LMCE will leverage the ability to generate very high divertor perpendicular heat flux q⊥ ~ 100MW/m 2 , extensive diagnostics, and liquid-metal-applicable infrastructure of NSTX-U. NSTX-U/LMCE will provide access to a high-confinement plasma core with majority self-driven plasma current, the flexibility to test a range of liquid metal divertor concepts, access to a range of separatrix collisionalities (from high to very low), and the ability to controllably vary the first-wall temperature to vary the plasma- wall interaction physics on liquid lithium components. Further, NSTX-U/LMCE will utilize more reactor-relevant high-Z refractory-metal PFC substrates. With these capabilities the NSTX-U/LMCE facility will explore the full continuum of core-edge solutions ranging from high core radiated power, to conditions with radiative losses concentrated in the scrape-off layer (SOL), and ultimately low recycling conditions. The low collisionality SOL that may be accessible in the low recycling regime is relatively unexplored and will require a kinetic treatment of the edge, which can be addressed theoretically, and with experiments in LTX-β. Additional smaller-scale preparatory R&D facilities will be required to reduce the risk of premature technical/engineering failure of liquid metal systems implemented in NSTX-U. The NSTX-U/LMCE facility aligns very well with recommendations in the FESAC Long-Range Plan and NASEM Pilot Plant reports and the Bold Decadal Vision, will be unique in the world program throughout the next decade, and is garnering private company interest in utilizing NSTX-U/LMCE for development of LM PFCs.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗