Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “failure analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

An ontology-based fault generation and fault propagation analysis approach for safety-critical computer systems at the design stage

Abstract Fault propagation analysis is a process used to determine the consequences of faults residing in a computer system. A typical computer system consists of diverse components (e.g., electronic and software components), thus, the faults contained in these components tend to possess diverse characteristics. How to describe and model such diverse faults, and further determine fault propagation through different components are challenging problems to be addressed in the fault propagation analysis. This paper proposes an ontology-based approach, which is an integrated method allowing for the generation, injection, and propagation through inference of diverse faults at an early stage of the design of a computer system. The results generated by the proposed framework can verify system robustness and identify safety and reliability risks with limited design level information. In this paper, we propose an ontological framework and its application to analyze an example safety-critical computer system. The analysis result shows that the proposed framework is capable of inferring fault propagation paths through software and hardware components and is effective in predicting the impact of faults.

97 MATHEMATICS AND COMPUTING↗

Determining Stress in Metallic Conducting Layers of Microelectronics Devices Using High Resolution Electron Backscatter Diffraction and Finite Element Analysis

Delayed failure due to stress voiding is a concern with some aging microelectronics, as these voids can grow large enough to cause an open circuit. Local measurements of stress in the metallic layers are crucial to understanding and predicting this failure, but such measurements are complicated by the fact that exposing the aluminum conducting lines will relieve most of their stress. In this study, we instead mechanically thin the device substrate and measure distortions on the thinned surface using high resolution electron backscatter diffraction (HREBSD). These measurements are then related to the stresses in the metallic layers through elastic simulations. This study found that in legacy components that had no obvious voids, the stresses were comparable to the theoretical stresses at the time of manufacture (≈300 MPa). Distortion fields in the substrate were also determined around known voids, which may be directly compared to stress voiding models. In conclusion, the technique presented here for stress determination, HREBSD coupled with finite element analysis to infer subsurface stresses, is a valuable tool for assessing failure in layered microelectronics devices.

HREBSD↗

Failure Mode and Effects Analysis for a Photovoltaic Inverter

While PV panel reliability continues to increase, PV inverters become the limiting factor for PV system reliability. Consequently, it is critical to have a generic tool from a third party for PV inverter reliability assessment to help 1) utilities/PV farm operators schedule maintenance in advance, and 2) inverter developers improve the next-generation design. However, these two things cannot be accomplished without first understanding the reasons behind inverter failure. Following this idea, as the first step, it is essential to identify and investigate the most failure-prone components within a PV inverter system. After all, any system is only as reliable as the components that are contained within it. This motivates the failure mode and effects analysis (FMEA) work presented for this workshop. The FMEA is conducted as follows: first, the overview of the methodology on the development of the FMEA is presented; then, based on a top-down approach starting from the PV inverter system, critical inverter components with high failure rates are identified and summarized; afterward, a thorough FMEA study at a component-level is performed and its results, including failure modes, failure mechanisms, and critical stressors, are tabulated; finally, according to three rankings (chance of occurrence, severity of occurrence, and ease of detection prior to failure) for each failure mechanism provided by the FMEA, risk priority numbers are calculated and the failure mechanisms along with the critical stressors are ranked in terms of their potentially detrimental effect on the PV inverter.

Brown, Buck↗

Facilitating Data Collection of Maintenance Events to Populate the Hydrogen Component Reliability Database (HyCReD)

The Hydrogen Component Reliability Database (HyCReD) is a collaborative project between the National Renewable Energy Laboratory, the University of Maryland, and hydrogen stakeholders to improve safety and reliability for hydrogen facilities by implementing component reliability data taxonomies that support hydrogen infrastructure failure rate analysis. The project aims to quantify failure rates of hydrogen components through high-quality data collection and analysis on root causes and maintenance needed. HyCReD provides a common database for cataloging hydrogen component failures which exists for reliability research in many other mature industries [2]. The database fills a gap for the hydrogen community by providing a scientifically rigorous approach to quantitative risk assessment (QRA), prognostic health management (PHM), and reliability-centered maintenance (RCM) analysis. High level results will be aggregated and anonymized to protect company sensitive information; detailed results will be used to help address issues of hydrogen components. These advanced analytics will support accelerated deployment of hydrogen infrastructure by enabling better: design and safety of projects (safety codes and standards development), infrastructure reliability and cost (component failure rates, maintenance protocols), and component R&D needs (robust supply chain). A key to a successful HyCReD implementation is facilitating the ease of reporting and data quality in the database that can be used for analysis. Maintenance data was a previously identified gap in initial efforts to populate and validate the database taxonomies [3]. Collection of maintenance data will be instrumental in identifying failure modes and rates, identifying incipient component failures or reduced performance, cataloging best practices for maintenance routines and methods for prognostic health management, and quantifying the risk and effect of different failure modes. Several key priorities are identified for streamlined data collection to achieve quality and detailed failure data: Applicability, Ease of Use, Accessibility, and Information Security. The HyCReD team has now begun deployment of the database to several companies and groups that have signed non-disclosure agreements to facilitate the data collection of failures in industry hydrogen refueling station infrastructure. This paper will provide an update into the process of HyCReD deployment including the development of a coding guide for facility personnel to reference and ensure data quality and consistency from one station to another as well as implementation of contextually dependent data fields of system taxonomy and formatted entries to provide ease of use. The goal is to communicate the lessons learned from the roll-out to technicians and engineers in the field, and the addition of need for high level of security to protect all stakeholders.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

General Failure Modes and Effects Analysis for Accelerator and Detector Magnet Design at JLab

The aim of this article is to develop a risk management procedure, which could be applied to the magnet design process, for both superconducting and normal magnets at the Jefferson Laboratory (JLab). This procedure allowed us to identify the key risks at each of the critical phases of design and propose procedures, tests, and checks to mitigate each risk. In this article, we present a qualitative and quantitative risk management procedure commonly referred to a “failure modes and effects analysis.” As part of this procedure, we calculated a risk priority number (RPN) for each activity of the process, identified the most critical activities and proposed mitigation activities, which in turn resulted in a revised RPN. Additionally, another benefit of this procedure was the identification of appropriate “control and hold” points within the design process, which allowed one to review and approve a particular outcome before proceeding to the next sequential activity.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Experimental and computational analysis of bending fatigue failure in chopped carbon fiber chip reinforced composites

With a better balance among good mechanical performance, high freedom of design, and low material and manufacturing cost, chopped carbon fiber chip reinforced sheet molding compound (SMC) composites show great potential in different engineering applications. Here in this paper, bending fatigue behaviors of SMC composites considering the heterogeneous fiber orientation distributions have been thoroughly investigated utilizing both experimental and computational methods. First, four-point bending fatigue tests are performed with designed SMC composites, and the local modulus is adopted as a metric to represent the local fiber orientation of two opposing sides. Interestingly, SMC composites with and without large discrepancy in local modulus of opposing sides show different fatigue behaviors. Interrupted tests are conducted to explore the bending fatigue failure mechanism, and the damage processes of valid specimens are also closely examined. We find that the fatigue failure of SMC composites under four-point bending is governed by crack propagation instead of crack initiation. Because of this, the heterogeneous local fiber orientations of both sides of the specimen influence fatigue life. The microstructure of the lower side shows a direct influence while that of the upper side also exhibiting influence which becomes more prominent for high cycle fatigue cases. Furthermore, a hybrid micro–macro computational model is proposed to efficiently study the cyclic bending behavior of SMC composites. The region of interest is reconstructed with a modified random sequential absorption algorithm to conserve all the microstructural details including the heterogeneous fiber orientation, while the rest of the regions are modeled as homogenized macro-scale continua. Combined with a framework to capture the progressive fatigue damage under cyclic bending, the bending fatigue behaviors of SMC composites are accurately captured by the hybrid computational model comparing with our experimental analysis.

36 MATERIALS SCIENCE↗

Numerical and experimental analysis of mechanically induced failure in electric vehicle battery modules

Mitigating thermal runaway and cell-to-cell propagation is essential for improving the safety of electric and hybrid vehicles. Enhancing digital twin capabilities to predict battery mechanical abuse is particularly critical for automotive and aerospace applications, where crashworthiness is a key concern. Understanding failure conditions and propagation in battery modules during mechanical abuse is complex due to interactions between structural deformation, heat transfer, electrochemical processes, exothermic reactions and mechanical fracture. While prior studies have focused on modeling cell-level behavior, extending these models to module or pack level is necessary for a system level understating of electric vehicle safety. This study develops coupled large deformation finite element models that simultaneously solve for electrochemistry, material failure, internal short circuit and thermal runaway propagation. The models account for mechanical and thermal interactions between lithium-ion cells and other battery components while the contact interfaces are evolving with time. Model-predicted voltage, temperature and force responses are compared with experimental data for validation. The results demonstrate that the approach captures key failure mechanisms, including thermal propagation through heat transfer, electrical propagation from short circuits in parallel-connected cells, and mechanical propagation via penetration and crack formation. These findings show that computational models are valuable tools for understanding battery module failure and providing insight that can reduce the need for extensive experimental testing.

25 ENERGY STORAGE↗

Setting Priorities for Photovoltaic Reliability Research Using Criticality Analysis

A forward-looking research opportunity number (RON) is defined for photovoltaic reliability researchers. The RON enables researchers to prioritize their efforts toward the highest impact. For a given degradation mode, the RON is based on three factors: the effect on levelized cost of electricity, the susceptibility of future module products, and the maturity of accelerated tests that can detect and quantify the mode. Reporting bias is avoided because the RON does not rely on polls. The RON is derived for three example cases: light and elevated temperature degradation, backsheet cracking, and antireflective coating abrasion. Finally, these examples demonstrate that targeted research has reduced the risk for these modes over the last several years.

14 SOLAR ENERGY↗

Safety Hazards of Batteries and Hydrogen Storage Systems in Proximity

This report addresses the safety concerns and mitigations for battery failures and their impact on hydrogen storage systems. Through an analysis of failure modes, this report highlights the risks posed by thermal runaway and chemical emissions caused by batteries. Although rare, battery thermal failure events may prompt the opening of the relief valve on the hydrogen tank. Strategies such as battery management systems, thermal management systems, and multiple thermally activated pressure relief devices can mitigate these risks. Potential simulations and experiments to better quantify the unique risks posed by lithium-ion batteries near tanks are suggested. Improving safety standards will enable integration of batteries and hydrogen storage systems in various energy storage technologies.

08 HYDROGEN↗

Orchestrating Fault Prediction with Live Migration and Checkpointing

Checkpoint/Restart (C/R) is widely used to provide fault tolerance on High-Performance Computing (HPC) systems. However, Parallel File System (PFS) overhead and failure uncertainty cause significant application overhead. This paper develops an adaptive multi-level C/R model that incorporates a failure prediction and analysis model, which orchestrates failure prediction, checkpointing, checkpoint frequency, and proactive live migration along with the additional benefit of Burst Buffers (BB). It effectively reduces the overheads due to failures, checkpointing, and recovery. Simulation results for the Summit supercomputer yield a reduction of ~20%-86% in application overhead due to BBs, orchestrated failure prediction, and migration. We also observe a ~29% decrease in checkpoint writes to BBs, which can increase the longevity of the BB storage devices.

Behera, Subhendu↗

Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis

HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc.

crossanalysis↗

Hazard analysis for identifying common cause failures of digital safety systems using a redundancy-guided systems-theoretic approach

Replacing the existing aging analog instrumentation and control (I&C) systems with modern safety control and protection, digital technology offers one of the foremost means of performance improvements and cost reductions for the existing nuclear power plants (NPPs). However, the qualification of digital I&C systems remains a challenge, especially considering the issue of software common-cause failures (CCFs), which are difficult to address. With the application and upgrades of advanced digital I&C systems, software CCFs have become a potential threat to plant safety because most redundant designs use similar digital platforms or software in the operating and application systems. With complex designs of multilayer redundancy to meet the single-failure criterion, digital I&C safety systems (e.g., engineered safety-features actuation system [ESFAS]) are of a particular concern in the U.S. Nuclear Regulatory Commission (NRC) licensing procedures. Here, this paper applies a modularized approach to conduct redundancy-guided systems-theoretic hazard analysis for an advanced digital ESFAS with multilevel redundancy designs. Systematic methods and risk-informed tools are incorporated to address both hardware and software CCFs, which provide guidance to eliminate the causal factors of potential single points of failure in the design of digital safety systems in advanced plant designs.

42 ENGINEERING↗

Advanced Power Electronics and Electric Machines

The advanced power electronics and electric machines (APEEM) research group at the National Renewable Energy Laboratory (NREL) has developed world-class experimental and modeling capabilities for designing and evaluating efficient and reliable power electronics and electric machines thermal management systems. They also design, fabricate and characterize advanced power electronics packaging, and are developing state-of-health monitoring techniques. These researchers deliver safe, reliable, high performing, power-dense components that allow seamless integration between renewable energy sources, electric transportation, and the grid, helping to make widespread electric vehicle (EV) adoption and greenhouse gas emissions reduction more feasible. This document outlines the group's major capabilities in the areas of power electronics; module development and characterization; thermal modeling and management; thermomechanical reliability analysis of devices, modules, inverters/converters, and electric machines; physics-of-failure-based reliability analysis; and microelectronics. It also overviews the group's state-of-the-art equipment for fluid-based thermal management; thermal measurement & characterization; thermomechanical reliability analysis; micro- and power electronics measurement & characterization; and prototype fabrication, as well as the group's world-class modeling and simulation capabilities.

advanced gate drivers↗

Development and Integration of a Stochastic Clad Damage Propagation Model into PRONGHORN-SC Subchannel Analysis Code

The failure of fuel pins in nuclear reactors is intrinsically stochastic. Typically, a combination of variation in manufacturing that affects the material characteristics and the fuel assembly dimensions, variation in operating conditions, such as local power, coolant flow rate, and irradiation induced changes in material properties lead to a large uncertainty in failure margin of the fuel pins. Failure, therefore, may occur in exceptional pins with adverse combinations of these variations. Upon a metal fuel pin (U-Pu-Zr/HT9) failure, depressurization of the fuel pin takes place by release of fission gas, liquid sodium bond, and potentially solid fuel particles or molten/eutectic fuel droplets through the hole in cladding. The effect of a fission gas jet on neighbor fuel pins and possible propagation of a clad damage during normal operation was studied experimentally in 1970s and it was found that the post-failure fission gas jet insulates the jet impingement area of the target fuel pin surface and could increase the target pin’s surface temperature by as much as 100 – 200 K during the failed pin depressurization. It was concluded that the effect should not lead to fuel pin failure propagation during normal operation. In accident scenarios of sodium and lead fast reactors such as Unprotected Loss-Of-Flow (ULOF) or Unprotected Transient Over Power (UTOP), the fuel pins can be subjected to higher clad temperatures and fuel pin pressures or fuel clad mechanical/chemical interaction where thermal creep margin becomes significantly lower compared to the normal operation conditions. Therefore, possible stochastic failure and the post-failure fission gas/fuel jet impingement could be critical in order to predict fuel pin failure propagation. Pin depressurization due to fission gas release may degrade the heat transfer by formation of a gas blanket on a neighboring pin surface, which is a local phenomenon, and by causing coolant flow deceleration and starvation, which could affect a surrounding region as well. Furthermore, the potential presence of solid fuel particles or molten fuel at the time of clad failure could boost post-failure jet induced degradation even further. The present study models the U-Pu-Zr/HT9 metal fuel pin failure and stochastic clad damage propagation by biased sampling based on a Cumulative Damage Fraction (CDF) type clad failure criterion and the normal distribution of fuel failure probability density as a function of logarithm of Cumulative Damage Fraction. In addition, the effect of post-failure fission gas jet on heat transfer degradation is modeled for the target pins. This model is called stochastic Clad Damage Propagation (CDAP). The CDAP model is now fully integrated into developmental version of PRONGHORN-SC subchannel analysis code, allowing for modeling local failures and its propagation potential. Section 2 describes the components of the CDAP models. Section 3 describes the model implementation to PRONGHORN-SC and input specifications. Section 4 describes the CDAP model validation coupled to PRONGHORN-SC.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Floating Wind Array Ontology and Modeling Framework

While there are many tools for designing and modeling a single floating turbine, array level design and modeling has much more to consider. Designing floating wind arrays requires a coupled approach considering many variables, from bathymetry to installation and maintenance to failure and risk analysis. With all of these considerations, an array-level modeling tool is needed to quickly evaluate array designs. The Floating Array Model (FAModel) tool developed at the National Renewable Energy Laboratory was created to fill this gap in low-fidelity array modeling. FAModel is a python framework created to streamline holistic low-fidelity floating wind modeling for array-level analysis. FAModel integrates site data and models with a variety of open-source modeling tools developed by NREL, including FLORIS, RAFT, MoorPy, and anchor capacity models. The integration of these tools allows users to quickly and holistically design an array by considering forces, area analysis, visualization, annual energy production, failure modeling, and component costs.

17 WIND ENERGY↗

Fitness For Service Assessment of a Corroded Heat Exchanger

Within the Fermi National Accelerator complex, there exist various water systems that support accelerator operations. One of these systems is extremely vital to the operation of the machine; that is the cooling system. The cooling system consists of nine relatively large heat exchangers that take untreated pond water and use it to cool the process fluid that further cools machine components. Over the 30 years these heat exchangers have been in operation, they have undergone significant material loss on the channels. This material loss, due to various forms of corrosion such as galvanic and microbiologically influenced corrosion (MIC) and possibly others, has deteriorated more than 80% of the nominal wall thickness of some of the exchangers and placed them in a questionable state. ASME FFS-1 (API 579) has been applied to address the condition of the heat exchangers due to their noncompliance with the governing code, BPVC Sec. VIII Div. 1. The assessments encompassed ASME FFS-1 parts 4: General Metal Loss and 9: Crack Like Flaw using level 1, 2, and 3 analysis techniques based on inspection data obtained by API 510 inspections. Level 1 and 2 assessments were deemed unfit for the corroded regions due to their location relative to a major structural discontinuity (channel to tube-sheet joint), so a level 3 analysis was conducted according to ASME Sec. VIII Div. 2 (design by analysis) rules for pressure vessels. Supplemental information included pond water tests to determine an accurate future corrosion allowance due to lacking inspection history. A leak before break (LBB) route was chosen to evaluate the possibility of leaking prior to the onset of failure. The analysis of one heat exchanger shows that the possibility the channel will develop a pinhole leak over 2.5 more years of operation should not be overlooked, but burst was unlikely from operation. The use of fracture mechanics show, that if a through-wall crack were to develop, it would not propagate further than the channel geometry and cause a leak not greater than 35 GPM. Using ASME Section XI Code Case N-705-1, allowing us to operate with a leak until the next outage given certain operating conditions and developing a leak mitigation procedure, this heat exchanger is deemed fit-for-service.

Humenik, Alex [Fermilab]↗

Nanometer Scale Imaging to Develop Quantitative Descriptors of Bipolar Membrane Junction Structure

Swings in pH can be achieved by electrically polarizing a bipolar membrane (BPM) to drive water dissociation at the BPM junction for electrochemical conversion and separation processes. BPM junction design is critical to tailor performance for specific applications; however, characterization techniques capable of resolving the nanometer scale physical structure of the junction are limited. We present sample preparation, imaging, and analysis workflows that are adaptable to a variety of BPM junction architectures. Atomic force microscopy produces BPM junction images with nanometer scale lateral resolution for samples with and without a graphene oxide water dissociation catalyst in the junction. Subsequent image segmentation and analysis quantify line edge roughness and catalyst layer thickness as descriptors of junction structure. Comparison of pre- and post-electrodialysis junctions suggests electric field-induced alignment of catalyst particles during electrodialysis. This characterization workflow can inform manufacturing protocols, computational modeling, and failure mode analysis for next-generation BPMs.

97 MATHEMATICS AND COMPUTING↗