Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Failure Rate”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Nuclear imaging to diagnose and correct target-driver registration at high-repetition-rate for improved reactor efficiency

Inertial Confinement Fusion produces energy from a burning plasma lasting a fraction of a nanosecond. Power plant designs based on Inertial Fusion Energy (IFE) will need to ignite targets 1-10 times a second, fired as projectiles into a chamber and delivering the driver to the target location. Driver asymmetry is known to impact ICF experiments at gain near unity and remains a candidate for primary yield degradation, and therefore fusion power plant energy output, for high-gain target designs. For a Fusion Power Plant (FPP), continuous and real-time monitoring of target performance provides an opportunity to stabilize or correct the target-driver registration. This requires x-ray and neutron imaging with a large field-of-view, sufficiently high resolution, fast analysis and to subtend a minimal solid angle. We introduce design criteria for such an imaging system that uses a coded aperture and time-gated, lens-coupled scintillators as a viable solution and outline the research steps required to field such a system. Integrating the imaging system into an IFE power plant as part of an active feedback loop could increase average power output by reducing the failure rate due to mis-aligned drivers with respect to the target.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Assessment of Crew Time for Maintenance and Repair Activities for Lunar Surface Missions

NASA is currently evaluating different methods to predict how much time crewmembers will spend conducting repair and maintenance activities on future space missions. As mission scope and spacecraft architectures change, understanding how crew repair and maintenance timelines are impacted by mission operations and technology changes is vital for future mission planning. Past work has been done using historical International Space Station (ISS) data to accurately predict crew habitation and operation timelines, resulting in the development of NASA’s Exploration Crew Time Model (ECTM). However, understanding crew maintenance and repair requirements has posed a unique challenge due to the complexity of available datasets, the probabilistic nature of sub-system failures, and the impacts of reliability growth on failure rates. This paper presents a methodology to collect and condition empirical repair and maintenance time data from available datasets, to extrapolate from that data to estimate projected maintenance and repair times for a lunar Surface Habitat (SH), and to assess how uncertainty in repair time could impact utilization time on the lunar surface. NASA ISS maintenance and crew time data are logged into two central databases: the Maintenance Data Collection (MDC) and the Operations Planning Timeline Integration System (OPTimIS). Separately, each of these two datasets capture only portions of the complete set of data required to generate an accurate assessment of crew time spent on maintenance activities at a sub-system level. To create a more useful crew time estimate for maintenance timelines, the authors developed a methodology to capture relevant data from each set and combine and utilize that data by linking crew time requirements to specific components. The authors compare the failure logs in the MDC to crew activity logs pulled from OPTimIS and then process the data to estimate required repair time for each failure and repair event. The entire maintenance activity dataset is then categorized based on the class of failed component to ensure a significant sample size for each class and accurate crew time estimates for any components lacking relevant data. This resultant component repair time data can be used in the future to generate Mean Time to Repair (MTTR) estimates and confidence intervals for each class of component based on a probabilistic distribution of documented maintenance events. These improved MTTR values can then be applied to candidate element sub-system architectures, along with component Mean Time Between Failure (MTBF) data to generate distributions for potential required system crew repair time estimates for a given mission. The authors applied these modeling methods to a case study of a crewed mission to the planned SH and produced expected corrective maintenance crew time distributions. The results produced an expected corrective maintenance crew time at over 24 hours per mission, and a maintenance crew time distribution that reflects the importance of planning for sufficient maintenance requirements each mission. Repair time distributions can then be used to develop more accurate crew schedules and to assess potential available utilization time.

Crew Time↗

Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis

HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc.

crossanalysis↗

Size of metallic and polyethylene debris particles in failed cemented total hip replacements

Reports of differing failure rates of total hip prostheses made of various metals prompted us to measure the size of metallic and polyethylene particulate debris around failed cemented arthroplasties. We used an isolation method, in which metallic debris was extracted from the tissues, and a non-isolation method of routine preparation for light and electron microscopy. Specimens were taken from 30 cases in which the femoral component was of titanium alloy (10), cobalt-chrome alloy (10), or stainless steel (10). The mean size of metallic particles with the isolation method was 0.8 to 1.0 microns by 1.5 to 1.8 microns. The non-isolation method gave a significantly smaller mean size of 0.3 to 0.4 microns by 0.6 to 0.7 microns. For each technique the particle sizes of the three metals were similar. The mean size of polyethylene particles was 2 to 4 microns by 8 to 13 microns. They were larger in tissue retrieved from failed titanium-alloy implants than from cobalt-chrome and stainless-steel implants. Our results suggest that factors other than the size of the metal particles, such as the constituents of the alloy, and the amount and speed of generation of debris, may be more important in the failure of hip replacements.

Non-NASA Center↗

Critical review and analysis of hydrogen safety data collection tools

The wider adoption of hydrogen in multiple sectors of the economy requires that safety and risk issues be rigorously investigated. Quantitative Risk Assessment (QRA) is an important tool for enabling safe deployment of hydrogen fueling stations and is increasingly embedded in the permitting process. QRA requires reliability data, and currently hydrogen QRA is limited by the lack of hydrogen specific reliability data, thereby hindering the development of necessary safety codes and standards [1]. Four tools have been identified that collect hydrogen system safety data: H2Tools Lessons Learned, Hydrogen Incidents and Accidents Database (HIAD), National Renewable Energy Lab's (NREL) Composite Data Products (CDPs), and the Center for Hydrogen Safety (CHS) Equipment and Component Failure Rate Data Submission Form. This work critically reviews and analyzes these tools for their quality and usability in QRA. It is determined that these tools lay a good foundation, however, the data collected by these tools needs improvement for use in QRA. Areas in which these tools can be improved are highlighted, and can be used to develop a path towards adequate reliability data collection for hydrogen systems.

08 HYDROGEN↗

Low Activity Waste Glass Optimization with Property Models from Machine Learning, Part 2: Experimental Validation and Active Learning

The United States Department of Energy is responsible for managing legacy nuclear waste stored in underground tanks at the Hanford Site. To treat the waste, it is planned as the current baseline to separately vitrify low-activity waste (LAW) and high-level waste fractions. Previously, machine learning (ML) based glass property models (e.g., chemical durability, viscosity, electrical conductivity and SO3 solubility) were developed with prediction uncertainties. A waste glass optimization approach was then established to enable the capability of using these ML models in LAW glass formulation. In this study, the previous ML models were first experimentally validated, and the results were incorporated back into the database to update the ML models. The updated models and formulations showed increased waste loading while reducing the failure rate, demonstrating improved predictive accuracy, reduced uncertainties, and the effectiveness of active learning in guiding high-dimensional, nonlinear LAW glass design. This represents the first experimental validation of ML based LAW glass formulation, with practical benefits such as higher waste loading, shorter mission duration, and lower operational risk.

Lu, Xiaonan (ORCID:0000000179708148)↗

Radio frequency mixing modules for superconducting qubit room temperature control systems

As the number of qubits in nascent quantum processing units increases, the connectorized RF (radio frequency) analog circuits used in first generation experiments become exceedingly complex. The physical size, cost, and electrical failure rate all become limiting factors in the extensibility of control systems. We have developed a series of compact RF mixing boards to address this challenge by integrating I/Q quadrature mixing, intermediate frequency/LO (local oscillator)/RF power level adjustments, and direct current bias fine tuning on a 40 × 80 mm 2 four-layer printed circuit board with electromagnetic interference shielding. The RF mixing module is designed to work with RF and LO frequencies between 2.5 and 8.5 GHz. The typical image rejection and adjacent channel isolation are measured to be ~27 dBc and ~50 dB. By scanning the drive phase in a loopback test, the module short-term amplitude and phase linearity are typically measured to be 5 ×10 -4 (V pp /V mean ) and 1 ×10 -3 radian (pk-pk). The operation of the RF mixing board was validated by integrating it into the room temperature control system of a superconducting quantum processor and executing randomized benchmarking characterization of single and two qubit gates. We measured a single-qubit process infidelity of 9.3(3) × 10 -4 and a two-qubit process infidelity of 2.7(1) × 10 -2 .

47 OTHER INSTRUMENTATION↗

Radiological HEPA Filter 10-year Lifetime Evaluation in Research Facilities

High-efficiency particulate air (HEPA) filters are widely employed by nuclear facilities to remove radiological particulate matter from their effluent exhaust streams. The purpose of this study is to evaluate the relationships between the 10-year HEPA filter lifetime deployment and its other performance indicators. This 10-year-long endeavor to collect and analyze data regarding the service life of HEPA filters at the Pacific Northwest National Laboratory began in 2010. A set of HEPA filters were selected and have been surveyed and analyzed at least annually to verify compliance with permit conditions. The study suggests the frequency of filter replacement should be based on the actual operational requirements, such as fume hood face velocity and/or efficiency test results, instead of on the prescribed filter “age limit” of 10 years from the date of manufacture (e.g., birth date) when operating under dry conditions. The study has now been completed, and over the past decade all the HEPA filters have been replaced, due to either technical issues as listed in this report or the previously recommended filter “age limit” of 10 years as prescribed by the oversight bodies. Experimentally determined failure rates are also determined from the data set and can be used to estimate the chances of HEPA filters surviving 15, 20, or even 30 years.

61 RADIATION PROTECTION AND DOSIMETRY↗

Quantitative Evaluation of Reliability Improvement: Case Study on a Self-healing Distribution System

In this work, we develop a methodology and tool to quantitatively evaluate the reliability of a self-healing system that considers practical distribution system features such as the distributed energy resources, microgrids, and service restoration strategies. Also, this paper addresses various practical issues when being applied to an actual Duke Energy distribution system, including the design of feasible and practical service restoration strategies that are used to identify the customer interruptions after a fault, and the incorporation of the utility’s historical reliability indices that are used to calibrate the failure rate and repair time of distribution system components such as overhead lines and underground cables. This case study demonstrates the effectiveness of the proposed method.

Dong, Jiaojiao↗

Adaptive stabilization of quantum circuits executed on unstable devices

Conventional computers have evolved to device components that demonstrate failure rates of 10 −17 or less, while current quantum computing devices typically exhibit error rates of 10 −2 or greater. This raises concerns about the reliability and reproducibility of the results obtained from quantum computers. The problem is highlighted by experimental observation that today’s NISQ devices are inherently unstable. Remote quantum cloud servers typically do not provide users with an ability to calibrate the device themselves. Using inaccurate characterization data for error mitigation can have devastating impact on reproducibility. In this study, we investigate if one can infer the critical channel parameters dynamically from the noisy binary output of the executed quantum circuit and use it to improve program stability. An open question however is how well does this methodology scale. We discuss the efficacy and efficiency of our adaptive algorithm using canonical quantum circuits such as the uniform superposition circuit. Our metric of performance is the Hellinger distance between the post-stabilization observations and the reference (ideal) distribution.

Dasgupta, Samudra↗

Remote Monitoring and Diagnostics of Pitch-Bearing Defects in an MW-Scale Wind Turbine Using Pitch Symmetrical-Component Analysis

Recently, multiple wind turbine failure databases have reviewed that the pitch system is one of the subassemblies with the highest failure rates and largest contributors to the overall downtime. Therefore, there has been an increasing interest to provide remote health monitoring for wind turbine pitch system. While most of the research articles are discussing pitch actuation system (hydraulic or electric actuator) faults only, there is very limited research on pitch-bearing-defect detection. This article provides a remote and hardware-free solution to monitor multiaxis pitch-bearing health condition called pitch symmetrical-component analysis. It leverages readily available low-resolution (100 Hz) electrical measurements, mechanical measurements, and control signals from the existing pitch control platform, and innovatively applies symmetrical-component analysis in multiphase ac system to multiaxis pitch control system and introduces multiaxis pitch-bearing degradation trending curves. This hardware-free solution can be directly applied to the existing wind turbines and successfully give the wind farm operator an early warning before multiaxis pitch bearing fails. It has been proved to be accurate, low cost, and has minimum impacts on turbine normal operation, and has been validated by field data from several North America MW-scale wind farms. This approach turns out to be the first hardware-free (no additional hardware needed) method to remotely monitor and diagnose multiaxis wind turbine pitch-bearing condition.

17 WIND ENERGY↗

Reliability Analysis of Power Grids Considering Component Failures of Variable Energy Resources

This paper proposes an improved model for the reliability assessment of power systems considering component failures of variable energy resources (VER). The inherent intermittency of VER such as solar photovoltaic (PV) and wind farms, along with their susceptibility to component failures, present significant challenges to reliable system operation. These issues, combined with power grid operation and network constraints, complicate the reliable operation of VER-integrated power systems. Here, to address these concerns, this paper introduces a reliability assessment framework that considers VER input variability, its impact on component availability, and their resulting impact on overall system reliability. Stochastic models based on discrete Markov processes are developed to incorporate variable irradiance, wind speeds, and their effects on PV and wind component failure rates. A next-event and state transition-based approach is then developed to integrate the stochastic models into a mixed-timing sequential Monte Carlo simulation framework for composite reliability assessment. Case studies on the RTS-GMLC system demonstrate the effectiveness of the proposed model in evaluating the reliability of VER-integrated systems.

Pandit, Dilip [Sandia National Laboratories (SNL-N↗

Underwater unexploded ordnance discrimination based on intrinsic target polarizabilities – A case study

Seabed unexploded ordnance that resulted partly from the high failure rate among munitions from more than 80 years ago and from decades of military training and testing of weapons systems poses an increasing concern all around the world. Although existing magnetic systems can detect clusters of debris, they are not able to tell whether a munition is still intact requiring special removal (e.g. in situ detonation) or is harmless scrap metal. The marine environment poses unique challenges, and transferring knowledge and approaches from land to a marine environment has not been easy and straightforward. On land, the background soil conductivity is much lower than the conductivity of the unexploded ordnance and the electromagnetic response of a target is essentially the same as that in free space. For those frequencies required for target characterization in the marine environment, the seawater response must be accounted for and removed from the measurements. The system developed for this study uses fields from three orthogonal transmitters to illuminate the target and four three-component receivers to measure the signal arranged in a configuration that inherently cancels the system's response due to the enclosing seawater, the sea–bottom interface and the air–sea interface for shallow deployments. The system was tested as a cued system on land and underwater in San Francisco Bay – it was mounted on a simple platform on top of a support structure that extended 1 m below and allowed the diver to place metal objects to a specific location even in low-visibility conditions. The measurements were stable and repeatable. Furthermore, target responses estimated from marine measurements matched those from land acquisition, confirming that the seawater and air–sea interface responses were removed successfully. Thirty-six channels of normalized induction responses were used for the classification, which was done by estimating the target principal dipole polarizabilities. Our results demonstrated that the system can resolve the intrinsic polarizabilities of the target, with clear distinctions between those of symmetric intact unexploded ordnance and irregular scrap metal. The prototype system was able to classify an object based on its size, shape and metal content and correctly estimate its location and orientation.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

A Multi-Objective Bayesian Optimized Human Assessed Multi-Target Generated Spectral Recommender System for Rapid Pareto Discoveries of Material Properties

Optimization for different tasks like material characterization, synthesis, and functional properties for desired applications over multi-dimensional control parameter and function spaces need a rapid strategic search through active learning. However, in all cases prior to optimization, the target material properties are assumed known and fixed, which mostly deviates from real-world scenarios in material synthesis. This can be critical for running expensive experiments on new materials, when the experimental results are fuzzy for any scientific outcomes due to improper target setting, ultimately wasting time and cost. The failure rate and cost are even higher over exploring on multi-target space, where we want to learn the pareto among multiple properties, to jointly optimize during material synthesis for desired applications. To address the challenge, here we introduce the human-operator attempt flexibility in the active learning based automated experiment framework, with generating multiple human assessed targets through a voting-based recommender system during real-time microscope measurements over the large material image space, sequentially learn/update multiple desired targets through a weighting system, and adaptively search in multiple material properties functional space for non-dominated pareto discoveries to maximize the custom structural similarity based acquisition function. We term this a multi-objective Bayesian optimized human assessed multi-target generated spectral recommender systems (MOBO-HAM-SRS). The approach has been demonstrated to peizoresponse force spectroscopy of a ferroelectric thin film, exploring with different kernels and acquisition functions. This work shows an advancement towards human-AI collaborated automated experiments, steering optimization trajectories through human overpowering AI at the early stage when uncertainty is high and AI overpowering human at the later stage with rapid exploration towards optimal goal, following human-assessed multiple targets properties.

Biswas, Arpan↗

Understanding GPU Memory Corruption at Extreme Scale: The Summit Case Study

GPU memory corruption and in particular double-bit errors (DBEs) remain one of the least understood aspects of HPC system reliability. Albeit rare, their occurrences always lead to job termination and can potentially cost thousands of node-hours, either from wasted computations or as the overhead from regular checkpointing needed to minimize the losses. As supercomputers and their components simultaneously grow in scale, density, failure rates, and environmental footprint, the efficiency of HPC operations becomes both an imperative and a challenge. We examine DBEs using system telemetry data and logs collected from the Summit supercomputer, equipped with 27,648 Tesla V100 GPUs with 2nd-generation high-bandwidth memory (HBM2). Using exploratory data analysis and statistical learning, we extract several insights about memory reliability in such GPUs. We find that GPUs with prior DBE occurrences are prone to experience them again due to otherwise harmless factors, correlate this phenomenon with GPU placement, and suggest manufacturing variability as a factor. On the general population of GPUs, we link DBEs to short- and long-term high power consumption modes while finding no significant correlation with higher temperatures. We also show that the workload type can be a factor in memory’s propensity to corruption.

Oles, Vlad↗

North Carolina Water Utility Builds Resilience with Distributed Energy Resources

As the frequency and duration of grid outages increase, backup power systems are becoming more important for ensuring that critical infrastructure continues to provide essential services. Most facilities rely on diesel generators, which may be ineffective during long outages owing to limited fuel supplies and high generator failure rates. Distributed energy resources such as solar, storage, and combined-heat-and-power systems, coupled with on-site biofuel production, offer an alternative source of on-site generation that can provide both cost savings and resilience (i.e., the ability to respond to catastrophic events with longer-term consequences). A mixed-integer linear program minimizes costs and maximizes resilience at a wastewater treatment plant in Wilmington, North Carolina. We find that the plant can reduce life-cycle energy costs by 3.1% through the installation of a hybrid combined-heat-and-power, photovoltaic, and storage system. When paired with existing diesel generators, this system can sustain full load for seven days while saving $664,000 over 25 years and reducing diesel fuel use by 48% compared with the diesel-only solution. This analysis informed a decision by the Cape Fear Public Utility Authority to allocate funds for the implementation of a combined-heat-and-power system at the wastewater treatment plant in fiscal year 2023. Finally, the benefits of deploying hybrid combined-heat-and-power technologies and the utilization of on-site biofuel production extend, on a national scale, to thousands of wastewater treatment facilities and other types of critical infrastructure.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Benchmark for Fuel Shuffling and Depletion for Pebble-Bed Reactors

Pebble bed reactors have specific operational characteristics when their fuel-cycle and fueling operations are considered. They are specifically distinguished by other type of nuclear reactor designs by their online fuel recycling scheme, where the fuel elements that have not yet reached discharge burnup can be reloaded and recycled continuously during normal operation. The fuel in a pebble bed reactor is not stationary and stochastically moves through the core once or several times during its lifetime, which allows them to operate without requiring a large excess reactivity hold for the burnup. However, this characteristic of pebble bed reactors introduces challenges in simulation, as each pebble can take many different trajectories through the core, its composition depends on the details of the irradiation history that is unique to its aggregated path through the core. For predicting the safety performance characteristics, such as source term, maximum fuel temperatures and fuel failure rates, etc., it is important to accurately incorporate the movement of pebbles through the core during their lifetime in a multi-physics simulation together with other phenomena. The equilibrium core analysis for pebble bed reactors are performed with multi-physics tools including fuel depletion in a multi pass reload coupled to the fuel movement. Currently, there are only a few legacy multi-physics simulation tools that can implement the pebble flow characteristics and perform equilibrium core analysis for pebble bed reactors. However, there are development efforts on-going under Department of Energy's Nuclear Energy Advanced Modelling and Simulation program and also in private industry for including these capabilities into their modelling and simulation tools. Any new development in the modelling and simulation tools needs to be validated by using tools such as experiments, analytical solutions or code-to-code benchmarks. In this work, a code-to-code benchmark for the equilibrium core analysis capability of pebble bed reactors was developed. Multiple cases were identified to capture different fuel cycle strategies that can be used in PBRs. The results of each case are presented in terms of overall equilibrium core characteristics: the discharge burnup; spatial burnup distribution; spatial isotopic distributions; axial and radial neutron flux distributions and power history of fuel elements per pass through the core for both a prototypical pebble bed High Temperature Gas-cooled Reactor and a prototypical pebble bed Fluoride-salt cooled High temperature Reactor.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗